Hadoop deployment and performance on Gordon data intensive supercomputer

The Hadoop framework is extensively used for scalable distributed processing of large datasets. This extended abstract provides information on the optimization of the Hadoop deployment on the Gordon data intensive supercomputer, at the San Diego Supercomputer Center (SDSC) at the University of California San Diego, using the myHadoop software. The details of the system configuration, the storage and network options (1 Gig-E, IPOIB, and UDA), tuning options considered, results using the TestDFSIO, TeraSort benchmarks, and bulk copy tests with distcp are presented in this extended abstract.