DocumentCode
2194324
Title
MapReduce for Data Intensive Scientific Analyses
Author
Ekanayake, Jaliya ; Pallickara, Shrideep ; Fox, Geoffrey
Author_Institution
Dept. of Comput. Sci., Indiana Univ. Bloomington, Bloomington, IN
fYear
2008
fDate
7-12 Dec. 2008
Firstpage
277
Lastpage
284
Abstract
Most scientific data analyses comprise analyzing voluminous data collected from various instruments. Efficient parallel/concurrent algorithms and frameworks are the key to meeting the scalability and performance requirements entailed in such scientific data analyses. The recently introduced MapReduce technique has gained a lot of attention from the scientific community for its applicability in large parallel data analyses. Although there are many evaluations of the MapReduce technique using large textual data collections, there have been only a few evaluations for scientific data analyses. The goals of this paper are twofold. First, we present our experience in applying the MapReduce technique for two scientific data analyses: (i) high energy physics data analyses; (ii) K-means clustering. Second, we present CGL-MapReduce, a streaming-based MapReduce implementation and compare its performance with Hadoop.
Keywords
data analysis; parallel algorithms; parallel programming; pattern clustering; physics computing; CGL-mapreduce technique; K-means clustering; data intensive scientific analyses; high energy physics data analyses; parallel algorithm; parallel programming; Astronomy; Biology computing; Clustering algorithms; Data analysis; Hardware; High energy physics instrumentation computing; Large Hadron Collider; Quality of service; Robustness; Scalability; MapReduce; Message passing; Parallel processing; Scientific Data Analysis;
fLanguage
English
Publisher
ieee
Conference_Titel
eScience, 2008. eScience '08. IEEE Fourth International Conference on
Conference_Location
Indianapolis, IN
Print_ISBN
978-1-4244-3380-3
Electronic_ISBN
978-0-7695-3535-7
Type
conf
DOI
10.1109/eScience.2008.59
Filename
4736768
Link To Document