DocumentCode
3717253
Title
Record-aware compression for big textual data analysis acceleration
Author
Dapeng Dong;John Herbert
Author_Institution
Mobile and Internet Systems Laboratory, University College Cork. Ireland
fYear
2015
Firstpage
1183
Lastpage
1190
Abstract
Big data analysis technologies are becoming more widely used in industry. The ever-increasing data volume, however, puts data analytic platforms such as Hadoop under constant pressure. Several compression methods have been made available on the Hadoop platform to effectively reduce data size and efficiently deliver data between cluster nodes. In the Hadoop context, compressed data can be categorized as splittable or non-splittable. Working with non-splittable data conflicts with the goal of parallelism. In addition, the current realization of splittable data by indexing is potentially harmful to the data locality property. To this end, we introduce the Record-aware Compression (RaC) scheme that makes the compressed contents splittable, uses a lightweight Hadoop Record Reader, and preserves the parallelism and data locality properties as much as possible. We evaluate RaC using a set of classical MapReduce jobs with a collection of well-known datasets from companies such as Google, Yahoo!, and Amazon. The experimental results show an average 24% improvement on analysis performance and up to 75% data size reduction.
Keywords
"Context","Compressors","Indexing","Distributed databases","Dictionaries","Big data","Data analysis"
Publisher
ieee
Conference_Titel
Big Data (Big Data), 2015 IEEE International Conference on
Type
conf
DOI
10.1109/BigData.2015.7363872
Filename
7363872
Link To Document