• DocumentCode
    3717253
  • Title

    Record-aware compression for big textual data analysis acceleration

  • Author

    Dapeng Dong;John Herbert

  • Author_Institution
    Mobile and Internet Systems Laboratory, University College Cork. Ireland
  • fYear
    2015
  • Firstpage
    1183
  • Lastpage
    1190
  • Abstract
    Big data analysis technologies are becoming more widely used in industry. The ever-increasing data volume, however, puts data analytic platforms such as Hadoop under constant pressure. Several compression methods have been made available on the Hadoop platform to effectively reduce data size and efficiently deliver data between cluster nodes. In the Hadoop context, compressed data can be categorized as splittable or non-splittable. Working with non-splittable data conflicts with the goal of parallelism. In addition, the current realization of splittable data by indexing is potentially harmful to the data locality property. To this end, we introduce the Record-aware Compression (RaC) scheme that makes the compressed contents splittable, uses a lightweight Hadoop Record Reader, and preserves the parallelism and data locality properties as much as possible. We evaluate RaC using a set of classical MapReduce jobs with a collection of well-known datasets from companies such as Google, Yahoo!, and Amazon. The experimental results show an average 24% improvement on analysis performance and up to 75% data size reduction.
  • Keywords
    "Context","Compressors","Indexing","Distributed databases","Dictionaries","Big data","Data analysis"
  • Publisher
    ieee
  • Conference_Titel
    Big Data (Big Data), 2015 IEEE International Conference on
  • Type

    conf

  • DOI
    10.1109/BigData.2015.7363872
  • Filename
    7363872