• DocumentCode
    669932
  • Title

    HcBench: Methodology, development, and characterization of a customer usage representative big data/Hadoop benchmark

  • Author

    Saletore, Vikram A. ; Krishnan, Karthikeyan ; Viswanathan, V. ; Tolentino, Matthew E.

  • Author_Institution
    Intel Corp., USA
  • fYear
    2013
  • fDate
    22-24 Sept. 2013
  • Firstpage
    77
  • Lastpage
    86
  • Abstract
    Big Data analytics using Map-Reduce over Hadoop has become a leading edge paradigm for distributed programming over large server clusters. The Hadoop platform is used extensively for interactive and batch analytics in ecommerce, telecom, media, retail, social networking, and being actively evaluated for use in other areas. However, to date no industry standard or customer representative benchmarks exist to measure and evaluate the true performance of a Hadoop cluster. Current Hadoop micro-benchmarks such as HiBench-2, GridMix-3, Terasort, etc. are narrow functional slices of applications that customers run to evaluate their Hadoop clusters. However, these benchmarks fail to capture the real usages and performance in a datacenter environment. Given that typical datacenter deployments of Hadoop process a wide variety of analytic interactive and query jobs in addition to batch transform jobs under strict Service Level Agreement (SLA) requirements, performance benchmarks used to evaluate clusters must capture the effects of concurrently running such diverse job types in production environments. In this paper, we present the methodology and the development of a customer datacenter usage representative Hadoop benchmark "HcBench" which includes a mix of large number of customer representative interactive, query, machine learning, and transform jobs, a variety of data sizes, and includes compute, storage 110, and network intensive jobs, with inter-job arrival times as in a typical datacenter environment. We present the details of this benchmark and discuss application level, server and cluster level performance characterization collected on an Intel Sandy Bridge Xeon Processor Hadoop cluster.
  • Keywords
    data analysis; distributed processing; file servers; learning (artificial intelligence); multiprocessing systems; query processing; GridMix-3; Hadoop microbenchmarks; HcBench; HiBench-2; Intel Sandy Bridge Xeon Processor Hadoop cluster; Map-Reduce; SLA; Terasort; analytic interactive jobs; big data analytics; customer datacenter usage representative Hadoop benchmark; customer usage representative big data; distributed programming; interjob arrival times; machine learning; network intensive jobs; query jobs; server clusters; service level agreement requirements; transform jobs; Artificial neural networks; Benchmark testing; Lead; Big Data; Hadoop Benchmark; Map-Reduce; Performance; Workload Characterization;
  • fLanguage
    English
  • Publisher
    ieee
  • Conference_Titel
    Workload Characterization (IISWC), 2013 IEEE International Symposium on
  • Conference_Location
    Portland, OR
  • Print_ISBN
    978-1-4799-0553-9
  • Type

    conf

  • DOI
    10.1109/IISWC.2013.6704672
  • Filename
    6704672