• DocumentCode
    3657145
  • Title

    Model-Driven Autoscaling for Hadoop Clusters

  • Author

    Anshul Gandhi;Parijat Dube;Andrzej Kochut;Li Zhang

  • fYear
    2015
  • fDate
    7/1/2015 12:00:00 AM
  • Firstpage
    155
  • Lastpage
    156
  • Abstract
    In this paper, we present the design and implementation of a model-driven auto scaling solution for Hadoop clusters. We first develop novel performance models for Hadoop workloads that relate job completion times to various workload and system parameters such as input size and resource allocation. We then employ statistical techniques to tune the models for specific workloads, including Terasort and K-means. Finally, we employ the tuned models to determine the resources required to successfully complete the Hadoop jobs as per the user-specified response time SLA. We implement our solution on an Open Stack-based cloud cluster running Hadoop. Our experimental results across different workloads demonstrate the auto scaling capabilities of our solution, and enable significant resource savings without compromising performance.
  • Keywords
    "Data models","Load modeling","Dynamic scheduling","Resource management","Monitoring","Cloud computing","Training data"
  • Publisher
    ieee
  • Conference_Titel
    Autonomic Computing (ICAC), 2015 IEEE International Conference on
  • Type

    conf

  • DOI
    10.1109/ICAC.2015.50
  • Filename
    7266955