• DocumentCode
    2791191
  • Title

    Accelerating Checkpoint Operation by Node-Level Write Aggregation on Multicore Systems

  • Author

    Ouyang, Xiangyong ; Gopalakrishnan, Karthik ; Panda, Dhabaleswar K.

  • Author_Institution
    Dept. of Comput. Sci. & Eng., Ohio State Univ., Columbus, OH, USA
  • fYear
    2009
  • fDate
    22-25 Sept. 2009
  • Firstpage
    34
  • Lastpage
    41
  • Abstract
    Clusters and applications continue to grow in size while their mean time between failure (MTBF) is getting smaller. Checkpoint/restart is becoming increasingly important for large scale parallel jobs. However, the performance of the checkpoint/restart mechanism does not scale well with increasing job size due to constraints within the file system. Furthermore, with the advent of multi-core architecture, the situation is aggravated due to larger number of processes running on the same node, trying to checkpoint simultaneously. This results in increased number of file writes at the time of checkpointing which leads to performance degradation. As a result, deployment of checkpoint/restart mechanisms for large scale parallel applications is limited. In this work, we explore the checkpoint/restart mechanism in MVAPICH2, which uses BLCR as the checkpointing library. Our profiling of the checkpoints for the NAS parallel benchmarks revealed a large number of small file writes interspersed with large writes. Based on these observation we propose to optimize checkpoint creation by classifying checkpoint file writes into small writes, medium writes and large writes based on their size of data to write, and use write aggregation to optimize the small and medium writes. At the aggregation threshold of 512 KB, the implementation of our design in BLCR shows improvements from 27% to 32% over the original BLCR in terms of time cost to checkpoint an MPI application.
  • Keywords
    checkpointing; message passing; multiprocessing systems; parallel processing; MPI application; MVAPICH2; NAS parallel benchmarks; checkpoint operation; checkpoint/restart mechanism; checkpointing library; file system; mean time between failure; multicore architecture; multicore systems; node-level write aggregation; parallel jobs; performance degradation; Acceleration; Application software; Checkpointing; Computer science; File systems; Large-scale systems; Libraries; Multicore processing; Parallel processing; Sun;
  • fLanguage
    English
  • Publisher
    ieee
  • Conference_Titel
    Parallel Processing, 2009. ICPP '09. International Conference on
  • Conference_Location
    Vienna
  • ISSN
    0190-3918
  • Print_ISBN
    978-1-4244-4961-3
  • Electronic_ISBN
    0190-3918
  • Type

    conf

  • DOI
    10.1109/ICPP.2009.73
  • Filename
    5361799