• DocumentCode
    451087
  • Title

    A Checkpointing Strategy for Scalable Recovery on Distributed Parallel Systems

  • Author

    Naik, Vijay K. ; Midkiff, Samuel P. ; Moreira, Jose E.

  • Author_Institution
    IBM T. J. Watson Research Center
  • fYear
    1997
  • fDate
    15-21 Nov. 1997
  • Firstpage
    32
  • Lastpage
    32
  • Abstract
    We describe a checkpoint/recovery scheme suitable for message-passing parallel applications. The novelty of our scheme is that checkpointed applications can be restored, from their checkpointed state, in reconfigured forms. Using this scheme, applications can quickly recover from partial system failures. A key component of our implementation is the distribution- independent representation of application array data structures in persistent storage. To further optimize the performance, we provide parallel array section streaming operations for distributed arrays. We compare the performance of the reconfigurable checkpoint/restart of parallel applications with that of conventional forms of checkpointing.
  • Keywords
    Checkpointing; Computer crashes; Context modeling; Data structures; Dynamic programming; Fault tolerant systems; Mission critical systems; Parallel programming; Protection; Runtime;
  • fLanguage
    English
  • Publisher
    ieee
  • Conference_Titel
    Supercomputing, ACM/IEEE 1997 Conference
  • Print_ISBN
    0-89791-985-8
  • Type

    conf

  • DOI
    10.1109/SC.1997.10039
  • Filename
    1592613