Title :
A Checkpointing Strategy for Scalable Recovery on Distributed Parallel Systems
Author :
Naik, Vijay K. ; Midkiff, Samuel P. ; Moreira, Jose E.
Author_Institution :
IBM T. J. Watson Research Center
Abstract :
We describe a checkpoint/recovery scheme suitable for message-passing parallel applications. The novelty of our scheme is that checkpointed applications can be restored, from their checkpointed state, in reconfigured forms. Using this scheme, applications can quickly recover from partial system failures. A key component of our implementation is the distribution- independent representation of application array data structures in persistent storage. To further optimize the performance, we provide parallel array section streaming operations for distributed arrays. We compare the performance of the reconfigurable checkpoint/restart of parallel applications with that of conventional forms of checkpointing.
Keywords :
Checkpointing; Computer crashes; Context modeling; Data structures; Dynamic programming; Fault tolerant systems; Mission critical systems; Parallel programming; Protection; Runtime;
Conference_Titel :
Supercomputing, ACM/IEEE 1997 Conference
Print_ISBN :
0-89791-985-8
DOI :
10.1109/SC.1997.10039