DocumentCode
451087
Title
A Checkpointing Strategy for Scalable Recovery on Distributed Parallel Systems
Author
Naik, Vijay K. ; Midkiff, Samuel P. ; Moreira, Jose E.
Author_Institution
IBM T. J. Watson Research Center
fYear
1997
fDate
15-21 Nov. 1997
Firstpage
32
Lastpage
32
Abstract
We describe a checkpoint/recovery scheme suitable for message-passing parallel applications. The novelty of our scheme is that checkpointed applications can be restored, from their checkpointed state, in reconfigured forms. Using this scheme, applications can quickly recover from partial system failures. A key component of our implementation is the distribution- independent representation of application array data structures in persistent storage. To further optimize the performance, we provide parallel array section streaming operations for distributed arrays. We compare the performance of the reconfigurable checkpoint/restart of parallel applications with that of conventional forms of checkpointing.
Keywords
Checkpointing; Computer crashes; Context modeling; Data structures; Dynamic programming; Fault tolerant systems; Mission critical systems; Parallel programming; Protection; Runtime;
fLanguage
English
Publisher
ieee
Conference_Titel
Supercomputing, ACM/IEEE 1997 Conference
Print_ISBN
0-89791-985-8
Type
conf
DOI
10.1109/SC.1997.10039
Filename
1592613
Link To Document