DocumentCode :
451087
Title :
A Checkpointing Strategy for Scalable Recovery on Distributed Parallel Systems
Author :
Naik, Vijay K. ; Midkiff, Samuel P. ; Moreira, Jose E.
Author_Institution :
IBM T. J. Watson Research Center
fYear :
1997
fDate :
15-21 Nov. 1997
Firstpage :
32
Lastpage :
32
Abstract :
We describe a checkpoint/recovery scheme suitable for message-passing parallel applications. The novelty of our scheme is that checkpointed applications can be restored, from their checkpointed state, in reconfigured forms. Using this scheme, applications can quickly recover from partial system failures. A key component of our implementation is the distribution- independent representation of application array data structures in persistent storage. To further optimize the performance, we provide parallel array section streaming operations for distributed arrays. We compare the performance of the reconfigurable checkpoint/restart of parallel applications with that of conventional forms of checkpointing.
Keywords :
Checkpointing; Computer crashes; Context modeling; Data structures; Dynamic programming; Fault tolerant systems; Mission critical systems; Parallel programming; Protection; Runtime;
fLanguage :
English
Publisher :
ieee
Conference_Titel :
Supercomputing, ACM/IEEE 1997 Conference
Print_ISBN :
0-89791-985-8
Type :
conf
DOI :
10.1109/SC.1997.10039
Filename :
1592613
Link To Document :
بازگشت