مرکز منطقه ای اطلاع رساني علوم و فناوري - A Checkpointing Strategy for Scalable Recovery on Distributed Parallel Systems

DocumentCode :

451087

Title :

A Checkpointing Strategy for Scalable Recovery on Distributed Parallel Systems

Author :

Naik, Vijay K. ; Midkiff, Samuel P. ; Moreira, Jose E.

Author_Institution :

IBM T. J. Watson Research Center

fYear :

1997

fDate :

15-21 Nov. 1997

Firstpage :

Lastpage :

Abstract :

We describe a checkpoint/recovery scheme suitable for message-passing parallel applications. The novelty of our scheme is that checkpointed applications can be restored, from their checkpointed state, in reconfigured forms. Using this scheme, applications can quickly recover from partial system failures. A key component of our implementation is the distribution- independent representation of application array data structures in persistent storage. To further optimize the performance, we provide parallel array section streaming operations for distributed arrays. We compare the performance of the reconfigurable checkpoint/restart of parallel applications with that of conventional forms of checkpointing.

Keywords :

Checkpointing; Computer crashes; Context modeling; Data structures; Dynamic programming; Fault tolerant systems; Mission critical systems; Parallel programming; Protection; Runtime;

fLanguage :

English

Publisher :

ieee

Conference_Titel :

Supercomputing, ACM/IEEE 1997 Conference

Print_ISBN :

0-89791-985-8

Type :

conf

DOI :

10.1109/SC.1997.10039

Filename :

1592613

Link To Document :

https://search.ricest.ac.ir/dl/search/defaultta.aspx?DTC=49&DC=451087