DocumentCode :
2121075
Title :
Data-Preservation in Scientific Workflow Middleware
Author :
Liu, D.T. ; Abdulla, G.M. ; Franklin, M.J. ; Garlick, J. ; Miller, M.
Author_Institution :
EECS Dept., California Univ., Berkeley, CA
fYear :
2006
fDate :
3-5 July 2006
Firstpage :
49
Lastpage :
58
Abstract :
This paper investigates data-preservation, a feature of scientific workflow middleware (SWM) useful for supporting data provenance and "smart recomputation." We observe that in order for an SWM supporting data preservation to achieve decent performance, it should execute on top of copy-on-write file systems. Unfortunately, most file systems in-use at scientific computing facilities were designed without copy-on-write semantics. In response, we design, implement and evaluate a middleware-level solution that is based on user-provided hints and parallelization. The solution can be deployed on top of current file systems and is able to scale almost arbitrarily. Our validation is based on real use-cases from astrophysics and experiments on a cluster with 4 file systems
Keywords :
distributed databases; middleware; natural sciences computing; network operating systems; astrophysics; copy-on-write semantics; data provenance; data-preservation; file systems; scientific computing; scientific workflow middleware; smart recomputation; Astrophysics; Collaborative work; Concurrent computing; File systems; Grid computing; Information retrieval; Laboratories; Middleware; Productivity; Scientific computing;
fLanguage :
English
Publisher :
ieee
Conference_Titel :
Scientific and Statistical Database Management, 2006. 18th International Conference on
Conference_Location :
Vienna
ISSN :
1551-6393
Print_ISBN :
0-7695-2590-3
Type :
conf
DOI :
10.1109/SSDBM.2006.18
Filename :
1644297
Link To Document :
بازگشت