Data-Preservation in Scientific Workflow Middleware

This paper investigates data-preservation, a feature of scientific workflow middleware (SWM) useful for supporting data provenance and "smart recomputation." We observe that in order for an SWM supporting data preservation to achieve decent performance, it should execute on top of copy-on-write file systems. Unfortunately, most file systems in-use at scientific computing facilities were designed without copy-on-write semantics. In response, we design, implement and evaluate a middleware-level solution that is based on user-provided hints and parallelization. The solution can be deployed on top of current file systems and is able to scale almost arbitrarily. Our validation is based on real use-cases from astrophysics and experiments on a cluster with 4 file systems

[1]  Adam Zemla,et al.  LGA: a method for finding 3D similarities in protein structures , 2003, Nucleic Acids Res..

[2]  John May,et al.  Parallel I/O for High Performance Computing , 2000 .

[3]  Michael J. Franklin,et al.  GridDB: a relational interface for the grid , 2003, SIGMOD '03.

[4]  Armin Rest,et al.  A Next Generation Microlensing Survey of the LMC , 2001 .

[5]  Jennifer Widom,et al.  Lineage tracing in data warehouses , 2001 .

[6]  Michael Pilato Version Control with Subversion , 2004 .

[7]  Erez Zadok,et al.  A Versatile and User-Oriented Versioning File System , 2004, FAST.

[8]  Sanjeev Khanna,et al.  Why and Where: A Characterization of Data Provenance , 2001, ICDT.

[9]  Randal C. Burns,et al.  Ext3cow: a time-shifting file system for regulatory compliance , 2005, TOS.

[10]  Yong Zhao,et al.  Chimera: a virtual data system for representing, querying, and automating data derivation , 2002, Proceedings 14th International Conference on Scientific and Statistical Database Management.

[11]  Jennifer Widom,et al.  Trio: A System for Integrated Management of Data, Accuracy, and Lineage , 2004, CIDR.

[12]  David Maier,et al.  Managing the Forecast Factory , 2006, 22nd International Conference on Data Engineering Workshops (ICDEW'06).

[13]  Rafael Hiriart,et al.  Object-Based Photometry Pipeline for SuperMACHO Project , 2003 .

[14]  Dror G. Feitelson,et al.  Mpi-io: a parallel file i/o interface for mpi , 1995 .

[15]  Norman W. Paton,et al.  Contextualised Workflow Execution in MyGrid , 2005, EGC.