Abstract
Scientific distributed applications have an increasing need to process and move large amounts of data across wide area networks. Existing systems either closely couple computation and data movement, or they require substantial human involvement during the end-to-end process. We propose a framework that enables scientists to build reliable and efficient data transfer and processing pipelines. Our framework provides a universal interface to different data transfer protocols and storage systems. It has sophisticated flow control and recovers automatically from network, storage system, software and hardware failures. We successfully used data pipelines to replicate and process three terabytes of DPOSS astronomy image dataset and several terabytes of WCER educational video dataset. In both cases, the entire process was performed without any human intervention and the data pipeline recovered automatically from various failures.
| Original language | English |
|---|---|
| Pages (from-to) | 609-620 |
| Number of pages | 12 |
| Journal | Concurrency and Computation: Practice and Experience |
| Volume | 18 |
| Issue number | 6 |
| DOIs | |
| State | Published - May 2006 |
Keywords
- Data intensive computing
- Data pipelines
- Data transfer
- Distributed systems
- Fault tolerance
- Grid computing
- Workflows
Fingerprint
Dive into the research topics of 'Building reliable and efficient data transfer and processing pipelines'. Together they form a unique fingerprint.Cite this
- APA
- Author
- BIBTEX
- Harvard
- Standard
- RIS
- Vancouver