Skip to main navigation Skip to search Skip to main content

Data pipelines: Enabling large scale multi-protocol data transfers

  • University of Wisconsin-Madison

Research output: Chapter in Book/Report/Conference proceedingConference contributionpeer-review

6 Scopus citations

Abstract

Collaborating users need to move terabytes of data among their sites, often involving multiple protocols. This process is very fragile and involves considerable human involvement to deal with failures. In this work, we propose data pipelines, an automated system for transferring data among collaborating sites. It speaks multiple protocols, has sophisticated flow control and recovers automatically from network, storage system, software and hardware failures. We successfully used data pipelines to transfer three terabytes of DPOSS data from SRB mass storage server at San Diego Supercomputing Center to UniTree mass storage at NCSA. The whole process did not require any human intervention and the data pipeline recovered automatically from various network, storage system, software and hardware failures.

Original languageEnglish
Title of host publicationProceedings of the 2nd Workshop on Middleware for Grid Computing, MGC '04
Pages63-68
Number of pages6
DOIs
StatePublished - 2004
Event2nd Workshop on Middleware for Grid Computing, MGC '04 - Toronto, ON, Canada
Duration: Oct 18 2004Oct 22 2004

Publication series

NameACM International Conference Proceeding Series
Volume76

Conference

Conference2nd Workshop on Middleware for Grid Computing, MGC '04
Country/TerritoryCanada
CityToronto, ON
Period10/18/0410/22/04

Keywords

  • Bulk data transfers
  • Data pipelines
  • Distributed systems
  • Fault-tolerance
  • Grid
  • Mass storage systems
  • Replication

Fingerprint

Dive into the research topics of 'Data pipelines: Enabling large scale multi-protocol data transfers'. Together they form a unique fingerprint.

Cite this