TY - GEN
T1 - Dynamically tuning level of parallelism in wide area data transfers
AU - Yildirim, Esma
AU - Balman, Mehmet
AU - Kosar, Tevfik
PY - 2008
Y1 - 2008
N2 - Using multiple parallel streams for wide area data transfers may yield much better performance than using a single stream, but overwhelming the network by opening too many streams may have an inverse effect. The congestion created by excess number of streams may cause a drop down in the throughput achieved. Hence, it is important to decide on the optimal number of streams without congesting the network. Predicting this 'magic' number is not straightforward, since it depends on many parameters specific to each individual transfer. Generic models that try to predict this number either rely too much on historical information or fail to achieve accurate predictions. In this paper, we present a set of new models which aim to approximate the optimal number with least history information and lowest prediction overhead. We measure the feasibility and accuracy of these models by comparing to actual GridFTP data transfers. We also discuss how these models can be used by a data scheduler to increase the overall performance of the incoming transfer requests.
AB - Using multiple parallel streams for wide area data transfers may yield much better performance than using a single stream, but overwhelming the network by opening too many streams may have an inverse effect. The congestion created by excess number of streams may cause a drop down in the throughput achieved. Hence, it is important to decide on the optimal number of streams without congesting the network. Predicting this 'magic' number is not straightforward, since it depends on many parameters specific to each individual transfer. Generic models that try to predict this number either rely too much on historical information or fail to achieve accurate predictions. In this paper, we present a set of new models which aim to approximate the optimal number with least history information and lowest prediction overhead. We measure the feasibility and accuracy of these models by comparing to actual GridFTP data transfers. We also discuss how these models can be used by a data scheduler to increase the overall performance of the incoming transfer requests.
KW - Data transfer, file transfer, gridftp, parallel tcp, parallelism
UR - https://www.scopus.com/pages/publications/57349092390
U2 - 10.1145/1383519.1383524
DO - 10.1145/1383519.1383524
M3 - Conference contribution
AN - SCOPUS:57349092390
SN - 9781605581545
T3 - International Symposium on High Performance Distributed Computing, HPDC 2008 - Proceedings of the 2008 International Workshop on Data-aware Distributed Computing 2008, DADC'08
SP - 39
EP - 47
BT - International Symposium on High Performance Distributed Computing, HPDC 2008 - Proceedings of the 2008 International Workshop on Data-aware Distributed Computing 2008, DADC'08
T2 - 2008 International Workshop on Data-aware Distributed Computing 2008, DADC'08
Y2 - 24 June 2008 through 24 June 2008
ER -