TY - GEN
T1 - PigOut
T2 - 2nd IEEE International Conference on Big Data, Big Data 2014
AU - Jeon, Kyungho
AU - Chandrashekhara, Sharath
AU - Shen, Feng
AU - Mehra, Shikhar
AU - Kennedy, Oliver
AU - Ko, Steven Y.
N1 - Publisher Copyright:
© 2014 IEEE.
PY - 2014
Y1 - 2014
N2 - This paper presents PigOut, a system that enables federated data processing over multiple Hadoop clusters. Using PigOut, a user (such as a data analyst) can write a single script in a high-level language to efficiently use multiple Hadoop clusters. There is no need to manually write multiple scripts and coordinate the execution for different clusters. PigOut accomplishes this by automatically partitioning a single, user-supplied script into multiple scripts that run on different clusters. Additionally, PigOut generates workflow descriptions to coordinate execution across clusters. In doing so, PigOut leverages existing tools built around Hadoop, avoiding extra effort required from users or administrators. For example, PigOut uses Pig Latin, a popular query language for Hadoop MapReduce, in a (virtually) unmodified form. Through our evaluation with PigMix, the standard benchmark for Pig, we demonstrate that PigOut's automatically-generated scripts and workflow definitions have comparable performance to manual, hand-tuned ones. We also report our experience with manually writing multiple scripts for a set of federated clusters, and compare the process with PigOut's automated approach.
AB - This paper presents PigOut, a system that enables federated data processing over multiple Hadoop clusters. Using PigOut, a user (such as a data analyst) can write a single script in a high-level language to efficiently use multiple Hadoop clusters. There is no need to manually write multiple scripts and coordinate the execution for different clusters. PigOut accomplishes this by automatically partitioning a single, user-supplied script into multiple scripts that run on different clusters. Additionally, PigOut generates workflow descriptions to coordinate execution across clusters. In doing so, PigOut leverages existing tools built around Hadoop, avoiding extra effort required from users or administrators. For example, PigOut uses Pig Latin, a popular query language for Hadoop MapReduce, in a (virtually) unmodified form. Through our evaluation with PigMix, the standard benchmark for Pig, we demonstrate that PigOut's automatically-generated scripts and workflow definitions have comparable performance to manual, hand-tuned ones. We also report our experience with manually writing multiple scripts for a set of federated clusters, and compare the process with PigOut's automated approach.
UR - https://www.scopus.com/pages/publications/84921806676
U2 - 10.1109/BigData.2014.7004218
DO - 10.1109/BigData.2014.7004218
M3 - Conference contribution
AN - SCOPUS:84921806676
T3 - Proceedings - 2014 IEEE International Conference on Big Data, IEEE Big Data 2014
SP - 100
EP - 109
BT - Proceedings - 2014 IEEE International Conference on Big Data, Big Data 2014
A2 - Lin, Jimmy
A2 - Pei, Jian
A2 - Hu, Xiaohua Tony
A2 - Chang, Wo
A2 - Nambiar, Raghunath
A2 - Aggarwal, Charu
A2 - Cercone, Nick
A2 - Honavar, Vasant
A2 - Huan, Jun
A2 - Mobasher, Bamshad
A2 - Pyne, Saumyadipta
PB - Institute of Electrical and Electronics Engineers Inc.
Y2 - 27 October 2014 through 30 October 2014
ER -