TY - GEN
T1 - Towards Performant Workflows, Monitoring and Measuring
AU - Sperhac, Jeanette
AU - Deleon, Robert L.
AU - White, Joseph P.
AU - Jones, Matthew
AU - Bruno, Andrew E.
AU - Ivey, Renette Jones
AU - Furlani, Thomas R.
AU - Bard, Jonathan E.
AU - Chaudhary, Vipin
N1 - Publisher Copyright:
© 2020 IEEE.
PY - 2020/8
Y1 - 2020/8
N2 - As part of the U.S. National Science Foundation (NSF) funded XD Metrics Service project, we are developing tools and techniques for the audit and analysis of High Performance Computing (HPC) and Cloud infrastructure. This includes a suite of tools for the analysis of HPC jobs, based on performance metrics collected from compute nodes. To date, we have developed two closely related utilities: XDMoD, which was designed to monitor usage and performance of NSF's innovative HPC resources (known as XSEDE), and Open XDMoD, which was designed to monitor usage and performance in academic, governmental or commercial cyberinfrastructures. Considerable effort has been made to continually improve XDMoD, in order to capture the most important aspects of modern research computing.One area in which XDMoD is lacking is in tracking workflows, which are broadly designated as containing the elements of data transfer/input and one to many computational steps. As data sets have become larger, data movement has become more time and resource intensive, and hence more important to characterize. In addition, multiple step workflows, in which one input spawns a complex series of processes, are becoming more common. Although XDMoD currently captures some of the information required to properly track complex workflows, there are clearly some key data that are missing. In this paper, we discuss the existing state of workflow monitoring, and suggest strategies to improve on the information captured.
AB - As part of the U.S. National Science Foundation (NSF) funded XD Metrics Service project, we are developing tools and techniques for the audit and analysis of High Performance Computing (HPC) and Cloud infrastructure. This includes a suite of tools for the analysis of HPC jobs, based on performance metrics collected from compute nodes. To date, we have developed two closely related utilities: XDMoD, which was designed to monitor usage and performance of NSF's innovative HPC resources (known as XSEDE), and Open XDMoD, which was designed to monitor usage and performance in academic, governmental or commercial cyberinfrastructures. Considerable effort has been made to continually improve XDMoD, in order to capture the most important aspects of modern research computing.One area in which XDMoD is lacking is in tracking workflows, which are broadly designated as containing the elements of data transfer/input and one to many computational steps. As data sets have become larger, data movement has become more time and resource intensive, and hence more important to characterize. In addition, multiple step workflows, in which one input spawns a complex series of processes, are becoming more common. Although XDMoD currently captures some of the information required to properly track complex workflows, there are clearly some key data that are missing. In this paper, we discuss the existing state of workflow monitoring, and suggest strategies to improve on the information captured.
KW - Computer Performance
KW - Data Processing
UR - https://www.scopus.com/pages/publications/85093824849
U2 - 10.1109/ICCCN49398.2020.9209647
DO - 10.1109/ICCCN49398.2020.9209647
M3 - Conference contribution
AN - SCOPUS:85093824849
T3 - Proceedings - International Conference on Computer Communications and Networks, ICCCN
BT - ICCCN 2020 - 29th International Conference on Computer Communications and Networks
PB - Institute of Electrical and Electronics Engineers Inc.
T2 - 29th International Conference on Computer Communications and Networks, ICCCN 2020
Y2 - 3 August 2020 through 6 August 2020
ER -