TY - GEN
T1 - Federating XDMoD to Monitor Affiliated Computing Resources
AU - Sperhac, Jeanette
AU - Plessinger, Benjamin D.
AU - Palmer, Jeffrey T.
AU - Chakraborty, Rudra
AU - Dean, Gregary
AU - Innus, Martins
AU - Rathsam, Ryan
AU - Simakov, Nikolay
AU - White, Joseph P.
AU - Furlani, Thomas R.
AU - Gallo, Steven M.
AU - DeLeon, Robert L.
AU - Jones, Matthew D.
AU - Cornelius, Cynthia
AU - Patra, Abani
N1 - Publisher Copyright:
© 2018 IEEE.
PY - 2018/10/29
Y1 - 2018/10/29
N2 - The XD Metrics on Demand (XDMoD) tool was designed to collect and present detailed utilization data from the XSEDE network of supercomputers. XDMoD displays a wealth of information on computational resources. Its metrics report accounting and performance figures for computational jobs, including resources consumed, wait times, and quality of service. XDMoD can aggregate this information over an ensemble of jobs, or report it for individual jobs. Its web-based interface supports charting, exploration, and reporting for any time range, across all computing resources. Over the last eight years, we have open-sourced XDMoD, generalizing and packaging it to produce Open XDMoD. Open XDMoD may be installed, configured, and run on any computing cluster, and we have extended and refined it to include support for storage, cloud resources, science gateways, and more. In this paper, we describe the federation of XDMoD instances. This new functionality makes it possible to associate and monitor networks of related resources, using the full power of XDMoD to aggregate and display data from disparate XDMoD installations. Charting and reports from XDMoD federations can encompass anything ranging from a single resource to the entire network. The federation capability is flexible; federated resources can be coupled loosely or tightly, and a variety of different information sharing and authentication paradigms are supported. We also highlight new functionality that incorporates storage and cloud facilities into XDMoD federations. Federation will enable XDMoD users to monitor and manage ever more diverse computing resources.
AB - The XD Metrics on Demand (XDMoD) tool was designed to collect and present detailed utilization data from the XSEDE network of supercomputers. XDMoD displays a wealth of information on computational resources. Its metrics report accounting and performance figures for computational jobs, including resources consumed, wait times, and quality of service. XDMoD can aggregate this information over an ensemble of jobs, or report it for individual jobs. Its web-based interface supports charting, exploration, and reporting for any time range, across all computing resources. Over the last eight years, we have open-sourced XDMoD, generalizing and packaging it to produce Open XDMoD. Open XDMoD may be installed, configured, and run on any computing cluster, and we have extended and refined it to include support for storage, cloud resources, science gateways, and more. In this paper, we describe the federation of XDMoD instances. This new functionality makes it possible to associate and monitor networks of related resources, using the full power of XDMoD to aggregate and display data from disparate XDMoD installations. Charting and reports from XDMoD federations can encompass anything ranging from a single resource to the entire network. The federation capability is flexible; federated resources can be coupled loosely or tightly, and a variety of different information sharing and authentication paradigms are supported. We also highlight new functionality that incorporates storage and cloud facilities into XDMoD federations. Federation will enable XDMoD users to monitor and manage ever more diverse computing resources.
KW - Cloud metrics
KW - Federation
KW - High performance computing
KW - Storage metrics
KW - XDMoD
UR - https://www.scopus.com/pages/publications/85057286588
U2 - 10.1109/CLUSTER.2018.00074
DO - 10.1109/CLUSTER.2018.00074
M3 - Conference contribution
AN - SCOPUS:85057286588
T3 - Proceedings - IEEE International Conference on Cluster Computing, ICCC
SP - 580
EP - 589
BT - Proceedings - 2018 IEEE International Conference on Cluster Computing, CLUSTER 2018
PB - Institute of Electrical and Electronics Engineers Inc.
T2 - 2018 IEEE International Conference on Cluster Computing, CLUSTER 2018
Y2 - 10 September 2018 through 13 September 2018
ER -