Skip to main navigation Skip to search Skip to main content

System-level scalable checkpoint-restart for petascale computing

  • Jiajun Cao
  • , Kapil Arya
  • , Rohan Garg
  • , Shawn Matott
  • , Dhabaleswar K. Panda
  • , Hari Subramoni
  • , Jerome Vienne
  • , Gene Cooperman
  • Northeastern University
  • Mesosphere, Inc.
  • Ohio State University
  • University of Texas at Austin

Research output: Chapter in Book/Report/Conference proceedingConference contributionpeer-review

28 Scopus citations

Abstract

Fault tolerance for the upcoming exascale generation has long been an area of active research. One of the components of a fault tolerance strategy is checkpointing. Petascale-level checkpointing is demonstrated through a new mechanism for virtualization of the InfiniBand UD (unreliable datagram) mode, and for updating the remote address on each UD-based send, due to lack of a fixed peer. Note that InfiniBand UD is required to support modern MPI implementations. An extrapolation from the current results to future SSD-based storage systems provides evidence that the current approach will remain practical in the exascale generation. This transparent checkpointing approach is evaluated using a framework of the DMTCP checkpointing package. Results are shown for HPCG (linear algebra), NAMD (molecular dynamics), and the NAS NPB benchmarks. In tests up to 32,752 MPI processes on 32,752 CPU cores, checkpointing of a computation with a 38 TB memory footprint in 11 minutes is demonstrated. Runtime overhead is reduced to less than 1%. The approach is also evaluated across three widely used MPI implementations.

Original languageEnglish
Title of host publicationProceedings - 22nd IEEE International Conference on Parallel and Distributed Systems, ICPADS 2016
EditorsXiaofei Liao, Robert Lovas, Xipeng Shen, Ran Zheng
PublisherIEEE Computer Society
Pages932-941
Number of pages10
ISBN (Electronic)9781509044573
DOIs
StatePublished - Jul 2 2016
Event22nd IEEE International Conference on Parallel and Distributed Systems, ICPADS 2016 - Wuhan, Hubei, China
Duration: Dec 13 2016Dec 16 2016

Publication series

NameProceedings of the International Conference on Parallel and Distributed Systems - ICPADS
Volume0
ISSN (Print)1521-9097

Conference

Conference22nd IEEE International Conference on Parallel and Distributed Systems, ICPADS 2016
Country/TerritoryChina
CityWuhan, Hubei
Period12/13/1612/16/16

Keywords

  • Checkpoint-Restart
  • Fault tolerance
  • InfiniBand
  • MPI
  • Supercomputing

Fingerprint

Dive into the research topics of 'System-level scalable checkpoint-restart for petascale computing'. Together they form a unique fingerprint.

Cite this