Search papers, labs, and topics across Lattice.
This paper introduces a framework for assessing the uptime efficiency of Exascale-class scientific computers, particularly in scenarios with significant application failure rates. By shifting the focus from time-based metrics to usage-based metrics (e.g., node-hours), the authors provide a more relevant evaluation of failure impacts on computational resources. The updated framework, which builds on Daly's 2006 work, allows for the optimization of checkpointing intervals to minimize resource losses, demonstrating its application with real data from the Frontier supercomputer.
Rethinking failure metrics in scientific computing could drastically enhance resource efficiency in Exascale systems.
We present a framework for evaluating uptime efficiency of Exascale-class scientific computers when application failure rates are appreciable. This is the situation that confronts current leadership-class scientific computing platforms and large AI training installations. What distinguishes scientific computing platforms is the heterogeneity of their applications. We argue that this diversity requires that failure rates and mean intervals between failures should be specified in terms of \emph{usage} (e.g. node-hours) rather than time, as is currently customary. We consider the usage loss terms due to failures, to checkpointing, and to restart costs, and update the framework of Daly (2006) allowing users to specify optimal checkpointing usage intervals that minimize such losses. We derive the machine computational efficiency, which specifies the expected fractional resource allocation that is available for scientific computation. We illustrate the methodology using one year of production runtime data from the \emph{Frontier} supercomputer at Oak Ridge National Laboratory.