Search NASASearch

SEARCH · Search NASA

Results for “Checkpoint”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

Reducing space overhead for independendent checkpointing

The main disadvantages of independent checkpointing are the possible domino effect and the associated storage space overhead for maintaining multiple checkpoints. In most previous work, it has been assumed that only the checkpoints older than the current global recovery line can be discarded. Here, we generalize a notion of recovery line to potential recovery line. Only the checkpoints belonging to at least one of the potential recovery lines cannot be discarded. By using the model of maximum-sized antichains on a partially ordered set, an efficient algorithm is developed for finding all non-discardable checkpoints, and we show that the number of non-discardable checkpoints cannot exceed N(N+1)/2, where N is the number of processors. Communication trace driven simulation for several hypercube programs is performed to show the benefit of the proposed algorithm for real applications.

Wang, Yi-Min

Lazy checkpoint coordination for bounding rollback propagation

Independent checkpointing allows maximum process autonomy but suffers from potential domino effects. Coordinated checkpointing eliminates the domino effect by sacrificing a certain degree of process autonomy. In this paper, we propose the technique of lazy checkpoint coordination which preserves process autonomy while employing communication-induced checkpoint coordination for bounding rollback propagation. The introduction of the notion of laziness allows a flexible trade-off between the cost for checkpoint coordination and the average rollback distance. Worst-case overhead analysis provides a means for estimating the extra checkpoint overhead. Communication trace-driven simulation for several parallel programs is used to evaluate the benefits of the proposed scheme for real applications.

Wang, Yi-Min

Compiler-assisted static checkpoint insertion

This paper describes a compiler-assisted approach for static checkpoint insertion. Instead of fixing the checkpoint location before program execution, a compiler enhanced polling mechanism is utilized to maintain both the desired checkpoint intervals and reproducible checkpoint 1ocations. The technique has been implemented in a GNU CC compiler for Sun 3 and Sun 4 (Sparc) processors. Experiments demonstrate that the approach provides for stable checkpoint intervals and reproducible checkpoint placements with performance overhead comparable to a previously presented compiler assisted dynamic scheme (CATCH) utilizing the system clock.

Long, Junsheng

An Efficient Checkpointing System for Large Machine Learning Model Training

As machine learning models increase in size and complexity rapidly, the cost of checkpointing in ML training became a bottleneck in storage and performance (time). For example, the latest GPT-4 model has massive parameters at the scale of 1.76 trillion. It is highly time and storage consuming to frequently writes the model to checkpoints with more than 1 trillion floating point values to storage. This work aims to understand and attempt to mitigate this problem. First, we characterize the checkpointing interface in a collection of representative large machine learning/language models with respect to storage consumption and performance overhead. Second, we propose the two optimizations: i) A periodic cleaning strategy that periodically cleans up outdated checkpoints to reduce the storage burden; ii) A data staging optimization that coordinates checkpoints between local and shared file systems for performance improvement.

machine learning, artificial intelligence

Physics-aware adaptive checkpointing with shadow systems for nonlinear PDE simulations

Large-scale simulations of nonlinear partial differential equations (PDEs) that exhibit strongly transient behavior and pattern-forming dynamics produce enormous amounts of data, which, even with modern storage systems, cannot be stored for later curation. Current I/O strategies either write dense time series of snapshots, which is often prohibitive in I/O and storage, or store a few checkpoints that enable restart but incur expensive recomputation cost and provide no control over post-restart error growth, especially when lossy compression is used. Moreover, most, if not all, existing strategies take no account of the actual physical state of the system. Here, we present a simple physics-aware I/O framework in which a low-cost shadow system adaptively triggers lossy checkpoints when the shadow system deviates from the fine-scale simulation. The shadow system can be a coarsened replica of the fine-scale simulation that evolves concurrently. This means that checkpoints are taken based on the physical state of the system: fewer checkpoints are triggered when the system is quiescent while more are taken when the system undergoes a rapid change. This type of behavior is observed in many systems such as Brusselator and FitzHugh–Nagumo. We illustrate that our framework maintains stable restarts, keeps fine-scale restart errors bounded by shadow errors, and reconstructs the time history with significantly lower error and storage than interpolating fixed-interval snapshots, with low-cost shadow replay and modest online synchronization overhead.

Gong, Qian [ORNL] (ORCID:0000000235704142)

AutoCheck: Automatically Identifying Variables for Checkpointing by Data Dependency Analysis

Checkpoint/Restart (C/R) has been widely deployed in numerous HPC systems, Clouds, and industrial data centers, which are typically operated by system engineers. Nevertheless, there is no existing approach that helps system engineers without domain expertise and domain scientists without system fault tolerance knowledge identify those critical variables accounted for correct application execution restoration in a failure for C/R. To address this problem, we propose an analytical model and a tool (AutoCheck) that can automatically identify critical variables to checkpoint for C/R. AutoCheck relies on first, analytically tracking and optimizing data dependency between variables and other application execution state, and second, a set of heuristics that identify critical variables for checkpointing from the refined data dependency graph (DDG). AutoCheck allows programmers to pinpoint critical variables to checkpoint quickly within a few minutes. We evaluate AutoCheck on 13 representative HPC benchmarks, demonstrating that AutoCheck can efficiently identify correct critical variables to checkpoint.

HPC

Scrutinizing Variables for Checkpoint Using Automatic Differentiation

Checkpoint/Restart (C/R) saves the running state of the programs periodically, which consumes considerable time and system resources. We observe that not every piece of data is involved in the computation in typical HPC applications; such unused data should be excluded from checkpointing for better storage and compute efficiency. We propose a systematic approach that leverages automatic differentiation (AD) to scrutinize every element within variables (e.g., arrays) necessary for checkpointing. This allows us to identify critical and uncritical elements and eliminate uncritical elements from checkpointing. Specifically, we inspect every single element within a variable necessary for checkpointing with an AD tool to determine whether the element has an impact on the application output or not. We validate our approach with all benchmarks from the NPB suite. We visualize the distribution of critical and uncritical elements within a variable with respect to its binary impact (yes or no) on the application output.

Huang, Xin [Kobe University]

Optimal checkpointing of real-time tasks

Analytical models for the design and evaluation of checkpointing of real-time tasks are developed. First, the execution of a real-time task is modeled under a common assumption of perfect coverage of on-line detection mechanisms (which is termed a basic model). Then, the model is generalized (to an extended model) to include more realistic cases, i.e., imperfect coverages of on-line detection mechanisms and acceptance tests. Finally, an optimal placement of checkpoints is determined to minimize the mean task execution time while the probability of an unreliable result (or lack of confidence) is kept below a specified level. In the basic model, it is shown that equidistant intercheckpoint intervals are optimal, whereas this is not necessarily true in the extended model. An algorithm for calculating the optimal number of checkpoints and intercheckpoint intervals is presented with some numerical examples for the extended model.

Shin, Kang G.

Selected 3D Flash-X Checkpoints for the Long-Time Evolution of a 9.6 Solar-Mass Core-Collapse Supernova Model

This dataset contains selected 3D Flash-X checkpoint files from the long-time evolution of a low-energy core-collapse supernova explosion of a 9.6 Msun zero-metallicity, low-mass iron-core progenitor. The checkpoints span the shock-breakout phase through the young-remnant phase, ending at approximately 3 yr after core bounce. The dataset is intended for follow-up analysis and post-processing, especially radiation-transport calculations using the hydrodynamic and compositional structure of the ejecta. For a full description of the numerical setup, physical assumptions, limitations, and interpretation of the simulation, users should refer to the associated paper.

79 ASTRONOMY AND ASTROPHYSICS

Pan‐Cancer Survival Impact of Immune Checkpoint Inhibitors in a National Healthcare System

ABSTRACT Background The cumulative, health system‐wide survival benefit of immune checkpoint inhibitors (ICIs) is unclear, particularly among real‐world patients with limited life expectancies and among subgroups poorly represented on clinical trials. We sought to determine the health system‐wide survival impact of ICIs. Methods We identified all patients receiving PD‐1/PD‐L1 or CTLA‐4 inhibitors from 2010 to 2023 in the national Veterans Health Administration (VHA) system (ICI cohort) and all patients who received non‐ICI systemic therapy in the years before ICI approval (historical control). ICI and historical control cohorts were matched on multiple cancer‐related prognostic factors, comorbidities, and demographics. The effect of ICI on overall survival was quantified with Cox regression incorporating matching weights. Cumulative life‐years gained system‐wide were calculated from the difference in adjusted 5‐year restricted mean survival times. Results There were 27,322 patients in the ICI cohort and 69,801 patients in the historical control cohort. Among ICI patients, the most common cancer types were NSCLC (46%) and melanoma (10%). ICI demonstrated a large OS benefit in most cancer types with heterogeneity across cancer types (NSCLC: adjusted HR [aHR] 0.56, 95% confidence interval [CI] 0.54–0.58,p < 0.001; urothelial: aHR 0.91, 95% CI 0.83–1.01,p = 0.066). The relative benefit of ICI was stable across patient age, comorbidity, and self‐reported race subgroups. Across VHA, 15,859 life‐years gained were attributable to ICI within 5‐years of treatment, with NSCLC contributing the most life‐years gained. Conclusion We demonstrated substantial increase in survival due to ICIs across a national health system, including in patient subgroups poorly represented on clinical trials.

Oncology

Implementing forward recovery using checkpointing in distributed systems

The paper describes the implementation of a forward recovery scheme using checkpoints and replicated tasks. The implementation is based on the concept of lookahead execution and rollback validation. In the experiment, two tasks are selected for the normal execution and one for rollback validation. It is shown that the recovery strategy has nearly error-free execution time and an average redundancy lower than TMR.

Long, Junsheng

Checkpoint and restart procedures for single and multi-stage structural model analysis in NASTRAN/COSMIC on a CDC 176

The Underwater Explosions Research Division (UERD) of the David Taylor Naval Ship Research and Development Center makes extensive use of NASTRAN/COSMIC on a CDC 176 to evaluate the structural response of ship structures subjected to underwater explosion shock loadings in the time domain. As relatively new users, UERD engineers have experienced difficulties with the checkpoint/restart feature because of the vague instructions in the user manual. Working procedures for the application of the checkpoint/restart feature to the transient analysis using NASTRAN/COSMIC are illustrated.

Camp, George H.

Aviation security screening optimizer for risk and throughput (ASSORT)

The increasing number of air travelers each year presents a challenge as many airports are near their capacity in terms of resources and space for passenger screening. Fortunately, advancements in technologies like next-generation millimeter wave scanning offer solutions to ease this strain. The focus remains on managing risk while enhancing the passenger experience for the traveling public. The risk model presented in this paper known as the Aviation Security Screening Optimizer for Risk and Throughput (ASSORT) is designed to assess risk-based approaches for passenger screening and checkpoint operations. Additionally, ASSORT is exploring various traveler categories — general, trusted, and trusted-plus — along with different checkpoint screening Concept of Operations tailored to each traveler type. For instance, travelers with a higher trust level may experience fewer screening technologies, resulting in quicker processing times at the checkpoint. The output of ASSORT provides a risk score for predefined threat scenarios, as well as the overall risk to the checkpoint, aircraft, and airport by traveler type. In conclusion, benefits of using this tool include assessing the trade-offs between the overall risk associated with checkpoints and the throughput rate of passengers screened. We show for example the impact that different passenger volumes at the checkpoint can have on risk.

99 GENERAL AND MISCELLANEOUS

Recoverable distributed shared virtual memory - Memory coherence and storage structures

This paper examines the problem of implementing rollback recovery in multicomputer distributed shared virtual memory environments, in which the shared memory is implemented in software and exists only virtually. A user-transparent checkpointing recovery scheme and new twin-page disk storage management are presented to implement a recoverable distributed shared virtual memory. The checkpointing scheme is integrated with the shared virtual memory management. The twin-page disk approach allows incremental checkpointing without an explicit undo at the time of recovery. A single consistent checkpoint state is maintained on stable disk storage. The recoverable distributed shared virtual memory allows the system to restart computation from a previous checkpoint due to a processor failure without a global restart.

Wu, Kun-Lung