Search NASA⌕ Search

SEARCH · Search NASA

Results for “application resilience prediction”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

PARIS: Predicting application resilience using machine learning

The traditional method to study application resilience to errors in HPC applications uses fault injection (FI), a time-consuming approach. Furthermore, while analytical models have been built to overcome the inefficiencies of FI, they lack accuracy. In this paper, we present PARIS, a machine-learning method to predict application resilience that avoids the time-consuming process of random FI and provides higher prediction accuracy than analytical models. PARIS captures the implicit relationship between application characteristics and application resilience, which is difficult to capture using most analytical models. We overcome many technical challenges for feature construction, extraction, and selection to use machine learning in our prediction approach. Our evaluation on 16 HPC benchmarks shows that PARIS achieves high prediction accuracy. PARIS is up to 450x faster than random FI (49x on average). Compared to the state-of-the-art analytical model, PARIS is at least 63% better in terms of accuracy and has comparable execution time on average.

97 MATHEMATICS AND COMPUTING↗

HAPPA: A Modular Platform for HPC Application Resilience Analysis with LLMs Embedded

High-performance computing (HPC) systems are increasingly vulnerable to soft errors, which pose significant challenges in maintaining computational accuracy and reliability. Predicting the resilience of HPC applications to these errors is crucial for robust code protection and detailed resilience analysis. In this study, we present HAppA, a modular platform designed for HPC Application Resilience Analysis. Embedding Large Language Models (LLMs), HAppA addresses understanding the context information of long code sequences typical in HPC applications. HAppA implements a novel code representation module that chunks the code into fixed-size segments and aggregates the embeddings of these segments. Three aggregation methods have been explored: MeanPooling, MaxPooling, and LSTM-based techniques. We built a DAtaset for REsilience analysis using Fault Injection (FI), named DARE. Using our DARE dataset, HAppA is trained for regression prediction tasks. Our evaluation results demonstrate the predictive accuracy of HAppA compared to other models, particularly noting that the LSTM-based aggregation method -- HAppA-LSTM -- achieves a mean squared error (MSE) of 0.078 for SDC prediction, surpassing the existing state-of-the-art PARIS model, which recorded an MSE of 0.1172. Additionally, HAppA with the KeyBERT model extracts a list of keywords representing the source code. A comprehensive importance analysis of these keywords further elucidates the code patterns contributing to the error rate. These findings highlight the effectiveness of HAppA in analyzing the resilience of HPC applications and establish a new benchmark for predictive accuracy in resilience.

Jiang, Hailong [Kent State University]↗

Proton radiation resilience of CdSeTe photovoltaics: High predicted end-of-life performance for space applications

Two types of cadmium selenide telluride (CdSeTe) photovoltaic devices have been exposed to high-energy (150–1500 keV) protons with fluences ranging from 1 × 10 11 to 9 × 10 13 cm −2 . Pre- and post-irradiation current density vs voltage characteristic data were collected and analyzed. Arsenic-doped CdSeTe devices retained 80% of the power conversion efficiency (PCE) relative to control devices after exposure to 10 12 cm −2 650 keV protons, while copper-doped CdSeTe devices retained about 95% of the control PCE under the same irradiation condition. Displacement damage dose analysis coupled with simulations for duration-dependent performance in a medium Earth orbit space mission revealed superior PCE remaining factors, indicating greater resilience to proton bombardment than state-of-the-art multijunction III-V based space photovoltaic technologies.

CdTe solar cells↗

Resilience and fault tolerance in high-performance computing for numerical weather and climate prediction

Progress in numerical weather and climate prediction accuracy greatly depends on the growth of the available computing power. As the number of cores in top computing facilities pushes into the millions, increased average frequency of hardware and software failures forces users to review their algorithms and systems in order to protect simulations from breakdown. This report surveys hardware, application-level and algorithm-level resilience approaches of particular relevance to time-critical numerical weather and climate prediction systems. A selection of applicable existing strategies is analysed, featuring interpolation-restart and compressed checkpointing for the numerical schemes, in-memory checkpointing, user-level failure mitigation and backup-based methods for the systems. Numerical examples showcase the performance of the techniques in addressing faults, with particular emphasis on iterative solvers for linear systems, a staple of atmospheric fluid flow solvers. The potential impact of these strategies is discussed in relation to current development of numerical weather prediction algorithms and systems towards the exascale. Trade-offs between performance, efficiency and effectiveness of resiliency strategies are analysed and some recommendations outlined for future developments.

54 ENVIRONMENTAL SCIENCES↗

Enhancing Power Distribution System Resilience with Fusion-GNN: A Dynamic Graph Representation Learning Approach

This paper explores the applications of Fusion Graph Neural Network (FuGNN) on power distribution systems. FuGNN effectively models dynamic networks with evolving topology and features. Applied to power system network reconfiguration, FuGNN demonstrates its feasibility in optimizing switch configurations to minimize unserved loads and operational costs during extreme events. Additionally, FuGNN supports various downstream tasks, such as node feature prediction, further enhancing its versatility and applicability in power system resilience.

Liu, Boming↗

Active Swarm Resiliency in the HelioSwarm Mission

Designed to observe plasma turbulence dynamics in solar wind over a distributed volume of space, the HelioSwarm mission comprises a primary chief spacecraft and eight smaller deputy satellites in uniquely assigned “loops” of periodic relative motion in a P/2 lunar resonant orbit. If one or more deputies fail, this multi-satellite architecture facilitates resiliency for science goals through repositioning of satellites to contingency loops. This strategy of Active Swarm Resiliency mitigates risk by modeling quantitative results ahead of time for mission operators to make informed decisions. Responsive actions meet minimum science objectives based on past and predicted system performance, an approach with applications to future missions with similar architecture and requirements.

Fault Management↗

Active Swarm Resiliency in the HelioSwarm Mission

Designed to observe plasma turbulence dynamics in solar wind over a distributed volume of space, the HelioSwarm mission comprises a primary chief spacecraft and eight smaller deputy satellites in uniquely assigned “loops” of periodic relative motion in a P/2 lunar resonant orbit. If one or more deputies fail, this multi-satellite architecture facilitates resiliency for science goals through repositioning of satellites to contingency loops. This strategy of Active Swarm Resiliency mitigates risk by modeling quantitative results ahead of time for mission operators to make informed decisions. Responsive actions meet minimum science objectives based on past and predicted system performance, an approach with applications to future missions with similar architecture and requirements.

Fault Management↗

Multigene engineering in plants: Technologies, applications, and future prospects

The emerging bioeconomy presents a promising solution to both economic and environmental challenges. Within the bioeconomy, plants serve as a renewable, sustainable, and cost-effective source of foods, fuels, chemicals, and materials. However, traditional breeding and single-gene engineering approaches fall short in addressing complex traits (e.g., drought tolerance, disease resistance, yield, nutrient use efficiency) which are controlled by multiple genes. The complexity of plant biology often necessitates the use of multigene engineering (MGE), which involves simultaneous ectopic expression, up/down-regulation, or editing of multiple genes, to enhance plant traits relevant to the bioeconomy. These genes may be associated with distinct traits or function as components of specific metabolic and regulatory pathways. This review summarizes current technologies for MGE within the synthetic biology-driven Design-Build-Test-Learn (DBTL) framework, detailing its four key stages: Design – gene construct development; Build – DNA assembly and plant transformation; Test – the molecular, biochemical, and physiological characterization of engineered plants; and Learn – computational modeling to refine, multiplex and iterate the process. Despite good progress in the applications of MGE in biofortification, metabolic engineering, and stress resilience, challenges remain in construct stability, coordinated gene expression, and regulatory predictability. We identified optimization paths and future directions to accelerate MGE deployment in sustainable agriculture, with possible societal benefits including reduced production costs, increased yield, and improved food and nutritional security.

AI-aided plant engineering↗

Physics-based hybrid machine learning for critical heat flux prediction with uncertainty quantification

Critical heat flux (CHF) is a key quantity in nuclear system modeling due to its impact on heat transfer, safety margins, and reactor performance. This study develops and validates an uncertainty-aware hybrid modeling approach that combines machine learning with physics-based models to predict CHF in cases of dryout. The Biasi and Bowring empirical correlations were paired with three ML uncertainty quantification (UQ) techniques: deep neural network (DNN) ensembles, Bayesian neural networks (BNNs), and deep Gaussian processes (DGPs). A pure ML model without a base model was evaluated for comparison. Model performance was assessed under plentiful (7,350 points) and limited (9 points) training data scenarios using parity, uncertainty distributions, and calibration curves. Results show that the Biasi hybrid DNN ensemble achieved the best overall performance, with a mean absolute relative error of 1.846%, and well-calibrated uncertainty estimates. The BNN-based hybrids showed slightly higher error (2.14%) but superior uncertainty calibration. DGP models underperformed, with over 6% error and poor uncertainty calibration. All hybrid models outperformed pure machine learning configurations, demonstrating resistance against data scarcity. These findings indicate that hybrid modeling significantly improves predictive accuracy, interpretability, and resilience to data scarcity. The integration of uncertainty awareness provides actionable confidence in CHF predictions, which is vital for safety-critical decisions in nuclear applications. This hybrid approach offers a viable pathway for deploying ML models in reactor analysis tools while preserving domain knowledge and physical consistency.

22 GENERAL STUDIES OF NUCLEAR REACTORS↗

Uncertainty Analysis in Multi‐Sector Systems: Considerations for Risk Analysis, Projection, and Planning for Complex Systems

Abstract Simulation models of multi‐sector systems are increasingly used to understand societal resilience to climate and economic shocks and change. However, multi‐sector systems are also subject to numerous uncertainties that prevent the direct application of simulation models for prediction and planning, particularly when extrapolating past behavior to a nonstationary future. Recent studies have developed a combination of methods to characterize, attribute, and quantify these uncertainties for both single‐ and multi‐sector systems. Here, we review challenges and complications to the idealized goal of fully quantifying all uncertainties in a multi‐sector model and their interactions with policy design as they emerge at different stages of analysis: (a) inference and model calibration; (b) projecting future outcomes; and (c) scenario discovery and identification of risk regimes. We also identify potential methods and research opportunities to help navigate the tradeoffs inherent in uncertainty analyses for complex systems. During this discussion, we provide a classification of uncertainty types and discuss model coupling frameworks to support interdisciplinary collaboration on multi‐sector dynamics (MSD) research. Finally, we conclude with recommendations for best practices to ensure that MSD research can be properly contextualized with respect to the underlying uncertainties.

54 ENVIRONMENTAL SCIENCES↗

Generative Design for Resilience of Interdependent Network Systems

Abstract Interconnected complex systems usually undergo disruptions due to internal uncertainties and external negative impacts such as those caused by harsh operating environments or regional natural disaster events. To maintain the operation of interconnected network systems under both internal and external challenges, design for resilience research has been conducted from both enhancing the reliability of the system through better designs and improving the failure recovery capabilities. As for enhancing the designs, challenges have arisen for designing a robust system due to the increasing scale of modern systems and the complicated underlying physical constraints. To tackle these challenges and design a resilient system efficiently, this study presents a generative design method that utilizes graph learning algorithms. The generative design framework contains a performance estimator and a candidate design generator. The generator can intelligently mine good properties from existing systems and output new designs that meet predefined performance criteria while the estimator can efficiently predict the performance of the generated design for a fast iterative learning process. Case studies results based on synthetic supply chain networks and power systems from the IEEE dataset have illustrated the applicability of the developed method for designing resilient interdependent network systems.

Engineering↗

An Inventory of AI-ready Benchmark Data for US Fires, Heatwaves, and Droughts

Extreme weather events, including fires, heatwaves, and droughts, have significant impacts on earth, environmental, and energy systems. Mechanistic and predictive understanding, as well as probabilistic risk assessment of these extreme weather events, are crucial for detecting, planning for, and responding to these extremes. Records of extreme weather events provide an important data source for understanding present and future extremes, but the existing data needs preprocessing before it can be used for analysis. Moreover, there are many nonstandard metrics defining the levels of severity or impacts of extremes. In this study, we compile a comprehensive benchmark data inventory of extreme weather events, including fires, heatwaves, and droughts. The dataset covers the period from 2001 to 2020 with a daily temporal resolution and a spatial resolution of 0.5°×0.5° (~55km×55km) over the continental United States (CONUS), and a spatial resolution of 1km × 1km over the Pacific Northwest (PNW) region, together with the co-located and relevant meteorological variables. By exploring and summarizing the spatial and temporal patterns of these extremes in various forms of marginal, conditional, and joint probability distributions, we gain a better understanding of the characteristics of climate extremes. The resulting AI/ML-ready data products can be readily applied to ML-based research, fostering and encouraging AI/ML research in the field of extreme weather. This study can contribute significantly to the advancement of extreme weather research, aiding researchers, policymakers, and practitioners in developing improved preparedness and response strategies to protect communities and ecosystems from the adverse impacts of extreme weather events. Usage Notes We presented a long term (2001-2020) and comprehensive data inventory of historical extreme events with daily temporal resolution covering the separate spatial extents of CONUS (0.5°×0.5°) and PNW(1km×1km) for various applications and studies. The dataset with 0.5°×0.5° resolution for CONUS can be used to help build more accurate climate models for the entire CONUS, which can help in understanding long-term climate trends, including changes in the frequency and intensity of extreme events, predicting future extreme events as well as understanding the implications of extreme events on society and the environment. The data can also be applied for risk accessment of the extremes. For example, ML/AI models can be developed to predict wildfire risk or forecast HWs by analyzing historical weather data, and past fires or heateave , allowing for early warnings and risk mitigation strategies. Using this dataset, AI-driven risk assessment models can also be built to identify vulnerable energy and utilities infrastructure, imrpove grid resilience and suggest adaptations to withstand extreme weather events. The high-resolution 1km×1km dataset ove PNW are advantageous for real-time, localized and detailed applications. It can enhance the accuracy of early warning systems for extreme weather events, helping authorities and communities prepare for and respond to disasters more effectively. For example, ML models can be developed to provide localized HW predictions for specific neighborhoods or cities, enabling residents and local emergency services to take targeted actions; the assessment of drought severity in specific communities or watersheds within the PNW can help local authorities manage water resources more effectively.

Lin, Xinming↗

Pulse duration dependent effects of ultrafast laser induced damage on a 1030 nm multi-layer dielectric mirror for high repetition rate, high average power laser systems

High repetition rate, high peak, and average power laser systems are crucial for next-generation particle accelerators, inertial confinement fusion, and secondary particle sources. These applications demand durable laser optics, particularly interference coatings on optics lasting millions of shots at high fluence. This study focuses on designing, testing, and simulating multi-layer dielectric (MLD) mirrors for pulse durations of 260 fs, 77 fs, and 25 fs at 1030 nm wavelength and 45-degree incidence angle with p -polarization. S-on-1 laser-induced damage thresholds (LIDT) for varying pulse numbers were determined, with single-shot LIDT values of 0.98 Jcm -2 , 1.63 Jcm -2 , and 2.3 Jcm -2 for 25 fs, 77 fs, and 260 fs respectively. A strong correlation between blister shape and local fluence was observed, implying that the layer expansion in a blister depends on local fluence. We have also examined mechanisms responsible for laser-induced stress generation and energy release rates in blister formation. Damage mechanisms are further explored by finite-difference time-domain (FDTD) simulations, incorporating Keldysh strong field ionization, whose predictions were in excellent agreement with the onset of damage determined experimentally. These findings offer insights for enhancing MLD coating technology, promising more efficient and resilient laser systems for diverse scientific and industrial applications.

Noor, Mohamed Yaseen (ORCID:000000021036644X)↗

Digital Twin Applications in the Water Sector: A Review

As cities develop and resource demands rise, the water sector faces crucial challenges to deliver reliable, sustainable, and efficient services. Digital Twins (DTs), virtual replicas of physical systems, offer a promising tool to transform how we manage water infrastructure. Originally developed in the aerospace industry, DTs are now gaining traction in the water sector, enabling real-time monitoring, simulation, and predictive control of water and wastewater treatment, collection and distribution networks, and water reclamation and reuse systems. While still emerging in the water sector, DTs have shown potential to enhance operational efficiency, reduce environmental impacts, and support smarter, more resilient water management. This review study provides a comprehensive overview of current DT applications in the water sector, highlighting successful case studies, technical challenges, and knowledge gaps. It also explores how DTs can help bridge the water–energy nexus by optimizing resources utilized across interconnected systems. By synthesizing recent advances and identifying future research directions, this paper illustrates how DTs can play a central role in building sustainable, adaptive, and digitally-enabled water infrastructure.

digital twin↗

Diaspora: Resilience-Enabling Services for Real-Time Distributed Workflows

The need for real-time processing to enable automated decision making and experimental steering has driven a shift from high-performance computing workflows on a centralized system to a distributed approach that integrates remote data sources, edge devices, and diverse compute facilities. Under this paradigm, data can be processed close to the source where it is generated, thus reducing latency and bandwidth usage. System resilience is thus a key challenge, requiring distributed workflows to survive component failures and to meet stringent quality-of-service requirements, which results in the need to mitigate anomalies such as congestion and low availability of resources. To address these challenges, we propose Diaspora, a unified resilience framework that is inspired by event-driven communication patterns used in public clouds. Specifically, we propose an event fabric that extends across sites, facilities, and computations to provide timely, reliable, and accurate information about data, application, and resource status. On top of the event fabric, we build resilience-enabling services that combine QoS-aware data streaming, resilient data views, resilient compute and data resources, and anomaly detection and prediction, all of which collectively enhance workflow resilience for these scientific cases.

Rao, Nageswara↗

Fault Diagnosis of Power Components with Reliability Assessment in Extraterrestrial Microgrids

This research investigates the possible failures caused by aging and other environmental and external factors that could significantly impact the performance of extraterrestrial power systems. Additionally, it presents a reliability assessment model for the space microgrid based on fault tree analysis (FTA). The reliability assessment model developed in this paper represents a tool that can be used by engineers to harden the system design for operational and economic benefits. To improve the reliability of the system, this work provides a broad review of the different fault detection and diagnosis (FDD) algorithms used for power microgrids and space applications. Using data sets from the Habitat Simulator developed through the NASA-funded Resilient Extraterrestrial Habitat Institute, this paper compares the applicability and accuracy of the different FDD methods. The primary FDD approach proposed and assessed in this work is based on the Markov reliability model. It predicts and detects future faults in the space microgrids by using past data samples and categorizing them into different classes. Data-driven-based models such as artificial neural networks are also investigated, tested, and evaluated using simulation data sets. According to the simulation results and the broad FDD algorithm comparison, this study provides the crew or maintenance engineers with a clear methodology to detect and localize power system failures.

Leila Chebbo↗

PROTEUS: Machine Learning Driven Resilience for Extreme-scale Systems

The objective of this project is to design, develop, and evaluate scalable software to enhance resilience, data checkpointing, program restart, and analysis. The proposed tasks are to 1) develop scalable machine learning techniques to learn temporal change patterns in a scalable and in-situ manner, and to minimize data movement and maximize learning locally closest to data; 2) design a concise data representation and indexing mechanism to capture the distribution of changes in data that can guarantee point-wise user-defined tolerable errors while reducing the data storage requirements by an order of magnitude or more; 3) develop data reduction techniques as library modules; 4) exploit local SSD for minimizing data movement in storage hierarchy; 5) develop anomaly detection algorithms that can predict corruptions based on learning of emerging patterns; 6) develop software libraries to be incorporated within widely used data formats and APIs; and 7) evaluate the proposed software using DOE scientific applications. The outcomes of the proposed work are to satisfy many synergistic data reduction and resilience requirements for large-scale data intensive applications executed on extreme-scale computing systems. The developed mechanism for error-bound data approximation is directly applicable to existing scientific applications. Through machine learning from historical events and change distribution, this work will enable anomaly detection for DOE computer facility.

97 MATHEMATICS AND COMPUTING↗

Material Resilience in Harsh Service Conditions

Resilience describes the attributes of a material that allow it to withstand or resist detrimental environmental effects degrading properties and performance. In service, materials may experience harsh or extreme conditions, but even modest thermal or load conditions experienced over a long period can degrade performance. Thus, the National Nuclear Security Administration mission requires predictive understanding of materials performance in harsh and extreme conditions over long periods. This performance is particularly relevant for applications in which replacement is impractical, impossible, or costly. This area of leadership addresses the evolution of material properties in environments that include static and dynamic stress, radiation, and chemical or thermal extremes. A particular focus is on situations when environments coexist or for which collection of experimental data is challenging or impossible. The capability to predict and control the nature and evolution of properties to allow designing resilience is a crucial aspect of mission success in national nuclear, global, and energy security.

36 MATERIALS SCIENCE↗