Search NASASearch

SEARCH · Search NASA

Results for “software failure”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 37 records · Page 2

wa-hls4ml: A GNN Surrogate Model for hls4ml

Recent advancements in use of machine learning techniques on field-programmable gate arrays (FPGAs) have allowed for implementation of embedded neural networks with extremely low latency. This is invaluable for particle detectors at the Large Hadron Collider, where latency and used area must be strictly bounded. The hls4ml framework is a procedure for converting from trained machine learning model software, to a synthesis result that can be used on an FPGA. However, running the pipeline is a time-consuming procedure, and there is a strong risk of failure. In particular, it is possible that the model is unable to be converted into a synthesis result, or that the resource consumption of the model will exceed the resources of the target FPGA. To aid with this development, we introduce wa-hls4ml, a surrogate model which uses a graph neural network to emulate the structure of the source models. The goal is to estimate the chance of success and resource consumption of an arbitrary model when passed through the hls4ml procedure, without the time consumption of actually running the pipeline.

43 PARTICLE ACCELERATORS

A Graph Neural Network Surrogate Model for hls4ml

Recent advancements in use of machine learning (ML) techniques on field-programmable gate arrays (FPGAs) have allowed for the implementation of embedded neural networks with extremely low latency. This is invaluable for particle detectors at the Large Hadron Collider, where latency and used area are strictly bounded. The hls4ml framework is a procedure that converts trained ML model software to a synthesis result to can be used on an FPGA. However, running the pipeline is a time-consuming procedure, and there is a strong risk of failure. In particular, it may not be possible to successfully convert a model into a synthesis result, or the resource consumption of the model may exceed the resources of the target FPGA. To aid with this development, we introduce wa-hls4ml, a surrogate model using a graph neural network to emulate the structure of the source models. The goal is to estimate the chance of success and resource consumption of a given model when passed through the hls4ml pipeline, without needing to run the pipeline.

Plotnikov, Dennis

Seismic Contingency Auto Generator

This code takes in premade earthquake scenario XML files from USGS, power grid data, and converts them into a contingency file (.con file) that can be used by power grid solvers. Within the .con file are a number (Specified by the user) of contingencies that have randomly failed power transformers based on their likelihood of failure and peak ground acceleration (PGA) value around the transformer. The transformers' likelihood of failure was calculated based on a variety of finite element modeling on various transformer designed for specific transformer voltage classes. Parameters from these FEM were used to create generic fragility curves for transformers within a specific voltage class, which correspond with earthquake PGA values to produced a probability of failure for a given earthquake scenario. More refined versions of this process, such as specifying specific transformer design categories within a voltage class, could also be applied in future iterations of the software.

Vaagensmith, Bjorn [Idaho National Laboratory (INL

A Standardized Analysis Process Using Digital Image Correlation to Calculate In Situ Cladding Strain from Modified Burst Tests for Fuel Performance Code Validation

Historical data collection on nuclear fuel cladding materials has focused on generating a statistically significant amount of data to assess the material and its failure behavior. Furthermore, data generated to support material model and failure criteria development were previously posttest evaluations, so a large number of tests was required to gain new understanding. A way to expedite this process is to develop techniques capable of generating large, high-fidelity data sets from a single test with lower uncertainty or quantified uncertainty. One such example of this approach is Oak Ridge National Laboratory’s use of modified burst tests (MBTs) to analyze the mechanical behavior and failure conditions of cladding during a simulated reactivity-initiated accident (RIA). Each test incorporates digital image correlation (DIC) analysis techniques that are used to assess the accumulated strain in situ, as well as eventual cladding failure. This work has been fruitful in defining strain-to-failure conditions for materials like silicon carbide (SiC) fiber–reinforced/SiC matrix composite tubes (SiC/SiC), iron-chromium-aluminum (FeCrAl) alloy tubes, and chromium-coated Zircaloy-4 tubes. However, there are numerous DIC software available, including open-source and proprietary software. The different DIC software use various algorithms to process images and calculate displacement values. Using these different software and algorithms can lead to varying results, and perhaps larger-than-expected uncertainties. In the present study, previously published MBT data encompassing a variety of test conditions were reanalyzed with two different DIC software to assess the variance in the calculated strain results. The data consisted of SiC/SiC, FeCrAl, and chromium-coated Zircaloy-4 tubes. Plots of the calculated strains during the transient revealed good agreement between the two DIC software. The average root-mean-square errors between the two software was 0.20% strain, which is slightly larger than a previously reported error value for these tests. In conclusion, this variance in results is low enough that this analysis method can be used for code validation.

Reactivity-initiated accident

srlife : A software tool for estimating the life of high temperature concentrating solar receivers. Part II – Ceramic receivers

As Concentrating Solar Power (CSP) technologies aim for higher operating temperatures to enhance efficiency and meet industrial process heat demands, high-temperature metallic materials, including nickel-based superalloys, face challenges in maintaining structural integrity. Advanced ceramics offer a promising alternative due to their superior high-temperature strength. However, accurately assessing the performance of ceramic components requires a fundamentally different approach from that used for metallic components. This Part II of a two-part paper describes the integration of ceramic statistical failure models within srlife – an open-source tool for predicting the life of high-temperature CSP receivers. These models account for the inherent variability in ceramic strength, as well as the effects of subcritical crack growth (SCG) under high temperature cyclic loads. Here, the paper includes an example problem that demonstrate the process of evaluating ceramic receivers using srlife. Part I details the life estimation process for metallic receivers (i.e. creep-fatigue life) along with input and output data structure, thermohydraulic analysis, and structural analysis. The complete tool is available as open-source software at https://github.com/srlife-project/srlife and can be installed via the PyPi package manager (https://pypi.org). By supporting both ceramic and metallic receiver analyses, srlife facilitates fair comparisons between competing metallic and ceramic designs, enabling accurate evaluations of plant efficiency and the economic benefits of ceramic solar receivers and other components.

High temperature ceramic receivers

Generic and ML Workloads in an HPC Datacenter: Node Energy, Job Failures, and Node-Job Analysis

HPC datacenters offer a backbone to the modern digital society. Increasingly, they run Machine Learning (ML) jobs next to generic, compute-intensive workloads, supporting science, business, and other decision-making processes. However, understanding how ML jobs impact the operation of HPC datacenters, relative to generic jobs, remains desirable but understudied. In this work, we leverage long-term operational data, collected from a national-scale production HPC datacenter, and statistically compare how ML and generic jobs can impact the performance, failures, resource utilization, and energy consumption of HPC datacenters. Our study provides key insights, e.g., ML-related power usage causes GPU nodes to run into temperature limitations, median/mean runtime and failure rates are higher for ML jobs than for generic jobs, both ML and generic jobs exhibit highly variable arrival processes and resource demands, significant amounts of energy are spent on unsuccessfully terminating jobs, and concurrent jobs tend to terminate in the same state. We open-source our cleaned-up data traces on Zenodo (https://doi. org/10.5281/zenodo.13685426), and provide our analysis toolkit as software hosted on GitHub (https://github.com/atlarge-research/2024-icpads-hpc-workload-characterization). This study offers multiple benefits for data center administrators, who can improve operational efficiency, and for researchers, who can further improve system designs, scheduling techniques, etc.

crossanalysis

SWARM: Reimagining scientific workflow management systems in a distributed world

Modern scientific workflows process massive amounts of data from diverse instruments and sensors, leveraging geographically distributed, heterogeneous compute and storage resources—from leadership-class systems to edge devices—connected by high-performance networks. The diversity of resources introduces challenges in harnessing their full potential, with resilience issues arising across applications, system software, networks, storage, and hardware. Today, workflow management systems (WMS) coordinate the execution of computation and data management tasks across target resources. However, WMS’s centralized nature makes them vulnerable to faults and scalability issues that may result in failures of entire computational campaigns. In conclusion, this paper introduces a novel agentic framework for workflow management, fully distributing and decentralizing the WMS functions and modeling them as swarm intelligence agents infused with advanced artificial intelligence solutions and traditional distributed computing algorithms that can make coordinated decisions in the presence of failures of the underlying cyberinfrastructure.

Swarm intelligence

Roadrunner

SAND2026-17073O Roadrunner software provides a comprehensive platform for simulating the mechanical behavior of crystalline materials under various loading conditions, allowing users to investigate the effects of dislocation slip hardening and damage evolution. Developed as a fork of the Multiphysics Object Oriented Simulation Environment (MOOSE) software from Idaho National Laboratory, Roadrunner is optimized for high-performance computing and can simulate large-scale problems, enabling researchers to explore complex scenarios. Its applications include material design and optimization in aerospace and automotive industries, investigation of failure mechanisms in structural materials, and development of predictive models for crystalline materials under various loading conditions. Sandia National Laboratories is a multimission laboratory managed and operated by National Technology & Engineering Solutions of Sandia, LLC, a wholly owned subsidiary of Honeywell International Inc., for the U.S. Department of Energy’s National Nuclear Security Administration under contract DE-NA0003525.

Lim, Hojun [Sandia National Lab. (SNL-CA), Livermo

Hydrogen Plus Other Alternative Fuels Risk Assessment Models (HyRAM+) Technical Reference Manual (V.6.0)

The HyRAM+ software is an open-source toolkit that provides publicly available models and default input values to enable straightforward and consistent safety assessments for hydrogen and other alternative fuel systems, such as natural gas and propane. The HyRAM+ quantitative risk assessment calculation incorporates annual likelihood of leaks or failures for both compressed gaseous and liquefied flammable fuels, as well as probabilistic models for the effects of heat flux and overpressure. HyRAM

08 HYDROGEN

FGMS-poster

Idaho National Laboratory (INL) performs post irradiation examination (PIE) of tri-structural isotropic (TRISO)-coated particle fuel to help qualify it for high temperature gas cooled reactors. TRISO fuel compacts are re-irradiated in the Neutron Radiography Reactor (NRAD) to generate the short lived fission products needed for fission product release testing. The Fuel Accident Condition Simulator (FACS) furnace and newly added Screen Neutron Irradiated Fuel for Failure (SNIFF) furnace heat the compacts in helium to temperatures of up to 2,000°C, prompting fission product release—predominantly gaseous xenon and krypton isotopes and condensable products such as cesium—from failed particles. These released isotopes are transported to a fission gas monitoring system (FGMS 1 or FGMS 3), where they accumulate in cryogenic cold traps and are quantified using high-purity germanium (HPGe) detectors. The addition of SNIFF and FGMS 3 increases throughput by enabling simultaneous testing of multiple compacts. Furthermore, automated INL developed software provides continuous, near-real time monitoring of fission product inventories and manages the liquid nitrogen cooling of the traps. These system enhancements improve the efficiency, data quality, and testing capacity of TRISO fuel performance evaluations.

07 - ISOTOPES AND RADIATION SOURCES

Programmable Digital Devices used in Advanced Reactors

This paper introduces the concepts of common cause failure, diversity, and defense-in-depth used by the nuclear industry to analyze resilience in reactors. A survey of publicly traded and private companies building advanced reactors and their licensing status is presented. Safety and non-safety systems found in the NuScale Power design are summarized and the likely hardware and software categories used by those systems are enumerated. The importance of industry partners is highlighted. This paper also identifies an alternate path forward without industry partners to advance the knowledge needed to use artificial intelligence to analyze HBOMs and SBOMs to better understand reactor resiliency.

cybersecurity

Quantitative insights for diagnosing performance bottlenecks in lithium–sulfur batteries

Lithium–sulfur (Li–S) batteries hold significant promise for electric vehicles and aviation due to their high energy density and cost-effectiveness. However, understanding the root causes of performance degradation remains a formidable challenge, as the interplay of multiple factors obscures key failure mechanisms. A major limitation has been the inability to quantify soluble sulfur species within practical detection limits accurately and to correlate electrochemical processes with associated physical inventory changes. Here, we introduce the high-performance liquid chromatography-ultraviolet spectroscopy and gas chromatography sequential characterization (HUGS) toolkit, capable of precisely quantifying seven distinct sulfur and polysulfide species at concentrations as low as 40 ppb. HUGS has been successfully applied to practical coin and pouch cells without requiring cell modification. Furthermore, our self-developed software, Dr HUGS, enhanced the data analysis speed by over 30 times, enabling multi-source data integration and delivering comprehensive analysis results within minutes. Using HUGS, we identify significant capacity losses from inactive lithium and sulfur during initial cycles and sulfide-rich solid–electrolyte interphase (SEI) formation on the anode during later cycles. Notably, our findings reveal that soluble polysulfides have minimal contributions to capacity loss, challenging long-standing assumptions. Moreover, HUGS demonstrates that constant-pressure setups in Li–S pouch cells improve compositional uniformity compared to constant-gap configurations. For sulfurized polyacrylonitrile (SPAN) cathodes, unique issues such as non-sulfide SEI formation and lithium pulverization are observed, which can be mitigated through localized high-concentration electrolytes to enhance lithium inventory retention. By enabling precise quantification of critical inventory components, HUGS provides transformative insights into failure mechanisms across various electrolytes and cathode chemistries, guiding rational design strategies for next-generation energy storage systems.

25 ENERGY STORAGE

Using a Large Language Model for Accurate Technical Language Generation in the Predictive Maintenance of Circulating Water Systems in Nuclear Power Plants

Machine learning (ML) methods for predictive maintenance (PdM) are emerging as effective proactive strategies for diagnosing equipment degradation and enabling effective decision-making. However, explainability and trustworthiness of artificial intelligence are two salient challenges that need to be addressed for wider deployment of these technologies in nuclear power plants (NPPs). Large language models (LLMs) offer a unique approach to tackle these challenges by explaining PdM, work orders, diagnosis results, and ML algorithms to users, who may not be familiar with ML and PdM in general. Moreover, by dynamically retrieving relevant information from technical documents and evaluating factuality of LLM generation, the accuracy and relevance of LLM generations can be improved. This work demonstrates using LLMs to explain the causes and consequences of circulating water system failures based on multiyear NPP work orders. This work tests the capability of multimodal LLM approaches in explaining the differences in the circulating water system from both the Salem and Hope Creek NPPs using both text and image resources. This work also demonstrates the use of multimodal LLMs in describing the diagnosis tab of a predictive maintenance software named VIsualization for PrEdictive maintenance Recommendation (VIPER) to users.

22 - GENERAL STUDIES OF NUCLEAR REACTORS

Software Quality Assurance Plan ANSYS LSDYNA Version 2023R1

ANSYS Inc. develops and markets engineering simulation software and services used in the aerospace, automotive, manufacturing, electronics, biomedical, energy, defense, and many other industries. ANSYS is dedicated to engineering simulation and is the world’s leading software provider. ANSYS was founded in 1970 and is headquartered in Canonsburg, Pennsylvania. ANSYS provides an engineering analysis tool combining structural, thermal, computational fluid dynamics, acoustic and electromagnetic simulation capabilities. ANSYS LS-DYNA is the most used explicit simulation program capable of simulating the response of materials to short periods of severe loading. Its many elements, contact formulations, material models, and other controls can be used to simulate complex models with control over all the details of the problem. ANSYS LS-DYNA has a vast array of capabilities to simulate extreme deformation problems using its explicit solver. Engineers can tackle simulations involving material failure and look at how the failure progresses through a part or through a system. Models with large amounts of parts or surfaces interacting with each other are also easily handled, and the interactions and load passing between complex behaviors are modeled accurately. Using computers with higher numbers of CPU cores can drastically reduce solution times. In addition, many consulting firms and hundreds of universities use ANSYS for analysis, research, and educational purposes. ANSYS is recognized worldwide as one of the most widely used and capable programs of its type. ANSYS has successfully passed over 100 customer quality system audits against American Society of Mechanical Engineers (ASME) NQA-1 and 10 CFR Part 50, Appendix B, since the company was founded, over 60 of which have been since 1997. ANSYS has successfully passed over 100 International Organization for Standardization (ISO) 9001 assessments. ANSYS design analysis software is the first created within a quality system with ISO 9001 certification, which is the internationally accepted quality standard. Product development, testing, maintenance, and support processes also meet the US Nuclear Regulatory Commission’s (NRC’s) quality requirements, as they have for nearly four decades. ANSYS staff perform more than 60,000 software verification tests before releasing each new product. ASME NQA-1-2012 (Subpart 2.7 is specific to software) is the industry- and NRC-accepted approach (consensus standard) for meeting 10 CFR Part 50, Appendix B, requirements.

11 NUCLEAR FUEL CYCLE AND FUEL MATERIALS

Time-Dependent Failure Assessment of Ceramic Receivers

The outlet temperature targets for Gen 3 Concentrating Solar Power (CSP) systems pose a significant challenge to the structural reliability of high temperature metallic components, including those manufactured from nickel-based superalloys. Advanced ceramics present a potential solution due to their excellent high-temperature strength. However, accurate assessment of ceramic components requires an entirely different approach compared to metallic components. This paper describes the implementation of time-dependent reliability analysis of ceramic components in srlife – an open-source software package for estimating the life of high temperature CSP components. This new capability will allow high temperature CSP designers to make fair comparisons between competing metallic and ceramic designs and accurately assess the performance of different ceramic materials for CSP receivers and other components. The current version of the tool is available at https://github.com/Argonne-National-Laboratory/srlife.

Barua, Bipul (ORCID:0000000247184113)

The ePIC Simulation Campaign Workflow on the Open Science Grid

The ePIC collaboration is realizing the first experiment of the future Electron-Ion Collider (EIC) at the Brookhaven National Laboratory that will allow for a precision study of the nucleons and the nucleus at the scale of sea quarks and gluons through the study of electron-proton/ion collisions. This paper will discuss the current workflow for running centralized simulation campaigns for ePIC on the Open Science Grid (OSG) infrastructure. This involves monthly releases of ePIC software and container deployments to CVMFS, generation of input datasets in HepMC format according to collaboration-defined policy, using Snakemake in CI/CD for validation and benchmarking, and submitting jobs to the OSG condor scheduler for opportunistic running on available resources. File transfers utilize XrootD, and Rucio is used for data management. The workflow is continuously refined to improve daily throughput (currently 50-100k core hours per day) and minimize job failures. Since May 2023, monthly simulation campaigns employing the workflow have cumulatively used over 20 million core hours on the OSG and produced over 350 TB of simulation data. The campaigns incorporate simulations for the broad science program of the EIC and are actively used for the detector and physics studies in preparation of the Technical Design Report (TDR).

73 NUCLEAR PHYSICS AND RADIATION PHYSICS

Resilient Entanglement Distribution in a Multihop Quantum Network

The evolution of quantum networking requires architectures capable of dynamically reconfigurable entanglement distribution to meet diverse user needs and ensure tolerance against transmission disruptions. We introduce multihop quantum networks to improve network reach and resilience by enabling quantum communications across intermediate nodes, thus broadening network connectivity and increasing scalability. We present multihop two-qubit polarization-entanglement distribution within a quantum network at the Oak Ridge National Laboratory campus. Our system uses wavelength-selective switches for adaptive bandwidth management on a software-defined quantum network that integrates a quantum data plane with classical data and control planes, creating a flexible, reconfigurable mesh. Our network distributes entanglement across six nodes within three subnetworks, each located in a separate building, optimizing quantum state fidelity and transmission rate through adaptive resource management. Additionally, we demonstrate the network's resilience by implementing a link recovery approach that monitors and reroutes quantum resources to maintain service continuity despite link failures—paving the way for scalable and reliable quantum networking infrastructures.

Alshowkan, Muneer [Oak Ridge National Laboratory (

Status of Multiple Channel Fuel Performance Capabilities Within the SAS4A/SASSYS-1 Safety Analysis Software

SAS4A/SASSYS-1 (SAS) is a fast-running simulation tool used to perform deterministic analysis of anticipated events as well as design basis and beyond design basis accidents for advanced liquid-metal-cooled nuclear reactors. It is a critical element of safety analysis capabilities for the U.S. Department of Energy and is utilized within industry to perform the transient safety analyses required to support the licensing of Liquid Metal-cooled Fast Reactors (LMFRs). Although SAS is exceptionally fast for most transient scenarios, fuel performance calculations, along with the associated pre-transient characterization of the fuel pin, may be required for transient scenarios where fuel pin failure is hypothesized. Both the pre-transient characterization and the transient fuel performance calculation are necessary to properly quantify margins to potential fuel failure and assess the time spent potentially exceeding such margins during events. While safety analysis calculations with fuel performance models provide a more detailed characterization of the reactor during a transient, the pre-transient characterization can be time-consuming and computationally expensive. Often, large numbers of fuel pins have been exposed to similar pre-transient irradiation conditions. Similarly, the same pre-transient fuel characterization may be applicable to numerous transient conditions. This provides an opportunity to optimize the SAS computational framework such that pre-transient fuel characterization can be shared across multiple channels (fuel pins) and across multiple simulations, thus dramatically reducing overall computational costs. This report summarizes progress toward enhancing the SAS computational framework to support shared, multiple channel fuel performance characterizations intended to significantly reduce computational costs. Preliminary testing has shown that the computational time saved by using the pre-transient sharing capability is approximately equal to the time it takes to perform the pre-transient characterization.

22 GENERAL STUDIES OF NUCLEAR REACTORS