Search NASA⌕ Search

SEARCH · Search NASA

Results for “failure detection”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 37 records · Page 2

System Engineers and Decisions: It?s All about Knowledge

In order to guarantee that a system meets adequate levels of reliability and availability, system performances are continuously monitored and analyzed thanks to the technological advancements driving the Industry 4.0 revolution. An Industry 4.0 approach is typically based on advanced statistical, big data mining, machine learning, and internet-of-things methods designed to detect anomalies in the behavior of system, detect the most likely failure modes, and provide indications to system engineers on when maintenance activities should be performed before system performance are deemed unacceptable (which can be generated by diagnostic and prognostic methods). However, these analyses, which are designed to automatize and increase the efficacy of the system maintenance program, require large amount of data which can come in various forms: numeric, textual, images, sounds etc. Such data constitutes the historic knowledge benchmark to track system performances and support system engineer decisions. Here we claim that data is not sufficient to support this kind of analyses when applied to systems characterized by complex architectures and behaviors. Robust system engineer decisions require the ability to understand the system operational context that lies behind the observed data elements. In this respect, system models are in fact necessary to “put data in context” and capture relationships between data elements. Industry 4.0 methods require in fact contextual knowledge as a basis upon which hypotheses can be generated and assumptions tested. In our view, for complex systems, model-based system engineering (MBSE) models can afford this contextual knowledge, as they are typically used to describe systems architecture and dynamic behaviors. System knowledge is here intended as the blending of collected data and system architecture which takes the form of a “knowledge graph”. A knowledge graph is a database which consists of a large set of nodes (in our case an entity can be either a data or an MBSE element) which are linked to each other. The types of nodes and links follow a pre-defined topology, sometimes also refers as an ontology, that is designed to fit the actual decisions that needs to be performed. We show here how a knowledge graph can be defined to support system engineer maintenance decisions and how the same graph can be built based on system MBSE models and pre-processed data from numeric (through anomaly detections and diagnostic methods) and textual elements (through technical language processing TLP).

97 - MATHEMATICS AND COMPUTING↗

Live State of Health Monitoring of Inverter Subsystems

This presentation talks about different failure modes and corresponding detection schemes of PV panels, PV inverters, electric machines such as motors and power converter. Several novel techniques have been presented that are capable of measuring degradation as well as detecting faults in various location in a PV based power system.

ENGINEERING,SOLAR ENERGY↗

Non-destructive electrochemical diagnosis of failure mechanisms in aqueous zinc batteries

The early detection of secondary reactions that affect the life and performance of zinc manganese oxide batteries requires a shift from conventional time-consuming and often destructive procedures to rapid lifetime-predictive techniques. In this work, an electrochemical approach is employed to elucidate independent signatures for four common types of failure mechanisms in zinc manganese dioxide (Zn||MnO2) batteries—namely, the loss of zinc inventory, the loss of active material at the cathode, electrolyte depletion, and increased cell impedance. Our findings, specific to coin cell configurations, reveal that each induced failure mechanism can be distinctively modeled and identified based on responses from the rest voltage and columbic-efficiency data for prompt detection. For instance, electrolyte depletion response manifests a distinctive abrupt (>80 %) decrease in columbic efficiency (CE) and charge-rest voltage (Vc) while the discharge-rest voltage remained constant at ~1.3 V. Furthermore, electrolyte rejuvenation of the cell increased the CE to >95 % and restored Vc from ~0.3 to >1.7 V. Recovery experiments and reference performance tests demonstrated consistency between electrochemical descriptors and their associated failure mechanisms. Further, the outcomes of this work provide valuable insights and data models for some of the dominant failure mechanisms present in zinc manganese battery chemistries, which are beneficial to accelerated early-lifetime diagnosis and advancement of Zn batteries development.

25 ENERGY STORAGE↗

Automating Anomaly Detection for Target systems at Spallation Neutron Source

The Spallation Neutron Source (SNS) at Oak Ridge National Laboratory, produces the world’s most intense pulse neutrons beams. An accelerated proton beam is directed into a mercury target to generate neutrons via spallation. The target system accounted for over 40% of the overall downtime of the facility in 2022. Thus, early detection in anomalies in the target systems can enable taking corrective actions to avoid failures and reduce downtime. Fault prognostics and anomaly detection in accelerators, both at SNS and outside, has largely focused on the beam side. This paper presents one the first studies exploring leveraging machine learning to automate the detection of anomalies in the target system. The target system consists of over 30 different interconnected subsystems, and the present work focuses on the mercury process system as a use case. Analyzing data from 28 process variables from 2022 and 2023, tree-based and reconstruction-based algorithms are employed to detect anomalies in archived data. The algorithms detected previously unreported anomalies, several of which were deemed alert worthy by human experts, particularly those found by reconstruction-based algorithms. Using data from each production run in the accelerator increased the generalizability of the models in time. Efforts are now underway to implement a workflow for incorporating human feedback to update the models and evaluating performance on unseen data. The models will eventually be integrated into the existing System Tracking and Reliability system with a web interface for automated anomaly detection and reporting along with a pathway for incorporating human feedback for model updates.

Raj, Anant [ORNL] (ORCID:0000000306711244)↗

Management of Risks Associated with Application of Novel Materials in Novel Operating Environments in Novel Reactor Designs

There is currently no widely agreed, detailed general method for licensing a novel plant incorporating novel materials (or materials being deployed in novel environments); in many such situations, there are no directly applicable engineering code cases for decision-makers (including regulators) to rely on. This paper discusses a framework for solving this problem that is based on the Reliability and Integrity Management (RIM) approach delineated in ASME BPVC Section XI Division 2. NRC Regulatory Guide 1.246, Rev. 0, endorses, with conditions, the subject portion of the ASME Code. The proposed framework is meant to support development of a licensing case by addressing certain technical challenges. The framework discussed here is compatible with the Licensing Modernization Project, but applying it in a specific case will call for advances in the state of practice, if not the state of the art. The RIM approach calls for applicants to (a) allocate reliability targets to plant structures, systems, and components (SSCs), (b) show that they are able to relate the currently observed physical condition of each SSC in the program to its failure probability well enough to determine whether the target reliability allocations are being satisfied, allowing for uncertainty related to the novelty of the materials/designs/operating environments, and (c) be able to demonstrate that the proposed program of surveillances will reliably detect unacceptable degradation of an SSC before SSC failure occurs. These challenges are discussed in the paper, and a potentially applicable modeling approach based on cumulative damage rather than failure rates is briefly illustrated.

22 GENERAL STUDIES OF NUCLEAR REACTORS↗

Collection of Disk Failure Events from Alpine, the Parallel File System for Summit Supercomputer

This dataset contains disk (HDD) failure events collected from the Alpine storage system of the Summit supercomputer, hosted at OLCF, spanning from January 4, 2019, to December 21, 2023 (a total of 4 years, 11 months, and 18 days), covering 89% of its operational lifetime. It includes 3,766 disk failure events, each recorded with its detection timestamp (in ISO 8601 format) and detailed by its location within the storage system - rack, enclosure, and drive slot number.

97 MATHEMATICS AND COMPUTING↗

Wasserstein normalized autoencoder for anomaly detection

A novel anomaly detection algorithm is presented. The Wasserstein normalized autoencoder (WNAE) is a normalized probabilistic model that minimizes the Wasserstein distance between the learned probability distribution—a Boltzmann distribution where the energy is the reconstruction error of the autoencoder (AE)—and the distribution of the training data. This algorithm has been developed and applied to the identification of semivisible jets—conical sprays of visible standard model (SM) particles and invisible dark matter states—with the CMS experiment at the CERN LHC. Trained on jets of particles from simulated SM processes, the WNAE is shown to learn the probability distribution of the input data in a fully unsupervised fashion, such that it effectively identifies new physics jets as anomalies. The model exhibits stable, convergent training and recovers strong classification performance for a wide range of signals against the selected background process, for which a standard AE fails because of outlier reconstruction. In addition, the model improves upon standard normalized autoencoders while remaining fully agnostic to the signal. The WNAE directly tackles the problem of outlier reconstruction, a common failure mode of autoencoders in anomaly detection tasks.

Hayrapetyan, Aram [Yerevan Phys. Inst.]↗

Delamination-informed lifecycle decisions: A dielectric and machine learning framework for composite sorting and recycling

Composite materials are widely used in aerospace, marine, and automotive sectors due to their high strength-to-weight ratio and durability. However, their long-term reliability can be compromised by damage accumulation. Specifically, delamination initiation serves as a precursor to structural failure, which is often difficult to detect during damage inspection. Identifying and sorting delamination initiation in samples not only increases operational safety while providing critical information for end-of-life decisions, which influences both the service life extension value and the efficiency of fiber extraction during recycling. This research addresses two challenges: (1) developing a nondestructive, ex-situ framework to sort composite materials based on damage severity, particularly delamination, and (2) understanding how damage in composites influences resin removal during pyrolysis. Both experimental work and finite element analysis were performed to predict critical stress levels that are associated with delamination onset. Based on these results, three loading levels 50 %, 75 %, and 90 % of maximum stress, were selected for controlled experiments, generating composite samples with varying extents of damage for machine learning model training. Microscopic imaging of these samples confirmed the damage progression from matrix cracking to delamination, validating the computational predictions. We explored supervised machine learning using dielectric measurements to classify damage states. Preliminary results show an artificial neural network can identify early delamination which is a potential precursor to failure, with 94.44 % accuracy on our dataset. A parallel investigation into the effect of damage severity on pyrolysis recycling showed that heavily delaminated samples required significantly less energy for comparable matrix removal than undamaged samples.

dielectric variables↗

Machine Learning for Anomaly Detection in Neural Network Security and SRF Cavities

This dissertation explores the development and deployment of machine learning approaches to address critical challenges in anomaly detection across two distinct domains: neural network security in federated learning settings and cavity behavior analysis in particle accelerator operations at Jefferson Lab in Newport News, Virginia. Anomaly detection identifies deviations from expected patterns, safeguarding systems in cybersecurity, industry, and research against malicious activities and failures. This dissertation demonstrates how our machine learning approaches enhance detection accuracy and efficiency in both neural network security and industrial applications. First, we investigate vulnerabilities in deep neural networks deployed in federated learning. Although federated learning preserves user privacy by training models locally, it remains vulnerable to backdoor attacks, in which malicious participants embed hidden triggers that induce targeted misbehavior. We propose a self-supervised contrastive learning framework to detect and mitigate such backdoor attacks. In our experiments, this method achieves higher detection accuracy and lower false positive rates than existing defenses, while operating without access to local model updates or original training data and thus preserving the privacy guarantees of the federated setting. Second, we address the operational reliability of superconducting radio-frequency (SRF) cavities at the Continuous Electron Beam Accelerator Facility (CEBAF). Our research leverages an unsupervised learning approach, combined with Principal Component Analysis (PCA) and k-means clustering, to identify anomalous behaviors in SRF cavities. Our method detects subtle anomalous behavior by analyzing SRF signal data. This knowledge allows for the early detection and resolution of potential faults, significantly improving the efficiency and reliability of operations. Third, we extend these insights to time-series anomaly detection more broadly. We design a contrastive-learning based model tailored to increasingly dynamic environments and academic research. This model improves detection accuracy in settings that require real-time monitoring and predictive maintenance. Our research underscores the broader applicability and impact of advanced machine learning techniques in anomaly detection. By extracting meaningful patterns from complex data, machine learning can significantly enhance security in distributed neural networks and improve the efficiency of particle accelerator operations. This dissertation serves as a stepping stone for future investigations into the vast possibilities of anomaly detection, inspiring further exploration and development of machine learning techniques in this field.

Ferguson, Hal [Old Dominion University]↗

ACCELERATED DEPLOYMENT OF NOVEL MATERIALS BASED ON RELIABILITY INTEGRITY MANAGEMENT USING CUMULATIVE DAMAGE MODELING

There is currently no widely agreed, detailed general method for licensing a novel plant incorporating novel materials (or materials being deployed in novel environments); in many such situations, there are no directly applicable engineering code cases for decision-makers (including regulators) to rely on. This paper discusses a framework for solving this problem that is based on the Reliability and Integrity Management (RIM) approach delineated in ASME BPVC Section XI Division 2. NRC Regulatory Guide 1.246, Rev. 0, endorses, with conditions, the subject portion of the 2019 ASME Code. The proposed framework is meant to support development of a licensing case by addressing certain remaining technical challenges. The framework discussed here is compatible with the Licensing Modernization Project, but applying it in a specific case will call for advances in the state of practice, if not the state of the art. The RIM approach calls for applicants to (a) allocate reliability targets to plant structures, systems, and components (SSCs), (b) show that they are able to relate the currently observed physical condition of each SSC in the program to its failure probability well enough to determine whether the target reliability allocations are being satisfied, allowing for uncertainty related to the novelty of the materials/designs/operating environments, and (c) be able to demonstrate that the proposed program of surveillances will reliably detect unacceptable degradation of an SSC before SSC failure occurs. A modeling approach potentially applicable to item (b), based on cumulative damage modeling rather than failure rates, is briefly illustrated.

22 - GENERAL STUDIES OF NUCLEAR REACTORS↗

Monitoring Methods for Early Detection of Inadvertent Fission Product Release at the Advanced Test Reactor

Isotope effluent data obtained during three instances of experiment failures at the Advanced Test Reactor (ATR) are analyzed to provide an overview of the methods used to detect initial signs of unintended fission product release. The data is contextualized with the operational experience, including means of identification and subsequent mitigation strategies, gained during these events. General trends as well as variations in isotopic behavior between the three failures are explored. Background on the Real Time Monitor, a High Purity Germanium detector, and other fission product monitoring systems utilized at the Advanced Test Reactor is also provided. The presented analysis was used to establish administrative action levels which are currently utilized by ATR for early detection of experiment fission product release. Early identification provides time to make programmatic decisions before approaching safety and environmental limits.

21 - SPECIFIC NUCLEAR REACTORS AND ASSOCIATED PLAN↗

Radiation portal monitor data file format for comprehensive background radiation monitoring

Radiation portal monitors (RPMs) are widely used at border security checkpoints to detect the presence of radioactive materials in people, vehicles, and cargo. Typically, RPM detection systems consist of two pillars equipped with gamma and neutron detectors. To improve detection efficiency, RPMs employ techniques such as a limited energy window, dynamic alarm thresholds, and lead shielding. However, without continuous monitoring of background radiation, signal interpretation can be compromised, because environmental factors and mechanical failures can cause fluctuations. Here, we introduce a daily file format that logs gamma background and neutron background radiation levels continuously over a 24 h period; this format is different from traditional formats that record data only when the RPM is active or occupied. The approach enables RPM operators and analysts to (1) identify and diagnose malfunctioning components, (2) adjust system settings to account for dynamic environmental factors, and (3) use the recorded data to characterize outer space phenomena. Continuous background reporting is essential for identifying issues such as faulty connections, voltage divider failures, and errors in background updates. Continuous background reporting also enables the detection of external influences, including nearby X-ray scanners, temperature fluctuations, rainfall, cosmic radiation, and lunar phase changes. These data files are designed to be easily evaluated and parsed using common tools, and a quick review by an expert is often sufficient for problem diagnosis. We anticipate that continuous background radiation monitoring and these new strategies will significantly improve the accuracy and reliability of RPM systems, reducing the rate of false alarms and enhancing overall system performance.

Background radiation monitoring↗

Federated Learning for Efficient Condition Monitoring and Anomaly Detection in Industrial Cyber-Physical Systems

Detecting and localizing anomalies in cyber-physical systems (CPS) has become increasingly challenging as systems grow in complexity, particularly due to varying sensor reliability and node failures in distributed environments. While federated learning (FL) offers a foundation for distributed model training, existing approaches lack mechanisms to handle these CPS-specific challenges. This paper presents an enhanced FL framework that introduces three key innovations: adaptive model aggregation based on sensor reliability, dynamic node selection for resource optimization, and Weibull-based checkpointing for fault tolerance. Our framework enables reliable condition monitoring while addressing the computational and reliability challenges of industrial CPS deployments. Experiments on NASA Bearing and Hydraulic System Datasets demonstrate superior performance over state-of-the-art FL methods, achieving 99.5% AUC-ROC in anomaly detection and maintaining accuracy under node failures. Statistical validation using Mann-Whitney (U) test confirms significant improvements (p < 0.05) in both detection accuracy and computational efficiency across diverse operational scenarios.1

Marfo, William [University of Texas at El Paso,Dep↗

Physical and Operational Status of the Photovoltaic Systems in the Municipality of Pinotepa de Don Luis, Oaxaca, One Year after Their Installation [Estado fisico y operativo de los sistemas fotovoltaicos en el municipio de Pinotepa de Don Luis, Oaxaca un ano despues de su instalacion]

In November 1990, an inspection visit was made to the residential lighting photovoltaic systems that had been installed approximately one year before in the rural communities of the Municipality of Pinotepa de Don Luis, in the State of Oaxaca. The main purpose of the visit was to evaluate the physical and operational status of these systems. The systems typically consist of a photovoltaic module with a capacity of 30 or 35 Watts, an automotive type battery, an indicator of the status of charge of the battery, 3 fluorescent lamps and the corresponding installation, including cables, connection or outlet box and switches. It can be seen that the systems are in good physical condition and do not show any signs of intentionally caused damage. However, the incidence of failures is higher than what had been expected; therefore the operational status of these systems is considered to be below a level that could be accepted, with a tendency to getting worse. The main problem that has been detected is the lack of automatic charge controllers. This problem has caused other important failures. Serious deficiencies were appreciated in the installations, which show evidence of the poor workmanship of the installers. The lack of system standardization is evident, both regarding the components and the installation. As a consequence, there are technical factors which indicate that the installed systems will keep deteriorating at a rate faster than normal, and will stop operating in a short time unless urgent preventive and corrective measures are taken. The attitude of the users with respect to the systems is generally positive. However, there are elements that make the correct operation and maintenance of these systems a very difficult task. The users did not receive any training regarding the best way to operate and take care of their systems, nor any other information in relation to their technical limitations.

14 SOLAR ENERGY↗

Anomaly Detection for Online Monitoring of Thermocouple Sensors in the Advanced Test Reactor

This study explores data-driven anomaly detection methods to analyze sensor fail- ures in the Advanced Gas Reactor (AGR) nuclear fuel irradiation experiments. Specifically, we examine failures of thermocouples (TCs), which are critical for mon- itoring and controlling in-reactor temperatures during operation. Failures were pri- marily observed during abrupt power transitions and manifested as sensor drop-outs, drifts, or unexplained behavior. We applied three time-series analysis techniques— rolling mean smoothing, matrix profile, and vector auto-regression (VAR)—to de- tect anomalies in TC data prior to failure events. The rolling mean method effec- tively highlighted deviations aligned with reported failures, while the matrix profile provided partial early warning but sometimes flagged normal fluctuations during power-down periods. VAR shows potential in capturing multivariate dependencies but requires further calibration. A rare case of TC drift was also documented, which did not result in failure, underscoring the challenge of building predictive models with sparse positive examples. Our findings demonstrate that traditional statistical tools can aid anomaly detection but have limited predictive power without richer training data. We propose future directions including synthetic data generation, real- time surrogate modeling, and multi-modal feature integration. This work provides a foundation for applying robust anomaly detection frameworks to mission-critical sensor systems in experimental settings.

11 - NUCLEAR FUEL CYCLE AND FUEL MATERIALS↗

LTAU-FF: Loss Trajectory Analysis for Uncertainty in atomistic Force Fields

Model ensembles are effective tools for estimating prediction uncertainty in deep learning atomistic force fields. However, their widespread adoption is hindered by high computational costs and overconfident error estimates. In this work, we address these challenges by leveraging distributions of per-sample errors obtained during training and employing a distance-based similarity search in the model latent space. Our method, which we call LTAU (Loss Trajectory Analysis for Uncertainty), efficiently estimates the full probability distribution function of errors for any test point using the logged training errors, achieving speeds that are 2–3 orders of magnitudes faster than typical ensemble methods and allowing it to be used for tasks where training or evaluating multiple models would be infeasible. We apply LTAU towards estimating parametric uncertainty in atomistic force fields (LTAU-FF), demonstrating that it produces well-calibrated confidence intervals and predicts errors that correlate strongly with the true errors for data near the training domain. Furthermore, we show that the errors predicted by LTAU-FF can be used in practical applications for detecting out-of-domain data, tuning model performance, and predicting failure during simulations. We believe that LTAU will be a valuable tool for uncertainty quantification in atomistic force fields and is a promising method that should be further explored in other domains of machine learning.

97 MATHEMATICS AND COMPUTING↗

Power Module Precursors and Prognostics

The ongoing work involves 1) the identification and evaluation of various electric signals as indicators to module failure and precursors to aging, and 2) the development of corresponding detection techniques. In the current phase, 1) the increase of on-state resistance is selected as the precursor to the power cycle aging of silicon carbide power MOSFET modules based on review and rest data. 2) A detection method leveraging the gate voltage fluctuations magnitude during commutation transients to inform the on-state resistance increase is investigated which has the benefit of being able to be integrated into the gate driver. A dedicated detection circuit for implementation is designed and simulated and is currently in fabrication.

ADVANCED PROPULSION SYSTEMS↗

Artificial Intelligence Thermostat to Detect Faults

Residential air conditioners and heat pumps often experience faults due to inadequate maintenance, which can severely reduce efficiency or even cause system failure. Common issues include dirty or clogged air filters and refrigerant leaks. These problems degrade performance and increase energy use and operating costs. This study presents a smart thermostat with embedded artificial intelligence to detect such faults and alert homeowners when maintenance is needed. The thermostat uses low-cost measurements—including return-air temperature, relative humidity, supply-air temperature, outdoor-air temperature, and condenser subcooling—to identify abnormal operations. Because different faults produce distinct response patterns, tailored algorithms are developed to recognize characteristic fault signatures. The investigation is built on a detailed co-simulation platform that couples EnergyPlus with the DOE/ORNL Heat Pump Design Model (HPDM). EnergyPlus represents the building’s dynamic environment, while HPDM is a high-fidelity, hardware-based model that can simulate fault-free performance as well as a wide range of faults, including gradual degradation such as minor refrigerant leakage. This platform provides a virtual training and testing environment that helps distinguish fault-induced behavior from normal operation and supports development of robust diagnostic algorithms. Using this framework, a Dynamic Bayesian Network was developed to identify two common faults—gradual refrigerant charge loss and indoor airflow blockage—and the AI-embedded thermostat was verified through annual building simulations.

Shen, Bo [ORNL] (ORCID:0000000336600393)↗