Search NASA⌕ Search

SEARCH · Search NASA

Results for “anomaly detection system”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 37 records · Page 2

Automating Anomaly Detection for Target systems at Spallation Neutron Source

The Spallation Neutron Source (SNS) at Oak Ridge National Laboratory, produces the world’s most intense pulse neutrons beams. An accelerated proton beam is directed into a mercury target to generate neutrons via spallation. The target system accounted for over 40% of the overall downtime of the facility in 2022. Thus, early detection in anomalies in the target systems can enable taking corrective actions to avoid failures and reduce downtime. Fault prognostics and anomaly detection in accelerators, both at SNS and outside, has largely focused on the beam side. This paper presents one the first studies exploring leveraging machine learning to automate the detection of anomalies in the target system. The target system consists of over 30 different interconnected subsystems, and the present work focuses on the mercury process system as a use case. Analyzing data from 28 process variables from 2022 and 2023, tree-based and reconstruction-based algorithms are employed to detect anomalies in archived data. The algorithms detected previously unreported anomalies, several of which were deemed alert worthy by human experts, particularly those found by reconstruction-based algorithms. Using data from each production run in the accelerator increased the generalizability of the models in time. Efforts are now underway to implement a workflow for incorporating human feedback to update the models and evaluating performance on unseen data. The models will eventually be integrated into the existing System Tracking and Reliability system with a web interface for automated anomaly detection and reporting along with a pathway for incorporating human feedback for model updates.

Raj, Anant [ORNL] (ORCID:0000000306711244)↗

Development of a Computer Architecture to Support the Optical Plume Anomaly Detection (OPAD) System

The NASA OPAD spectrometer system relies heavily on extensive software which repetitively extracts spectral information from the engine plume and reports the amounts of metals which are present in the plume. The development of this software is at a sufficiently advanced stage where it can be used in actual engine tests to provide valuable data on engine operation and health. This activity will continue and, in addition, the OPAD system is planned to be used in flight aboard space vehicles. The two implementations, test-stand and in-flight, may have some differing requirements. For example, the data stored during a test-stand experiment are much more extensive than in the in-flight case. In both cases though, the majority of the requirements are similar. New data from the spectrograph is generated at a rate of once every 0.5 sec or faster. All processing must be completed within this period of time to maintain real-time performance. Every 0.5 sec, the OPAD system must report the amounts of specific metals within the engine plume, given the spectral data. At present, the software in the OPAD system performs this function by solving the inverse problem. It uses powerful physics-based computational models (the SPECTRA code), which receive amounts of metals as inputs to produce the spectral data that would have been observed, had the same metal amounts been present in the engine plume. During the experiment, for every spectrum that is observed, an initial approximation is performed using neural networks to establish an initial metal composition which approximates as accurately as possible the real one. Then, using optimization techniques, the SPECTRA code is repetitively used to produce a fit to the data, by adjusting the metal input amounts until the produced spectrum matches the observed one to within a given level of tolerance. This iterative solution to the original problem of determining the metal composition in the plume requires a relatively long period of time to execute the software in a modern single-processor workstation, and therefore real-time operation is currently not possible. A different number of iterations may be required to perform spectral data fitting per spectral sample. Yet, the OPAD system must be designed to maintain real-time performance in all cases. Although faster single-processor workstations are available for execution of the fitting and SPECTRA software, this option is unattractive due to the excessive cost associated with very fast workstations and also due to the fact that such hardware is not easily expandable to accommodate future versions of the software which may require more processing power. Initial research has already demonstrated that the OPAD software can take advantage of a parallel computer architecture to achieve the necessary speedup. Current work has improved the software by converting it into a form which is easily parallelizable. Timing experiments have been performed to establish the computational complexity and execution speed of major components of the software. This work provides the foundation of future work which will create a fully parallel version of the software executing in a shared-memory multiprocessor system.

Katsinis, Constantine↗

Lessons learned from the introduction of autonomous monitoring to the EUVE science operations center

The University of California at Berkeley's (UCB) Center for Extreme Ultraviolet Astrophysics (CEA), in conjunction with NASA's Ames Research Center (ARC), has implemented an autonomous monitoring system in the Extreme Ultraviolet Explorer (EUVE) science operations center (ESOC). The implementation was driven by a need to reduce operations costs and has allowed the ESOC to move from continuous, three-shift, human-tended monitoring of the science payload to a one-shift operation in which the off shifts are monitored by an autonomous anomaly detection system. This system includes Eworks, an artificial intelligence (AI) payload telemetry monitoring package based on RTworks, and Epage, an automatic paging system to notify ESOC personnel of detected anomalies. In this age of shrinking NASA budgets, the lessons learned on the EUVE project are useful to other NASA missions looking for ways to reduce their operations budgets. The process of knowledge capture, from the payload controllers for implementation in an expert system, is directly applicable to any mission considering a transition to autonomous monitoring in their control center. The collaboration with ARC demonstrates how a project with limited programming resources can expand the breadth of its goals without incurring the high cost of hiring additional, dedicated programmers. This dispersal of expertise across NASA centers allows future missions to easily access experts for collaborative efforts of their own. Even the criterion used to choose an expert system has widespread impacts on the implementation, including the completion time and the final cost. In this paper we discuss, from inception to completion, the areas where our experiences in moving from three shifts to one shift may offer insights for other NASA missions.

Lewis, M.↗

Reliable statistics-based detection and investigation of anomalies in a SMART valve system

Reliable anomaly detection and diagnosis are critical for the safe operation of complex engineered systems. This study presents a unified framework that integrates statistical, model-based, and data-driven techniques for anomaly detection and investigation, demonstrated on SMART valve systems in hybrid energy applications. Four detection methods—mean deviation, seasonal extreme studentized deviate, ARIMA forecasting, and matrix profiling—were implemented and compared. Matrix profiling was particularly effective in revealing subtle deviations and hidden relationships among variables. Anomaly investigation was performed by analyzing variable-level and grouped signal profiles, with system topology incorporated to distinguish primary faults from propagated effects. Grouping signals by type enhanced interpretability, enabling accurate localization of anomalies across multi-dimensional datasets. Experimental results confirmed the framework's capability to consistently detect and isolate anomalies while providing actionable insights into system interdependencies. The proposed methodology offers a robust, interpretable, and scalable solution for condition monitoring, with potential applications in safety-critical domains such as nuclear energy, aerospace, and process industries.

ARIMA models↗

Communications and tracking expert systems study

The original objectives of the study consisted of five broad areas of investigation: criteria and issues for explanation of communication and tracking system anomaly detection, isolation, and recovery; data storage simplification issues for fault detection expert systems; data selection procedures for decision tree pruning and optimization to enhance the abstraction of pertinent information for clear explanation; criteria for establishing levels of explanation suited to needs; and analysis of expert system interaction and modularization. Progress was made in all areas, but to a lesser extent in the criteria for establishing levels of explanation suited to needs. Among the types of expert systems studied were those related to anomaly or fault detection, isolation, and recovery.

Leibfried, T. F.↗

Space Shuttle Main Engine: Advanced Health Monitoring System

The main gola of the Space Shuttle Main Engine (SSME) Advanced Health Management system is to improve flight safety. To this end the new SSME has robust new components to improve the operating margen and operability. The features of the current SSME health monitoring system, include automated checkouts, closed loop redundant control system, catastropic failure mitigation, fail operational/ fail-safe algorithms, and post flight data and inspection trend analysis. The features of the advanced health monitoring system include: a real time vibration monitor system, a linear engine model, and an optical plume anomaly detection system. Since vibration is a fundamental measure of SSME turbopump health, it stands to reason that monitoring the vibration, will give some idea of the health of the turbopumps. However, how is it possible to avoid shutdown, when it is not necessary. A sensor algorithm has been developed which has been exposed to over 400 test cases in order to evaluate the logic. The optical plume anomaly detection (OPAD) has been developed to be a sensitive monitor of engine wear, erosion, and breakage.

Singer, Chirs↗

SSME Advanced Health Management: Project Overview

This document is the viewgraphs from a presentation concerning the development of the Health Management system for the Space Shuttle Main Engine (SSME). It reviews the historical background of the SSME Advanced Health Management effort through the present final Health management configuration. The document includes reviews of three subsystems to the Advanced Health Management System: (1) the Real-Time Vibration Monitor System, (2) the Linear Engine Model, and (3) the Optical Plume Anomaly Detection system.

Plowden, John↗

Integrated System Health Management Development Toolkit

This software toolkit is designed to model complex systems for the implementation of embedded Integrated System Health Management (ISHM) capability, which focuses on determining the condition (health) of every element in a complex system (detect anomalies, diagnose causes, and predict future anomalies), and to provide data, information, and knowledge (DIaK) to control systems for safe and effective operation.

Figueroa, Jorge↗

NASA Stennis Space Center Integrated System Health Management Test Bed and Development Capabilities

Integrated System Health Management (ISHM) is a capability that focuses on determining the condition (health) of every element in a complex System (detect anomalies, diagnose causes, prognosis of future anomalies), and provide data, information, and knowledge (DIaK)-not just data-to control systems for safe and effective operation. This capability is currently done by large teams of people, primarily from ground, but needs to be embedded on-board systems to a higher degree to enable NASA's new Exploration Mission (long term travel and stay in space), while increasing safety and decreasing life cycle costs of spacecraft (vehicles; platforms; bases or outposts; and ground test, launch, and processing operations). The topics related to this capability include: 1) ISHM Related News Articles; 2) ISHM Vision For Exploration; 3) Layers Representing How ISHM is Currently Performed; 4) ISHM Testbeds & Prototypes at NASA SSC; 5) ISHM Functional Capability Level (FCL); 6) ISHM Functional Capability Level (FCL) and Technology Readiness Level (TRL); 7) Core Elements: Capabilities Needed; 8) Core Elements; 9) Open Systems Architecture for Condition-Based Maintenance (OSA-CBM); 10) Core Elements: Architecture, taxonomy, and ontology (ATO) for DIaK management; 11) Core Elements: ATO for DIaK Management; 12) ISHM Architecture Physical Implementation; 13) Core Elements: Standards; 14) Systematic Implementation; 15) Sketch of Work Phasing; 16) Interrelationship Between Traditional Avionics Systems, Time Critical ISHM and Advanced ISHM; 17) Testbeds and On-Board ISHM; 18) Testbed Requirements: RETS AND ISS; 19) Sustainable Development and Validation Process; 20) Development of on-board ISHM; 21) Taxonomy/Ontology of Object Oriented Implementation; 22) ISHM Capability on the E1 Test Stand Hydraulic System; 23) Define Relationships to Embed Intelligence; 24) Intelligent Elements Physical and Virtual; 25) ISHM Testbeds and Prototypes at SSC Current Implementations; 26) Trailer-Mounted RETS; 27) Modeling and Simulation; 28) Summary ISHM Testbed Environments; 29) Data Mining - ARC; 30) Transitioning ISHM to Support NASA Missions; 31) Feature Detection Routines; 32) Sample Features Detected in SSC Test Stand Data; and 33) Health Assessment Database (DIaK Repository).

Figueroa, Fernando↗

ISHM Implementation for Constellation Systems

Integrated System Health Management (ISHM) is a capability that focuses on determining the condition (health) of every element in a complex System (detect anomalies, diagnose causes, prognosis of future anomalies), and provide data, information, and knowledge (DIaK) "not just data" to control systems for safe and effective operation. This capability is currently done by large teams of people, primarily from ground, but needs to be embedded on-board systems to a higher degree to enable NASA's new Exploration Mission (long term travel and stay in space), while increasing safety and decreasing life cycle costs of systems (vehicles; platforms; bases or outposts; and ground test, launch, and processing operations). This viewgraph presentation reviews the use of ISHM for the Constellation system.

Figueroa, Fernando↗

Intelligent Integrated Health Management for a System of Systems

An intelligent integrated health management system (IIHMS) incorporates major improvements over prior such systems. The particular IIHMS is implemented for any system defined as a hierarchical distributed network of intelligent elements (HDNIE), comprising primarily: (1) an architecture (Figure 1), (2) intelligent elements, (3) a conceptual framework and taxonomy (Figure 2), and (4) and ontology that defines standards and protocols. Some definitions of terms are prerequisite to a further brief description of this innovation: A system-of-systems (SoS) is an engineering system that comprises multiple subsystems (e.g., a system of multiple possibly interacting flow subsystems that include pumps, valves, tanks, ducts, sensors, and the like); 'Intelligent' is used here in the sense of artificial intelligence. An intelligent element may be physical or virtual, it is network enabled, and it is able to manage data, information, and knowledge (DIaK) focused on determining its condition in the context of the entire SoS; As used here, 'health' signifies the functionality and/or structural integrity of an engineering system, subsystem, or process (leading to determination of the health of components); 'Process' can signify either a physical process in the usual sense of the word or an element into which functionally related sensors are grouped; 'Element' can signify a component (e.g., an actuator, a valve), a process, a controller, an actuator, a subsystem, or a system; The term Integrated System Health Management (ISHM) is used to describe a capability that focuses on determining the condition (health) of every element in a complex system (detect anomalies, diagnose causes, prognosis of future anomalies), and provide data, information, and knowledge (DIaK) not just data to control systems for safe and effective operation. A major novel aspect of the present development is the concept of intelligent integration. The purpose of intelligent integration, as defined and implemented in the present IIHMS, is to enable automated analysis of physical phenomena in imitation of human reasoning, including the use of qualitative methods. Intelligent integration is said to occur in a system in which all elements are intelligent and can acquire, maintain, and share knowledge and information. In the HDNIE of the present IIHMS, an SoS is represented as being operationally organized in a hierarchical-distributed format. The elements of the SoS are considered to be intelligent in that they determine their own conditions within an integrated scheme that involves consideration of data, information, knowledge bases, and methods that reside in all elements of the system. The conceptual framework of the HDNIE and the methodologies of implementing it enable the flow of information and knowledge among the elements so as to make possible the determination of the condition of each element. The necessary information and knowledge is made available to each affected element at the desired time, satisfying a need to prevent information overload while providing context-sensitive information at the proper level of detail. Provision of high-quality data is a central goal in designing this or any IIHMS. In pursuit of this goal, functionally related sensors are logically assigned to groups denoted processes. An aggregate of processes is considered to form a system. Alternatively or in addition to what has been said thus far, the HDNIE of this IIHMS can be regarded as consisting of a framework containing object models that encapsulate all elements of the system, their individual and relational knowledge bases, generic methods and procedures based on models of the applicable physics, and communication processes (Figure 2). The framework enables implementation of a paradigm inspired by how expert operators monitor the health of systems with the help of (1) DIaK from various sources, (2) software tools that assist in rapid visualization of the condition of the system, (3) analical software tools that assist in reasoning about the condition, (4) sharing of information via network communication hardware and software, and (5) software tools that aid in making decisions to remedy unacceptable conditions or improve performance.

Smith, Harvey↗

Federated Learning for Efficient Condition Monitoring and Anomaly Detection in Industrial Cyber-Physical Systems

Detecting and localizing anomalies in cyber-physical systems (CPS) has become increasingly challenging as systems grow in complexity, particularly due to varying sensor reliability and node failures in distributed environments. While federated learning (FL) offers a foundation for distributed model training, existing approaches lack mechanisms to handle these CPS-specific challenges. This paper presents an enhanced FL framework that introduces three key innovations: adaptive model aggregation based on sensor reliability, dynamic node selection for resource optimization, and Weibull-based checkpointing for fault tolerance. Our framework enables reliable condition monitoring while addressing the computational and reliability challenges of industrial CPS deployments. Experiments on NASA Bearing and Hydraulic System Datasets demonstrate superior performance over state-of-the-art FL methods, achieving 99.5% AUC-ROC in anomaly detection and maintaining accuracy under node failures. Statistical validation using Mann-Whitney (U) test confirms significant improvements (p < 0.05) in both detection accuracy and computational efficiency across diverse operational scenarios.1

Marfo, William [University of Texas at El Paso,Dep↗

Method for Real-Time Model Based Structural Anomaly Detection

A system and methods for real-time model based vehicle structural anomaly detection are disclosed. A real-time measurement corresponding to a location on a vehicle structure during an operation of the vehicle is received, and the real-time measurement is compared to expected operation data for the location to provide a modeling error signal. A statistical significance of the modeling error signal to provide an error significance is calculated, and a persistence of the error significance is determined. A structural anomaly is indicated, if the persistence exceeds a persistence threshold value.

Smith, Timothy A.↗

Automated Scoring of Morphological Changes in Images of Pentaerythritol Tetranitrate

Recent advances in characterization techniques that generate large datasets of material microstructure images require robust, automated image-processing. We applied an unsupervised anomaly detection method called feature anomaly detection system (FADS) to automatically detect and quantify microstructure changes in images of the explosive pentaerythritol tetranitrate (PETN) aged at various temperatures. We demonstrated the FADS approach on two-dimensional images extracted from computed tomography scans, but the same technique can be readily applied to other imaging modalities. FADS calculates anomaly scores on the basis of differences in filter activations of nominal and test data in pretrained convolutional neural networks. The FADS scores successfully differentiated between pristine PETN and PETN aged at a temperature where material coarsening occurred. Morphological metric analysis of segmented images verified observed trends in FADS scores as a function of aging temperature and aging time, specifically by calculating volume fractions, specific boundary lengths, two-point correlation functions, and local thicknesses. Here, the FADS technique has two important advantages compared to traditional morphological analysis: First, it uses grayscale images as input, rather than images that are segmented to separate the appropriate phases; and second, FADS scores capture any type of changes among image sets, rather than requiring prior knowledge or selection of a relevant set of metrics.

Accelerated aging↗

Reducing Communication Overhead in Federated Learning for Network Anomaly Detection with Adaptive Client Selection

Communication overhead in federated learning (FL) poses a significant challenge for network anomaly detection systems, where the myriad of client configurations and network conditions can severely impact system efficiency and detection accuracy. While existing approaches attempt to address this through individual optimization techniques, they often fail to maintain the delicate balance between reduced overhead and detection performance. This paper presents an adaptive FL framework that dynamically combines batch size optimization, client selection, and asynchronous updates to achieve efficient anomaly detection. Through extensive profiling and experimental analysis on two distinct datasets-UNSW-NBIS for general network traffic and ROAD for automotive networks-our framework reduces communication overhead by 97.6%; (from 700.0s to 16.8s) compared to synchronous baseline approaches while maintaining comparable detection accuracy (95.10%; vs. 95.12%;). Statistical validation using Mann-Whitney U test confirms significant improvements (p < 0.05) over existing FL approaches across both datasets, demonstrating the framework's adaptability to different network security contexts. Detailed profiling analysis reveals the efficiency gains through dramatic reductions in GPU operations and memory transfers while maintaining robust detection performance under varying client conditions.

Marfo, William [University of Texas at El Paso]↗

Cybersecurity Challenges in Low-Inertia Power-Electronics-Dominated Grids

Here, the low inertia characteristics of the power electronics dominated grid (PEDG) introduces challenges while restoring voltage and frequency to their nominal values. These stability challenges create new cybersecurity vulnerabilities that are not thoroughly discussed in the literature. Cyber events such as false data injection (FDI), denial of service (DoS), man-in-the-middle attacks, stealthy attacks, and advanced persistent threats target PEDG to disrupt grid stability or gain financial benefits. The low inertia of PEDG (< 2s) compared to traditional grids (~10s) exacerbates these vulnerabilities. In response to stealthy attacks on state variables that supervisory layers cannot detect until significant harm occurs, the low inertia characteristics of PEDG offer substantial stealthy attack surfaces. To counteract such threats, PEDG must be equipped with ultra-fast real-time anomaly detection system and trajectory prediction mechanism to achieve effective cyberattack resiliency.

24 POWER TRANSMISSION AND DISTRIBUTION↗

Exploring AI/ML-based Real-time Anomaly Detection in DUNE for Supernova Burst Neutrinos

The Deep Underground Neutrino Experiment (DUNE) is currently under construction with far detectors consisting of 4 liquid argon time projection chamber (LArTPC) modules at SURF (South Dakota Underground Research Facility) and a near detector complex with neutrino beam production at Fermilab to unambiguously determine neutrino mass ordering, to discover and precisely measure Charge-Parity (CP) violation phase in leptonic sector, to search for Beyond Stand Model (BSM) physics, and to study solar and supernova burst neutrinos. Anomalies in this project are classified in three categories: new physics signals, supernova burst neutrinos, and detector malfunction. We report here on promising early studies toward an Artificial Intelligence/Machine Learning-based real-time anomaly detection system, using a prototype autoencoder model currently under development. Additionally, the current status of an improved model and its performance will be presented. The model will be evaluated not only for its sensitivity to supernova neutrinos, but also to BSM physics signals and detector malfunctions. We will also consider how such a real-time algorithm might be used in DUNE.

de Jonge, Anselm [Kirchhoff Inst. Phys.] (ORCID:00↗

Accurate and Fast Anomaly Detection in Additive Composite-Based Manufacturing using Thermal Cameras

Today, large-scale additive manufacturing with plastics and composite materials requires continuous monitoring by experienced staff to prevent, detect and correct anomalous events affecting the performance of the printed part. We address the complexity of this demanding task by designing a camera-based anomaly detection system utilizing probabilistic principal component analysis (PPCA). This is a machine learning technique is trained with thermal images collected during normal operation of the large-scale printer (Cincinnati BAAM). This technique is advantageous for practical applications as there is no need to artificially introduce anomalous conditions into model training. During deployment, we challenge this model by introducing deliberate variations of the extruder speed. We reduce extrusion speed to a lower level, between 70 and 95% of the nominal value to collected test images. Our results show that images are easily identified as anomalous for extruder speeds at or below 85% of the nominal speed, meaning that an anomalous reduction of the material deposition rate can be detected within seconds of its onset. We show that our results are robust to (a) camera-to-camera variability and (b) print-to-print variability.

Pike, John [ORNL]↗