Search NASA⌕ Search

SEARCH · Search NASA

Results for “statistical fault detection”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

Toward Statistical Real-Time Power Fault Detection

We propose statistical fault detection methodology based on high-frequency data streams that are becoming available in modern power grids. Our approach can be treated as an online (sequential) change point monitoring methodology. However, due to the mostly unexplored and very nonstandard structure of high-frequency power grid streaming data, substantial new statistical development is required to make this methodology practically applicable. The paper includes development of scalar detectors based on multichannel data streams, determination of data-driven alarm thresholds and investigation of the performance and robustness of the new tools. Due to a reasonably large database of faults, we can calculate frequencies of false and correct fault signals, and recommend implementations that optimize these empirical success rates.

bolted faults↗

Fault detection and isolation

Erroneous measurements in multisensor navigation systems must be detected and isolated. A recursive estimator can find fast growing errors; a least squares batch estimator can find slow growing errors. This process is called fault detection. A protection radius can be calculated as a function of time for a given location. This protection radius can be used to guarantee the integrity of the navigation data. Fault isolation can be accomplished using either a snapshot method or by examining the history of the fault detection statistics.

Bernath, Greg↗

Automated Monitoring with a BCP Fault-Decision Test

The Bayesian conditional probability (BCP) technique is a statistical fault-decision technique that is suitable as the mathematical basis of the fault-manager module in the automated-monitoring system and method described in the immediately preceding article. Within the automated-monitoring system, the fault-manager module operates in conjunction with the fault-detector module, which can be based on any one of several fault-detection techniques; examples include a threshold-limit-comparison technique or the BSP or SPRT technique mentioned in the preceding article. The present BCP technique is used to evaluate a series of one or more fault-detection events for the purpose of filtering out occasional false alarms produced by many types of statistical fault-detection procedures. The BCP technique increases the probability that an automated monitoring system produces a correct decision regarding the presence or absence of a fault. Because occasional false alarms are an inevitable consequence of the SPRT, BSP, or any other statistically based fault-detection test, there is a need for a logical procedure to distinguish between true and false alarms. Heretofore, it has been common practice to make a fault decision on an ad hoc basis for example by following a multiple-observation voting strategy in which a signal is declared to be indicative of a fault if m of the last n observations produced a fault-detection alarm. The BCP technique was developed to obtain results more reliable than those afforded by a voting strategy. The BCP technique involves a test in which one applies Bayesian inference techniques to a series of one or more single-observation alarms produced by a fault-detection test. One considers the last n decisions generated by a fault-detection test in order to evaluate the conditional probability that a failure is indicated (see figure). Each new decision reached by a fault-detection test is treated as a new piece of evidence about the state of the monitored asset, and the conditional probability of failure for the system is updated on the basis of this new evidence. The conditional probability of failure is compared with a predefined limit. For a probability below the limit, the asset is declared to be healthy. For a probability above the limit, the asset is declared to be faulty.

Bickford, Randall L.↗

Image-based novel fault detection with deep learning classifiers using hierarchical labels

One important characteristic of modern fault classification systems is the ability to flag the system when faced with previously unseen fault types. This work considers the unknown fault detection capabilities of deep neural network-based fault classifiers. Specifically, we propose a methodology on how, when available, labels regarding the fault taxonomy can be used to increase unknown fault detection performance without sacrificing model performance. To achieve this, we propose to utilize soft label techniques to improve the state-of-the-art deep novel fault detection techniques during the training process and novel hierarchically consistent detection statistics for online novel fault detection. Lastly, we demonstrated increased detection performance on novel fault detection in inspection images from the hot steel rolling process, with results well replicated across multiple scenarios and baseline detection methods.

42 ENGINEERING↗

Advanced information processing system: Fault injection study and results

The objective of the AIPS program is to achieve a validated fault tolerant distributed computer system. The goals of the AIPS fault injection study were: (1) to present the fault injection study components addressing the AIPS validation objective; (2) to obtain feedback for fault removal from the design implementation; (3) to obtain statistical data regarding fault detection, isolation, and reconfiguration responses; and (4) to obtain data regarding the effects of faults on system performance. The parameters are described that must be varied to create a comprehensive set of fault injection tests, the subset of test cases selected, the test case measurements, and the test case execution. Both pin level hardware faults using a hardware fault injector and software injected memory mutations were used to test the system. An overview is provided of the hardware fault injector and the associated software used to carry out the experiments. Detailed specifications are given of fault and test results for the I/O Network and the AIPS Fault Tolerant Processor, respectively. The results are summarized and conclusions are given.

Burkhardt, Laura F.↗

Emulation and detection of physical faults and cyber-attacks on building energy systems through real-time hardware-in-the-loop experiments

The increasing use of remote or mobile access, integrated wearable technologies, data exchange, and cloud-based data analytics in modern smart buildings is steering the building industry towards open communication technologies. The increased connectivity and accessibility could lead to more cyber-attacks in smart buildings. On the other hand, physical faults (e.g., HVAC -heating, ventilation, and air-conditioning faults) may have similar adverse impacts as those from the cyber-attacks on building energy systems, such as occupant discomfort, energy wastage, and equipment downtime. However, current physical behavior-based anomaly detection methods fail to differentiate between cyber-attacks and physical faults in building energy systems. Moreover, the challenge in collecting real-world threat data with ground truth has led researchers to rely on numerical models with user-defined assumptions, which may not accurately reflect real-world conditions due to the lack of in-situ experimental datasets. To address these challenges and gaps, this paper presents a flexible hardware-in-the-loop (HIL) testbed for generating cyber-attack and physical fault datasets and demonstrating threat detection algorithms in a real building automation system (BAS) environment. This testbed combines hardware (i.e., real BAS with local HVAC controllers and a physical network) with software (i.e., high-fidelity models to represent behaviors of building envelope and HVAC energy systems), enabling emulations of realistic threats. Five HIL experiments, including one baseline without any threats, two with physical faults, and two with cyber-attacks, were conducted to generate datasets containing detailed network traffic and system states. A joint classification framework, incorporating a network analyzer and a physical HVAC fault detector, was proposed to automatically detect cyber-physical abnormalities on BAS at both the network and the physical HVAC levels. The network analyzer comprises a conditional random fields (CRF) based command validator and a statistics-based detection strategy. The fault detector employs a weather and schedule-based pattern matching and feature-based principal component analysis (WPM-FPCA) method. Evaluation of the classification using four metrics from the multi-class confusion matrix revealed an average accuracy of 90.2%, recall of 89.7%, precision of 88.5% and F1-score of 89.2%. Finally, these results demonstrate that the proposed joint classification framework can effectively differentiate between specific types of cyber-attacks (e.g., device reinitialization attack, network Denial-of-Service attack) and physical faults (e.g., air handling unit operational fault, cooling coil valve stuck) in real time for improved building energy management.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗

Supply Current Diagnosis in VLSI

This paper presents a technique based upon the power supply current signature (cd) which allows for the testing of mixed-signal systems, in situ. Through experiments with a microprocessor, the cd is shown to contain important information concerning the operational status of the system which may be easily extracted using approaches based on statistical signal detection theory. The fault-detection performance of these techniques is compared to that achieved through auto-regressive modeling of the cd.

Frenzel, J. F.↗

Discrete Data Qualification System and Method Comprising Noise Series Fault Detection

A Sensor Data Qualification (SDQ) function has been developed that allows the onboard flight computers on NASA s launch vehicles to determine the validity of sensor data to ensure that critical safety and operational decisions are not based on faulty sensor data. This SDQ function includes a novel noise series fault detection algorithm for qualification of the output data from LO2 and LH2 low-level liquid sensors. These sensors are positioned in a launch vehicle s propellant tanks in order to detect propellant depletion during a rocket engine s boost operating phase. This detection capability can prevent the catastrophic situation where the engine operates without propellant. The output from each LO2 and LH2 low-level liquid sensor is a discrete valued signal that is expected to be in either of two states, depending on whether the sensor is immersed (wet) or exposed (dry). Conventional methods for sensor data qualification, such as threshold limit checking, are not effective for this type of signal due to its discrete binary-state nature. To address this data qualification challenge, a noise computation and evaluation method, also known as a noise fault detector, was developed to detect unreasonable statistical characteristics in the discrete data stream. The method operates on a time series of discrete data observations over a moving window of data points and performs a continuous examination of the resulting observation stream to identify the presence of anomalous characteristics. If the method determines the existence of anomalous results, the data from the sensor is disqualified for use by other monitoring or control functions.

Fulton, Christopher↗

Characterization of fault recovery through fault injection on FTMP

The development of fault-injection procedures and statistical analysis techniques to characterize the fault recovery of fault-tolerant systems is described. Pin-level fault-injection was conducted on a fault-tolerant microprocessor computer in order to generate data to assess the utility of current fault-injection sampling methods. The validity of common reliability-modeling assumptions concerning the statistical distribution of recovery times is investigated. A multiple comparison analysis for detecting behavior variations, and a distribution fitting for determining the best fit for the data were conducted. It is observed that the detection behavior is not homogeneous across all data sets, and that none of the factors under experimental control can account for the observed groupings of behavior. It is determined that no single distribution fits all the data sets, and that stratified random sampling and statistically robust parameter-estimation techniques are required to characterize fault detection time.

Finelli, George B.↗

Detecting Latent Faults In Digital Flight Controls

Report discusses theory, conduct, and results of tests involving deliberate injection of low-level faults into digital flight-control system. Part of study of effectiveness of techniques for detection of and recovery from faults, based on statistical assessment of inputs and outputs of parts of control systems. Offers exceptional new capability to establish reliabilities of critical digital electronic systems in aircraft.

Mcgough, John↗

Sensor Selection and Data Validation for Reliable Integrated System Health Management

For new access to space systems with challenging mission requirements, effective implementation of integrated system health management (ISHM) must be available early in the program to support the design of systems that are safe, reliable, highly autonomous. Early ISHM availability is also needed to promote design for affordable operations; increased knowledge of functional health provided by ISHM supports construction of more efficient operations infrastructure. Lack of early ISHM inclusion in the system design process could result in retrofitting health management systems to augment and expand operational and safety requirements; thereby increasing program cost and risk due to increased instrumentation and computational complexity. Having the right sensors generating the required data to perform condition assessment, such as fault detection and isolation, with a high degree of confidence is critical to reliable operation of ISHM. Also, the data being generated by the sensors needs to be qualified to ensure that the assessments made by the ISHM is not based on faulty data. NASA Glenn Research Center has been developing technologies for sensor selection and data validation as part of the FDDR (Fault Detection, Diagnosis, and Response) element of the Upper Stage project of the Ares 1 launch vehicle development. This presentation will provide an overview of the GRC approach to sensor selection and data quality validation and will present recent results from applications that are representative of the complexity of propulsion systems for access to space vehicles. A brief overview of the sensor selection and data quality validation approaches is provided below. The NASA GRC developed Systematic Sensor Selection Strategy (S4) is a model-based procedure for systematically and quantitatively selecting an optimal sensor suite to provide overall health assessment of a host system. S4 can be logically partitioned into three major subdivisions: the knowledge base, the down-select iteration, and the final selection analysis. The knowledge base required for productive use of S4 consists of system design information and heritage experience together with a focus on components with health implications. The sensor suite down-selection is an iterative process for identifying a group of sensors that provide good fault detection and isolation for targeted fault scenarios. In the final selection analysis, a statistical evaluation algorithm provides the final robustness test for each down-selected sensor suite. NASA GRC has developed an approach to sensor data qualification that applies empirical relationships, threshold detection techniques, and Bayesian belief theory to a network of sensors related by physics (i.e., analytical redundancy) in order to identify the failure of a given sensor within the network. This data quality validation approach extends the state-of-the-art, from red-lines and reasonableness checks that flag a sensor after it fails, to include analytical redundancy-based methods that can identify a sensor in the process of failing. The focus of this effort is on understanding the proper application of analytical redundancy-based data qualification methods for onboard use in monitoring Upper Stage sensors.

Garg, Sanjay↗

Long-Term Statistical Process Monitoring of an Ultrafiltration Water Treatment Process

As water treatment technology has improved, the amount of available process data has substantially increased, making real-time, data-driven fault detection a reality. One shortcoming of the fault detection literature is that methods are usually evaluated by comparing their performance on hand-picked, short-term case studies, which yields no insight into long-term performance. In this work, we first evaluate multiple statistical and machine learning approaches for detrending process data. Then, we evaluate the performance of a PCA-based fault detection approach, applied to the detrended data, to monitor influent water quality, filtrate quality, and membrane fouling of an ultrafiltration membrane system for indirect potable reuse. Based on two short case studies, the adaptive lasso detrending method is selected, and the performance of the multivariate approach is evaluated over more than a year. The method is tested for different sets of three critical tuning parameters, and we find that for long-term, autonomous monitoring to be successful, these parameters should be carefully evaluated. However, in comparison with industry standards of simpler, univariate monitoring or daily pressure decay tests, multivariate monitoring produces substantial benefits in long-term testing.

ammonia↗

Neural net diagnostics for VLSI test

This paper discusses the application of neural network pattern analysis algorithms to the IC fault diagnosis problem. A fault diagnostic is a decision rule combining what is known about an ideal circuit test response with information about how it is distorted by fabrication variations and measurement noise. The rule is used to detect fault existence in fabricated circuits using real test equipment. Traditional statistical techniques may be used to achieve this goal, but they can employ unrealistic a priori assumptions about measurement data. Our approach to this problem employs an adaptive pattern analysis technique based on feedforward neural networks. During training, a feedforward network automatically captures unknown sample distributions. This is important because distributions arising from the nonlinear effects of process variation can be more complex than is typically assumed. A feedforward network is also able to extract measurement features which contribute significantly to making a correct decision. Traditional feature extraction techniques employ matrix manipulations which can be particularly costly for large measurement vectors. In this paper we discuss a software system which we are developing that uses this approach. We also provide a simple example illustrating the use of the technique for fault detection in an operational amplifier.

Lin, T.↗

Artemis II Inertial Sensor Failure & Abort Condition Study

Given Artemis II as the first crewed lunar mission in decades, significant effort is directed towards understanding Space Launch System (SLS) response to failures of the three inertial sensors: two Ring-Laser Gyro Assemblies (RGAs) and the Redundant Inertial Navigation Unit (RINU). Parametric studies identify Loss of Mission (LOM) or Loss of Crew (LOC) sensitive mission phases. Fault Detection & Isolation (FDI) algorithms’ thresholds’ performance is assessed with respect to detecting LOM or LOC causing sensor failures. Monte Carlo studies are performed targeting specific trajectory phases for insight into re-optimizing FDI thresholds and improving knowledge of SLS sensor failure response statistics.

SLS↗

Machine Learning–Based Condition Monitoring of a Circulating Water System of a Canadian Nuclear Plant

With the need to maintain long-term reliable energy using nuclear power plants, there is an underlying demand to ensure that the maintenance of plant components and systems is also done in an efficient and cost-effective manner. One way to achieve this is by moving from time-based maintenance to condition-based maintenance. The research presented in this paper focuses on applying statistical and machine-learning-based methods to capture anomalies within data for fault detection to further develop into condition monitoring. This paper focuses on system data for a circulating water system (CWS) of a pressurized heavy-water reactor for detecting anomalies. The different methodologies used for detecting and capturing anomalies in the CWS data are matrix profile, density-based spatial clustering of applications with noise (DBSCAN), and support vector machines (SVMs). Matrix profile and DBSCAN are used to distinguish between normal data and anomalous data. This paper presents a hybrid method using DBSCAN and SVM when a portion of the data is used for DBSCAN to generate clusters. This portion of data is then used to train the SVM along with the clusters generated by DBSCAN as output. SVM is then tested on unseen data as a predictive tool, which can work in real time to categorize data points as either normal or anomalous. This paper presents results that show the high accuracies of DBSCAN and SVM in capturing anomalies within the data for a CWS for fault detection. Thus, the maintenance plan would be focused on component condition rather than a time-based schedule by switching to an automated system to identify and predict faults within a CWS.

21 SPECIFIC NUCLEAR REACTORS AND ASSOCIATED PLANTS↗

Fault detection using a two-model test for changes in the parameters of an autoregressive time series

This article describes an investigation of a statistical hypothesis testing method for detecting changes in the characteristics of an observed time series. The work is motivated by the need for practical automated methods for on-line monitoring of Deep Space Network (DSN) equipment to detect failures and changes in behavior. In particular, on-line monitoring of the motor current in a DSN 34-m beam waveguide (BWG) antenna is used as an example. The algorithm is based on a measure of the information theoretic distance between two autoregressive models: one estimated with data from a dynamic reference window and one estimated with data from a sliding reference window. The Hinkley cumulative sum stopping rule is utilized to detect a change in the mean of this distance measure, corresponding to the detection of a change in the underlying process. The basic theory behind this two-model test is presented, and the problem of practical implementation is addressed, examining windowing methods, model estimation, and detection parameter assignment. Results from the five fault-transition simulations are presented to show the possible limitations of the detection method, and suggestions for future implementation are given.

Scholtz, P.↗

Strategy Developed for Selecting Optimal Sensors for Monitoring Engine Health

Sensor indications during rocket engine operation are the primary means of assessing engine performance and health. Effective selection and location of sensors in the operating engine environment enables accurate real-time condition monitoring and rapid engine controller response to mitigate critical fault conditions. These capabilities are crucial to ensure crew safety and mission success. Effective sensor selection also facilitates postflight condition assessment, which contributes to efficient engine maintenance and reduced operating costs. Under the Next Generation Launch Technology program, the NASA Glenn Research Center, in partnership with Rocketdyne Propulsion and Power, has developed a model-based procedure for systematically selecting an optimal sensor suite for assessing rocket engine system health. This optimization process is termed the systematic sensor selection strategy. Engine health management (EHM) systems generally employ multiple diagnostic procedures including data validation, anomaly detection, fault-isolation, and information fusion. The effectiveness of each diagnostic component is affected by the quality, availability, and compatibility of sensor data. Therefore systematic sensor selection is an enabling technology for EHM. Information in three categories is required by the systematic sensor selection strategy. The first category consists of targeted engine fault information; including the description and estimated risk-reduction factor for each identified fault. Risk-reduction factors are used to define and rank the potential merit of timely fault diagnoses. The second category is composed of candidate sensor information; including type, location, and estimated variance in normal operation. The final category includes the definition of fault scenarios characteristic of each targeted engine fault. These scenarios are defined in terms of engine model hardware parameters. Values of these parameters define engine simulations that generate expected sensor values for targeted fault scenarios. Taken together, this information provides an efficient condensation of the engineering experience and engine flow physics needed for sensor selection. The systematic sensor selection strategy is composed of three primary algorithms. The core of the selection process is a genetic algorithm that iteratively improves a defined quality measure of selected sensor suites. A merit algorithm is employed to compute the quality measure for each test sensor suite presented by the selection process. The quality measure is based on the fidelity of fault detection and the level of fault source discrimination provided by the test sensor suite. An inverse engine model, whose function is to derive hardware performance parameters from sensor data, is an integral part of the merit algorithm. The final component is a statistical evaluation algorithm that characterizes the impact of interference effects, such as control-induced sensor variation and sensor noise, on the probability of fault detection and isolation for optimal and near-optimal sensor suites.

Source record↗

Topographic Slant Range Modeling and Fault Detection for Precision Planetary Landing

This work presents a novel landing site relative topographic measurement model for aslant range sensor being utilized for precision planetary landing operations. The measurement model accounts for the local terrain the slant range sensor captures and leverages knowledge of the estimated landing site provided by the navigation filter. Notably, in contrast to previous works, the new model does not rely on surface normal approximation, reducing the model’s sensitivity to noisy digital elevation maps which represent the local topography. In addition to the measurement model, this work introduces a novel fault detection method, denoted the probabilistic inspection of topographic filter altitude likelihood (PITFAL) algorithm, that implements a statistical outlier rejection algorithm. PITFAL is designed for multi-beam slant range sensors, such as the Navigation Doppler LIDAR (NDL), and identifies statistically inconsistent range estimates through a consensus check on the set of apparent altitudes computed for each individual beam. These models are numerically validated by the Safe and Precise Landing Capability Evolution (SPLICE) project’s high-fidelity terrestrial and lunar lander simulations.

Davis W Adams↗