Search NASA⌕ Search

SEARCH · Search NASA

Results for “Fault Injection”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 109 records · Page 6

Development and validation of techniques for improving software dependability

A collection of document abstracts are presented on the topic of improving software dependability through NASA grant NAG-1-1123. Specific topics include: modeling of error detection; software inspection; test cases; Magnetic Stereotaxis System safety specifications and fault trees; and injection of synthetic faults into software.

Knight, John C.↗

Injecting Artificial Memory Errors Into a Running Computer Program

Single-event upsets (SEUs) or bitflips are computer memory errors caused by radiation. BITFLIPS (Basic Instrumentation Tool for Fault Localized Injection of Probabilistic SEUs) is a computer program that deliberately injects SEUs into another computer program, while the latter is running, for the purpose of evaluating the fault tolerance of that program. BITFLIPS was written as a plug-in extension of the open-source Valgrind debugging and profiling software. BITFLIPS can inject SEUs into any program that can be run on the Linux operating system, without needing to modify the program s source code. Further, if access to the original program source code is available, BITFLIPS offers fine-grained control over exactly when and which areas of memory (as specified via program variables) will be subjected to SEUs. The rate of injection of SEUs is controlled by specifying either a fault probability or a fault rate based on memory size and radiation exposure time, in units of SEUs per byte per second. BITFLIPS can also log each SEU that it injects and, if program source code is available, report the magnitude of effect of the SEU on a floating-point value or other program variable.

Bornstein, Benjamin J.↗

A quantitative risk assessment framework for fault reactivation in underground hydrogen storage: Coupled simulation and deep learning approach

Underground hydrogen storage (UHS) is emerging as a critical solution for large-scale energy storage. However, like all subsurface fluid injection activities, UHS poses the risk of injection-induced fault reactivation. Accurate risk assessment is essential to ensuring the safety and efficiency of UHS operations. This study presents the development of deep-learning surrogate models for fault reactivation prediction in UHS, trained on a comprehensive database of fully coupled fluid flow-geomechanics simulations. Our findings reveal that analytical models often yield unreliable estimates, with errors up to 54% in the allowable injection pressure, potentially leading to a 40% reduction in UHS operational capacity. The developed surrogate models were incorporated into a quantitative risk assessment (QRA) framework, enabling probabilistic evaluation of fault reactivation risk while accounting for uncertainties in the input variables. Site-specific features, such as horizontal stress gradients, fault’s dip and strike angles, and operational parameters like bottom-hole injection pressure and well-fault distance, were identified as the primary drivers of fault reactivation across various stress regimes. Whereas other hydraulic, geological, and poroelastic reservoir properties were found to have a secondary impact. Notably, we observed that the risk of fault reactivation for a critically oriented fault with a static friction coefficient greater than 0.55 remains below 10% in a normal faulting stress regime. However, the risk significantly increases as the stress regime transitions from normal to strike-slip and ultimately to reverse faulting conditions. These findings underscore the importance of rigorous site characterization and comprehensive QRA evaluations to optimize UHS performance and minimize geomechanical risks.

25 ENERGY STORAGE↗

Predeployment validation of fault-tolerant systems through software-implemented fault insertion

Fault injection-based automated testing (FIAT) environment, which can be used to experimentally characterize and evaluate distributed realtime systems under fault-free and faulted conditions is described. A survey is presented of validation methodologies. The need for fault insertion based on validation methodologies is demonstrated. The origins and models of faults, and motivation for the FIAT concept are reviewed. FIAT employs a validation methodology which builds confidence in the system through first providing a baseline of fault-free performance data and then characterizing the behavior of the system with faults present. Fault insertion is accomplished through software and allows faults or the manifestation of faults to be inserted by either seeding faults into memory or triggering error detection mechanisms. FIAT is capable of emulating a variety of fault-tolerant strategies and architectures, can monitor system activity, and can automatically orchestrate experiments involving insertion of faults. There is a common system interface which allows ease of use to decrease experiment development and run time. Fault models chosen for experiments on FIAT have generated system responses which parallel those observed in real systems under faulty conditions. These capabilities are shown by two example experiments each using a different fault-tolerance strategy.

Czeck, Edward W.↗

Subsurface Energy Systems Mapping Inquiry Tool (MapIT)

The Subsurface Energy Systems Mapping Inquiry Tool (MapIT) is an online web mapping tool designed to help users discover available public-sourced data to facilitate data exploration for subsurface energy exploration and characterization efforts for resource identification (e.g. critical minerals, hydrocarbons, geothermal) as well as injection of geologic sequestration of carbon dioxide (e.g. enhanced oil recovery, saline storage, etc.). Modules within the tool curate data related to geology, faults, fractures, injection and confining zones, hydrologic information, groundwater, groundwater wells, geomechanical and petrophysical data, and geochemical data. User documentation on how to use the tool is also provided. Data have been collected from authoritative national, state, and local sources and made available in this tool. The data is also available as a data catalog and Esri Geodatabase at: https://edx.netl.doe.gov/dataset/mapit-database Disclaimer: There is no guarantee of completeness or appropriateness for individual user’s requirements. Use of this tool is solely at the discretion of the user. See full Federal Disclaimer for further information (https://netl.doe.gov/home/disclaimer). This project was funded by the United States Department of Energy, National Energy Technology Laboratory, in part, through a site support contract. Neither the United States Government nor any agency thereof, nor any of their employees, nor the support contractor, nor any of their employees, makes any warranty, express or implied, or assumes any legal liability or responsibility for the accuracy, completeness, or usefulness of any information, apparatus, product, or process disclosed, or represents that its use would not infringe privately owned rights. Reference herein to any specific commercial product, process, or service by trade name, trademark, manufacturer, or otherwise does not necessarily constitute or imply its endorsement, recommendation, or favoring by the United States Government or any agency thereof. The views and opinions of authors expressed herein do not necessarily state or reflect those of the United States Government or any agency thereof. https://www.netl.doe.gov/home/disclaimer

Carbon Sequestration↗

Impact of device level faults in a digital avionic processor

This paper describes an experimental analysis of the impact of gate and device-level faults in the processor of a flight control system. Via mixed mode simulation faults were injected both at the gate (stuck-at) and at the transistor levels, and their propagation through the chip to the output pins was measured. The results show that there is little correspondence between a stuck-at and a device-level fault model insofar as error activity or detection within a functional unit is concerned. Insofar as error activity outside the injected unit and at the output pins are concerned, the stuck-at and device models track each other, although the stuck-at model overestimates, by over one hundred percent, the probability of fault propagation to the output pins. The stuck-at model significantly underestimates the impact of an internal chip fault on the output pins.

Kim, S.↗

Impact of device level faults in a digital avionic processor

This study describes an experimental analysis of the impact of gate and device-level faults in the processor of a Bendix BDX-930 flight control system. Via mixed mode simulation, faults were injected at the gate (stuck-at) and at the transistor levels and, their propagation through the chip to the output pins was measured. The results show that there is little correspondence between a stuck-at and a device-level fault model, as far as error activity or detection within a functional unit is concerned. In so far as error activity outside the injected unit and at the output pins are concerned, the stuck-at and device models track each other. The stuck-at model, however, overestimates, by over 100 percent, the probability of fault propagation to the output pins. An evaluation of the Mean Error Durations and the Mean Time Between Errors at the output pins shows that the stuck-at model significantly underestimates (by 62 percent) the impact of an internal chip fault on the output pins. Finally, the study also quantifies the impact of device fault by location, both internally and at the output pins.

Suk, Ho Kim↗

Validation environment for AIPS/ALS: Implementation and results

The work is presented which was performed in porting the Fault Injection-based Automated Testing (FIAT) and Programming and Instrumentation Environments (PIE) validation tools, to the Advanced Information Processing System (AIPS) in the context of the Ada Language System (ALS) application, as well as an initial fault free validation of the available AIPS system. The PIE components implemented on AIPS provide the monitoring mechanisms required for validation. These mechanisms represent a substantial portion of the FIAT system. Moreover, these are required for the implementation of the FIAT environment on AIPS. Using these components, an initial fault free validation of the AIPS system was performed. The implementation is described of the FIAT/PIE system, configured for fault free validation of the AIPS fault tolerant computer system. The PIE components were modified to support the Ada language. A special purpose AIPS/Ada runtime monitoring and data collection was implemented. A number of initial Ada programs running on the PIE/AIPS system were implemented. The instrumentation of the Ada programs was accomplished automatically inside the PIE programming environment. PIE's on-line graphical views show vividly and accurately the performance characteristics of Ada programs, AIPS kernel and the application's interaction with the AIPS kernel. The data collection mechanisms were written in a high level language, Ada, and provide a high degree of flexibility for implementation under various system conditions.

Segall, Zary↗

Advanced Diagnostic System on Earth Observing One

In this infusion experiment, the Livingstone 2 (L2) model-based diagnosis engine, developed by the Computational Sciences division at NASA Ames Research Center, has been uploaded to the Earth Observing One (EO-1) satellite. L2 is integrated with the Autonomous Sciencecraft Experiment (ASE) which provides an on-board planning capability and a software bridge to the spacecraft's 1773 data bus. Using a model of the spacecraft subsystems, L2 predicts nominal state transitions initiated by control commands, monitors the spacecraft sensors, and, in the case of failure, isolates the fault based on the discrepant observations. Fault detection and isolation is done by determining a set of component modes, including most likely failures, which satisfy the current observations. All mode transitions and diagnoses are telemetered to the ground for analysis. The initial L2 model is scoped to EO-1's imaging instruments and solid state recorder. Diagnostic scenarios for EO-1's nominal imaging timeline are demonstrated by injecting simulated faults on-board the spacecraft. The solid state recorder stores the science images and also hosts: the experiment software. The main objective of the experiment is to mature the L2 technology to Technology Readiness Level (TRL) 7. Experiment results are presented, as well as a discussion of the challenging technical issues encountered. Future extensions may explore coordination with the planner, and model-based ground operations.

Hayden, Sandra C.↗

Experimental evaluation of a COTS system for space applications

The use of COTS-based systems in space missions for scientific data processing is very attractive, as their ratio of performance to power consumption of commercial components can be an order of magnitude greater than that of radiation hardened components, and the price differential is even higher.

fault injection on-board processing cluster commut↗

An experimental evaluation of the REE SIFT environment for spaceborne applications

This paper presents an experimental evaluation of a software-implemented fault tolerance environment built around a set of self-checking ARMOR proceses running on different machines that provide error detection and recovery services to themselves and to spaceborne scientific applications.

fault injection on-board processing cluster commut↗

Portable Health Algorithms Test System

A document discusses the Portable Health Algorithms Test (PHALT) System, which has been designed as a means for evolving the maturity and credibility of algorithms developed to assess the health of aerospace systems. Comprising an integrated hardware-software environment, the PHALT system allows systems health management algorithms to be developed in a graphical programming environment, to be tested and refined using system simulation or test data playback, and to be evaluated in a real-time hardware-in-the-loop mode with a live test article. The integrated hardware and software development environment provides a seamless transition from algorithm development to real-time implementation. The portability of the hardware makes it quick and easy to transport between test facilities. This hard ware/software architecture is flexible enough to support a variety of diagnostic applications and test hardware, and the GUI-based rapid prototyping capability is sufficient to support development execution, and testing of custom diagnostic algorithms. The PHALT operating system supports execution of diagnostic algorithms under real-time constraints. PHALT can perform real-time capture and playback of test rig data with the ability to augment/ modify the data stream (e.g. inject simulated faults). It performs algorithm testing using a variety of data input sources, including real-time data acquisition, test data playback, and system simulations, and also provides system feedback to evaluate closed-loop diagnostic response and mitigation control.

Melcher, Kevin J.↗

A Cryogenic Fluid System Simulation in Support of Integrated Systems Health Management

Simulations serve as important tools throughout the design and operation of engineering systems. In the context of sys-tems health management, simulations serve many uses. For one, the underlying physical models can be used by model-based health management tools to develop diagnostic and prognostic models. These simulations should incorporate both nominal and faulty behavior with the ability to inject various faults into the system. Such simulations can there-fore be used for operator training, for both nominal and faulty situations, as well as for developing and prototyping health management algorithms. In this paper, we describe a methodology for building such simulations. We discuss the design decisions and tools used to build a simulation of a cryogenic fluid test bed, and how it serves as a core technology for systems health management development and maturation.

cryogenics↗

Cryogenic Fuel Valve Testbed Development

The goal for this project is to update the cryogenic valve testbed program in LabVIEW to schedule and automate tests and experiments. By using an automated system, tens or hundreds of tests may be performed. This will ensure that accurate data is being collected for testing of the remaining useful life and end of life predictions. From the data obtained, new diagnostic and prognostic methods will be developed to manage or predict potential leaks which may occur in the future. The Cryogenic valve testbed injects controlled faults into the cryogenic fuel valve system in order to accurately determine failure behavior.

Prognostics↗

Validation of an SEU simulation technique for a complex processor: PowerPC7400

Published data on the processors sensitivites with respect to SEU is generally obtained from radiation ground testing during which the program is executed by the DUT consists in the sequential inspection of each of the processor memory cells accessible to the user, through the execution of a suitable instruction sequence. In such programs, so-called static tests, typically considered memory cells are general-purpose registers, special registers (program counter, stack pointer...) and internal memory. Nevertheless, the register use and duty cycle of the final application will be very different, including using instructions no in the static tests and disturbing other potential SEU targets. The ideal would be the use of the final application program for the radiation ground testing, but generally this program is either unknown or unavailable when the qualification testing is performed on candidate circuits to space projects.

Radiation↗

Uncovering Hazards Using a Multi-Objective Optimization to Explore the Faulty State-Space

Considering resilience when designing complex engineered systems is crucial to ensure the system is safe under unexpected hazardous scenarios. Traditional risk-based approaches, such as Failure Modes and Effects Analysis (FMEA) are useful for designing the system to mitigate hazardous scenarios that can be identified by the designer, but often require experience or prior knowledge of system failures to generate. More recently, researchers have developed simulation tools that enable the designer to model large sets of hazardous scenarios (driven by both internal faults and external factors) through simulation. While these tools enable a wider scope of fault modes to be evaluated (e.g., by injecting combined set of fault modes or injecting modes at different times), the resulting assessments (like FMEA) still require knowledge of the specific modes to be evaluated. However, failure to analyze a wide variety of fault scenarios can lead to an incomplete picture of the system resilience, especially to "surprise events'' which may be difficult for the designer to identify and predict beforehand. To overcome this challenge, previous work developed a fault sampling approach for resilience simulations which would procedurally-generate a wide variety of potential faults by systematically perturbing the health states of the system. While the resulting fault modes generated covered a much larger space hazards than would be otherwise considered (and identified many unique failure trajectories which would not have otherwise been identified), it also significantly increased the computational cost of the analysis and resulted in the simulation and analysis of a large set of essentially duplicate scenarios. Additionally, as the number of dimensions in the faulty state-space increases, the full elaboration of possible modes becomes computationally infeasible, justifying the use of a more targeted search. To resolve this limitation, this work proposes the use of a multiobjective optimization algorithm to search the health state space for potential fault modes that are both (1) hazardous and (2) unique. To solve this type of problem, this work proposes the use of a cooperative co-evolutionary algorithm. To demonstrate this approach, it will be applied to a model of an autonomous rover which uses line markings to navigate, focusing on potential hazards in the drive system which could cause the rover to crash. To determine the merit of the approach, it will further be compared with the previously-presented range elaboration approach and a random mode generation approach on the basis of computational efficiency and found modes.

Resilience↗

Fault-tolerance experiments with the JPL STAR computer.

Results of fault-tolerance experiments performed using an experimental computer with dynamic (standby) redundancy, including replaceable subsystems and a 'program rollback' provision to eliminate transient-caused errors. After a brief review of the specification of fault-tolerance with respect to transient faults, including a description of the method of injection of transient faults in software and system tests, fault-tolerance experiments carried out with this computer with regard to the determination of fault classes, software verification, system verification, and recovery stability are summarized. A test and repair processor is described which constitutes a special monitor unit of the computer and is used to obtain information for fault detection in the other subsystems of the computer and to ensure that proper recovery occurs when a fault is detected.

Avizienis, A.↗