Search NASA⌕ Search

SEARCH · Search NASA

Results for “Fault Injection”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 91 records · Page 5

Validation of the Mars 2020 Fault Protection Design: Navigating the Infinity of the Off-Nominal

On July 30th 2020, the Mars 2020 mission successfully launched out of Cape Canaveral, Florida, passed through the Earth’s shadow, and began its short cruise to Mars. Less than seven months later, the Perseverance rover touched down safely in Jezero Crater to begin its ambitious mission that includes looking for signs of ancient life and collecting samples for future return to Earth. Getting to the successful landing, or “Tango Delta Nominal,” could not have been achieved without also considering the off-nominal. One of the teams supporting this ambitious mission is the fault protection (FP) team. This team is tasked with assessing the various failures, or faults, that could prevent mission success and with ensuring that the autonomous behaviors built into the software and hardware can detect faults and recover the vehicle to a safe state. As part of its charter, the FP team designed a test campaign to provide confidence in the system’s robustness to off-nominal scenarios across all of Mars 2020’s mission phases. The greatest challenge associated with designing such a validation campaign was reducing the infinite number of anomalous scenarios into a finite test suite. In addition, the tests needed to be executed efficiently in order to utilize the team’s limited test venue access, but still needed to maintain a level of rigor that guaranteed confidence in the test outcomes. Given that each test scenario generated massive amounts of data, the team also developed methods for quickly ascertaining whether the autonomous fault protection behaviors maintained vehicle safety in the presence of an anomaly. This paper summarizes the processes that the Mars 2020 fault protection team employed to execute its off-nominal validation campaign. It captures both the methods of generating a suite of off-nominal tests, as well as reducing it to a subset that can be realistically executed within schedule and resource constraints. It also describes the various processes and philosophies that the team utilized to execute the tests efficiently, including creating a standardized procedure template, keeping the test cases modular so that they could be easily interchanged, and capturing common fault injections in a change-controlled database. Finally, it will describe the tools and processes for assessing the test data, focusing in particular on a tool that evaluated vehicle state using “secondary” sources of data to validate that the software had truly configured the spacecraft to the expected safe state.

Morantz, Chaz↗

Adaptive Independent Verification and Validation (IV&V) Reduces Risk of Software Impacting Safety in Artemis Missions

The National Aeronautics and Space Administration (NASA) is asking more of its human spaceflight programs than ever before through the collective Artemis Missions. The NASA Independent Verification and Validation (IV&V) Program contributes to NASA’s human spaceflight goals by providing IV&V services for NASA’s critical spacecraft and ground software. The IV&V Program is tasked with providing assurance from both individual and integrated mission software perspectives. The Artemis IV&V organization is actively supporting six distinct development efforts: Orion, the Space Launch System (SLS), Exploration Ground Systems (EGS), Mission Control Center (MCC), the Lunar Gateway, and the Human Landing System (HLS), representing a wide diversity of developer organizations, management structures, and development approaches. With much of this extremely complex flight and ground software being essential to human safety both on the ground and in space, Artemis IV&V is likewise challenged to provide more value-added assurance to future Artemis missions within a constrained budget. To meet this challenge, Artemis IV&V employs a variety of novel and evolving “Adaptive IV&V” approaches for planning and executing IV&V analysis to increase both the efficiency and effectiveness of the IV&V Program’s assurance activities, and to address the difficulties imposed by assuring software for a large, highly integrated, multi-mission enterprise managed and executed by physically and organizationally distinct programs. Instilling agile principles like iterative planning cycles, self-organizing teams, and regular retrospectives, into IV&V planning and execution has led to a more rapid turnaround of a minimum viable assurance product and allowed for increased alignment of assurance activities with development progress. Adopting an assurance case methodology has led to greater consistency and clearer communication of assurance design and provided a foundation for long-term maintenance of assurance plans, products, and results across missions. The IV&V-developed Assurance / Safety Case Analytical Network (A-SCAN) framework and tool has enabled the quantification and tracking of system/software risk and confidence. These confidence measures provide a means to repeatedly express the impact of planned and completed assurance work and the remaining residual risk. Applied as part of a “Follow-the-Risk” organizational ethos, this allows consistent rightsizing of analysis rigor and intensity commensurate with the perceived risk of defects, as well as appropriate targeting of the highest risk areas of the software to find safety issues before they can manifest. Finally, the development of the IV&V Advanced Risk Reduction Integrated Software Test and Operations Tri-program Lightweight Environment (ARRISTOTLE), an integrated software-only simulation of Orion, SLS, and EGS systems, has made it possible to independently test integrated pad and flight scenarios and inject faults to observe how the Artemis multi-program, mission software behaves in degraded modes and in response to hazards. These adaptive IV&V investments have enabled Artemis IV&V to become more efficient and effective in IV&V planning and execution and respond more readily to changes in the risk landscape, increasing the breadth and depth of risk reduction possible within the available resources. Residual risk tracking allows IV&V to communicate more effectively with stakeholders, both internal and external at all levels, and inform key decision-making personnel. This evolving assurance design approach provides IV&V surety that work is performed in the highest risk, most value-added areas of the software, to keep our astronauts and ground crews safe and ensure mission success.

Gerek Whitman↗

Design for dependability: A simulation-based approach

This research addresses issues in simulation-based system level dependability analysis of fault-tolerant computer systems. The issues and difficulties of providing a general simulation-based approach for system level analysis are discussed and a methodology that address and tackle these issues is presented. The proposed methodology is designed to permit the study of a wide variety of architectures under various fault conditions. It permits detailed functional modeling of architectural features such as sparing policies, repair schemes, routing algorithms as well as other fault-tolerant mechanisms, and it allows the execution of actual application software. One key benefit of this approach is that the behavior of a system under faults does not have to be pre-defined as it is normally done. Instead, a system can be simulated in detail and injected with faults to determine its failure modes. The thesis describes how object-oriented design is used to incorporate this methodology into a general purpose design and fault injection package called DEPEND. A software model is presented that uses abstractions of application programs to study the behavior and effect of software on hardware faults in the early design stage when actual code is not available. Finally, an acceleration technique that combines hierarchical simulation, time acceleration algorithms and hybrid simulation to reduce simulation time is introduced.

Goswami, Kumar K.↗

Load flows and faults considering dc current injections

The authors present novel methods for incorporating current injection sources into dc power flow computations and determining network fault currents when electronic devices limit fault currents. Combinations of current and voltage sources into a single network are considered in a general formulation. An example of relay coordination is presented. The present study is pertinent to the development of the Space Station Freedom electrical generation, transmission, and distribution system.

Kusic, G. L.↗

Measurement and analysis of workload effects on fault latency in real-time systems

The authors demonstrate the need to address fault latency in highly reliable real-time control computer systems. It is noted that the effectiveness of all known recovery mechanisms is greatly reduced in the presence of multiple latent faults. The presence of multiple latent faults increases the possibility of multiple errors, which could result in coverage failure. The authors present experimental evidence indicating that the duration of fault latency is dependent on workload. A synthetic workload generator is used to vary the workload, and a hardware fault injector is applied to inject transient faults of varying durations. This method makes it possible to derive the distribution of fault latency duration. Experimental results obtained from the fault-tolerant multiprocessor at the NASA Airlab are presented and discussed.

Woodbury, Michael H.↗

Viscosity determinations of some frictionally generated silicate melts: Implications for slip zone rheology during impact-induced faulting

Analytical scanning electron microscopy, using combined energy dispersive and wavelength dispersive spectrometry, was used to determine the major-element compositions of some natural and artificial glasses and their crystalline equivalents derived by the frictional melting of acid to intermediate protoliths. The major-element compositions are used to calculate the viscosities of their melt precursors using the model of Shaw at temperatures of 800-1400 C, with Fe(2+)/Fe(tot) = 0.5 and for 1-3 wt percent H2O. These results are then modified to account for suspension effects in order to determine viscosities. The results have implications for the generation of pseudotachylitic breccias as seen in the basement lithologies of the Sudbury and Vredefort structures and possibly certain dimict lunar breccias. Many of these breccias show similarities with the more commonly developed pseudotachylite fault and injection veins seen in endogenic fault zones that typically occur in thicknesses of a few centimeters or less. The main difference is one of scale: Impact-induced pseudotachylite breccias can attain several meters in thickness. This would suggest that they were generated under exceptionally high slip rates and hence high strain rates and that the friction melts generated possessed extremely low viscosities.

Spray, John G.↗

Development and evaluation of a Fault-Tolerant Multiprocessor (FTMP) computer. Volume 3: FTMP test and evaluation

The experimental test and evaluation of the Fault-Tolerant Multiprocessor (FTMP) is described. Major objectives of this exercise include expanding validation envelope, building confidence in the system, revealing any weaknesses in the architectural concepts and in their execution in hardware and software, and in general, stressing the hardware and software. To this end, pin-level faults were injected into one LRU of the FTMP and the FTMP response was measured in terms of fault detection, isolation, and recovery times. A total of 21,055 stuck-at-0, stuck-at-1 and invert-signal faults were injected in the CPU, memory, bus interface circuits, Bus Guardian Units, and voters and error latches. Of these, 17,418 were detected. At least 80 percent of undetected faults are estimated to be on unused pins. The multiprocessor identified all detected faults correctly and recovered successfully in each case. Total recovery time for all faults averaged a little over one second. This can be reduced to half a second by including appropriate self-tests.

Lala, J. H.↗

Graphics enhanced computer emulation for improved timing-race and fault tolerance control system analysis

A computer simulation system has been developed for the Space Shuttle's advanced Centaur liquid fuel booster rocket, in order to conduct systems safety verification and flight operations training. This simulation utility is designed to analyze functional system behavior by integrating control avionics with mechanical and fluid elements, and is able to emulate any system operation, from simple relay logic to complex VLSI components, with wire-by-wire detail. A novel graphics data entry system offers a pseudo-wire wrap data base that can be easily updated. Visual subsystem operations can be selected and displayed in color on a six-monitor graphics processor. System timing and fault verification analyses are conducted by injecting component fault modes and min/max timing delays, and then observing system operation through a red line monitor.

Szatkowski, G. P.↗

Development and validation of techniques for improving software dependability

A collection of document abstracts are presented on the topic of improving software dependability through NASA grant NAG-1-1123. Specific topics include: modeling of error detection; software inspection; test cases; Magnetic Stereotaxis System safety specifications and fault trees; and injection of synthetic faults into software.

Knight, John C.↗

Injecting Artificial Memory Errors Into a Running Computer Program

Single-event upsets (SEUs) or bitflips are computer memory errors caused by radiation. BITFLIPS (Basic Instrumentation Tool for Fault Localized Injection of Probabilistic SEUs) is a computer program that deliberately injects SEUs into another computer program, while the latter is running, for the purpose of evaluating the fault tolerance of that program. BITFLIPS was written as a plug-in extension of the open-source Valgrind debugging and profiling software. BITFLIPS can inject SEUs into any program that can be run on the Linux operating system, without needing to modify the program s source code. Further, if access to the original program source code is available, BITFLIPS offers fine-grained control over exactly when and which areas of memory (as specified via program variables) will be subjected to SEUs. The rate of injection of SEUs is controlled by specifying either a fault probability or a fault rate based on memory size and radiation exposure time, in units of SEUs per byte per second. BITFLIPS can also log each SEU that it injects and, if program source code is available, report the magnitude of effect of the SEU on a floating-point value or other program variable.

Bornstein, Benjamin J.↗

Predeployment validation of fault-tolerant systems through software-implemented fault insertion

Fault injection-based automated testing (FIAT) environment, which can be used to experimentally characterize and evaluate distributed realtime systems under fault-free and faulted conditions is described. A survey is presented of validation methodologies. The need for fault insertion based on validation methodologies is demonstrated. The origins and models of faults, and motivation for the FIAT concept are reviewed. FIAT employs a validation methodology which builds confidence in the system through first providing a baseline of fault-free performance data and then characterizing the behavior of the system with faults present. Fault insertion is accomplished through software and allows faults or the manifestation of faults to be inserted by either seeding faults into memory or triggering error detection mechanisms. FIAT is capable of emulating a variety of fault-tolerant strategies and architectures, can monitor system activity, and can automatically orchestrate experiments involving insertion of faults. There is a common system interface which allows ease of use to decrease experiment development and run time. Fault models chosen for experiments on FIAT have generated system responses which parallel those observed in real systems under faulty conditions. These capabilities are shown by two example experiments each using a different fault-tolerance strategy.

Czeck, Edward W.↗

Impact of device level faults in a digital avionic processor

This paper describes an experimental analysis of the impact of gate and device-level faults in the processor of a flight control system. Via mixed mode simulation faults were injected both at the gate (stuck-at) and at the transistor levels, and their propagation through the chip to the output pins was measured. The results show that there is little correspondence between a stuck-at and a device-level fault model insofar as error activity or detection within a functional unit is concerned. Insofar as error activity outside the injected unit and at the output pins are concerned, the stuck-at and device models track each other, although the stuck-at model overestimates, by over one hundred percent, the probability of fault propagation to the output pins. The stuck-at model significantly underestimates the impact of an internal chip fault on the output pins.

Kim, S.↗

Impact of device level faults in a digital avionic processor

This study describes an experimental analysis of the impact of gate and device-level faults in the processor of a Bendix BDX-930 flight control system. Via mixed mode simulation, faults were injected at the gate (stuck-at) and at the transistor levels and, their propagation through the chip to the output pins was measured. The results show that there is little correspondence between a stuck-at and a device-level fault model, as far as error activity or detection within a functional unit is concerned. In so far as error activity outside the injected unit and at the output pins are concerned, the stuck-at and device models track each other. The stuck-at model, however, overestimates, by over 100 percent, the probability of fault propagation to the output pins. An evaluation of the Mean Error Durations and the Mean Time Between Errors at the output pins shows that the stuck-at model significantly underestimates (by 62 percent) the impact of an internal chip fault on the output pins. Finally, the study also quantifies the impact of device fault by location, both internally and at the output pins.

Suk, Ho Kim↗

Validation environment for AIPS/ALS: Implementation and results

The work is presented which was performed in porting the Fault Injection-based Automated Testing (FIAT) and Programming and Instrumentation Environments (PIE) validation tools, to the Advanced Information Processing System (AIPS) in the context of the Ada Language System (ALS) application, as well as an initial fault free validation of the available AIPS system. The PIE components implemented on AIPS provide the monitoring mechanisms required for validation. These mechanisms represent a substantial portion of the FIAT system. Moreover, these are required for the implementation of the FIAT environment on AIPS. Using these components, an initial fault free validation of the AIPS system was performed. The implementation is described of the FIAT/PIE system, configured for fault free validation of the AIPS fault tolerant computer system. The PIE components were modified to support the Ada language. A special purpose AIPS/Ada runtime monitoring and data collection was implemented. A number of initial Ada programs running on the PIE/AIPS system were implemented. The instrumentation of the Ada programs was accomplished automatically inside the PIE programming environment. PIE's on-line graphical views show vividly and accurately the performance characteristics of Ada programs, AIPS kernel and the application's interaction with the AIPS kernel. The data collection mechanisms were written in a high level language, Ada, and provide a high degree of flexibility for implementation under various system conditions.

Segall, Zary↗

Advanced Diagnostic System on Earth Observing One

In this infusion experiment, the Livingstone 2 (L2) model-based diagnosis engine, developed by the Computational Sciences division at NASA Ames Research Center, has been uploaded to the Earth Observing One (EO-1) satellite. L2 is integrated with the Autonomous Sciencecraft Experiment (ASE) which provides an on-board planning capability and a software bridge to the spacecraft's 1773 data bus. Using a model of the spacecraft subsystems, L2 predicts nominal state transitions initiated by control commands, monitors the spacecraft sensors, and, in the case of failure, isolates the fault based on the discrepant observations. Fault detection and isolation is done by determining a set of component modes, including most likely failures, which satisfy the current observations. All mode transitions and diagnoses are telemetered to the ground for analysis. The initial L2 model is scoped to EO-1's imaging instruments and solid state recorder. Diagnostic scenarios for EO-1's nominal imaging timeline are demonstrated by injecting simulated faults on-board the spacecraft. The solid state recorder stores the science images and also hosts: the experiment software. The main objective of the experiment is to mature the L2 technology to Technology Readiness Level (TRL) 7. Experiment results are presented, as well as a discussion of the challenging technical issues encountered. Future extensions may explore coordination with the planner, and model-based ground operations.

Hayden, Sandra C.↗

Experimental evaluation of a COTS system for space applications

The use of COTS-based systems in space missions for scientific data processing is very attractive, as their ratio of performance to power consumption of commercial components can be an order of magnitude greater than that of radiation hardened components, and the price differential is even higher.

fault injection on-board processing cluster commut↗

An experimental evaluation of the REE SIFT environment for spaceborne applications

This paper presents an experimental evaluation of a software-implemented fault tolerance environment built around a set of self-checking ARMOR proceses running on different machines that provide error detection and recovery services to themselves and to spaceborne scientific applications.

fault injection on-board processing cluster commut↗