Search NASASearch

SEARCH · Search NASA

Results for “Fault Injection”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 37 records · Page 2

Measuring fault tolerance with the FTAPE fault injection tool

This paper describes FTAPE (Fault Tolerance And Performance Evaluator), a tool that can be used to compare fault-tolerant computers. The major parts of the tool include a system-wide fault-injector, a workload generator, and a workload activity measurement tool. The workload creates high stress conditions on the machine. Using stress-based injection, the fault injector is able to utilize knowledge of the workload activity to ensure a high level of fault propagation. The errors/fault ratio, performance degradation, and number of system crashes are presented as measures of fault tolerance.

Tsai, Timothy K.

A study of fault injection in multichannel spacecraft power systems

NASA/Marshall Space Flight Center proposes to implement fault injection into an electrical power system breadboard to study the reactions of the various control elements of this breadboard. Among the elements to be studied are the remote power controllers, the algorithms in the control computers, and the artificially intelligent control programs resident in this breadboard. To this end, a study of electrical power is being performed to yield a list of the most common power system faults. The results of this study are being applied to a multichannel high-voltage DC spacecraft power system called the Large Autonomous Spacecraft Electrical Power System Breadboard. Some of the reactions of the breadboard to some of the faults which have been encountered are presented along with the results of this study.

Dugal-Whitehead, Norma R.

Fault Injection Techniques and Tools

Dependability evaluation involves the study of failures and errors. The destructive nature of a crash and long error latency make it difficult to identify the causes of failures in the operational environment. It is particularly hard to recreate a failure scenario for a large, complex system. To identify and understand potential failures, we use an experiment-based approach for studying the dependability of a system. Such an approach is applied not only during the conception and design phases, but also during the prototype and operational phases. To take an experiment-based approach, we must first understand a system's architecture, structure, and behavior. Specifically, we need to know its tolerance for faults and failures, including its built-in detection and recovery mechanisms, and we need specific instruments and tools to inject faults, create failures or errors, and monitor their effects.

Hsueh, Mei-Chen

Experimental analysis of computer system dependability

This paper reviews an area which has evolved over the past 15 years: experimental analysis of computer system dependability. Methodologies and advances are discussed for three basic approaches used in the area: simulated fault injection, physical fault injection, and measurement-based analysis. The three approaches are suited, respectively, to dependability evaluation in the three phases of a system's life: design phase, prototype phase, and operational phase. Before the discussion of these phases, several statistical techniques used in the area are introduced. For each phase, a classification of research methods or study topics is outlined, followed by discussion of these methods or topics as well as representative studies. The statistical techniques introduced include the estimation of parameters and confidence intervals, probability distribution characterization, and several multivariate analysis methods. Importance sampling, a statistical technique used to accelerate Monte Carlo simulation, is also introduced. The discussion of simulated fault injection covers electrical-level, logic-level, and function-level fault injection methods as well as representative simulation environments such as FOCUS and DEPEND. The discussion of physical fault injection covers hardware, software, and radiation fault injection methods as well as several software and hybrid tools including FIAT, FERARI, HYBRID, and FINE. The discussion of measurement-based analysis covers measurement and data processing techniques, basic error characterization, dependency analysis, Markov reward modeling, software-dependability, and fault diagnosis. The discussion involves several important issues studies in the area, including fault models, fast simulation techniques, workload/failure dependency, correlated failures, and software fault tolerance.

Iyer, Ravishankar, K.

Understanding Mixed Precision GEMM with MPGemmFI: Insights into Fault Resilience

Emerging deep learning workloads urgently need fast general matrix multiplication (GEMM). Thus, one of the critical features of machine-learning-specific accelerators such as NVIDIA Tensor Cores, AMD Matrix Cores, and Google TPUs is the support of mixed-precision enabled GEMM. For DNN models, lower-precision FP data formats and computation offer acceptable correctness but significant performance, area, and memory footprint improvement. While promising, the mixed-precision computation on error resilience remains unexplored. To this end, we develop a fault injection framework that systematically injects fault into the mixed-precision computation results. We investigate how the faults affect the accuracy of machine learning applications. Based on the characteristics of error resilience, we offer lightweight error detection and correction solutions that significantly improve the overall model accuracy by 75% if the models experience hardware faults. The solutions can be efficiently integrated into the accelerator's pipelines.

Fang, Bo

Control Mechanisms for Self‐Sealing in Activated Clay‐Rich Faults Through Controlled Hydraulic Injection Experiment

Abstract In a high‐pressure injection fault activation experiment conducted at the Mont Terri underground research laboratory in Switzerland, the transmissivity of the Opalinus Clay fault significantly increased due to opening and shearing. The fluid injection, spanning a few hours, generated a 10 m radius fault activation patch. Subsequent pressure pulse tests conducted bi‐weekly for a year revealed the gradual return of fault transmissivity to its initial state. The study utilized fluid pressure decay analysis, optical fiber monitoring, continuous active source seismic measurements and borehole displacement sensors for measuring fault displacements. The fault zone exhibited a dilation of approximately 1.4 mm, associated with both normal and tangential movements during activation, resulting in a sudden transmissivity increase from 1 × 10 −12 to 3.2 × 10 −7 m 2 /s. Early post‐activation, transient compaction and the subsequent slow compaction were observed, transitioning to an extension regime. The pressure pulse tests demonstrated a rapid transmissivity drop by more than two orders of magnitude within the first 10 days, followed by a gradual and less pronounced decrease. Plastic shear and compaction dominated the transmissivity evolution until 70 days after injection ended, followed by a period where additional factors, such as clay mineral swelling, influenced the behavior. Extrapolation suggested a sealing process taking at least 50 years after the initial activation. Plain Language Summary A field‐scale fault activation experiment offers valuable insights into the elasto‐plastic processes governing the sealing of shale faults. The experiment reveals a rapid increase in the fault's transmissivity by approximately five orders of magnitude during activation. Subsequent observations show a gradual transmissivity decrease by about three orders of magnitude post‐activation, with slow long‐term plastic shear and compaction of the fault competing against secondary processes, notably clay mineral swelling. All conceptual models employed to interpret these field data converge on the estimation that the fault's return to its initial low transmissivity state would require a minimum of 50 years. Key Points High‐pressure injection fault activation experiment at the Mont Terri underground research laboratory Continuous transmissivity measurements record self‐sealing inside a clay‐rich fault zone Transmissivity undergoes a phase of domination by slow plastic compaction and shearing during the initial post‐activation period, with mineral swelling exerting its influence over the long term

Guglielmi, Yves

Fault recovery characteristics of the fault tolerant multi-processor

The fault handling performance of the fault tolerant multiprocessor (FTMP) was investigated. Fault handling errors detected during fault injection experiments were characterized. In these fault injection experiments, the FTMP disabled a working unit instead of the faulted unit once every 500 faults, on the average. System design weaknesses allow active faults to exercise a part of the fault management software that handles byzantine or lying faults. It is pointed out that these weak areas in the FTMP's design increase the probability that, for any hardware fault, a good LRU (line replaceable unit) is mistakenly disabled by the fault management software. It is concluded that fault injection can help detect and analyze the behavior of a system in the ultra-reliable regime. Although fault injection testing cannot be exhaustive, it has been demonstrated that it provides a unique capability to unmask problems and to characterize the behavior of a fault-tolerant system.

Padilla, Peter A.

Board Level Proton Testing Book of Knowledge for NASA Electronic Parts and Packaging Program

This book of knowledge (BoK) provides a critical review of the benefits and difficulties associated with using proton irradiation as a means of exploring the radiation hardness of commercial-off-the-shelf (COTS) systems. This work was developed for the NASA Electronic Parts and Packaging (NEPP) Board Level Testing for the COTS task. The fundamental findings of this BoK are the following. The board-level test method can reduce the worst case estimate for a board's single-event effect (SEE) sensitivity compared to the case of no test data, but only by a factor of ten. The estimated worst case rate of failure for untested boards is about 0.1 SEE/board-day. By employing the use of protons with energies near or above 200 MeV, this rate can be safely reduced to 0.01 SEE/board-day, with only those SEEs with deep charge collection mechanisms rising this high. For general SEEs, such as static random-access memory (SRAM) upsets, single-event transients (SETs), single-event gate ruptures (SEGRs), and similar cases where the relevant charge collection depth is less than 10 μm, the worst case rate for SEE is below 0.001 SEE/board-day. Note that these bounds assume that no SEEs are observed during testing. When SEEs are observed during testing, the board-level test method can establish a reliable event rate in some orbits, though all established rates will be at or above 0.001 SEE/board-day. The board-level test approach we explore has picked up support as a radiation hardness assurance technique over the last twenty years. The approach originally was used to provide a very limited verification of the suitability of low cost assemblies to be used in the very benign environment of the International Space Station (ISS), in limited reliability applications. Recently the method has been gaining popularity as a way to establish a minimum level of SEE performance of systems that require somewhat higher reliability performance than previous applications. This sort of application of the method suggests a critical analysis of the method is in order. This is also of current consideration because the primary facility used for this type of work, the Indiana University Cyclotron Facility (IUCF) (also known as the Integrated Science and Technology (ISAT) hall), has closed permanently, and the future selection of alternate test facilities is critically important. This document reviews the main theoretical work on proton testing of assemblies over the last twenty years. It augments this with review of reported data generated from the method and other data that applies to the limitations of the proton board-level test approach. When protons are incident on a system for test they can produce spallation reactions. From these reactions, secondary particles with linear energy transfers (LETs) significantly higher than the incident protons can be produced. These secondary particles, together with the protons, can simulate a subset of the space environment for particles capable of inducing single event effects (SEEs). The proton board-level test approach has been used to bound SEE rates, establishing a maximum possible SEE rate that a test article may exhibit in space. This bound is not particularly useful in many cases because the bound is quite loose. We discuss the established limit that the proton board-level test approach leaves us with. The remaining possible SEE rates may be as high as one per ten years for most devices. The situation is actually more problematic for many SEE types with deep charge collection. In cases with these SEEs, the limits set by the proton board-level test can be on the order of one per 100 days. Because of the limited nature of the bounds established by proton testing alone, it is possible that tested devices will have actual SEE sensitivity that is very low (e.g., fewer than one event in 1 × 10(exp 4) years), but the test method will only be able to establish the limits indicated above. This BoK further examines other benefits of proton board-level testing besides hardness assurance. The primary alternate use is the injection of errors. Error injection, or fault injection, is something that is often done in a simulation environment. But the proton beam has the benefit of injecting the majority of actual SEEs without risk of something being missed, and without the risk of simulation artifacts misleading the SEE investigation.

Guertin, Steven M.

Estimating the distribution of fault latency in a digital processor

Presented is a statistical approach to measuring fault latency in a digital processor. The method relies on the use of physical fault injection where the duration of the fault injection can be controlled. Although a specific fault's latency period is never directly measured, the method indirectly determines the distribution of fault latency.

Ellis, Erik L.

In-circuit fault injector user's guide

A fault injector system, called an in-circuit injector, was designed and developed to facilitate fault injection experiments performed at NASA-Langley's Avionics Integration Research Lab (AIRLAB). The in-circuit fault injector (ICFI) allows fault injections to be performed on electronic systems without special test features, e.g., sockets. The system supports stuck-at-zero, stuck-at-one, and transient fault models. The ICFI system is interfaced to a VAX-11/750 minicomputer. An interface program has been developed in the VAX. The computer code required to access the interface program is presented. Also presented is the connection procedure to be followed to connect the ICFI system to a circuit under test and the ICFI front panel controls which allow manual control of fault injections.

Padilla, Peter A.

Experimental evaluation of the certification-trail method

Certification trails are a recently introduced and promising approach to fault-detection and fault-tolerance. A comprehensive attempt to assess experimentally the performance and overall value of the method is reported. The method is applied to algorithms for the following problems: huffman tree, shortest path, minimum spanning tree, sorting, and convex hull. Our results reveal many cases in which an approach using certification-trails allows for significantly faster overall program execution time than a basic time redundancy-approach. Algorithms for the answer-validation problem for abstract data types were also examined. This kind of problem provides a basis for applying the certification-trail method to wide classes of algorithms. Answer-validation solutions for two types of priority queues were implemented and analyzed. In both cases, the algorithm which performs answer-validation is substantially faster than the original algorithm for computing the answer. Next, a probabilistic model and analysis which enables comparison between the certification-trail method and the time-redundancy approach were presented. The analysis reveals some substantial and sometimes surprising advantages for ther certification-trail method. Finally, the work our group performed on the design and implementation of fault injection testbeds for experimental analysis of the certification trail technique is discussed. This work employs two distinct methodologies, software fault injection (modification of instruction, data, and stack segments of programs on a Sun Sparcstation ELC and on an IBM 386 PC) and hardware fault injection (control, address, and data lines of a Motorola MC68000-based target system pulsed at logical zero/one values). Our results indicate the viability of the certification trail technique. It is also believed that the tools developed provide a solid base for additional exploration.

Sullivan, Gregory F.

Modeling injection-induced fault slip using long short-term memory networks

Stress changes due to changes in fluid pressure and temperature in a faulted formation may lead to the opening/shearing of the fault. This can be due to subsurface (geo)engineering activities such as fluid injections and geologic disposal of nuclear waste. Such activities are expected to rise in the future making it necessary to assess their short- and long-term safety. Here, a new machine learning (ML) approach to model pore pressure and fault displacements in response to high-pressure fluid injection cycles is developed. The focus is on fault behavior near the injection borehole. To capture the temporal dependencies in the data, long short-term memory (LSTM) networks are utilized. To prevent error accumulation within the forecast window, four critical measures to train a robust LSTM model for predicting fault response are highlighted: (i) setting an appropriate value of LSTM lag, (ii) calibrating the LSTM cell dimension, (iii) learning rate reduction during weight optimization, and (iv) not adopting an independent injection cycle as a validation set. Several numerical experiments were conducted, which demonstrated that the ML model can capture peaks in pressure and associated fault displacement that accompany an increase in fluid injection. The model also captured the decay in pressure and displacement during the injection shut-in period. Further, the ability of an ML model to highlight key changes in fault hydromechanical activation processes was investigated, which shows that ML can be used to monitor risk of fault activation and leakage during high pressure fluid injections.

58 GEOSCIENCES

Fail-Safe Logic Design Strategies Within Modern FPGA Architectures

Fail-safe computing refers to computing systems that revert to a non-operational safe state when a fault occurs. In this paper, we investigate a circuit level technique as mitigation for single event upsets (SEUs) and fault injection attacks on field programmable gate arrays (FPGAs), and analyze the effectiveness of the technique as a fail-safe monitor for an encryption algorithm. The propagation of fault effects through FPGA primitives including lookup tables (LUTs) and programmable interconnect points (PIPs) is assessed within an FPGA architecture created using an open source tool, and validated using fault injection experiments on an FPGA. The analysis reveals additional vulnerabilities exist within reconfigurable architectures over those in equivalent fail-safe application specific integrated circuit (ASIC), thus requiring a more elaborate network of redundant circuits and checking logic. The configuration memory bits (CMBs), which configure routing and designate logic functions within the LUTs of the FPGA, add complexity to fail-safe design strategies by introducing additional fault conditions and fault propagation paths. A resource-efficient fail-safe circuit design technique called DEsign for Fail-safe in reCONfigurable systems (DEFCON) is proposed. The benefits and limitations associated with DEFCON are described in the context of fault injection experiments carried out as simulations and in FPGA hardware.

Bhakta, Priya A. [Univ. of New Mexico, Albuquerque

Establishing Fault Tolerance for a Class of Systems by Experiment

A long-standing problem in system verification is establishing fault tolerance at the ultra-high level by experiment. It is considered impossible because of system complexity and the enormous number of trials needed. This paper considers the problem for a class of digital systems that use redundancy to achieve reliability. The class is the systems that operate for a period of time without maintenance followed by a maintenance check that replaces components identified as faulty. The paper considers simulating a natural life test where a natural life test observes a number of operating periods. If the system does not fail during the test, it can be said to have a certain reliability at a certain confidence level. The approach in this paper is to make the simulated life test more efficient while maintaining realism by integrating structural arguments, information on fault occurrence, and fault injection in the lab. The major result of this paper is constructing a global fault model using the failure rate of the components and proving theorems about the model that tell how many, what kind, when, and where to inject faults. A simple example illustrates applying the theorems.

design of experiments

Hierarchical Simulation to Assess Hardware and Software Dependability

This thesis presents a method for conducting hierarchical simulations to assess system hardware and software dependability. The method is intended to model embedded microprocessor systems. A key contribution of the thesis is the idea of using fault dictionaries to propagate fault effects upward from the level of abstraction where a fault model is assumed to the system level where the ultimate impact of the fault is observed. A second important contribution is the analysis of the software behavior under faults as well as the hardware behavior. The simulation method is demonstrated and validated in four case studies analyzing Myrinet, a commercial, high-speed networking system. One key result from the case studies shows that the simulation method predicts the same fault impact 87.5% of the time as is obtained by similar fault injections into a real Myrinet system. Reasons for the remaining discrepancy are examined in the thesis. A second key result shows the reduction in the number of simulations needed due to the fault dictionary method. In one case study, 500 faults were injected at the chip level, but only 255 propagated to the system level. Of these 255 faults, 110 shared identical fault dictionary entries at the system level and so did not need to be resimulated. The necessary number of system-level simulations was therefore reduced from 500 to 145. Finally, the case studies show how the simulation method can be used to improve the dependability of the target system. The simulation analysis was used to add recovery to the target software for the most common fault propagation mechanisms that would cause the software to hang. After the modification, the number of hangs was reduced by 60% for fault injections into the real system.

Ries, Gregory Lawrence

Model-Based Verification and Validation of Spacecraft Avionics

Verification and Validation (V&V) at JPL is traditionally performed on flight or flight-like hardware running flight software. For some time, the complexity of avionics has increased exponentially while the time allocated for system integration and associated V&V testing has remained fixed. There is an increasing need to perform comprehensive system level V&V using modeling and simulation, and to use scarce hardware testing time to validate models; the norm for thermal and structural V&V for some time. Our approach extends model-based V&V to electronics and software through functional and structural models implemented in SysML. We develop component models of electronics and software that are validated by comparison with test results from actual equipment. The models are then simulated enabling a more complete set of test cases than possible on flight hardware. SysML simulations provide access and control of internal nodes that may not be available in physical systems. This is particularly helpful in testing fault protection behaviors when injecting faults is either not possible or potentially damaging to the hardware. We can also model both hardware and software behaviors in SysML, which allows us to simulate hardware and software interactions. With an integrated model and simulation capability we can evaluate the hardware and software interactions and identify problems sooner. The primary missing piece is validating SysML model correctness against hardware; this experiment demonstrated such an approach is possible.

MBV&V

Fault detection, isolation and reconfiguration in FTMP Methods and experimental results

The Fault-Tolerant Multiprocessor (FTMP) is a highly reliable computer designed to meet a goal of 10 to the -10th failures per hour and built with the objective of flying an active-control transport aircraft. Fault detection, identification, and recovery software is described, and experimental results obtained by injecting faults in the pin level in the FTMP are presented. Over 21,000 faults were injected in the CPU, memory, bus interface circuits, and error detection, masking, and error reporting circuits of one LRU of the multiprocessor. Detection, isolation, and reconfiguration times were recorded for each fault, and the results were found to agree well with earlier assumptions made in reliability modeling.

Lala, J. H.

Formal Validation of Fault Management Design Solutions

The work presented in this paper describes an approach used to develop SysML modeling patterns to express the behavior of fault protection, test the model's logic by performing fault injection simulations, and verify the fault protection system's logical design via model checking. A representative example, using a subset of the fault protection design for the Soil Moisture Active-Passive (SMAP) system, was modeled with SysML State Machines and JavaScript as Action Language. The SysML model captures interactions between relevant system components and system behavior abstractions (mode managers, error monitors, fault protection engine, and devices/switches). Development of a method to implement verifiable and lightweight executable fault protection models enables future missions to have access to larger fault test domains and verifiable design patterns. A tool-chain to transform the SysML model to jpf-Statechart compliant Java code and then verify the generated code via model checking was established. Conclusions and lessons learned from this work are also described, as well as potential avenues for further research and development.

Statechart