Search NASA⌕ Search

SEARCH · Search NASA

Results for “transient faults”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

TFDA (Transient Fault Detection Algorithm) [SWR-24-23]

The Transient Fault Detection Algorithm (TFDA) is an algorithm to detect the non-fault and transient fault conditions, and classify the fault types. This is part of OEDI-SI project (https://data.openei.org/search?q=oedi%20si).

Hao, Jun↗

A preliminary transient-fault experiment on the SIFT computer system

This paper presents the results of a preliminary experiment to study the effectiveness of a fault-tolerant system's ability to handle transient faults. The primary goal of the experiment was to develop the techniques to measure the parameters needed for a reliability analysis of the SIFT computer system which includes th effects of transient faults. A key aspect of such an analysis is the determination of the effectiveness of the operating system's ability to discriminate between transient and permanent faults. A detailed description of the preliminary transient fault experiment along with the results from 297 transient fault injections are given. Although not enough data was obtained to draw statistically significant conclusions, the foundation has been laid for a large-scale transient fault experiment.

Butler, Ricky W.↗

Modeling transient faults in TMR computer systems

In this paper we report the development of a technique for modeling transient faults in redundant computer systems. Transient faults are characterized by their arrival rate and their duration. Fault detection, transient recovery, and the effect of permanent faults are included. A fault occurrence/recovery status state diagram is drawn to illustrate the operational status of the system while undergoing faults. The state diagram is used to formulate the equations for the mission failure probability. The techniques are then applied to a triple modular redundant computer system.

Merryman, P. M.↗

Transient Faults in Computer Systems

A powerful technique particularly appropriate for the detection of errors caused by transient faults in computer systems was developed. The technique can be implemented in either software or hardware; the research conducted thus far primarily considered software implementations. The error detection technique developed has the distinct advantage of having provably complete coverage of all errors caused by transient faults that affect the output produced by the execution of a program. In other words, the technique does not have to be tuned to a particular error model to enhance error coverage. Also, the correctness of the technique can be formally verified. The technique uses time and software redundancy. The foundation for an effective, low-overhead, software-based certification trail approach to real-time error detection resulting from transient fault phenomena was developed.

Masson, Gerald M.↗

Large transient fault current test of an electrical roll ring

The space station uses precision rotary gimbals to provide for sun tracking of its photoelectric arrays. Electrical power, command signals and data are transferred across the gimbals by roll rings. Roll rings have been shown to be capable of highly efficient electrical transmission and long life, through tests conducted at the NASA Lewis Research Center and Honeywell's Satellite and Space Systems Division in Phoenix, AZ. Large potential fault currents inherent to the power system's DC distribution architecture, have brought about the need to evaluate the effects of large transient fault currents on roll rings. A test recently conducted at Lewis subjected a roll ring to a simulated worst case space station electrical fault. The system model used to obtain the fault profile is described, along with details of the reduced order circuit that was used to simulate the fault. Test results comparing roll ring performance before and after the fault are also presented.

Yenni, Edward J.↗

Large transient fault current test of an electrical roll ring

The Space Station Freedom uses precision rotary gimbals to provide for sun tracking of its photoelectric arrays. Electrical power, command signals, and data are transferred across the gimbals by roll rings. Roll rings have been shown to be capable of highly efficient electrical transmission and long life, through tests conducted at the NASA Lewis Research Center and Honeywell's Satellite and Space Systems Division in Phoenix, AZ. Large potential fault currents inherent to the power system's DC distribution architecture have brought about the need to evaluate the effects of large transient fault currents on roll rings. A test recently conducted at Lewis subjected a roll ring to a simulated worst case space station electrical fault. The system model used to obtain the fault profile is described, along with details of the reduced order circuit that was used to simulate the fault. Test results comparing roll ring performance before and after the fault are also presented.

Yenni, Edward J.↗

The containment set approach to digital system tolerance of lightning-induced transient faults

A fault model and a system model are necessary to assess the tolerance or resilience of a digital system to lightning-induced transients. It is noted that these models are usually developed separately. A new approach is outlined here for this assessment problem which combines the fault and system models into an overall model. With this approach, referred to as the containment set approach, an assessment of the effects of lightning-induced transients can be made in terms of a state transition matrix. This matrix can be generated by means of fault injection experiments. In addition, certain nonredundancy-oriented design alternatives to the achievement of lightning-induced transient tolerance in digital systems are indicated by the containment set approach.

Masson, G. M.↗

Transient fault behavior in a microprocessor: A case study

An experimental analysis is described which studies the susceptibility of a microprocessor based jet engine controller to upsets caused by current and voltage transients. A design automation environment which allows the run time injection of transients and the tracing from their impact device to the pin level is described. The resulting error data are categorized by the charge levels of the injected transients by location and by their potential to cause logic upsets, latched errors, and pin errors. The results show a 3 picoCouloumb threshold, below which the transients have little impact. An Arithmetic and Logic Unit transient is most likely to result in logic upsets and pin errors (i.e., impact the external environment). The transients in the countdown unit are potentially serious since they can result in latched errors, thus causing latent faults. Suggestions to protect the processor against these errors, by incorporating internal error detection and transient suppression techniques, are also made.

Duba, Patrick↗

Rapid recovery from transient faults in the fault-tolerant processor with fault-tolerant shared memory

The Draper fault-tolerant processor with fault-tolerant shared memory (FTP/FTSM), which is designed to allow application tasks to continue execution during the memory alignment process, is described. Processor performance is not affected by memory alignment. In addition, the FTP/FTSM incorporates a hardware scrubber device to perform the memory alignment quickly during unused memory access cycles. The FTP/FTSM architecture is described, followed by an estimate of the time required for channel reintegration.

Harper, Richard E.↗

How earthquakes organize stress

Stress is not uniform in the Earth. Therefore, we must use natural experiments to measure the distribution of stresses and related quantities, rather than single values. For instance, dynamic triggering shows that faults are uniformly distributed over their loading cycles in Southern California. The probability that a fault ruptures across a barrier measures the in situ energy distribution. Fault roughness reflects the distribution of strength. These natural experiments produce observable distributions that are surprisingly consistent and suggest some degree of self-organization in the Earth’s crust. Once established, the functional form of the distributions can be used to track changes in response to earthquakes as well as to distinguish fundamentally different fault systems. Transient fault locking before stress release in laboratory experiments can be interpreted as a consequence of self-organization of fault stress. The robust self-organization of multiple variables in earthquake systems suggests that the most consequential mechanical outcome of earthquakes may be the redistribution of stress and the strain energy associated with it. The low friction on a fault during seismic slip as inferred by temperature measurements of the Tohoku earthquake is consistent with dissipation playing a secondary role to this redistribution process. Through stress redistribution and interaction, subduction zone faults tend to synchronize, perhaps due to their geometric simplicity, while the continental system of Southern California cannot synchronize, perhaps due to the complexity of the fault network. Earthquakes organize stress in the crust and produce a suite of well-defined, consistent distributions.

earthquakes↗

Review and verification of CARE 3 mathematical model and code

The CARE-III mathematical model and code verification performed by Boeing Computer Services were documented. The mathematical model was verified for permanent and intermittent faults. The transient fault model was not addressed. The code verification was performed on CARE-III, Version 3. A CARE III Version 4, which corrects deficiencies identified in Version 3, is being developed.

Rose, D. M.↗

Open Source Fault-tolerant Grid Frequency Measurement for Solar Inverters

The Discrete Fourier transform (DFT) based measurement algorithms are one of the most common measurement algorithms for grid parameter estimation such as rms, phase angle, frequency. Over the past few years, many DFT based algorithms have been developed to enhance its measurement accuracy under steady-state and/or dynamic grid conditions. For example, an adaptive band-pass filter utilizing exponential modulation filter has been proposed to reduce measurement errors at the presence of large frequency deviations. Measurement accuracy of different algorithms including FIR filter, extended Kalman filtering (EKF), and enhanced DFT method have been compared in detail under different grid conditions. Two artificial signals that have 90-degree phase difference were constructed by the Clarke transformation to address the frequency spectrum leakage of DFT. A multi-module approach was developed to enhance both steady-state and dynamic measurement accuracies, in which each module was developed to eliminate some specific errors. Besides DFT-based measurement algorithms, some signal model-based algorithms have been developed to further improve the accuracy under dynamic conditions. However, a key drawback of the state-of-the-art algorithms is that they cannot perform measurements accurately during system transient faults. In the Blue Cut Fire event, there was a phase angle jump of about 26 degrees in the voltage waveform during the transient fault. The phase angle jump fault will cause waveform discontinuity, and these algorithms will fail to provide reliable measurements during this period because they typically assume the waveform to be measured is continuous, no matter what method (DFT, PLL, EKF, FIR, or Taylor WLS) is used for estimation. In fact, the measurement errors during the system transient faults like phase-jump is not required in the IEEE Standard. As a result, although a measurement instrument can pass the strict IEEE Standard, it could still be the source of the problem in the future if we have similar system transient faults, which could happen again. Therefore, developing the fault-tolerant measurement technology is the key to solve the problem.

14 SOLAR ENERGY↗

Transient upsets in microprocessor controllers

The modeling and analysis of transient faults in microprocessor based controllers are discussed. Such controllers typically consist of a microprocessor, read only memory storing and application program, random access memory for data storage, and input/output devices for external communications. The effects of transient faults on the performance of the controller are reviewed. An instruction level perspective of performance is taken which is the basis of a useful high level program state description of the microprocessor controller. A transition matrix is defined which determines the controller's response to transient fault arrivals.

Glaser, R. E.↗

An experimental study of fault propagation in a jet-engine controller

An experimental analysis of the impact of transient faults on a microprocessor-based jet engine controller, used in the Boeing 747 and 757 aircrafts is described. A hierarchical simulation environment which allows the injection of transients during run-time and the tracing of their impact is described. Verification of the accuracy of this approach is also provided. A determination of the probability that a transient results in latch, pin or functional errors is made. Given a transient fault, there is approximately an 80 percent chance that there is no impact on the chip. An empirical model to depict the process of error exploration and degeneration in the target system is derived. The model shows that, if no latch errors occur within eight clock cycles, no significant damage is likely to happen. Thus, the overall impact of a transient is well contained. A state transition model is also derived from the measured data, to describe the error propagation characteristics within the chip, and to quantify the impact of transients on the external environment. The model is used to identify and isolate the critical fault propagation paths, the module most sensitive to fault propagation and the module with the highest potential of causing external pin errors.

Choi, Gwan Seung↗

Analysis of a hardware and software fault tolerant processor for critical applications

Computer systems for critical applications must be designed to tolerate software faults as well as hardware faults. A unified approach to tolerating hardware and software faults is characterized by classifying faults in terms of duration (transient or permanent) rather than source (hardware or software). Errors arising from transient faults can be handled through masking or voting, but errors arising from permanent faults require system reconfiguration to bypass the failed component. Most errors which are caused by software faults can be considered transient, in that they are input-dependent. Software faults are triggered by a particular set of inputs. Quantitative dependability analysis of systems which exhibit a unified approach to fault tolerance can be performed by a hierarchical combination of fault tree and Markov models. A methodology for analyzing hardware and software fault tolerant systems is applied to the analysis of a hypothetical system, loosely based on the Fault Tolerant Parallel Processor. The models consider both transient and permanent faults, hardware and software faults, independent and related software faults, automatic recovery, and reconfiguration.

Dugan, Joanne B.↗

Network Connectivity for Permanent, Transient, Independent, and Correlated Faults

This paper develops a method for the quantitative analysis of network connectivity in the presence of both permanent and transient faults. Even though transient noise is considered a common occurrence in networks, a survey of the literature reveals an emphasis on permanent faults. Transient faults introduce a time element into the analysis of network reliability. With permanent faults it is sufficient to consider the faults that have accumulated by the end of the operating period. With transient faults the arrival and recovery time must be included. The number and location of faults in the system is a dynamic variable. Transient faults also introduce system recovery into the analysis. The goal is the quantitative assessment of network connectivity in the presence of both permanent and transient faults. The approach is to construct a global model that includes all classes of faults: permanent, transient, independent, and correlated. A theorem is derived about this model that give distributions for (1) the number of fault occurrences, (2) the type of fault occurrence, (3) the time of the fault occurrences, and (4) the location of the fault occurrence. These results are applied to compare and contrast the connectivity of different network architectures in the presence of permanent, transient, independent, and correlated faults. The examples below use a Monte Carlo simulation, but the theorem mentioned above could be used to guide fault-injections in a laboratory.

White, Allan L.↗

A verified design of a fault-tolerant clock synchronization circuit: Preliminary investigations

Schneider demonstrates that many fault tolerant clock synchronization algorithms can be represented as refinements of a single proven correct paradigm. Shankar provides mechanical proof that Schneider's schema achieves Byzantine fault tolerant clock synchronization provided that 11 constraints are satisfied. Some of the constraints are assumptions about physical properties of the system and cannot be established formally. Proofs are given that the fault tolerant midpoint convergence function satisfies three of the constraints. A hardware design is presented, implementing the fault tolerant midpoint function, which is shown to satisfy the remaining constraints. The synchronization circuit will recover completely from transient faults provided the maximum fault assumption is not violated. The initialization protocol for the circuit also provides a recovery mechanism from total system failure caused by correlated transient faults.

Miner, Paul S.↗

Flight experience with a fail-operational digital fly-by-wire control system

The NASA Dryden Flight Research Center is flight testing a triply redundant digital fly-by-wire (DFBW) control system installed in an F-8 aircraft. The full-time, full-authority system performs three-axis flight control computations, including stability and command augmentation, autopilot functions, failure detection and isolation, and self-test functions. Advanced control law experiments include an active flap mode for ride smoothing and maneuver drag reduction. This paper discusses research being conducted on computer synchronization, fault detection, fault isolation, and recovery from transient faults. The F-8 DFBW system has demonstrated immunity from nuisance fault declarations while quickly identifying truly faulty components.

Brown, S. R.↗