Search NASA⌕ Search

SEARCH · Search NASA

Results for “physics of faulting”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

Experimental analysis of computer system dependability

This paper reviews an area which has evolved over the past 15 years: experimental analysis of computer system dependability. Methodologies and advances are discussed for three basic approaches used in the area: simulated fault injection, physical fault injection, and measurement-based analysis. The three approaches are suited, respectively, to dependability evaluation in the three phases of a system's life: design phase, prototype phase, and operational phase. Before the discussion of these phases, several statistical techniques used in the area are introduced. For each phase, a classification of research methods or study topics is outlined, followed by discussion of these methods or topics as well as representative studies. The statistical techniques introduced include the estimation of parameters and confidence intervals, probability distribution characterization, and several multivariate analysis methods. Importance sampling, a statistical technique used to accelerate Monte Carlo simulation, is also introduced. The discussion of simulated fault injection covers electrical-level, logic-level, and function-level fault injection methods as well as representative simulation environments such as FOCUS and DEPEND. The discussion of physical fault injection covers hardware, software, and radiation fault injection methods as well as several software and hybrid tools including FIAT, FERARI, HYBRID, and FINE. The discussion of measurement-based analysis covers measurement and data processing techniques, basic error characterization, dependency analysis, Markov reward modeling, software-dependability, and fault diagnosis. The discussion involves several important issues studies in the area, including fault models, fast simulation techniques, workload/failure dependency, correlated failures, and software fault tolerance.

Iyer, Ravishankar, K.↗

Using Markov Models of Fault Growth Physics and Environmental Stresses to Optimize Control Actions

A generalized Markov chain representation of fault dynamics is presented for the case that available modeling of fault growth physics and future environmental stresses can be represented by two independent stochastic process models. A contrived but representatively challenging example will be presented and analyzed, in which uncertainty in the modeling of fault growth physics is represented by a uniformly distributed dice throwing process, and a discrete random walk is used to represent uncertain modeling of future exogenous loading demands to be placed on the system. A finite horizon dynamic programming algorithm is used to solve for an optimal control policy over a finite time window for the case that stochastic models representing physics of failure and future environmental stresses are known, and the states of both stochastic processes are observable by implemented control routines. The fundamental limitations of optimization performed in the presence of uncertain modeling information are examined by comparing the outcomes obtained from simulations of an optimizing control policy with the outcomes that would be achievable if all modeling uncertainties were removed from the system.

Bole, Brian↗

DRS: Derivational Reasoning System

The high reliability requirements for airborne systems requires fault-tolerant architectures to address failures in the presence of physical faults, and the elimination of design flaws during the specification and validation phase of the design cycle. Although much progress has been made in developing methods to address physical faults, design flaws remain a serious problem. Formal methods provides a mathematical basis for removing design flaws from digital systems. DRS (Derivational Reasoning System) is a formal design tool based on advanced research in mathematical modeling and formal synthesis. The system implements a basic design algebra for synthesizing digital circuit descriptions from high level functional specifications. DRS incorporates an executable specification language, a set of correctness preserving transformations, verification interface, and a logic synthesis interface, making it a powerful tool for realizing hardware from abstract specifications. DRS integrates recent advances in transformational reasoning, automated theorem proving and high-level CAD synthesis systems in order to provide enhanced reliability in designs with reduced time and cost.

Bose, Bhaskar↗

Simulation-based reasoning about the physical propagation of fault effects

The research described deals with the effects of faults on complex physical systems, with particular emphasis on aircraft and spacecraft systems. Given that a malfunction has occurred and been diagnosed, the goal is to determine how that fault will propagate to other subsystems, and what the effects will be on vehicle functionality. In particular, the use of qualitative spatial simulation to determine the physical propagation of fault effects in 3-D space is described.

Feyock, Stefan↗

Estimating the distribution of fault latency in a digital processor

Presented is a statistical approach to measuring fault latency in a digital processor. The method relies on the use of physical fault injection where the duration of the fault injection can be controlled. Although a specific fault's latency period is never directly measured, the method indirectly determines the distribution of fault latency.

Ellis, Erik L.↗

A New On-Line Diagnosis Protocol for the SPIDER Family of Byzantine Fault Tolerant Architectures

This paper presents the formal verification of a new protocol for online distributed diagnosis for the SPIDER family of architectures. An instance of the Scalable Processor-Independent Design for Electromagnetic Resilience (SPIDER) architecture consists of a collection of processing elements communicating over a Reliable Optical Bus (ROBUS). The ROBUS is a specialized fault-tolerant device that guarantees Interactive Consistency, Distributed Diagnosis (Group Membership), and Synchronization in the presence of a bounded number of physical faults. Formal verification of the original SPIDER diagnosis protocol provided a detailed understanding that led to the discovery of a significantly more efficient protocol. The original protocol was adapted from the formally verified protocol used in the MAFT architecture. It required O(N) message exchanges per defendant to correctly diagnose failures in a system with N nodes. The new protocol achieves the same diagnostic fidelity, but only requires O(1) exchanges per defendant. This paper presents this new diagnosis protocol and a formal proof of its correctness using PVS.

Geser, Alfons↗

Implementation of an experimental fault-tolerant memory system

The experimental fault-tolerant memory system described in this paper has been designed to enable the modular addition of spares, to validate the theoretical fault-secure and self-testing properties of the translator/corrector, to provide a basis for experiments using the new testing and correction processes for recovery, and to determine the practicality of such systems. The hardware design and implementation are described, together with methods of fault insertion. The hardware/software interface, including a restricted single error correction/double error detection (SEC/DED) code, is specified. Procedures are carefully described which, (1) test for specified physical faults, (2) ensure that single error corrections are not miscorrections due to triple faults, and (3) enable recovery from double errors.

Carter, W. C.↗

General linear codes for fault-tolerant matrix operations on processor arrays

Various checksum codes have been suggested for fault-tolerant matrix computations on processor arrays. Use of these codes is limited due to potential roundoff and overflow errors. Numerical errors may also be misconstrued as errors due to physical faults in the system. In this a set of linear codes is identified which can be used for fault-tolerant matrix operations such as matrix addition, multiplication, transposition, and LU-decomposition, with minimum numerical error. Encoding schemes are given for some of the example codes which fall under the general set of codes. With the help of experiments, a rule of thumb for the selection of a particular code for a given application is derived.

Nair, V. S. S.↗

Columbia Accident Investigation Board Report

The Columbia Accident Investigation Board's independent investigation into the tragic February 1, 2003, loss of the Space Shuttle Columbia and its seven-member crew lasted nearly seven months and involved 13 Board members, approximately 120 Board investigators, and thousands of NASA and support personnel. Because the events that initiated the accident were not apparent for some time, the investigation's depth and breadth were unprecedented in NASA history. Further, the Board determined early in the investigation that it intended to put this accident into context. We considered it unlikely that the accident was a random event; rather, it was likely related in some degree to NASA's budgets, history, and program culture, as well as to the politics, compromises, and changing priorities of the democratic process. We are convinced that the management practices overseeing the Space Shuttle Program were as much a cause of the accident as the foam that struck the left wing. The Board was also influenced by discussions with members of Congress, who suggested that this nation needed a broad examination of NASA's Human Space Flight Program, rather than just an investigation into what physical fault caused Columbia to break up during re-entry. Findings and recommendations are in the relevant chapters and all recommendations are compiled in Chapter 11. Volume I is organized into four parts: The Accident; Why the Accident Occurred; A Look Ahead; and various appendices. To put this accident in context, Parts One and Two begin with histories, after which the accident is described and then analyzed, leading to findings and recommendations. Part Three contains the Board's views on what is needed to improve the safety of our voyage into space. Part Four is reference material. In addition to this first volume, there will be subsequent volumes that contain technical reports generated by the Columbia Accident Investigation Board and NASA, as well as volumes containing reference documentation and other related material.

Gehman, Harold W., Jr.↗

Analysis of the Radiated Field in an Electromagnetic Reverberation Chamber as an Upset-Inducing Stimulus for Digital Systems

Preliminary data analysis for a physical fault injection experiment of a digital system exposed to High Intensity Radiated Fields (HIRF) in an electromagnetic reverberation chamber suggests a direct causal relation between the time profile of the field strength amplitude in the chamber and the severity of observed effects at the outputs of the radiated system. This report presents an analysis of the field strength modulation induced by the movement of the field stirrers in the reverberation chamber. The analysis is framed as a characterization of the discrete features of the field strength waveform responsible for the faults experienced by a radiated digital system. The results presented here will serve as a basis to refine the approach for a detailed analysis of HIRF-induced upsets observed during the radiation experiment. This work offers a novel perspective into the use of an electromagnetic reverberation chamber to generate upset-inducing stimuli for the study of fault effects in digital systems.

Torres-Pomales, Wilfredo↗

Robust fault diagnosis of physical systems in operation

Ideas are presented and demonstrated for improved robustness in diagnostic problem solving of complex physical systems in operation, or operative diagnosis. The first idea is that graceful degradation can be viewed as reasoning at higher levels of abstraction whenever the more detailed levels proved to be incomplete or inadequate. A form of abstraction is defined that applies this view to the problem of diagnosis. In this form of abstraction, named status abstraction, two levels are defined. The lower level of abstraction corresponds to the level of detail at which most current knowledge-based diagnosis systems reason. At the higher level, a graph representation is presented that describes the real-world physical system. An incremental, constructive approach to manipulating this graph representation is demonstrated that supports certain characteristics of operative diagnosis. The suitability of this constructive approach is shown for diagnosing fault propagation behavior over time, and for sometimes diagnosing systems with feedback. A way is shown to represent different semantics in the same type of graph representation to characterize different types of fault propagation behavior. An approach is demonstrated that threats these different behaviors as different fault classes, and the approach moves to other classes when previous classes fail to generate suitable hypotheses. These ideas are implemented in a computer program named Draphys (Diagnostic Reasoning About Physical Systems) and demonstrated for the domain of inflight aircraft subsystems, specifically a propulsion system (containing two turbofan systems and a fuel system) and hydraulic subsystem.

Abbott, Kathy Hamilton↗

An application of synthetic seismicity in earthquake statistics - The Middle America Trench

The way in which seismicity calculations which are based on the concept of fault segmentation incorporate the physics of faulting through static dislocation theory can improve earthquake recurrence statistics and hone the probabilities of hazard is shown. For the Middle America Trench, the spread parameters of the best-fitting lognormal or Weibull distributions (about 0.75) are much larger than the 0.21 intrinsic spread proposed in the Nishenko Buland (1987) hypothesis. Stress interaction between fault segments disrupts time or slip predictability and causes earthquake recurrence to be far more aperiodic than has been suggested.

Ward, Steven N.↗

Making real-time reactive systems reliable

A reactive system is characterized by a control program that interacts with an environment (or controlled program). The control program monitors the environment and reacts to significant events by sending commands to the environment. This structure is quite general. Not only are most embedded real time systems reactive systems, but so are monitoring and debugging systems and distributed application management systems. Since reactive systems are usually long running and may control physical equipment, fault tolerance is vital. The research tries to understand the principal issues of fault tolerance in real time reactive systems and to build tools that allow a programmer to design reliable, real time reactive systems. In order to make real time reactive systems reliable, several issues must be addressed: (1) How can a control program be built to tolerate failures of sensors and actuators. To achieve this, a methodology was developed for transforming a control program that references physical value into one that tolerates sensors that can fail and can return inaccurate values; (2) How can the real time reactive system be built to tolerate failures of the control program. Towards this goal, whether the techniques presented can be extended to real time reactive systems is investigated; and (3) How can the environment be specified in a way that is useful for writing a control program. Towards this goal, whether a system with real time constraints can be expressed as an equivalent system without such constraints is also investigated.

Marzullo, Keith↗

Discovering operating modes in telemetry data from the Shuttle Reaction Control System

This paper addresses the problem of detecting and diagnosing faults in physical systems, for which suitable system models are not available. An architecture is proposed that integrates the on-line acquisition and exploitation of monitoring and diagnostic knowledge. The focus is on the component of the architecture that discovers classes of behaviors with similar characteristics by observing a system in operation. A characterization of behaviors based on best fitting approximation models is investigated. An experimental prototype has been implemented to test it. Preliminary results in diagnosing faults of the reaction control system of the space shuttle are presented. The merits and limitations of the approach are identified and directions for future work are set.

Manganaris, Stefanos↗

Towards a machine learning framework for acquiring and exploiting monitoring and diagnostic knowledge

In this paper we address the problem of detecting and diagnosing faults in physical systems, for which neither prior expertise for the task nor suitable system models are available. We propose an architecture that integrates the on-line acquisition and exploitation of monitoring and diagnostic knowledge. The focus of the paper is on the component of the architecture that discovers classes of behaviors with similar characteristics by observing a system in operation. We investigate a characterization of behaviors based on best fitting approximation models. An experimental prototype has been implemented to test it. We present preliminary results in diagnosing faults of the Reaction Control System of the Space Shuttle. The merits and limitations of the approach are identified and directions for future work are set.

Manganaris, Stefanos↗

Intelligent Engine Systems Work Element 1.3: Sub System Health Management

The objectives of this program were to develop health monitoring systems and physics-based fault detection models for engine sub-systems including the start, lubrication, and fuel. These models will ultimately be used to provide more effective sub-system fault identification and isolation to reduce engine maintenance costs and engine down-time. Additionally, the bearing sub-system health is addressed in this program through identification of sensing requirements, a review of available technologies and a demonstration of a demonstration of a conceptual monitoring system for a differential roller bearing. This report is divided into four sections; one for each of the subtasks. The start system subtask is documented in section 2.0, the oil system is covered in section 3.0, bearing in section 4.0, and the fuel system is presented in section 5.0.

Ashby, Malcolm↗

Recommendations for Enabling Manual Component Level Electronic Repair for Future Space Missions

Long duration missions to the Moon and Mars pose a number of challenges to mission designers, controllers, and the crews. Among these challenges are planning for corrective maintenance actions which often require a repair. Current repair strategies on the International Space Station (ISS) rely primarily on the use of Orbital Replacement Units (ORUs), where a faulty unit is replaced with a spare, and the faulty unit typically returns to Earth for analysis and possible repair. The strategy of replace to repair has posed challenges even for the ISS program. Repairing faulty hardware at lower levels such as the component level can help maintain system availability in situations where no spares exist and potentially reduce logistic resupply mass.This report provides recommendations to help enable manual replacement of electronics at the component-level for future manned space missions. The recommendations include hardware, tools, containment options, and crew training. The recommendations are based on the work of the Component Level Electronics Assembly Repair (CLEAR) task of the Exploration Technology Development Program from 2006 to 2009. The recommendations are derived based on the experience of two experiments conducted by the CLEAR team aboard the International Space Station as well as a group of experienced Miniature/Microminiature (2M) electronics repair technicians and instructors from the U.S. Navy 2M Project Office. The emphasis of the recommendations is the physical repair. Fault diagnostics and post-repair functional test are discussed in other CLEAR reports.

Struk, Peter M.↗

Physics Based Model for Online Fault Detection in Autonomous Cryogenic Loading System

We report the progress in the development of the chilldown model for rapid cryogenic loading system developed at KSC. The nontrivial characteristic feature of the analyzed chilldown regime is its active control by dump valves. The two-phase flow model of the chilldown is approximated as one-dimensional homogeneous fluid flow with no slip condition for the interphase velocity. The model is built using commercial SINDAFLUINT software. The results of numerical predictions are in good agreement with the experimental time traces. The obtained results pave the way to the application of the SINDAFLUINT model as a verification tool for the design and algorithm development required for autonomous loading operation.

Cryogenics↗