Search NASA⌕ Search

SEARCH · Search NASA

Results for “fault analysis”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 361 records · Page 20

An experiment in software reliability

The results of a software reliability experiment conducted in a controlled laboratory setting are reported. The experiment was undertaken to gather data on software failures and is one in a series of experiments being pursued by the Fault Tolerant Systems Branch of NASA Langley Research Center to find a means of credibly performing reliability evaluations of flight control software. The experiment tests a small sample of implementations of radar tracking software having ultra-reliability requirements and uses n-version programming for error detection, and repetitive run modeling for failure and fault rate estimation. The experiment results agree with those of Nagel and Skrivan in that the program error rates suggest an approximate log-linear pattern and the individual faults occurred with significantly different error rates. Additional analysis of the experimental data raises new questions concerning the phenomenon of interacting faults. This phenomenon may provide one explanation for software reliability decay.

Dunham, J. R.↗

Experiments in software reliability - Life-critical applications

The paper discusses four reliability data gathering experiments which were conducted using a small sample of programs for two problems having ultrareliability requirements, n-version programming for fault detection, and repetitive run modeling for failure and fault rate estimation. The experimental results agree with those of Nagel and Skrivan in that the program error rates suggest an approximate log-linear pattern and the individual faults occurred with significantly different error rates. Additional analysis of the experimental data raises new questions concerning the phenomenon of interacting faults. This phenomenon may provide one explanation for software reliability decay. The fourth experiment underscored the difficulty in distinguishing between observations of deficiencies in the design of the algorithm and observations of software faults for real-time process control software. These experiments are a part of a program of serial experiments being pursued by the System Validation Methods of NASA-Langley Research Center to find a means of credibly performing reliability evaluations of flight control software.

Dunham, J. R.↗

Design and fabrication of prototype system for early warning of impending bearing failure

A test program was conducted with the objective of developing a method and equipment for on-line monitoring of installed ball bearings to detect deterioration or impending failure of the bearings. The program was directed at the spin-axis bearings of a control moment gyro. The bearings were tested at speeds of 6000 and 8000 rpm, thrust loads from 50 to 1000 pounds, with a wide range of lubrication conditions, with and without a simulated fatigue spall implanted in the inner race ball track. It was concluded that a bearing monitor system based on detection and analysis of modulations of a fault indicating bearing resonance frequency can provide a low threshold of sensitivity.

Meacher, J.↗

Theory of reliable systems

Research is reported in the program to refine the current notion of system reliability by identifying and investigating attributes of a system which are important to reliability considerations, and to develop techniques which facilitate analysis of system reliability. Reliability analysis, and on-line fault diagnosis are discussed.

Meyer, J. F.↗

Synchronization and fault-masking in redundant real-time systems

A real time computer may fail because of massive component failures or not responding quickly enough to satisfy real time requirements. An increase in redundancy - a conventional means of improving reliability - can improve the former but can - in some cases - degrade the latter considerably due to the overhead associated with redundancy management, namely the time delay resulting from synchronization and voting/interactive consistency techniques. The implications of synchronization and voting/interactive consistency algorithms in N-modular clusters on reliability are considered. All these studies were carried out in the context of real time applications. As a demonstrative example, we have analyzed results from experiments conducted at the NASA Airlab on the Software Implemented Fault Tolerance (SIFT) computer. This analysis has indeed indicated that in most real time applications, it is better to employ hardware synchronization instead of software synchronization and not allow reconfiguration.

Krishna, C. M.↗

MTK: An AI tool for model-based reasoning

A 1988 goal for the Systems Autonomy Demonstration Project Office of the NASA Ames Research Center is to apply model-based representation and reasoning techniques in a knowledge-based system that will provide monitoring, fault diagnosis, control and trend analysis of the space station Thermal Management System (TMS). A number of issues raised during the development of the first prototype system inspired the design and construction of a model-based reasoning tool called MTK, which was used in the building of the second prototype. These issues are outlined, along with examples from the thermal system to highlight the motivating factors behind them. An overview of the capabilities of MTK is given.

Erickson, William K.↗

MTK: An AI tool for model-based reasoning

A 1988 goal for the Systems Autonomy Demonstration Project Office of the NASA Ames Research Office is to apply model-based representation and reasoning techniques in a knowledge-based system that will provide monitoring, fault diagnosis, control, and trend analysis of the Space Station Thermal Control System (TCS). A number of issues raised during the development of the first prototype system inspired the design and construction of a model-based reasoning tool called MTK, which was used in the building of the second prototype. These issues are outlined here with examples from the thermal system to highlight the motivating factors behind them, followed by an overview of the capabilities of MTK, which was developed to address these issues in a generic fashion.

Erickson, William K.↗

Automated generation of reliability models

The abstract semi-Markov specification interface to the SURE (Semi-Markov Range Evaluator) tool (ASSIST) program allows the user to describe the Markov model in a high-level language. Instead of listing the individual states of the model, the user specifies the rules governing the behavior of the system, and these are used to automatically generate the model. A small number of statements in the abstract language can describe a large, complex model. Becuase no assumptions are made about the system being modeled, ASSIST can be used to generate models describing the behavior of any type of system. The abstract model definition and the automatic model generation strategy are described. Analysis of an example fault-tolerant architecture, a triad of processor with cold spare processors, shows how the behavior of a system can be captured by a few general rules. The syntax of the ASSIST input language is then described and demonstrated by creating a model to describe the fault behavior of the example architecture. The flexibility of the abstract language is demonstrated by expanding the example to model multiple triads of processors sharing a pool of cold spare processors.

Johnson, Sally C.↗

Modeling and measurement of error propagation in a multimodule computing system

An error propagation model has been developed for multimodule computing systems in which the main parameters are the distribution functions of error propagation times. A digraph model is used to represent a multimodule computing system, and error propagation in the system is modeled by general distributions of error propagation times between all pairs of modules. Two algorithms are developed to compute systematically and efficiently the distributions of error propagation times. Experiments are also conducted to measure the distributions of error propagation times with the fault-tolerant microprocessor (FTMP). Statistical analysis of experimental data shows that the error propagation times in FTMP do not follow a well-known distribution, thus justifying the use of general distributions in the present model.

Shin, Kang G.↗

Semi-Markov Unreliability Range Evaluator (SURE)

Analysis tool for reconfigurable, fault-tolerant systems, SURE provides efficient way to calculate accurate upper and lower bounds for death state probabilities for large class of semi-Markov models. Calculated bounds close enough for use in reliability studies of ultrareliable computer systems. Written in PASCAL for interactive execution and runs on DEC VAX computer under VMS.

Butler, R. W.↗

Generating Semi-Markov Models Automatically

Abstract Semi-Markov Specification Interface to SURE Tool (ASSIST) program developed to generate semi-Markov model automatically from description in abstract, high-level language. ASSIST reads input file describing failure behavior of system in abstract language and generates Markov models in format needed for input to Semi-Markov Unreliability Range Evaluator (SURE) program (COSMIC program LAR-13789). Facilitates analysis of behavior of fault-tolerant computer. Written in PASCAL.

Johnson, Sally C.↗

Framework for a space shuttle main engine health monitoring system

A framework developed for a health management system (HMS) which is directed at improving the safety of operation of the Space Shuttle Main Engine (SSME) is summarized. An emphasis was placed on near term technology through requirements to use existing SSME instrumentation and to demonstrate the HMS during SSME ground tests within five years. The HMS framework was developed through an analysis of SSME failure modes, fault detection algorithms, sensor technologies, and hardware architectures. A key feature of the HMS framework design is that a clear path from the ground test system to a flight HMS was maintained. Fault detection techniques based on time series, nonlinear regression, and clustering algorithms were developed and demonstrated on data from SSME ground test failures. The fault detection algorithms exhibited 100 percent detection of faults, had an extremely low false alarm rate, and were robust to sensor loss. These algorithms were incorporated into a hierarchical decision making strategy for overall assessment of SSME health. A preliminary design for a hardware architecture capable of supporting real time operation of the HMS functions was developed. Utilizing modular, commercial off-the-shelf components produced a reliable low cost design with the flexibility to incorporate advances in algorithm and sensor technology as they become available.

Hawman, Michael W.↗

Networks for image acquisition, processing and display

The human visual system comprises layers of networks which sample, process, and code images. Understanding these networks is a valuable means of understanding human vision and of designing autonomous vision systems based on network processing. Ames Research Center has an ongoing program to develop computational models of such networks. The models predict human performance in detection of targets and in discrimination of displayed information. In addition, the models are artificial vision systems sharing properties with biological vision that has been tuned by evolution for high performance. Properties include variable density sampling, noise immunity, multi-resolution coding, and fault-tolerance. The research stresses analysis of noise in visual networks, including sampling, photon, and processing unit noises. Specific accomplishments include: models of sampling array growth with variable density and irregularity comparable to that of the retinal cone mosaic; noise models of networks with signal-dependent and independent noise; models of network connection development for preserving spatial registration and interpolation; multi-resolution encoding models based on hexagonal arrays (HOP transform); and mathematical procedures for simplifying analysis of large networks.

Ahumada, Albert J., Jr.↗

Complex structure of the Thaumasia region of Mars

The Thaumasia region was the first center of Tharsis tectonism, and it is the most complex and poorly understood. Therefore, a geologic map of the entire Thaumasia region (lat 15 deg to 50 deg S, long 50 deg to 115 deg), at 1:5,000,000 scale is being compiled. This region is mostly made up of the Thaumasia plateau, the highlands are fractured by Thaumasia, southern Claritas, Coracis, Melas, and Nectaris Fossae. Preliminary structural analysis of the most complexly faulted area in the region (the central part, at lat 30 deg to 45 deg S, long 80 deg to 100 deg) indicates that, unlike other regions of Mars, Thaumasia has undergone extensive deformation by both small- and large-scale extensional and compressional structures. These results indicate that the early (Noachian) style of tectonism commonly involved lithospheric-scale deformation, in contrast to most younger tectonism (which mainly affected the upper parts of the crust above mechanical discontinuities); this difference may be due to a weaker (and thus more readily deformable) early lithosphere in this region.

Tanaka, Kenneth L.↗

A Model-based Health Monitoring and Diagnostic System for the UH-60 Helicopter

Model-based reasoning techniques hold much promise in providing comprehensive monitoring and diagnostics capabilities for complex systems. We are exploring the use of one of these techniques, which utilizes multi-signal modeling and the TEAMS-RT real-time diagnostic engine, on the UH-60 Rotorcraft Aircrew Systems Concepts Airborne Laboratory (RASCAL) flight research aircraft. We focus on the engine and transmission systems, and acquire sensor data across the 1553 bus as well as by direct analog-to-digital conversion from sensors to the QHuMS (Qualtech health and usage monitoring system) computer. The QHuMS computer uses commercially available components and is rack-mounted in the RASCAL facility. A multi-signal model of the transmission and engine subsystems enables studies of system testability and analysis of the degree of fault isolation available with various instrumentation suites. The model and examples of these analyses will be described and the data architectures enumerated. Flight tests of this system will validate the data architecture and provide real-time flight profiles to be further analyzed in the laboratory.

Patterson-Hine, Ann↗

Stochastic Stability of Sampled Data Systems with a Jump Linear Controller

In this paper an equivalence between the stochastic stability of a sampled-data system and its associated discrete-time representation is established. The sampled-data system consists of a deterministic, linear, time-invariant, continuous-time plant and a stochastic, linear, time-invariant, discrete-time, jump linear controller. The jump linear controller models computer systems and communication networks that are subject to stochastic upsets or disruptions. This sampled-data model has been used in the analysis and design of fault-tolerant systems and computer-control systems with random communication delays without taking into account the inter-sample response. This paper shows that the known equivalence between the stability of a deterministic sampled-data system and the associated discrete-time representation holds even in a stochastic framework.

Gonzalez, Oscar R.↗

Stochastic Stability of Nonlinear Sampled Data Systems with a Jump Linear Controller

This paper analyzes the stability of a sampled- data system consisting of a deterministic, nonlinear, time- invariant, continuous-time plant and a stochastic, discrete- time, jump linear controller. The jump linear controller mod- els, for example, computer systems and communication net- works that are subject to stochastic upsets or disruptions. This sampled-data model has been used in the analysis and design of fault-tolerant systems and computer-control systems with random communication delays without taking into account the inter-sample response. To analyze stability, appropriate topologies are introduced for the signal spaces of the sampled- data system. With these topologies, the ideal sampling and zero-order-hold operators are shown to be measurable maps. This paper shows that the known equivalence between the stability of a deterministic, linear sampled-data system and its associated discrete-time representation as well as between a nonlinear sampled-data system and a linearized representation holds even in a stochastic framework.

Gonzalez, Oscar R.↗

Anomaly Resolution in the International Space Station

Topics include post flight 2A status, groundrules, anomaly resolution, Early Communications Subsystem anomaly and resolution, Logistics and Maintenance plan, case for obscuration, case for electrical short, and manual fault isolation, and post mission analysis. Photographs from flight 2A.1 are used to illustrate anomalies.

Evans, William A.↗