Search NASA⌕ Search

SEARCH · Search NASA

Results for “fault analysis”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 271 records · Page 15

Mapping of Landsat satellite and gravity lineaments in west Tennessee

The analysis of earthquake fault lineament patterns within the alluvial valley of west Tennessee, which is often made difficult by the presence of unconsolidated sediments, is presently undertaken through a synergistic use of Landsat satellite images in conjunction with gravity anomaly data, which were quantitatively analyzed and compared by means of two-dimensional histograms and rose diagrams. The northeastern trend revealed for the lineaments corresponds to faults and is in keeping with reactivation of the Reelfoot rift near the Mississippi River; this suggests that deeper features, perhaps at earthquake focal depth, may extend to the land surface as Landsat-detectable lineaments.

Argialas, Demetre P.↗

A HOL theory for voting

Central to fault-tolerant computing is redundancy management, and common to proofs of fault-tolerance is a maximum fault assumption. Typically a maximum fault assumption is rather restrictive. Usually, this is necessary to avoid assumptions about the behavior of faulty channels. A maximum fault assumption is useful because it allows reasoning about fault tolerance in the presence of arbitrarily malicious fault behavior. However, analysis of the architecture may establish certain scenarios in which the assumption may be weakened. Proofs comparing majority and plurality and proofs of simple reconfiguration strategies are presented in viewgraph form.

Miner, Paul S.↗

DEPEND - A design environment for prediction and evaluation of system dependability

The development of DEPEND, an integrated simulation environment for the design and dependability analysis of fault-tolerant systems, is described. DEPEND models both hardware and software components at a functional level, and allows automatic failure injection to assess system performance and reliability. It relieves the user of the work needed to inject failures, maintain statistics, and output reports. The automatic failure injection scheme is geared toward evaluating a system under high stress (workload) conditions. The failures that are injected can affect both hardware and software components. To illustrate the capability of the simulator, a distributed system which employs a prediction-based, dynamic load-balancing heuristic is evaluated. Experiments were conducted to determine the impact of failures on system performance and to identify the failures to which the system is especially susceptible.

Goswami, Kumar K.↗

Fault detection in digital and analog circuits using an i(DD) temporal analysis technique

An i(sub DD) temporal analysis technique which is used to detect defects (faults) and fabrication variations in both digital and analog IC's by pulsing the power supply rails and analyzing the temporal data obtained from the resulting transient rail currents is presented. A simple bias voltage is required for all the inputs, to excite the defects. Data from hardware tests supporting this technique are presented.

Beasley, J.↗

Adaptation of a Control Center Development Environment for Industrial Process Control

In the control center, raw telemetry data is received for storage, display, and analysis. This raw data must be combined and manipulated in various ways by mathematical computations to facilitate analysis, provide diversified fault detection mechanisms, and enhance display readability. A development tool called the Graphical Computation Builder (GCB) has been implemented which provides flight controllers with the capability to implement computations for use in the control center. The GCB provides a language that contains both general programming constructs and language elements specifically tailored for the control center environment. The GCB concept allows staff who are not skilled in computer programming to author and maintain computer programs. The GCB user is isolated from the details of external subsystem interfaces and has access to high-level functions such as matrix operators, trigonometric functions, and unit conversion macros. The GCB provides a high level of feedback during computation development that improves upon the often cryptic errors produced by computer language compilers. An equivalent need can be identified in the industrial data acquisition and process control domain: that of an integrated graphical development tool tailored to the application to hide the operating system, computer language, and data acquisition interface details. The GCB features a modular design which makes it suitable for technology transfer without significant rework. Control center-specific language elements can be replaced by elements specific to industrial process control.

Killough, Ronnie L.↗

Parallel NPARC: Implementation and Performance

Version 3 of the NPARC Navier-Stokes code includes support for large-grain (block level) parallelism using explicit message passing between a heterogeneous collection of computers. This capability has the potential for significant performance gains, depending upon the block data distribution. The parallel implementation uses a master/worker arrangement of processes. The master process assigns blocks to workers, controls worker actions, and provides remote file access for the workers. The processes communicate via explicit message passing using an interface library which provides portability to a number of message passing libraries, such as PVM (Parallel Virtual Machine). A Bourne shell script is used to simplify the task of selecting hosts, starting processes, retrieving remote files, and terminating a computation. This script also provides a simple form of fault tolerance. An analysis of the computational performance of NPARC is presented, using data sets from an F/A-18 inlet study and a Rocket Based Combined Cycle Engine analysis. Parallel speedup and overall computational efficiency were obtained for various NPARC run parameters on a cluster of IBM RS6000 workstations. The data show that although NPARC performance compares favorably with the estimated potential parallelism, typical data sets used with previous versions of NPARC will often need to be reblocked for optimum parallel performance. In one of the cases studied, reblocking increased peak parallel speedup from 3.2 to 11.8.

Townsend, S. E.↗

Developing An Autonomy Infusion Infrastructure for Robotic Exploration

Future robotic exploration missions will require autonomy in order to accomplish mission goals for operational efficiency and science return. For example, it will require three communication cycles for the Mars Exploration Rovers, Spirit and Opportunity, to place an instrument on a science target. Reducing this time necessitates highly accurate navigation, obstacle avoidance, target tracking, target analysis, manipulation, and fault diagnosis. Technologies to address these and other operational elements are currently being developed at NASA and within academia. However, infusion into missions has always been a difficult task for researchers. In order to keep risk down, mission managers are reluctant to include new technologies unless they have undergone extensive testing and verification under flight-realistic conditions. Furthermore, infusion of new technologies into missions is made more difficult by the variety of software frameworks under which these technologies are developed. Missions would like to see competing solutions demonstrated on a common platform so that they can compare performance and choose the solution best suited to their application.

Bualat, Maria G.↗

Electrical Engineering Medium Voltage Protection and Coordination

This presentation covers the basics of the medium voltage electrical system on the Kennedy Space Center. The material will include the process of selective coordination between protective devices and the analysis of a fault that occurred in the electrical system. I will discuss the most common causes of faults found and how the system reacts to the faults at different locations in the subsystem. Explain the importance of the sensitivity and coordination between the devices to protect the equipment it is supplying.

Wessner, Jason R.↗

Software fault tolerance in computer operating systems

This chapter provides data and analysis of the dependability and fault tolerance for three operating systems: the Tandem/GUARDIAN fault-tolerant system, the VAX/VMS distributed system, and the IBM/MVS system. Based on measurements from these systems, basic software error characteristics are investigated. Fault tolerance in operating systems resulting from the use of process pairs and recovery routines is evaluated. Two levels of models are developed to analyze error and recovery processes inside an operating system and interactions among multiple instances of an operating system running in a distributed environment. The measurements show that the use of process pairs in Tandem systems, which was originally intended for tolerating hardware faults, allows the system to tolerate about 70% of defects in system software that result in processor failures. The loose coupling between processors which results in the backup execution (the processor state and the sequence of events occurring) being different from the original execution is a major reason for the measured software fault tolerance. The IBM/MVS system fault tolerance almost doubles when recovery routines are provided, in comparison to the case in which no recovery routines are available. However, even when recovery routines are provided, there is almost a 50% chance of system failure when critical system jobs are involved.

Iyer, Ravishankar K.↗

Analysis of a Disturbance Event with Inverter-Based Resources Using EMT Simulations

Increasing penetration of inverter-based resources (IBRs) necessitates newer methods of planning and analysis of disturbances. The existing phasor-domain transient stability (TS) analysis may not capture the dynamics of IBRs during fault events. Here, in this paper, electromagnetic transient (EMT) simulations using high-fidelity detailed model of power grid and one of the affected photovoltaic (PV) plants during the Angeles Forest disturbance in 2018 are performed. In these simulations, the processes to develop EMT models of power grid from traditional phasor-domain TS data and PV plant from collected data are described. Thereafter, using these simulations, the response of the PV plant during the fault event in 2018 is replicated and a sensitivity analysis is performed. The sensitivity analysis consists of making changes to the components within the PV plant and in the power grid to evaluate the impact they have on the response observed by the PV plant during the fault event. This analysis provides an understanding of the components that impact the operation of a PV plant during fault events and provide guidance to system planners on the studies that need to be performed to maintain a reliable power grid as new IBR plants are integrated.

42 ENGINEERING↗

Analysis of the survivability of the shuttle (ALT) fault-tolerant avionics system

An extension of the Complementary-Analytic-Simulative Technique (CAST) is presented which is applicable to the Shuttle Data Processing Subsystem (DPS). A two step process was used. The first step provides models, both analytic and simulative, for analysis of the Approach-Landing Test (ALT) configuration. The ALT modeling and analysis are presented. Since CAST had already been shown to be multicomputer systems, the emphasis was placed on extending the CAST concept so it is applicable to computer systems including the multiplicity of input and output devices found in a real-time control system application. The DPS mission-critical survivability for a six-hour mission was determined to be 0.999863 for the Shuttle ALT baseline configuration. Thus it can be said that for ALT, the survivability is adequate. However, the fact that orbiting missions of up to 30 days are planned illustrates the necessity of extending the ALT work to be applicable to OFT and actual mission scenarios. The above analysis led to the evaluation of three selected options which identified two areas of possible improvement. These improvements would result from use of a recovery technique which combines roll ahead with memory copy, and increased TACAN fault detectability.

Source record↗

Modeling uncertainty in requirements engineering decision support

One inherent characteristic of requrements engineering is a lack of certainty during this early phase of a project. Nevertheless, decisions about requirements must be made in spite of this uncertainty. Here we describe the context in which we are exploring this, and some initial work to support elicitation of uncertain requirements, and to deal with the combination of such information from multiple stakeholders.

risk analysis↗

A Test Generation Framework for Distributed Fault-Tolerant Algorithms

Heavyweight formal methods such as theorem proving have been successfully applied to the analysis of safety critical fault-tolerant systems. Typically, the models and proofs performed during such analysis do not inform the testing process of actual implementations. We propose a framework for generating test vectors from specifications written in the Prototype Verification System (PVS). The methodology uses a translator to produce a Java prototype from a PVS specification. Symbolic (Java) PathFinder is then employed to generate a collection of test cases. A small example is employed to illustrate how the framework can be used in practice.

Goodloe, Alwyn↗

Post-carboniferous tectonics in the Anadarko Basin, Oklahoma: Evidence from side-looking radar imagery

The Anadarko Basin of western Oklahoma is a WNW-ESE elongated trough filled with of Paleozoic sediments. Most models call for tectonic activity to end in Pennsylvanian times. NASA Shuttle Imaging Radar revealed a distinctive and very straight lineament set extending virtually the entire length of the Anadarko Basin. The lineaments cut across the relatively flat-lying Permian units exposed at the surface. The character of these lineaments is seen most obviously as a tonal variation. Major streams, including the Washita and Little Washita rivers, appear to be controlled by the location of the lineaments. Subsurface data indicate the lineaments may be the updip expression of a buried major fault system, the Mountain View fault. Two principal conclusions arise from this analysis: (1) the complex Mountain View Fault system appears to extend southeast to join the Reagan, Sulphur, and/or Mill Creek faults of the Arbuckle Mountains, and (2) this fault system has been reactivated in Permian or younger times.

Nielsen, K. C.↗

An experimental study of fault propagation in a jet-engine controller

An experimental analysis of the impact of transient faults on a microprocessor-based jet engine controller, used in the Boeing 747 and 757 aircrafts is described. A hierarchical simulation environment which allows the injection of transients during run-time and the tracing of their impact is described. Verification of the accuracy of this approach is also provided. A determination of the probability that a transient results in latch, pin or functional errors is made. Given a transient fault, there is approximately an 80 percent chance that there is no impact on the chip. An empirical model to depict the process of error exploration and degeneration in the target system is derived. The model shows that, if no latch errors occur within eight clock cycles, no significant damage is likely to happen. Thus, the overall impact of a transient is well contained. A state transition model is also derived from the measured data, to describe the error propagation characteristics within the chip, and to quantify the impact of transients on the external environment. The model is used to identify and isolate the critical fault propagation paths, the module most sensitive to fault propagation and the module with the highest potential of causing external pin errors.

Choi, Gwan Seung↗

Analytical Approaches to Guide SLS Fault Management (FM) Development

Extensive analysis is needed to determine the right set of FM capabilities to provide the most coverage without significantly increasing the cost, reliability (FP/FN), and complexity of the overall vehicle systems. Strong collaboration with the stakeholders is required to support the determination of the best triggers and response options. The SLS Fault Management process has been documented in the Space Launch System Program (SLSP) Fault Management Plan (SLS-PLAN-085).

Patterson, Jonathan D.↗

Model-Based Investigation of Multi-Fault Interactions and Performance Degradation in Residential Heat Pump Systems

Faults in heat pump systems can significantly degrade performance, reduce efficiency, and accelerate component wear, leading to higher operating costs and maintenance demands. While numerous studies have investigated the impact of individual faults, the interactions between multiple concurrent faults remain insufficiently understood, despite their common occurrence in real-world operation. This study conducts a comprehensive simulation analysis of multiple simultaneous faults in a vapor compression heat pump using a validated heat pump design model (HPDM) tool. Detailed component-level modeling methods are implemented to examine performance sensitivity under combinations of refrigerant flow and heat exchanger faults. The results reveal complex fault interactions that can mask or amplify system deviations, challenging conventional diagnostic approaches. Findings from this work provide meaningful insights for the development of more robust fault detection and diagnosis algorithms, supporting improved reliability and energy efficiency in next-generation heat pump technologies.

Hu, Yifeng [ORNL] (ORCID:0000000242875185)↗

Hierarchical Simulation to Assess Hardware and Software Dependability

This thesis presents a method for conducting hierarchical simulations to assess system hardware and software dependability. The method is intended to model embedded microprocessor systems. A key contribution of the thesis is the idea of using fault dictionaries to propagate fault effects upward from the level of abstraction where a fault model is assumed to the system level where the ultimate impact of the fault is observed. A second important contribution is the analysis of the software behavior under faults as well as the hardware behavior. The simulation method is demonstrated and validated in four case studies analyzing Myrinet, a commercial, high-speed networking system. One key result from the case studies shows that the simulation method predicts the same fault impact 87.5% of the time as is obtained by similar fault injections into a real Myrinet system. Reasons for the remaining discrepancy are examined in the thesis. A second key result shows the reduction in the number of simulations needed due to the fault dictionary method. In one case study, 500 faults were injected at the chip level, but only 255 propagated to the system level. Of these 255 faults, 110 shared identical fault dictionary entries at the system level and so did not need to be resimulated. The necessary number of system-level simulations was therefore reduced from 500 to 145. Finally, the case studies show how the simulation method can be used to improve the dependability of the target system. The simulation analysis was used to add recovery to the target software for the most common fault propagation mechanisms that would cause the software to hang. After the modification, the number of hangs was reduced by 60% for fault injections into the real system.

Ries, Gregory Lawrence↗