Search NASA⌕ Search

SEARCH · Search NASA

Results for “faults”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 91 records · Page 5

Software-implemented fault insertion: An FTMP example

This report presents a model for fault insertion through software; describes its implementation on a fault-tolerant computer, FTMP; presents a summary of fault detection, identification, and reconfiguration data collected with software-implemented fault insertion; and compares the results to hardware fault insertion data. Experimental results show detection time to be a function of time of insertion and system workload. For the fault detection time, there is no correlation between software-inserted faults and hardware-inserted faults; this is because hardware-inserted faults must manifest as errors before detection, whereas software-inserted faults immediately exercise the error detection mechanisms. In summary, the software-implemented fault insertion is able to be used as an evaluation technique for the fault-handling capabilities of a system in fault detection, identification and recovery. Although the software-inserted faults do not map directly to hardware-inserted faults, experiments show software-implemented fault insertion is capable of emulating hardware fault insertion, with greater ease and automation.

Czeck, Edward W.↗

A Unified Nonlinear Adaptive Approach for Detection and Isolation of Engine Faults

A challenging problem in aircraft engine health management (EHM) system development is to detect and isolate faults in system components (i.e., compressor, turbine), actuators, and sensors. Existing nonlinear EHM methods often deal with component faults, actuator faults, and sensor faults separately, which may potentially lead to incorrect diagnostic decisions and unnecessary maintenance. Therefore, it would be ideal to address sensor faults, actuator faults, and component faults under one unified framework. This paper presents a systematic and unified nonlinear adaptive framework for detecting and isolating sensor faults, actuator faults, and component faults for aircraft engines. The fault detection and isolation (FDI) architecture consists of a parallel bank of nonlinear adaptive estimators. Adaptive thresholds are appropriately designed such that, in the presence of a particular fault, all components of the residual generated by the adaptive estimator corresponding to the actual fault type remain below their thresholds. If the faults are sufficiently different, then at least one component of the residual generated by each remaining adaptive estimator should exceed its threshold. Therefore, based on the specific response of the residuals, sensor faults, actuator faults, and component faults can be isolated. The effectiveness of the approach was evaluated using the NASA C-MAPSS turbofan engine model, and simulation results are presented.

Tang, Liang↗

Application of Fault Containment Principles to EMC

A hardware fault in aerospace is an undesired response to a designed engineering function in hardware. Therefore, in aerospace systems fault management is of prime importance. An important aspect of hardware faults is fault propagation. Fault propagation (also known as failure propagation) is a condition where a fault will not only produce an undesired hardware response, at its location of origin, but the fault will also propagate to other interfaced hardware and cause additional faults, or failures. Fault containment is necessary to avoid fault propagation. A fault containment region is an electronic or electromechanical region within a given hardware assembly where a fault in that region will not physically propagate to other regions of the assembly, and beyond. Rather, the fault will cause a functional failure of the hardware, where the fault occurred, without causing any additional propagated failures. Electromagnetic Compatibility (EMC) is a desired state in aerospace hardware. Electromagnetic Interference (EMI) in aerospace hardware is a fault condition of EMC. Like other faults EMI can also propagate to other aerospace hardware unless it occurs inside an EMI containment region. The paper addresses the concepts of fault containment region and introduces the concept of EMI containment region with examples of both.

Perez, Reinaldo↗

Network Connectivity for Permanent, Transient, Independent, and Correlated Faults

This paper develops a method for the quantitative analysis of network connectivity in the presence of both permanent and transient faults. Even though transient noise is considered a common occurrence in networks, a survey of the literature reveals an emphasis on permanent faults. Transient faults introduce a time element into the analysis of network reliability. With permanent faults it is sufficient to consider the faults that have accumulated by the end of the operating period. With transient faults the arrival and recovery time must be included. The number and location of faults in the system is a dynamic variable. Transient faults also introduce system recovery into the analysis. The goal is the quantitative assessment of network connectivity in the presence of both permanent and transient faults. The approach is to construct a global model that includes all classes of faults: permanent, transient, independent, and correlated. A theorem is derived about this model that give distributions for (1) the number of fault occurrences, (2) the type of fault occurrence, (3) the time of the fault occurrences, and (4) the location of the fault occurrence. These results are applied to compare and contrast the connectivity of different network architectures in the presence of permanent, transient, independent, and correlated faults. The examples below use a Monte Carlo simulation, but the theorem mentioned above could be used to guide fault-injections in a laboratory.

White, Allan L.↗

Analysis of a hardware and software fault tolerant processor for critical applications

Computer systems for critical applications must be designed to tolerate software faults as well as hardware faults. A unified approach to tolerating hardware and software faults is characterized by classifying faults in terms of duration (transient or permanent) rather than source (hardware or software). Errors arising from transient faults can be handled through masking or voting, but errors arising from permanent faults require system reconfiguration to bypass the failed component. Most errors which are caused by software faults can be considered transient, in that they are input-dependent. Software faults are triggered by a particular set of inputs. Quantitative dependability analysis of systems which exhibit a unified approach to fault tolerance can be performed by a hierarchical combination of fault tree and Markov models. A methodology for analyzing hardware and software fault tolerant systems is applied to the analysis of a hypothetical system, loosely based on the Fault Tolerant Parallel Processor. The models consider both transient and permanent faults, hardware and software faults, independent and related software faults, automatic recovery, and reconfiguration.

Dugan, Joanne B.↗

Earthquake recurrence on the southern San Andreas modulated by fault-normal stress

Earthquake recurrence data from the Pallett Creek and Wrightwood paleoseismic sites on the San Andreas fault appear to show temporal variations in repeat interval. We investigate the interaction between strike-slip faults and auxiliary reverse and normal faults as a physical mechanism capable of producing such variations. Under the assumption that fault strength is a function of fault-normal stress (e.g. Byerlee's Law), failure of an auxiliary fault modifies the strength of the strike-slip fault, thereby modulating the recurrence interval for earthquakes. In our finite element model, auxiliary faults are driven by stress accumulation near restraining and releasing bends of a strike-slip fault. Earthquakes occur when fault strength is exceeded and are incorporated as a stress drop which is dependent on fault-normal stress. The model is driven by a velocity boundary condition over many earthquake cycles. Resulting synthetic strike-slip earthquake recurrence data display temporal variations similar to observed paleoseismic data within time windows surrounding auxiliary fault failures. Our simple model supports the idea that interaction between a strike-slip fault and auxiliary reverse or normal faults can modulate the recurrence interval of events on the strike-slip fault, possibly producing short term variations in earthquake recurrence interval.

Palmer, Randy↗

Automated Monitoring with a BCP Fault-Decision Test

The Bayesian conditional probability (BCP) technique is a statistical fault-decision technique that is suitable as the mathematical basis of the fault-manager module in the automated-monitoring system and method described in the immediately preceding article. Within the automated-monitoring system, the fault-manager module operates in conjunction with the fault-detector module, which can be based on any one of several fault-detection techniques; examples include a threshold-limit-comparison technique or the BSP or SPRT technique mentioned in the preceding article. The present BCP technique is used to evaluate a series of one or more fault-detection events for the purpose of filtering out occasional false alarms produced by many types of statistical fault-detection procedures. The BCP technique increases the probability that an automated monitoring system produces a correct decision regarding the presence or absence of a fault. Because occasional false alarms are an inevitable consequence of the SPRT, BSP, or any other statistically based fault-detection test, there is a need for a logical procedure to distinguish between true and false alarms. Heretofore, it has been common practice to make a fault decision on an ad hoc basis for example by following a multiple-observation voting strategy in which a signal is declared to be indicative of a fault if m of the last n observations produced a fault-detection alarm. The BCP technique was developed to obtain results more reliable than those afforded by a voting strategy. The BCP technique involves a test in which one applies Bayesian inference techniques to a series of one or more single-observation alarms produced by a fault-detection test. One considers the last n decisions generated by a fault-detection test in order to evaluate the conditional probability that a failure is indicated (see figure). Each new decision reached by a fault-detection test is treated as a new piece of evidence about the state of the monitored asset, and the conditional probability of failure for the system is updated on the basis of this new evidence. The conditional probability of failure is compared with a predefined limit. For a probability below the limit, the asset is declared to be healthy. For a probability above the limit, the asset is declared to be faulty.

Bickford, Randall L.↗

A distributed fault-detection and diagnosis system using on-line parameter estimation

The development of a model-based fault-detection and diagnosis system (FDD) is reviewed. The system can be used as an integral part of an intelligent control system. It determines the faults of a system from comparison of the measurements of the system with a priori information represented by the model of the system. The method of modeling a complex system is described and a description of diagnosis models which include process faults is presented. There are three distinct classes of fault modes covered by the system performance model equation: actuator faults, sensor faults, and performance degradation. A system equation for a complete model that describes all three classes of faults is given. The strategy for detecting the fault and estimating the fault parameters using a distributed on-line parameter identification scheme is presented. A two-step approach is proposed. The first step is composed of a group of hypothesis testing modules, (HTM) in parallel processing to test each class of faults. The second step is the fault diagnosis module which checks all the information obtained from the HTM level, isolates the fault, and determines its magnitude. The proposed FDD system was demonstrated by applying it to detect actuator and sensor faults added to a simulation of the Space Shuttle Main Engine. The simulation results show that the proposed FDD system can adequately detect the faults and estimate their magnitudes.

Guo, T.-H.↗

Optimization of Second Fault Detection Thresholds to Maximize Mission POS

In order to support manned spaceflight safety requirements, the Space Launch System (SLS) has defined program-level requirements for key systems to ensure successful operation under single fault conditions. To accommodate this with regards to Navigation, the SLS utilizes an internally redundant Inertial Navigation System (INS) with built-in capability to detect, isolate, and recover from first failure conditions and still maintain adherence to performance requirements. The unit utilizes multiple hardware- and software-level techniques to enable detection, isolation, and recovery from these events in terms of its built-in Fault Detection, Isolation, and Recovery (FDIR) algorithms. Successful operation is defined in terms of sufficient navigation accuracy at insertion while operating under worst case single sensor outages (gyroscope and accelerometer faults at launch). In addition to first fault detection and recovery, the SLS program has also levied requirements relating to the capability of the INS to detect a second fault, tracking any unacceptable uncertainty in knowledge of the vehicle's state. This detection functionality is required in order to feed abort analysis and ensure crew safety. Increases in navigation state error and sensor faults can drive the vehicle outside of its operational as-designed environments and outside of its performance envelope causing loss of mission, or worse, loss of crew. The criteria for operation under second faults allows for a larger set of achievable missions in terms of potential fault conditions, due to the INS operating at the edge of its capability. As this performance is defined and controlled at the vehicle level, it allows for the use of system level margins to increase probability of mission success on the operational edges of the design space. Due to the implications of the vehicle response to abort conditions (such as a potentially failed INS), it is important to consider a wide range of failure scenarios in terms of both magnitude and time. As such, the Navigation team is taking advantage of the INS's capability to schedule and change fault detection thresholds in flight. These values are optimized along a nominal trajectory in order to maximize probability of mission success, and reducing the probability of false positives (defined as when the INS would report a second fault condition resulting in loss of mission, but the vehicle would still meet insertion requirements within system-level margins). This paper will describe an optimization approach using Genetic Algorithms to tune the threshold parameters to maximize vehicle resilience to second fault events as a function of potential fault magnitude and time of fault over an ascent mission profile. The analysis approach, and performance assessment of the results will be presented to demonstrate the applicability of this process to second fault detection to maximize mission probability of success.

Anzalone, Evan↗

Correlation of data on strain accumulation adjacent to the San Andreas Fault with available models

Theoretical and numerical studies of deformation on strike slip faults were performed and the results applied to geodetic observations performed in the vicinity of the San Andreas Fault in California. The initial efforts were devoted to an extensive series of finite element calculations of the deformation associated with cyclic displacements on a strike-slip fault. Measurements of strain accumulation adjacent to the San Andreas Fault indicate that the zone of strain accumulation extends only a few tens of kilometers away from the fault. There is a concern about the tendency to make geodetic observations along the line to the source. This technique has serious problems for strike slip faults since the vector velocity is also along the fault. Use of a series of stations lying perpendicular to the fault whose positions are measured relative to a reference station are suggested to correct the problem. The complexity of faulting adjacent to the San Andreas Fault indicated that the homogeneous elastic and viscoelastic approach to deformation had serious limitations. These limitation led to the proposal of an approach that assumes a fault is composed of a distribution of asperities and barriers on all scales. Thus, an earthquake on a fault is treated as a failure of a fractal tree. Work continued on the development of a fractal based model for deformation in the western United States. In order to better understand the distribution of seismicity on the San Andreas Fault system a fractal analog was developed. The fractal concept also provides a means of testing whether clustering in time or space is a scale-invariant process.

Turcotte, Donald L.↗

Block rotations, fault domains and crustal deformation in the western US

The aim of the project was to develop a 3D model of crustal deformation by distributed fault sets and to test the model results in the field. In the first part of the project, Nur's 2D model (1986) was generalized to 3D. In Nur's model the frictional strength of rocks and faults of a domain provides a tight constraint on the amount of rotation that a fault set can undergo during block rotation. Domains of fault sets are commonly found in regions where the deformation is distributed across a region. The interaction of each fault set causes the fault bounded blocks to rotate. The work that has been done towards quantifying the rotation of fault sets in a 3D stress field is briefly summarized. In the second part of the project, field studies were carried out in Israel, Nevada and China. These studies combined both paleomagnetic and structural information necessary to test the block rotation model results. In accordance with the model, field studies demonstrate that faults and attending fault bounded blocks slip and rotate away from the direction of maximum compression when deformation is distributed across fault sets. Slip and rotation of fault sets may continue as long as the earth's crustal strength is not exceeded. More optimally oriented faults must form, for subsequent deformation to occur. Eventually the block rotation mechanism may create a complex pattern of intersecting generations of faults.

Nur, Amos↗

Stress field rotation or block rotation: An example from the Lake Mead fault system

The Coulomb criterion, as applied by Anderson (1951), has been widely used as the basis for inferring paleostresses from in situ fault slip data, assuming that faults are optimally oriented relative to the tectonic stress direction. Consequently if stress direction is fixed during deformation so must be the faults. Freund (1974) has shown that faults, when arranged in sets, must generally rotate as they slip. Nur et al., (1986) showed how sufficiently large rotations require the development of new sets of faults which are more favorably oriented to the principal direction of stress. This leads to the appearance of multiple fault sets in which older faults are offset by younger ones, both having the same sense of slip. Consequently correct paleostress analysis must include the possible effect of fault and material rotation, in addition to stress field rotation. The combined effects of stress field rotation and material rotation were investigated in the Lake Meade Fault System (LMFS) especially in the Hoover Dam area. Fault inversion results imply an apparent 60 degrees clockwise (CW) rotation of the stress field since mid-Miocene time. In contrast structural data from the rest of the Great Basin suggest only a 30 degrees CW stress field rotation. By incorporating paleomagnetic and seismic evidence, the 30 degrees discrepancy can be neatly resolved. Based on paleomagnetic declination anomalies, it is inferred that slip on NW trending right lateral faults caused a local 30 degrees counter-clockwise (CCW) rotation of blocks and faults in the Lake Mead area. Consequently the inferred 60 degrees CW rotation of the stress field in the LMFS consists of an actual 30 degrees CW rotation of the stress field (as for the entire Great Basin) plus a local 30 degrees CCW material rotation of the LMFS fault blocks.

Ron, Hagai↗

Software Evolution and the Fault Process

In developing a software system, we would like to estimate the way in which the fault content changes during its development, as well determine the locations having the highest concentration of faults. In the phases prior to test, however, there may be very little direct information regarding the number and location of faults. This lack of direct information requires developing a fault surrogate from which the number of faults and their location can be estimated. We develop a fault surrogate based on changes in the fault index, a synthetic measure which has been successfully used as a fault surrogate in previous work. We show that changes in the fault index can be used to estimate the rates at which faults are inserted into a system between successive revisions. We can then continuously monitor the total number of faults inserted into a system, the residual fault content, and identify those portions of a system requiring the application of additional fault detection and removal resources.

Nikora, Allen P.↗

Fault Management Techniques in Human Spaceflight Operations

This paper discusses human spaceflight fault management operations. Fault detection and response capabilities available in current US human spaceflight programs Space Shuttle and International Space Station are described while emphasizing system design impacts on operational techniques and constraints. Preflight and inflight processes along with products used to anticipate, mitigate and respond to failures are introduced. Examples of operational products used to support failure responses are presented. Possible improvements in the state of the art, as well as prioritization and success criteria for their implementation are proposed. This paper describes how the architecture of a command and control system impacts operations in areas such as the required fault response times, automated vs. manual fault responses, use of workarounds, etc. The architecture includes the use of redundancy at the system and software function level, software capabilities, use of intelligent or autonomous systems, number and severity of software defects, etc. This in turn drives which Caution and Warning (C&W) events should be annunciated, C&W event classification, operator display designs, crew training, flight control team training, and procedure development. Other factors impacting operations are the complexity of a system, skills needed to understand and operate a system, and the use of commonality vs. optimized solutions for software and responses. Fault detection, annunciation, safing responses, and recovery capabilities are explored using real examples to uncover underlying philosophies and constraints. These factors directly impact operations in that the crew and flight control team need to understand what happened, why it happened, what the system is doing, and what, if any, corrective actions they need to perform. If a fault results in multiple C&W events, or if several faults occur simultaneously, the root cause(s) of the fault(s), as well as their vehicle-wide impacts, must be determined in order to maintain situational awareness. This allows both automated and manual recovery operations to focus on the real cause of the fault(s). An appropriate balance must be struck between correcting the root cause failure and addressing the impacts of that fault on other vehicle components. Lastly, this paper presents a strategy for using lessons learned to improve the software, displays, and procedures in addition to determining what is a candidate for automation. Enabling technologies and techniques are identified to promote system evolution from one that requires manual fault responses to one that uses automation and autonomy where they are most effective. These considerations include the value in correcting software defects in a timely manner, automation of repetitive tasks, making time critical responses autonomous, etc. The paper recommends the appropriate use of intelligent systems to determine the root causes of faults and correctly identify separate unrelated faults.

O'Hagan, Brian↗

Identifiability of Additive Actuator and Sensor Faults by State Augmentation

A class of fault detection and identification (FDI) methods for bias-type actuator and sensor faults is explored in detail from the point of view of fault identifiability. The methods use state augmentation along with banks of Kalman-Bucy filters for fault detection, fault pattern determination, and fault value estimation. A complete characterization of conditions for identifiability of bias-type actuator faults, sensor faults, and simultaneous actuator and sensor faults is presented. It is shown that FDI of simultaneous actuator and sensor faults is not possible using these methods when all sensors have unknown biases. The fault identifiability conditions are demonstrated via numerical examples. The analytical and numerical results indicate that caution must be exercised to ensure fault identifiability for different fault patterns when using such methods.

Joshi, Suresh↗

Identifiability of Additive, Time-Varying Actuator and Sensor Faults by State Augmentation

Recent work has provided a set of necessary and sucient conditions for identifiability of additive step faults (e.g., lock-in-place actuator faults, constant bias in the sensors) using state augmentation. This paper extends these results to an important class of faults which may affect linear, time-invariant systems. In particular, the faults under consideration are those which vary with time and affect the system dynamics additively. Such faults may manifest themselves in aircraft as, for example, control surface oscillations, control surface runaway, and sensor drift. The set of necessary and sucient conditions presented in this paper are general, and apply when a class of time-varying faults affects arbitrary combinations of actuators and sensors. The results in the main theorems are illustrated by two case studies, which provide some insight into how the conditions may be used to check the theoretical identifiability of fault configurations of interest for a given system. It is shown that while state augmentation can be used to identify certain fault configurations, other fault configurations are theoretically impossible to identify using state augmentation, giving practitioners valuable insight into such situations. That is, the limitations of state augmentation for a given system and configuration of faults are made explicit. Another limitation of model-based methods is that there can be large numbers of fault configurations, thus making identification of all possible configurations impractical. However, the theoretical identifiability of known, credible fault configurations can be tested using the theorems presented in this paper, which can then assist the efforts of fault identification practitioners.

Upchurch, Jason M.↗

Faults on Skylab imagery of the Salton Trough area, Southern California

The author has identified the following significant results. Large segments of the major high angle faults in the Salton Trough area are readily identifiable in Skylab images. Along active faults, distinctive topographic features such as scarps and offset drainage, and vegetation differences due to ground water blockage in alluvium are visible. Other fault-controlled features along inactive as well as active faults visible in Skylab photography include straight mountain fronts, linear valleys, and lithologic differences producing contrasting tone, color or texture. A northwestern extension of a fault in the San Andreas set, is postulated by the regional alignment of possible fault-controlled features. The suspected fault is covered by Holocene deposits, principally windblown sand. A northwest trending tonal change in cultivated fields across Mexicali Valley is visible on Skylab photos. Surface evidence for faulting was not observed; however, the linear may be caused by differences in soil conditions along an extension of a segment of the San Jacinto fault zone. No evidence of faulting could be found along linears which appear as possible extensions of the Substation and Victory Pass faults, demonstrating that the interpretation of linears as faults in small scale photography must be corroborated by field investigations.

Merifield, P. M.↗

Measurement and application of fault latency

The time interval between the occurrence of a fault and the detection of the error caused by the fault is divided by the generation of that error into two parts: fault latency and error latency. Since the moment of error generation is not directly observable, all related works in the literature have dealt with only the sum of fault and error latencies, thereby making the analysis of their separate effects impossible. To remedy this deficiency, (1) a new methodology for indirectly measuring fault latency is presented; the distribution of fault latency is derived from the methodology; and (3) the knowledge of fault latency is applied to the analysis of two important examples. The proposed methodology has been implemented for measuring fault latency in the Fault-Tolerant Multiprocessor (FTMP) at the NASA Airlab. The experimental results show wide variations in the mean fault latencies of different function circuits within FTMP. Also, the measured distributions of fault latency are shown to have monotone hazard rates. Consequently, Gamma and Weibull distributions are selected for the least-squares fit as the distribution of fault latency.

Shin, K. G.↗