Search NASASearch

SEARCH · Search NASA

Results for “multiple faults”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 109 records · Page 6

The UCLA Design Diversity Experiment (DEDIX) system: A distributed testbed for multiple-version software

To establish a long-term research facility for experimental investigations of design diversity as a means of achieving fault-tolerant systems, a distributed testbed for multiple-version software was designed. It is part of a local network, which utilizes the Locus distributed operating system to operate a set of 20 VAX 11/750 computers. It is used in experiments to measure the efficacy of design diversity and to investigate reliability increases under large-scale, controlled experimental conditions.

Avizienis, A.

Operations management system advanced automation: Fault detection isolation and recovery prototyping

The purpose of this project is to address the global fault detection, isolation and recovery (FDIR) requirements for Operation's Management System (OMS) automation within the Space Station Freedom program. This shall be accomplished by developing a selected FDIR prototype for the Space Station Freedom distributed processing systems. The prototype shall be based on advanced automation methodologies in addition to traditional software methods to meet the requirements for automation. A secondary objective is to expand the scope of the prototyping to encompass multiple aspects of station-wide fault management (SWFM) as discussed in OMS requirements documentation.

Hanson, Matt

Partitioning in Avionics Architectures: Requirements, Mechanisms, and Assurance

Automated aircraft control has traditionally been divided into distinct "functions" that are implemented separately (e.g., autopilot, autothrottle, flight management); each function has its own fault-tolerant computer system, and dependencies among different functions are generally limited to the exchange of sensor and control data. A by-product of this "federated" architecture is that faults are strongly contained within the computer system of the function where they occur and cannot readily propagate to affect the operation of other functions. More modern avionics architectures contemplate supporting multiple functions on a single, shared, fault-tolerant computer system where natural fault containment boundaries are less sharply defined. Partitioning uses appropriate hardware and software mechanisms to restore strong fault containment to such integrated architectures. This report examines the requirements for partitioning, mechanisms for their realization, and issues in providing assurance for partitioning. Because partitioning shares some concerns with computer security, security models are reviewed and compared with the concerns of partitioning.

Rushby, John

A strainmeter array as the fulcrum of novel observatory sites along the Alto Tiberina Near Fault Observatory

Fault slip is a complex natural phenomenon involving multiple spatiotemporal scales from seconds to days to weeks. To understand the physical and chemical processes responsible for the full fault slip spectrum, a multidisciplinary approach is highly recommended. The Near Fault Observatories (NFOs) aim at providing high-precision and spatiotemporally dense multidisciplinary near-fault data, enabling the generation of new original observations and innovative scientific products. The Alto Tiberina Near Fault Observatory is a permanent monitoring infrastructure established around the Alto Tiberina fault (ATF), a 60 km long low-angle normal fault (mean dip 20°), located along a sector of the Northern Apennines (central Italy) undergoing an extension at a rate of about 3 mm yr –1 . The presence of repeating earthquakes on the ATF and a steep gradient in crustal velocities measured across the ATF by GNSS stations suggest large and deep (5–12 km) portions of the ATF undergoing aseismic creep. Both laboratory and theoretical studies indicate that any given patch of a fault can creep, nucleate slow earthquakes, and host large earthquakes, as also documented in nature for certain ruptures (e.g., Iquique in 2014, Tōhoku in 2011, and Parkfield in 2004). Nonetheless, how a fault patch switches from one mode of slip to another, as well as the interaction between creep, slow slip, and regular earthquakes, is still poorly documented by near-field observation. With the strainmeter array along the Alto Tiberina fault system (STAR) project, we build a series of six geophysical observatory sites consisting of 80–160 m deep vertical boreholes instrumented with strainmeters and seismometers as well as meteorological and GNSS antennas and additional seismometers at the surface. By covering the portions of the ATF that exhibits repeated earthquakes at shallow depth (above 4 km) with these new observatory sites, we aim to collect unique open-access data to answer fundamental questions about the relationship between creep, slow slip, dynamic earthquake rupture, and tectonic faulting.

58 GEOSCIENCES

R2U2 in Space: System and Software Health Management for Small Satellites

In order for small but complex systems like rovers, SmallSats, or Unmanned Aircraft (UAS) to operate autonomously, they must have a real-time solution for assessing their own system health. System and Software Health Management (SHM) enables better detection of faulty sensors and software problems, and enables better fault management including mitigation of unpredicted fault scenarios in the absence of a human on-board. In recent work, we have developed a Responsive, Realizable, Unobtrusive Unit (R2U2) for on-board SHM of autonomous UAS and demonstrated its ability to detect faults during flight time. These faults, from sensor failures, to software problems, to malicious security attacks, can present as transient temporal faults that even humans are challenged to find. An R2U2 congfiuration is a modular combination of multiple types of temporal logic runtime observers with fault-specic Bayesian Nets and sensor filters. R2U2 reasons about both on-board hardware and software components; R2U2 itself can be instantiated as an independent FPGA (Field-Programmable Gate Array)-based conguration or as a software component running independently from other software on-board. Small satellites, such as CubeSats, also require on-board SHM and failure mitigation, as limited telemetry bandwidth does not allow the transmission of the entire system state for ground-based health management. However, the autonomous operation of satellites brings a set of challenges different from UAS, including the effects of radiation on non-rad-hard, low-cost components, and the harsher environment of space. We surmise that a new extension of R2U2 could be adapted to help better detect, for example, radiation errors in cheaper COTS (Commercial Off the Shelf) (not rad-hard) components often used in small space systems. Since small satellites often operate in coordination, we will also examine new ways of distributed monitoring of their communication and cooperation and real-time detection of off-nominal situations utilizing multiple satellites. This talk will discuss preliminary work and ideas for building on terrestrial success of system and software health management for the harsher, and differently challenging, environment of space.

Runtime Verification & Validation

An Indirect Adaptive Control Scheme in the Presence of Actuator and Sensor Failures

The problem of controlling a system in the presence of unknown actuator and sensor faults is addressed. The system is assumed to have groups of actuators, and groups of sensors, with each group consisting of multiple redundant similar actuators or sensors. The types of actuator faults considered consist of unknown actuators stuck in unknown positions, as well as reduced actuator effectiveness. The sensor faults considered include unknown biases and outages. The approach employed for fault detection and estimation consists of a bank of Kalman filters based on multiple models, and subsequent control reconfiguration to mitigate the effect of biases caused by failed components as well as to obtain stability and satisfactory performance using the remaining actuators and sensors. Conditions for fault identifiability are presented, and the adaptive scheme is applied to an aircraft flight control example in the presence of actuator failures. Simulation results demonstrate that the method can rapidly and accurately detect faults and estimate the fault values, thus enabling safe operation and acceptable performance in spite of failures.

Sun, Joy Z.

A Systematic Framework for Tuning Open-Source Multifunctional IBR Models To Emulate OEM Black-Box Fault Dynamics

This paper presents a systematic framework to tune a generic IBR EMT model to match with an OEM provided balckbox inverter model based on the fault current responses. The key learnings and findings are summarized as follows: The tunable key parameters include inner control loops and current limiters to align the fault current magnitude, sequence content, and phase trajectories with the OEM models across diverse fault type and locations. The tuned model's fidelity is validated through comparative analysis with an OEM blackbox model, assessing both the fault current response and the responses of multiple relay elements. The results demonstrate the tuned generic model can trigger relay decision logic that is identical or near identical to that of the OEM model, thus generating very good match model for fault studies.

24 POWER TRANSMISSION AND DISTRIBUTION

Gyro-based Maximum-Likelihood Thruster Fault Detection and Identification

When building smaller, less expensive spacecraft, there is a need for intelligent fault tolerance vs. increased hardware redundancy. If fault tolerance can be achieved using existing navigation sensors, cost and vehicle complexity can be reduced. A maximum likelihood-based approach to thruster fault detection and identification (FDI) for spacecraft is developed here and applied in simulation to the X-38 space vehicle. The system uses only gyro signals to detect and identify hard, abrupt, single and multiple jet on- and off-failures. Faults are detected within one second and identified within one to five accords,

Wilson, Edward

A unified method for analyzing mission reliability for fault tolerant computer systems.

For fault-tolerant computer systems consisting of multiple classes of modules, a unified method for analyzing mission reliability is proposed and evaluated. The analysis proceeds by generalizing the notions of standby and N modular redundancy into a concept called hybrid-degraded redundancy. The probabilistic evaluation of the unified redundancy concept is then developed to yield, for a given modular class, the joint distribution of success and the number of nonfailed modules from that class, at special times. With this information, a Markov chain analysis gives the reliability of an entire sequence of phases (mission profile).

Bricker, J. L.

Tradeoffs in implementing primary-backup protocols

One way to implement a fault-tolerant service is by using multiple servers that fail independently. The state of the service is replicated and distributed among these servers, and updates are coordinated so that even when a subset of the servers fail, the service remains available. A common approach to structuring such replicated services is to designate one server as the primary and all the others as backups. Clients make requests by sending messages only to the primary. If the primary fails, then a failover occurs and one of the backups takes over. This service architecture is commonly called the primary-backup or the primary-copy approach. In most such primary-backup protocols, when the primary receives a client request, it informs the backups about the request, and then responds to the client. Informally, this primary-backup protocol is non-blocking if the primary does not wait for an acknowledgement from the backups before it sends the response; otherwise, it is blocking. Most of the existing protocols are blocking as non-blocking protocols cannot be constructed for some kinds of failures. However, it is shown that non-blocking protocols can be constructed for most of the process and communication failures that are expected to occur in the primary-backup systems of the future. Since non-blocking protocols can theoretically achieve the smallest possible response time, this paper analyzes these protocols under various system parameters. Two kinds of non-blocking protocols are analyzed: one in which the processes use point-to-point communication to exchange messages, and the other in which processes use hardware broadcasts.

Budhiraja, Navin

From Informal Safety-Critical Requirements to Property-Driven Formal Validation

Most of the efforts in formal methods have historically been devoted to comparing a design against a set of requirements. The validation of the requirements themselves, however, has often been disregarded, and it can be considered a largely open problem, which poses several challenges. The first challenge is given by the fact that requirements are often written in natural language, and may thus contain a high degree of ambiguity. Despite the progresses in Natural Language Processing techniques, the task of understanding a set of requirements cannot be automatized, and must be carried out by domain experts, who are typically not familiar with formal languages. Furthermore, in order to retain a direct connection with the informal requirements, the formalization cannot follow standard model-based approaches. The second challenge lies in the formal validation of requirements. On one hand, it is not even clear which are the correctness criteria or the high-level properties that the requirements must fulfill. On the other hand, the expressivity of the language used in the formalization may go beyond the theoretical and/or practical capacity of state-of-the-art formal verification. In order to solve these issues, we propose a new methodology that comprises of a chain of steps, each supported by a specific tool. The main steps are the following. First, the informal requirements are split into basic fragments, which are classified into categories, and dependency and generalization relationships among them are identified. Second, the fragments are modeled using a visual language such as UML. The UML diagrams are both syntactically restricted (in order to guarantee a formal semantics), and enriched with a highly controlled natural language (to allow for modeling static and temporal constraints). Third, an automatic formal analysis phase iterates over the modeled requirements, by combining several, complementary techniques: checking consistency; verifying whether the requirements entail some desirable properties; verify whether the requirements are consistent with selected scenarios; diagnosing inconsistencies by identifying inconsistent cores; identifying vacuous requirements; constructing multiple explanations by enabling the fault-tree analysis related to particular fault models; verifying whether the specification is realizable.

Cimatti, Alessandro

Aircraft Loss-of-Control Accident Prevention: Switching Control of the GTM Aircraft with Elevator Jam Failures

Switching control, servomechanism, and H2 control theory are used to provide a practical and easy-to-implement solution for the actuator jam problem. A jammed actuator not only causes a reduction of control authority, but also creates a persistent disturbance with uncertain amplitude. The longitudinal dynamics model of the NASA GTM UAV is employed to demonstrate that a single fixed reconfigured controller design based on the proposed approach is capable of accommodating an elevator jam failure with arbitrary jam position as long as the thrust control has enough control authority. This paper is a first step towards solving a more comprehensive in-flight loss-of-control accident prevention problem that involves multiple actuator failures, structure damages, unanticipated faults, and nonlinear upset regime recovery, etc.

Chang, Bor-Chin

Intelligent Wireless Sensor Networks for System Health Monitoring

Wireless sensor networks (WSN) based on the IEEE 802.15.4 Personal Area Network (PAN) standard are finding increasing use in the home automation and emerging smart energy markets. The network and application layers, based on the ZigBee 2007 Standard, provide a convenient framework for component-based software that supports customer solutions from multiple vendors. WSNs provide the inherent fault tolerance required for aerospace applications. The Discovery and Systems Health Group at NASA Ames Research Center has been developing WSN technology for use aboard aircraft and spacecraft for System Health Monitoring of structures and life support systems using funding from the NASA Engineering and Safety Center and Exploration Technology Development and Demonstration Program. This technology provides key advantages for low-power, low-cost ancillary sensing systems particularly across pressure interfaces and in areas where it is difficult to run wires. Intelligence for sensor networks could be defined as the capability of forming dynamic sensor networks, allowing high-level application software to identify and address any sensor that joined the network without the use of any centralized database defining the sensors characteristics. The IEEE 1451 Standard defines methods for the management of intelligent sensor systems and the IEEE 1451.4 section defines Transducer Electronic Datasheets (TEDS), which contain key information regarding the sensor characteristics such as name, description, serial number, calibration information and user information such as location within a vehicle. By locating the TEDS information on the wireless sensor itself and enabling access to this information base from the application software, the application can identify the sensor unambiguously and interpret and present the sensor data stream without reference to any other information. The application software is able to read the status of each sensor module, responding in real-time to changes of PAN configuration, providing the appropriate response for maintaining overall sensor system function, even when sensor modules fail or the WSN is reconfigured. The session will present the architecture and technical feasibility of creating fault-tolerant WSNs for aerospace applications based on our application of the technology to a Structural Health Monitoring testbed. The interim results of WSN development and testing including our software architecture for intelligent sensor management will be discussed in the context of the specific tradeoffs required for effective use. Initial certification measurement techniques and test results gauging WSN susceptibility to Radio Frequency interference are introduced as key challenges for technology adoption. A candidate Developmental and Flight Instrumentation implementation using intelligent sensor networks for wind tunnel and flight tests is developed as a guide to understanding key aspects of the aerospace vehicle design, test and operations life cycle.

networks

Evaluation of an Enhanced Bank of Kalman Filters for In-Flight Aircraft Engine Sensor Fault Diagnostics

In this paper, an approach for in-flight fault detection and isolation (FDI) of aircraft engine sensors based on a bank of Kalman filters is developed. This approach utilizes multiple Kalman filters, each of which is designed based on a specific fault hypothesis. When the propulsion system experiences a fault, only one Kalman filter with the correct hypothesis is able to maintain the nominal estimation performance. Based on this knowledge, the isolation of faults is achieved. Since the propulsion system may experience component and actuator faults as well, a sensor FDI system must be robust in terms of avoiding misclassifications of any anomalies. The proposed approach utilizes a bank of (m+1) Kalman filters where m is the number of sensors being monitored. One Kalman filter is used for the detection of component and actuator faults while each of the other m filters detects a fault in a specific sensor. With this setup, the overall robustness of the sensor FDI system to anomalies is enhanced. Moreover, numerous component fault events can be accounted for by the FDI system. The sensor FDI system is applied to a commercial aircraft engine simulation, and its performance is evaluated at multiple power settings at a cruise operating point using various fault scenarios.

Kobayashi, Takahisa

How earthquakes organize stress

Stress is not uniform in the Earth. Therefore, we must use natural experiments to measure the distribution of stresses and related quantities, rather than single values. For instance, dynamic triggering shows that faults are uniformly distributed over their loading cycles in Southern California. The probability that a fault ruptures across a barrier measures the in situ energy distribution. Fault roughness reflects the distribution of strength. These natural experiments produce observable distributions that are surprisingly consistent and suggest some degree of self-organization in the Earth’s crust. Once established, the functional form of the distributions can be used to track changes in response to earthquakes as well as to distinguish fundamentally different fault systems. Transient fault locking before stress release in laboratory experiments can be interpreted as a consequence of self-organization of fault stress. The robust self-organization of multiple variables in earthquake systems suggests that the most consequential mechanical outcome of earthquakes may be the redistribution of stress and the strain energy associated with it. The low friction on a fault during seismic slip as inferred by temperature measurements of the Tohoku earthquake is consistent with dissipation playing a secondary role to this redistribution process. Through stress redistribution and interaction, subduction zone faults tend to synchronize, perhaps due to their geometric simplicity, while the continental system of Southern California cannot synchronize, perhaps due to the complexity of the fault network. Earthquakes organize stress in the crust and produce a suite of well-defined, consistent distributions.

earthquakes

On reliability modeling and analysis of ultrareliable fault-tolerant digital systems.

The processes of protective redundancy, namely, standby replacement (SR) redundancy and hybrid redundancy (a combination of SR and multiple-line voting redundancy), find application in the architecture of fault-tolerant digital computers and enable them to be ultrareliable and self-repairing. The claims to ultrareliability lead to the challenge of quantitatively evaluating and assigning a value to the probability of survival as a function of the mission durations intended. This note presents various mathematical models, and derives and displays quantitative evaluations of system reliability as a function of various mission parameters of interest to the system designer.

Mathur, F. P.

Active multi-mode data analysis to improve fault diagnosis in AHUs

Faults in heating, ventilation and air conditioning systems can lead to increased energy consumption, occupant comfort issues, and reduced equipment lifetime. Commercial fault detection and diagnosis (FDD) tools has been increasingly deployed in U.S. commercial buildings. While they are helping to achieve energy efficiency and operational reliability, there remain gaps in their fault diagnostic capabilities. The diagnostic results often contain multiple distinct candidate root causes (CRCs) or offer no insight into CRCs. This study developed a novel active rule-based multi-mode data analysis method to enhance diagnostic resolution by applying proven rule sets and additional new rules to data from multiple known operational modes. The proposed method was demonstrated using enhanced air handling unit performance assessment rule sets and validated with the simulated data of two air handling units. New metrics, namely, reduced number of CRCs and improvement ratio, were developed to quantify the improvement of fault diagnostic resolution. The validation results showed that the proposed method effectively reduced the number of CRCs in contrast to analyzing data solely for a single mode of operation. It achieved a median improvement ratio of 80% in 19 test cases.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI

Automated Monitoring with a BCP Fault-Decision Test

The Bayesian conditional probability (BCP) technique is a statistical fault-decision technique that is suitable as the mathematical basis of the fault-manager module in the automated-monitoring system and method described in the immediately preceding article. Within the automated-monitoring system, the fault-manager module operates in conjunction with the fault-detector module, which can be based on any one of several fault-detection techniques; examples include a threshold-limit-comparison technique or the BSP or SPRT technique mentioned in the preceding article. The present BCP technique is used to evaluate a series of one or more fault-detection events for the purpose of filtering out occasional false alarms produced by many types of statistical fault-detection procedures. The BCP technique increases the probability that an automated monitoring system produces a correct decision regarding the presence or absence of a fault. Because occasional false alarms are an inevitable consequence of the SPRT, BSP, or any other statistically based fault-detection test, there is a need for a logical procedure to distinguish between true and false alarms. Heretofore, it has been common practice to make a fault decision on an ad hoc basis for example by following a multiple-observation voting strategy in which a signal is declared to be indicative of a fault if m of the last n observations produced a fault-detection alarm. The BCP technique was developed to obtain results more reliable than those afforded by a voting strategy. The BCP technique involves a test in which one applies Bayesian inference techniques to a series of one or more single-observation alarms produced by a fault-detection test. One considers the last n decisions generated by a fault-detection test in order to evaluate the conditional probability that a failure is indicated (see figure). Each new decision reached by a fault-detection test is treated as a new piece of evidence about the state of the monitored asset, and the conditional probability of failure for the system is updated on the basis of this new evidence. The conditional probability of failure is compared with a predefined limit. For a probability below the limit, the asset is declared to be healthy. For a probability above the limit, the asset is declared to be faulty.

Bickford, Randall L.