Search NASA⌕ Search

SEARCH · Search NASA

Results for “SYSTEM FAILURE”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 181 records · Page 10

An accumulation method for early fault warning and its application to wind turbine systems

Unexpected failures in engineering systems lead to expensive maintenance actions and should be avoided if at all possible. This is particularly true for wind turbine systems for which unexpected failures not only demand costly repairs but also cause long downtime. Motivated by this need, we present an accumulation method for fault early warning and failure anticipation. Here, our research shows that one critical element allowing the ability of early warning is to accumulate the small-magnitude symptoms resulting from gradual changes in an engineering system like wind turbines. Our idea is inspired by the classical cumulative sum method, or CUSUM, but we have to redesign the accumulation mechanism for tackling unique challenges in wind turbine data. The new accumulation method is applied to two real wind turbine datasets, one with gearbox failures and the other with generator failures, and demonstrates superior performance as compared with CUSUM.

17 WIND ENERGY↗

Direct Adaptive Control of Systems with Actuator Failures: State of the Art and Continuing Challenges

In this paper, the problem of controlling systems with failures and faults is introduced, and an overview of recent work on direct adaptive control for compensation of uncertain actuator failures is presented. Actuator failures may be characterized by some unknown system inputs being stuck at some unknown (fixed or varying) values at unknown time instants, that cannot be influenced by the control signals. The key task of adaptive compensation is to design the control signals in such a manner that the remaining actuators can automatically and seamlessly take over for the failed ones, and achieve desired stability and asymptotic tracking. A certain degree of redundancy is necessary to accomplish failure compensation. The objective of adaptive control design is to effectively use the available actuation redundancy to handle failures without the knowledge of the failure patterns, parameters, and time of occurrence. This is a challenging problem because failures introduce large uncertainties in the dynamic structure of the system, in addition to parametric uncertainties and unknown disturbances. The paper addresses some theoretical issues in adaptive actuator failure compensation: actuator failure modeling, redundant actuation requirements, plant-model matching, error system dynamics, adaptation laws, and stability, tracking, and performance analysis. Adaptive control designs can be shown to effectively handle uncertain actuator failures without explicit failure detection. Some open technical challenges and research problems in this important research area are discussed.

Tao, Gang↗

Error latency measurements in symbolic architectures

Error latency, the time that elapses between the occurrence of an error and its detection, has a significant effect on reliability. In computer systems, failure rates can be elevated during a burst of system activity due to increased detection of latent errors. A hybrid monitoring environment is developed to measure the error latency distribution of errors occurring in main memory. The objective of this study is to develop a methodology for gauging the dependability of individual data categories within a real-time application. The hybrid monitoring technique is novel in that it selects and categorizes a specific subset of the available blocks of memory to monitor. The precise times of reads and writes are collected, so no actual faults need be injected. Unlike previous monitoring studies that rely on a periodic sampling approach or on statistical approximation, this new approach permits continuous monitoring of referencing activity and precise measurement of error latency.

Young, L. T.↗

Analysis of Aircraft Clusters to Measure Sector-Independent Airspace Congestion

The Distributed Air/Ground Traffic Management (DAG-TM) concept of operations* permits appropriately equipped aircraft to conduct Free Maneuvering operations. These independent aircraft have the freedom to optimize their trajectories in real time according to user preferences; however, they also take on the responsibility to separate themselves from other aircraft while conforming to any local Traffic Flow Management (TFM) constraints imposed by the air traffic service provider (ATSP). Examples of local-TFM constraints include temporal constraints such as a required time of arrival (RTA), as well as spatial constraints such as regions of convective weather, special use airspace, and congested airspace. Under current operations, congested airspace typically refers to a sector(s) that cannot accept additional aircraft due to controller workload limitations; hence Dynamic Density (a metric that is indicative of controller workload) can be used to quantify airspace congestion. However, for Free Maneuvering operations under DAG-TM, an additional metric is needed to quantify the airspace congestion problem from the perspective of independent aircraft. Such a metric would enable the ATSP to prevent independent aircraft from entering any local areas of congestion in which the flight deck based systems and procedures may not be able to ensure separation. This new metric, called Gaggle Density, offers the ATSP a mode of control to regulate normal operations and to ensure safety and stability during rare-normal or off-normal situations (e.g., system failures). It may be difficult to certify Free Maneuvering systems for unrestricted operations, but it may be easier to certify systems and procedures for specified levels of Gaggle Density that could be monitored by the ATSP, and maintained through relatively minor flow-rate (RTA type) restrictions. Since flight deck based separation assurance is airspace independent, the challenge is to measure congestion independent of sector boundaries. Figure 1 , reproduced from Ref. 1, depicts an example traffic situation. When the situation is analyzed by sector boundaries (left side of figure), a Dynamic Density metric would identify excessive congestion in the central sector. When the same traffic situation is analyzed independent of sector boundaries (right side of figure), a Gaggle Density metric would identify congestion in two dynamically defined areas covering portions of several sectors. The first step towards measuring airspace-independent congestion is to identify aircraft clusters, i.e., groups of closely spaced aircraft. The objective of this work is to develop techniques to detect and classify clusters of aircraft.

Bilimoria, Karl D.↗

Lox/Gox related failures during Space Shuttle Main Engine development

Specific rocket engine hardware and test facility system failures are described which were caused by high pressure liquid and/or gaseous oxygen reactions. The failures were encountered during the development and testing of the space shuttle main engine. Failure mechanisms are discussed as well as corrective actions taken to prevent or reduce the potential of future failures.

Cataldo, C. E.↗

Update on N2O4 Molecular Sieving with 3A Material at NASA/KSC

During its operational life, the Shuttle Program has experienced numerous failures in the Nitrogen Tetroxide (N2O4) portion of Reaction Control System (RCS), many of which were attributed to iron-nitrate contamination. Since the mid-1980's, N2O4 has been processed through a molecular sieve at the N2O4 manufacturer's facility which results in an iron content typically less than 0.5 parts-per-million-by-weight (ppmw). In February 1995, a Tiger Team was formed to attempt to resolve the iron nitrate problem. Eighteen specific actions were recommended as possibly reducing system failures. Those recommended actions include additional N2O4 molecular sieving at the Shuttle launch site. Testing at NASA White Sands Test Facility (WSTF) determined an alternative molecular sieve material could also reduce the water-equivalent content (free water and HNO3) and thereby further reduce the natural production of iron nitrate in N2O4 while stored in iron-alloy storage tanks. Since April '96, NASA Kennedy Space Center (KSC) has been processing N2O4 through the alternative molecular sieve material prior to delivery to Shuttle launch pad N2O4 storage tanks. A new, much larger capacity molecular sieve unit has also been used. This paper will evaluate the effectiveness of N2O4 molecular sieving on a large-scale basis and attempt to determine if the resultant lower-iron and lower-water content N2O4 maintains this new purity level in pad storage tanks and shuttle flight systems.

Davis, Chuck↗

Reducing the cognitive workload: Trouble managing power systems

The complexity of space-based systems makes monitoring them and diagnosing their faults taxing for human beings. Mission control operators are well-trained experts but they can not afford to have their attention diverted by extraneous information. During normal operating conditions monitoring the status of the components of a complex system alone is a big task. When a problem arises, immediate attention and quick resolution is mandatory. To aid humans in these endeavors we have developed an automated advisory system. Our advisory expert system, Trouble, incorporates the knowledge of the power system designers for Space Station Freedom. Trouble is designed to be a ground-based advisor for the mission controllers in the Control Center Complex at Johnson Space Center (JSC). It has been developed at NASA Lewis Research Center (LeRC) and tested in conjunction with prototype flight hardware contained in the Power Management and Distribution testbed and the Engineering Support Center, ESC, at LeRC. Our work will culminate with the adoption of these techniques by the mission controllers at JSC. This paper elucidates how we have captured power system failure knowledge, how we have built and tested our expert system, and what we believe are its potential uses.

Manner, David B.↗

Reducing the cognitive workload - Trouble managing power systems

The complexity of space-based systems makes monitoring them and diagnosing their faults taxing for human beings. When a problem arises, immediate attention and quick resolution is mandatory. To aid humans in these endeavors we have developed an automated advisory system. Our advisory expert system, Trouble, incorporates the knowledge of the power system designers for Space Station Freedom. Trouble is designed to be a ground-based advisor for the mission controllers in the Control Center Complex at Johnson Space Center (JSC). It has been developed at NASA Lewis Research Center (LeRC) and tested in conjunction with prototype flight hardware contained in the Power Management and Distribution testbed and the Engineering Support Center, ESC, at LeRC. Our work will culminate with the adoption of these techniques by the mission controllers at JSC. This paper elucidates how we have captured power system failure knowledge, how we have built and tested our expert system, and what we believe its potential uses are.

Manner, David B.↗

Emergency Flight Control of a Twin-Jet Commercial Aircraft using Manual Throttle Manipulation

The Department of Homeland Security (DHS) created the PCAR (Propulsion-Controlled Aircraft Recovery) project in 2005 to mitigate the ManPADS (man-portable air defense systems) threat to the commercial aircraft fleet with near-term, low-cost proven technology. Such an attack could potentially cause a major FCS (flight control system) malfunction or other critical system failure onboard the aircraft, despite the extreme reliability of current systems. For the situations in which nominal flight controls are lost or degraded, engine thrust may be the only remaining means for emergency flight control [ref 1]. A computer-controlled thrust system, known as propulsion-controlled aircraft (PCA), was developed in the mid 1990s with NASA, McDonnell Douglas and Honeywell. PCA's major accomplishment was a demonstration of an automatic landing capability using only engine thrust [ref 11. Despite these promising results, no production aircraft have been equipped with a PCA system, due primarily to the modifications required for implementation. A minimally invasive option is TOC (throttles-only control), which uses the same control principles as PCA, but requires absolutely no hardware, software or other aircraft modifications. TOC is pure piloting technique, and has historically been utilized several times by flight crews, both military and civilian, in emergency situations stemming from a loss of conventional control. Since the 1990s, engineers at NASA Dryden Flight Research Center (DFRC) have studied TOC, in both simulation and flight, for emergency flight control with test pilots in numerous configurations. In general, it was shown that TOC was effective on certain aircraft for making a survivable landing. DHS sponsored both NASA Dryden Flight Research Center (Edwards, CA) and United Airlines (Denver, Colorado) to conduct a flight and simulation study of the TOC characteristics of a twin-jet commercial transport, and assess the ability of a crew to control an aircraft down to a survivable runway landing using TOC. The PCAR project objective was a set of pilot procedures for operation of a specific aircraft without hydraulics that (a) have been validated in both simulation and flight by relevant personnel, and (b) mesh well with existing commercial operations, maintenance, and training at a minimum cost. As a result of this study, a procedure has been developed to assist a crew in making a survivable landing using TOC. In a simulation environment, line pilots with little or no previous TOC experience performed survivable runway landings after a few practice TOC approaches. In-flight evaluations put line pilots in a simulated emergency situation where TOC was used to recover the aircraft, maneuver to a landing site, and perform an approach down to 200 feet AGL. The results of this research, including pilot observations, procedure comments, recommendations, future work and lessons learned, will he discussed. Flight data and video footage of TOC approaches may also be shown.

Cole, Jennifer H.↗

Designing Fault-Injection Experiments for the Reliability of Embedded Systems

This paper considers the long-standing problem of conducting fault-injections experiments to establish the ultra-reliability of embedded systems. There have been extensive efforts in fault injection, and this paper offers a partial summary of the efforts, but these previous efforts have focused on realism and efficiency. Fault injections have been used to examine diagnostics and to test algorithms, but the literature does not contain any framework that says how to conduct fault-injection experiments to establish ultra-reliability. A solution to this problem integrates field-data, arguments-from-design, and fault-injection into a seamless whole. The solution in this paper is to derive a model reduction theorem for a class of semi-Markov models suitable for describing ultra-reliable embedded systems. The derivation shows that a tight upper bound on the probability of system failure can be obtained using only the means of system-recovery times, thus reducing the experimental effort to estimating a reasonable number of easily-observed parameters. The paper includes an example of a system subject to both permanent and transient faults. There is a discussion of integrating fault-injection with field-data and arguments-from-design.

White, Allan L.↗

Extending the life and recycle capability of earth storable propellant systems.

Rocket propulsion systems for reusable vehicles will be required to operate reliably for a large number of missions with a minimum of maintenance and a fast turnaround. For the space shuttle reaction control system to meet these requirements, current and prior related system failures were examined for their impact on reuse and, where warranted, component design and/or system configuration changes were defined for improving system service life. It was found necessary to change the pressurization component arrangement used on many single-use applications in order to eliminate a prevalent check valve failure mode and to incorporate redundant expulsion capability in propellant tank designs to achieve the necessary system reliability. Material flaws in pressurant and propellant tanks were noted to have a significant effect on tank cycle life. Finally, maintenance considerations dictated a modularized systems approach, allowing the system to be removed from the vehicle for service and repair at a remote site.

Schweickert, T. F.↗

Method and system for detecting a failure or performance degradation in a dynamic system such as a flight vehicle

A method and system for detecting a failure or performance degradation in a dynamic system having sensors for measuring state variables and providing corresponding output signals in response to one or more system input signals are provided. The method includes calculating estimated gains of a filter and selecting an appropriate linear model for processing the output signals based on the input signals. The step of calculating utilizes one or more models of the dynamic system to obtain estimated signals. The method further includes calculating output error residuals based on the output signals and the estimated signals. The method also includes detecting one or more hypothesized failures or performance degradations of a component or subsystem of the dynamic system based on the error residuals. The step of calculating the estimated values is performed optimally with respect to one or more of: noise, uncertainty of parameters of the models and un-modeled dynamics of the dynamic system which may be a flight vehicle or financial market or modeled financial system.

Miller, Robert H.↗

Fault-tolerant system considerations for a redundant strapdown inertial measurement unit

The development and evaluation of a fault-tolerant system for the Redundant Strapdown Inertial Measurement Unit (RSDIMU) being developed and evaluated by the NASA Langley Research Center was continued. The RSDIMU consists of four two-degree-of-freedom gyros and accelerometers mounted on the faces of a semi-octahedron which can be separated into two halves for damage protection. Compensated and uncompensated fault-tolerant system failure decision algorithms were compared. An algorithm to compensate for sensor noise effects in the fault-tolerant system thresholds was evaluated via simulation. The effects of sensor location and magnitude of the vehicle structural modes on system performance were assessed. A threshold generation algorithm, which incorporates noise compensation and filtered parity equation residuals for structural mode compensation, was evaluated. The effects of the fault-tolerant system on navigational accuracy were also considered. A sensor error parametric study was performed in an attempt to improve the soft failure detection capability without obtaining false alarms. Also examined was an FDI system strategy based on the pairwise comparison of sensor measurements. This strategy has the specific advantage of, in many instances, successfully detecting and isolating up to two simultaneously occurring failures.

Motyka, P.↗

Advanced orbit transfer vehicle propulsion system study

A reuseable orbit transfer vehicle concept was defined and subsequent recommendations for the design criteria of an advanced LO2/LH2 engine were presented. The major characteristics of the vehicle preliminary design include a low lift to drag aerocapture capability, main propulsion system failure criteria of fail operational/fail safe, and either two main engines with an attitude control system for backup or three main engines to meet the failure criteria. A maintenance and servicing approach was also established for the advanced vehicle and engine concepts. Design tradeoff study conclusions were based on the consideration of reliability, performance, life cycle costs, and mission flexibility.

Cathcart, J. A.↗

Eucalyptus – An Analysis Suite for Fault Trees with Uncertainty Quantification

Eucalyptus is a novel code developed at Lawrence Livermore National Laboratory to incorporate uncertainty quantification into Fault Tree Analysis (FTA). This tool addresses the challenge of imperfect knowledge in “grey-box” systems by allowing analysts to incorporate and propagate uncertainty from component-level assessments to system-level effects. Eucalyptus facilitates a consistent evaluation of the impact of subject matter expert judgment and knowledge gaps on overall system response by Monte Carlo generation of possible system fault trees, sampling probabilities of the existence of subsystems and components. Here, the code supports the specification of fault trees through text and allows export to various formats, including auto-generated images, easing analysis and reducing errors. It has undergone extensive verification testing, demonstrating its reliability and readiness for deployment, and leverages on-node parallelism for rapid analysis. Example analyses are shown that include the identification of system failure paths and quantification of the value of further information about system components.

Fault Tree Analysis↗

PKI solar thermal plant evaluation at Capitol Concrete Products, Topeka, Kansas

A system feasibility test to determine the technical and operational feasibility of using a solar collector to provide industrial process heat is discussed. The test is of a solar collector system in an industrial test bed plant at Capitol Concrete Products in Topeka, Kansas, with an experiment control at Sandia National Laboratories, Albuquerque. Plant evaluation will occur during a year-long period of industrial utilization. It will include performance testing, operability testing, and system failure analysis. Performance data will be recorded by a data acquisition system. User, community, and environmental inputs will be recorded in logs, journals, and files. Plant installation, start-up, and evaluation, are anticipated for late November, 1981.

Hauger, J. S.↗

Use of Thermoregulatory Models to Enhance Space Shuttle and Space Station operations and Review of Human Thermoregulatory Control

Thermoregulation in the space environment is critical for survival, especially in off- nominal operations. In such cases, mathematical models of thermoregulation are frequently employed to evaluate safety-of-flight issues in various human mission scenarious. In this study, the 225-node Wissler model and the 41-Node Metabolic Man model are employed to evaluate the effects of such a scenario. Metabolic loads on astronauts wearing the advanced crew escape suit (ACES) and liquid cooled ventilation garment (LCVG) are imposed on astronauts exposed to elevated cabin temperatures resulting from a systems failure. The study indicates that the performance of the ACES/LCVG cooling system is marginal. Increases in workload and or cabin temperature above nominal will increase rectal temperature, stored heat load, heart rate, and sweating, which could lead to deficits in the performance of cognitive and motor tasks. This is of concern as the ACES/LCVG is employed during Shuttle decent when the likelihood of a safe landing may be compromised. The study indicates that the most effective mitigation strategy would be to decrease the LCVG inlet temperature.

Pisacane, V. L.↗