Search NASA⌕ Search

SEARCH · Search NASA

Results for “SYSTEM FAILURE”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 163 records · Page 9

Low-thrust mission risk analysis.

A computerized multi-stage failure process simulation procedure is used to evaluate the risk in a solar electric space mission. The procedure uses currently available thrust-subsystem reliability data and performs approximate simulations of the thrust subsystem burn operation, the system failure processes, and the retargetting operations. The application of the method is used to assess the risks in carrying out a 1980 rendezvous mission to Comet Encke. Analysis of the results and evaluation of the effects of various risk factors on the mission show that system component failure rates is the limiting factor in attaining a high mission reliability. But it is also shown that a well-designed trajectory and system operation mode can be used effectively to partially compensate for unreliable thruster performance.

Yen, C. L.↗

Development of an adaptive failure detection and identification system for detecting aircraft control element failures

A methodology for designing a failure detection and identification (FDI) system to detect and isolate control element failures in aircraft control systems is reviewed. An FDI system design for a modified B-737 aircraft resulting from this methodology is also reviewed, and the results of evaluating this system via simulation are presented. The FDI system performed well in a no-turbulence environment, but it experienced an unacceptable number of false alarms in atmospheric turbulence. An adaptive FDI system, which adjusts thresholds and other system parameters based on the estimated turbulence level, was developed and evaluated. The adaptive system performed well over all turbulence levels simulated, reliably detecting all but the smallest magnitude partially-missing-surface failures.

Bundick, W. Thomas↗

A three-failure-tolerant computer system.

Two basic factors influence the design of highly reliable computer systems: the amount of failures required to be tol erated and the reliability or MTBF required. A computer system designed to tolerate any three single failures in a fail-operational-fail-operational-fail-safe manner for a real-time control application is presented. The design approach uses adaptive majority voting in both hardware and software with a four-level redundant system. Various methods of performing the adaptive majority voting functions were evaluated with the selected approach using a special module termed a voter-comparator switch (VCS). The VCS module allows the computer system to be operated in a variety of redundant modes, depending on the failure tolerance required at any particular time.

Koczela, L. J.↗

Streaming Analytics for Anomaly Detection in Large-Scale Data

Anomalous behavior poses serious risks to assured performance and reliability of complex, high-consequence systems. For spaceborne assets and their state-of-health (SOH) telemetry, the challenges of high-dimensional data of varying data types are compounded by computational limitations from size, weight, and power (SWaP) constraints as well as data availability. Automated anomaly detection methods tend to perform poorly under these constraints, while current operational approaches can introduce delays in response time due to the manual, retrospective processes for understanding system failures. As a result, presently deployed space systems, and those deployed in the near future, face situations where mission operations might be delayed or only be able to operate under degraded capabilities. Here, we examine a near-term lightweight solution that provides real-time detection capabilities for rare events and assess state-of-the-art anomaly detection techniques against real SOH telemetry from space platforms. This report describes our methodology and research, which could support more automated capabilities for comprehensive space operations as well as for other resource-constrained edge applications.

97 MATHEMATICS AND COMPUTING↗

A Model-Based Expert System for Space Power Distribution Diagnostics

When engineers diagnose system failures, they often use models to confirm system operation. This concept has produced a class of advanced expert systems that perform model-based diagnosis. A model-based diagnostic expert system for the Space Station Freedom electrical power distribution test bed is currently being developed at the NASA Lewis Research Center. The objective of this expert system is to autonomously detect and isolate electrical fault conditions. Marple, a software package developed at TRW, provides a model-based environment utilizing constraint suspension. Originally, constraint suspension techniques were developed for digital systems. However, Marple provides the mechanisms for applying this approach to analog systems such as the test bed, as well. The expert system was developed using Marple and Lucid Common Lisp running on a Sun Sparc-2 workstation. The Marple modeling environment has proved to be a useful tool for investigating the various aspects of model-based diagnostics. This report describes work completed to date and lessons learned while employing model-based diagnostics using constraint suspension within an analog system.

Quinn, Todd M.↗

Microgravity Analogues of Herpes Virus Pathogenicity: Human Cytomegalovirus (hCMV) and Varicella Zoster (VZV) Infectivity in Human Tissue Like Assemblies (TLAs)

The old adage we are our own worst enemies may perhaps be the most profound statement ever made when applied to man s desire for extraterrestrial exploration and habitation of Space. Consider the immune system protects the integrity of the entire human physiology and is comprised of two basic elements the adaptive or circulating and the innate immune system. Failure of the components of the adaptive system leads to venerability of the innate system from opportunistic microbes; viral, bacteria, and fungal, which surround us, are transported on our skin, and commonly inhabit the human physiology as normal and imunosuppressed parasites. The fine balance which is maintained for the preponderance of our normal lives, save immune disorders and disease, is deregulated in microgravity. Thus analogue systems to study these potential Risks are essential for our progress in conquering Space exploration and habitation. In this study we employed two known physiological target tissues in which the reactivation of hCMV and VZV occurs, human neural and lung systems created for the study and interaction of these herpes viruses independently and simultaneously on the innate immune system. Normal human neural and lung tissue analogues called tissue like assemblies (TLAs) were infected with low MOIs of approximately 2 x 10(exp -5) pfu hCMV or VZV and established active but prolonged low grade infections which spanned .7-1.5 months in length. These infections were characterized by the ability to continuously produce each of the viruses without expiration of the host cultures. Verification and quantification of viral replication was confirmed via RT_PCR, IHC, and confocal spectral analyses of the respective essential viral genomes. All host TLAs maintained the ability to actively proliferate throughout the entire duration of the experiments as is analogous to normal in vivo physiological conditions. These data represent a significant advance in the ability to study the triggering mechanisms which surround Herpes vial reactivation and proliferation. Additionally, prolonged replication of these viruses will allow the tracking of viral genomic shift.

Goodwin, T. J.↗

ISS Regenerative Life Support: Challenges and Success in the Quest for Long-Term Habitability in Space

The International Space Station's (ISS) Regenerative Environmental Control and Life Support System (ECLSS) was launched in 2008 to continuously recycle urine and crew sweat into drinking water and oxygen using brand new technologies. This functionality was highly important to the ability of the ISS to transition to the long-term goal of 6-crew operations as well as being critical tests for long-term space habitability. Through the initial activation and long-term operations of these systems, important lessons were learned about the importance of system redundancy and operational workarounds that allow Systems Engineers to maintain functionality with limited on-orbit spares. This presentation will share some of these lessons learned including how to balance water through the different systems, store and use water for use in system failures and creating procedures to operate the systems in ways that they were not initially designed to do.

Bazley, Jesse↗

A unified method for evaluating real-time computer controllers: A case study

A real time control system consists of a synergistic pair, that is, a controlled process and a controller computer. Performance measures for real time controller computers are defined on the basis of the nature of this synergistic pair. A case study of a typical critical controlled process is presented in the context of new performance measures that express the performance of both controlled processes and real time controllers (taken as a unit) on the basis of a single variable: controller response time. Controller response time is a function of current system state, system failure rate, electrical and/or magnetic interference, etc., and is therefore a random variable. Control overhead is expressed as a monotonically nondecreasing function of the response time and the system suffers catastrophic failure, or dynamic failure, if the response time for a control task exceeds the corresponding system hard deadline, if any. A rigorous probabilistic approach is used to estimate the performance measures. The controlled process chosen for study is an aircraft in the final stages of descent, just prior to landing. First, the performance measures for the controller are presented. Secondly, control algorithms for solving the landing problem are discussed and finally the impact of the performance measures on the problem is analyzed.

Shin, K. G.↗

A unified method for evaluating real-time computer controllers and its application

A real time control system consists of a synergistic pair, that is, a controlled process and a controller computer. Performance measures for real time controller computers are defined on the basis of the nature of this synergistic pair. A case study of a typical critical controlled process is presented in the context of new performance measures that express the performance of both controlled processes and real time controllers (taken as a unit) on the basis of a single variable: controller response time. Controller response time is a function of current system state, system failure rate, electrical and/or magnetic interference, etc., and is therefore a random variable. Control overhead is expressed as a monotonically nondecreasing function of the response time and the system suffers catastrophic failure, or dynamic failure, if the response time for a control task exceeds the corresponding system hard deadline, if any. A rigorous probabilistic approach is used to estimate the performance measures. The controlled process chosen for study is an aircraft in the final stages of descent, just prior to landing. First, the performance measures for the controller are presented. Secondly, control algorithms for solving the landing problem are discussed and finally the impact of the performance measures on the problem is analyzed.

Shin, K. G.↗

An Efficient Approach for the Reliability Analysis of Phased-Mission Systems with Dependent Failures

We consider the reliability analysis of phased-mission systems with common-cause failures in this paper. Phased-mission systems (PMS) are systems supporting missions characterized by multiple, consecutive, and nonoverlapping phases of operation. System components may be subject to different stresses as well as different reliability requirements throughout the course of the mission. As a result, component behavior and relationships may need to be modeled differently from phase to phase when performing a system-level reliability analysis. This consideration poses unique challenges to existing analysis methods. The challenges increase when common-cause failures (CCF) are incorporated in the model. CCF are multiple dependent component failures within a system that are a direct result of a shared root cause, such as sabotage, flood, earthquake, power outage, or human errors. It has been shown by many reliability studies that CCF tend to increase a system's joint failure probabilities and thus contribute significantly to the overall unreliability of systems subject to CCF.We propose a separable phase-modular approach to the reliability analysis of phased-mission systems with dependent common-cause failures as one way to meet the above challenges in an efficient and elegant manner. Our methodology is twofold: first, we separate the effects of CCF from the PMS analysis using the total probability theorem and the common-cause event space developed based on the elementary common-causes; next, we apply an efficient phase-modular approach to analyze the reliability of the PMS. The phase-modular approach employs both combinatorial binary decision diagram and Markov-chain solution methods as appropriate. We provide an example of a reliability analysis of a PMS with both static and dynamic phases as well as CCF as an illustration of our proposed approach. The example is based on information extracted from a Mars orbiter project. The reliability model for this orbiter considers the various phases of Launch, Cruise, Mars Orbit Insertion, and Orbit. Some of the CCF for the orbiter in this mission include environmental effects, such as micrometeoroids, human operator errors, and software errors.

reliability analysis↗

A Comparative Study on Computation of Cumulative Distribution Function in Predicting Time of Failure of Engineering Systems

Estimating accurate Time-of-Failure (ToF) of a system is key in making the decisions that impact operational safety and optimize cost. In this context, it is interesting to note that different approaches have been explored to tackle the problem of estimating ToF. The difference is in part characterized by different definitions of the hazard zones. The conventional definition for the cumulative distribution function (CDF) calculation is assumed to have well-defined hazard zones, that is, hazard zones defined as a function of the system state trajectory. An alternate method suggests the use of hazard zones defined as a function of the system state at time , instead of hazard zones defined as a function of system state up to and including time k (Acuna and Orchard 2018, 2017). This paper explores these differences and their impact on ToF estimation. Results for the conventional CDF definition indicated that, (i) the cumulative distribution function is always an increasing function of time, even when realizations of the degradation process are not monotonic, (ii) the sum of all probabilities is always 1 and does not need to be normalized, and (iii) all probabilities are positive and less than or equal to 1. Similar results are not observed for CDF calculation with hazard zones defined as a function only of the system state at time . Results for ToF estimation using Acuna's definition differ, suggesting that there is an underlying assumption of independence in the hazard zone definition. Therefore, we present an alternate definition of hazard zone which guarantees the properties of a well-defined CDF with a more straightforward ToF definition.

Hazard Zone↗

An innovative design for autonomous backup attitude control of the Gamma Ray Observatory

The Gamma Ray Observatory is a NASA funded three-axis stabilized spacecraft which will carry four scientific instruments to observe gamma ray phenomena. The requirement to protect the scientific mission from system failures led to the attitude control and determination system design described in this paper. The design employs nine control modes with error detection, hardware substitution, and autonomous mode switching. The system architecture evolved to eliminate cross-dependence between the primary on-board computer (OBC) and the backup control processor electronics. Cross strapping of sensors and actuators and separation of the input/output electronics ensure that a reliable set of sensors and actuators will be available for backup mode operation. The OBC software includes failure detection, hardware reconfiguration, and mode switching logic which provide the ability to autonomously transfer, upon anomaly, to a reliable backup mode. Verification of this mode transition design is done in four test programs: at the unit level, by analytical simulation, by a hybrid breadboard electronics-simulation setup, and by a flight hardware-simulation test.

Tai, F.↗

Detecting and Characterizing Patterns of Failure in Complex Systems: An Ontology Development and Clustering Approach

While the causes of failures in complex engineered systems are often clear in hindsight, it can be challenging to predict failures proactively during the design of novel engineered products or systems. Identifying patterns can be useful for capturing common characteristics that may lead to failure. In this paper, we present a methodology for identifying patterns of failure from NASA’s publicly available Lessons Learned Information System (LLIS). We apply an ontology development and clustering approach to identify representative patterns leading to failures in historical lessons learned. A joint inductive-deductive approach reveals the key themes in lessons that lead to failure, which are formalized and recorded as an ontology of complex systems failure causes. Documents from the LLIS are manually tagged with relevant characteristics from the ontology. From the tagged set, clustering is used to capture co-occurring sets of characteristics that lead to failure. The primary contribution of this work is a method for extracting a set of generic failure patterns in complex engineered systems and characteristics for these patterns that can be identified at design time, knowledge of which can be used to plan mitigation strategies.

Systems Engineering↗

Concepts for Distributed Engine Control

Gas turbine engines for aero-propulsion systems are found to be highly optimized machines after over 70 years of development. Still, additional performance improvements are sought while reduction in the overall cost is increasingly a driving factor. Control systems play a vitally important part in these metrics but are severely constrained by the operating environment and the consequences of system failure. The considerable challenges facing future engine control system design have been investigated. A preliminary analysis has been conducted of the potential benefits of distributed control architecture when applied to aero-engines. In particular, reductions in size, weight, and cost of the control system are possible. NASA is conducting research to further explore these benefits, with emphasis on the particular benefits enabled by high temperature electronics and an open-systems approach to standardized communications interfaces.

Culley, Dennis E.↗

The Parable of the Boiled Safety Professional

Common and unique issues contribute to system failures. This paper touches on the concept of drift to failure as a cautionary message. Managers and leaders, design team members, fabricators and assemblers, analysis and assurance personnel, and others associated with operating and maintaining systems, need to pay attention to identify the manifestation of individual and collective behaviors that might indicate slips in rigor or focus or decisions that might eat away at safety margins as our system drifts to failure. Corrections to drift made during design and development phases may efficiently prevent or mitigate drift problems occurring in the operational phase.

Shivers, Charles H.↗

Flight Results of the NF-15B Intelligent Flight Control System (IFCS) Aircraft with Adaptation to a Longitudinally Destabilized Plant

Adaptive flight control systems have the potential to be resilient to extreme changes in airplane behavior. Extreme changes could be a result of a system failure or of damage to the airplane. The goal for the adaptive system is to provide an increase in survivability in the event that these extreme changes occur. A direct adaptive neural-network-based flight control system was developed for the National Aeronautics and Space Administration NF-15B Intelligent Flight Control System airplane. The adaptive element was incorporated into a dynamic inversion controller with explicit reference model-following. As a test the system was subjected to an abrupt change in plant stability simulating a destabilizing failure. Flight evaluations were performed with and without neural network adaptation. The results of these flight tests are presented. Comparison with simulation predictions and analysis of the performance of the adaptation system are discussed. The performance of the adaptation system is assessed in terms of its ability to stabilize the vehicle and reestablish good onboard reference model-following. Flight evaluation with the simulated destabilizing failure and adaptation engaged showed improvement in the vehicle stability margins. The convergent properties of this initial system warrant additional improvement since continued maneuvering caused continued adaptation change. Compared to the non-adaptive system the adaptive system provided better closed-loop behavior with improved matching of the onboard reference model. A detailed discussion of the flight results is presented.

Bosworth, John T.↗