Search NASA⌕ Search

SEARCH · Search NASA

Results for “System Unreliability”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

System applications of the fault tolerant memory

Conventional memory technologies currently employed in aerospace applications contribute at least fifty percent to system unreliability (where the system includes CPU, I/O and memory). A fault tolerant memory performs both error correction and memory replacement at the bit plane level. To determine the effects of system design of using a fault tolerant memory in space applications, analysis was performed to determine tradeable hardware configurations that meet the reliability goals of each program. The candidate configurations, which included redundant elements of the computer system with both conventional and fault tolerant memories, were then traded in terms of selection criteria of cost, weight, volume, and power. These trade studies demonstrated that a fault tolerant memory provided significant advantages in terms of cost, weight, and volume. The memory selected for this analysis was a recently developed five fault tolerant memory.

Murphy, L. J.↗

Design Methods and Practices for Fault Prevention and Management in Spacecraft

Integrated Systems Health Management (ISHM) is intended to become a critical capability for all space, lunar and planetary exploration vehicles and systems at NASA. Monitoring and managing the health state of diverse components, subsystems, and systems is a difficult task that will become more challenging when implemented for long-term, evolving deployments. A key technical challenge will be to ensure that the ISHM technologies are reliable, effective, and low cost, resulting in turn in safe, reliable, and affordable missions. To ensure safety and reliability, ISHM functionality, decisions and knowledge have to be incorporated into the product lifecycle as early as possible, and ISHM must be considered as an essential element of models developed and used in various stages during system design. During early stage design, many decisions and tasks are still open, including sensor and measurement point selection, modeling and model-checking, diagnosis, signature and data fusion schemes, presenting the best opportunity to catch and prevent potential failures and anomalies in a cost-effective way. Using appropriate formal methods during early design, the design teams can systematically explore risks without committing to design decisions too early. However, the nature of ISHM knowledge and data is detailed, relying on high-fidelity, detailed models, whereas the earlier stages of the product lifecycle utilize low-fidelity, high-level models of systems and their functionality. We currently lack the tools and processes necessary for integrating ISHM into the vehicle system/subsystem design. As a result, most existing ISHM-like technologies are retrofits that were done after the system design was completed. It is very expensive, and sometimes futile, to retrofit a system health management capability into existing systems. Last-minute retrofits result in unreliable systems, ineffective solutions, and excessive costs (e.g., Space Shuttle TPS monitoring which was considered only after 110 flights and the Columbia disaster). High false alarm or false negative rates due to substandard implementations hurt the credibility of the ISHM discipline. This paper presents an overview of the current state of ISHM design,and a review of formal design methods to make recommendations about possible approaches to enable the ISHM capabilities to be designed-in at the system-level, from the very beginning of the vehicle design process.

Tumer, Irem Y.↗

An Integrated Approach to Life Cycle Analysis

Life Cycle Analysis (LCA) is the evaluation of the impacts that design decisions have on a system and provides a framework for identifying and evaluating design benefits and burdens associated with the life cycles of space transportation systems from a "cradle-to-grave" approach. Sometimes called life cycle assessment, life cycle approach, or "cradle to grave analysis", it represents a rapidly emerging family of tools and techniques designed to be a decision support methodology and aid in the development of sustainable systems. The implementation of a Life Cycle Analysis can vary and may take many forms; from global system-level uncertainty-centered analysis to the assessment of individualized discriminatory metrics. This paper will focus on a proven LCA methodology developed by the Systems Analysis and Concepts Directorate (SACD) at NASA Langley Research Center to quantify and assess key LCA discriminatory metrics, in particular affordability, reliability, maintainability, and operability. This paper will address issues inherent in Life Cycle Analysis including direct impacts, such as system development cost and crew safety, as well as indirect impacts, which often take the form of coupled metrics (i.e., the cost of system unreliability). Since LCA deals with the analysis of space vehicle system conceptual designs, it is imperative to stress that the goal of LCA is not to arrive at the answer but, rather, to provide important inputs to a broader strategic planning process, allowing the managers to make risk-informed decisions, and increase the likelihood of meeting mission success criteria.

Chytka, T. M.↗

Procedure for Failure Mode, Effects, and Criticality Analysis (FMECA)

This document provides guidelines for the accomplishment of Failure Mode, Effects, and Criticality Analysis (FMECA) on the Apollo program. It is a procedure for analysis of hardware items to determine those items contributing most to system unreliability and crew safety problems.

Source record↗

An airline study of advanced technology requirements for advanced high speed commercial transport engines. 2: Engine preliminary design assessment

The advanced technology requirements for an advanced high speed commercial transport engine are presented. The results of the phase 2 study effort cover the following areas: (1) general review of preliminary engine designs suggested for a future aircraft, (2) presentation of a long range view of airline propulsion system objectives and the research programs in noise, pollution, and design which must be undertaken to achieve the goals presented, (3) review of the impact of propulsion system unreliability and unscheduled maintenance on cost of operation, (4) discussion of the reliability and maintainability requirements and guarantees for future engines.

Sallee, G. P.↗

The cost of software fault tolerance

The proposed use of software fault tolerance techniques as a means of reducing software costs in avionics and as a means of addressing the issue of system unreliability due to faults in software is examined. A model is developed to provide a view of the relationships among cost, redundancy, and reliability which suggests strategies for software development and maintenance which are not conventional.

Migneault, G. E.↗

Markov reliability models for digital flight control systems

The reliability of digital flight control systems can often be accurately predicted using Markov chain models. The cost of numerical solution depends on a model's size and stiffness. Acyclic Markov models, a useful special case, are particularly amenable to efficient numerical solution. Even in the general case, instantaneous coverage approximation allows the reduction of some cyclic models to more readily solvable acyclic models. After considering the solution of single-phase models, the discussion is extended to phased-mission models. Phased-mission reliability models are classified based on the state restoration behavior that occurs between mission phases. As an economical approach for the solution of such models, the mean failure rate solution method is introduced. A numerical example is used to show the influence of fault-model parameters and interphase behavior on system unreliability.

Mcgough, John↗

An Efficient Approach for the Reliability Analysis of Phased-Mission Systems with Dependent Failures

We consider the reliability analysis of phased-mission systems with common-cause failures in this paper. Phased-mission systems (PMS) are systems supporting missions characterized by multiple, consecutive, and nonoverlapping phases of operation. System components may be subject to different stresses as well as different reliability requirements throughout the course of the mission. As a result, component behavior and relationships may need to be modeled differently from phase to phase when performing a system-level reliability analysis. This consideration poses unique challenges to existing analysis methods. The challenges increase when common-cause failures (CCF) are incorporated in the model. CCF are multiple dependent component failures within a system that are a direct result of a shared root cause, such as sabotage, flood, earthquake, power outage, or human errors. It has been shown by many reliability studies that CCF tend to increase a system's joint failure probabilities and thus contribute significantly to the overall unreliability of systems subject to CCF.We propose a separable phase-modular approach to the reliability analysis of phased-mission systems with dependent common-cause failures as one way to meet the above challenges in an efficient and elegant manner. Our methodology is twofold: first, we separate the effects of CCF from the PMS analysis using the total probability theorem and the common-cause event space developed based on the elementary common-causes; next, we apply an efficient phase-modular approach to analyze the reliability of the PMS. The phase-modular approach employs both combinatorial binary decision diagram and Markov-chain solution methods as appropriate. We provide an example of a reliability analysis of a PMS with both static and dynamic phases as well as CCF as an illustration of our proposed approach. The example is based on information extracted from a Mars orbiter project. The reliability model for this orbiter considers the various phases of Launch, Cruise, Mars Orbit Insertion, and Orbit. Some of the CCF for the orbiter in this mission include environmental effects, such as micrometeoroids, human operator errors, and software errors.

reliability analysis↗

Algorithms for Multiple Fault Diagnosis With Unreliable Tests

In this paper, we consider the problem of constructing optimal and near-optimal multiple fault diagnosis (MFD) in bipartite systems with unreliable (imperfect) tests. It is known that exact computation of conditional probabilities for multiple fault diagnosis is NP-hard. The novel feature of our diagnostic algorithms is the use of Lagrangian relaxation and subgradient optimization methods to provide: (1) near optimal solutions for the MFD problem, and (2) upper bounds for an optimal branch-and-bound algorithm. The proposed method is illustrated using several examples. Computational results indicate that: (1) our algorithm has superior computational performance to the existing algorithms (approximately three orders of magnitude improvement), (2) the near optimal algorithm generates the most likely candidates with a very high accuracy, and (3) our algorithm can find the most likely candidates in systems with as many as 1000 faults.

Shakeri, Mojdeh↗

On-line replacement of program modules using AdaPT

One purpose of our research is the investigation of the effectiveness and expressiveness of AdaPT(1), a set of language extensions to Ada 83, for distributed systems. As a part of that effort, we are now investigating the subject of replacing, e.g., upgrading, software modules while the software system remains in operation. The AdaPT language extension provide a good basis for this investigation for several reasons: (1) they include the concept of specific, self-contained program modules which can be manipulated; (2) support for program configuration is included in the language; and (3) although the discussion will be in terms of the AdaPT language, the AdaPT to Ada 83 conversion methodology being developed as another part of this project will provide a basis for the application of our findings to Ada 83 systems. The purpose of this investigation is to explore the basic mechanisms to the replacement process. Thus, while replacement in the presence of real-time deadlines, heterogeneous systems, and unreliable networks is certainly a topic of interest, we will first gain an understanding of the basic processes in the absence of such concerns. The extension of the replacement process to more complex situations can be made later. This report will establish an overview of the on-line upgrade problem, and present a taxonomy of the various aspects of the replacement process.

Waldrop, Raymond S.↗

On-line upgrade of program modules using AdaPT

One purpose of our research is the investigation of the effectiveness and expressiveness of AdaPT, a set of language extensions to Ada 83, for distributed systems. As a part of that effort, we are now investigating the subject of replacing, e.g. upgrading, software modules while the software system remains in operation. The AdaPT language extensions provide a good basis for this investigation for several reasons: they include the concept of specific, self-contained program modules which can be manipulated; support for program configuration is included in the language; and although the discussion will be in terms of the AdaPT language, the AdaPT to Ada 83 conversion methodology being developed as another part of this project will provide a basis for the application of our findings to Ada 83 and Ada 9X systems. The purpose of this investigation is to explore the basic mechanisms of the replacement process. With this purpose in mind, we will avoid including issues whose presence would obscure these basic mechanisms by introducing additional, unrelated concerns. Thus, while replacement in the presence of real-time deadlines, heterogeneous systems, and unreliable networks is certainly a topic of interest, we will first gain an understanding of the basic processes in the absence of such concerns. The extension of the replacement process to more complex situations can be made later. A previous report established an overview of the module replacement problem, a taxonomy of the various aspects of the replacement process, and a solution to one case in the replacement taxonomy. This report provides solutions to additional cases in the replacement process taxonomy: replacement of partitions with state and replacement of nodes. The solutions presented here establish the basic principles for module replacement. Extension of these solutions to other more complicated cases in the replacement taxonomy is direct, though requiring substantial work beyond the available funding.

Waldrop, Raymond S.↗

Ti-48Al-2Cr-2Nb Evaluated Under Fretting Conditions

Material parameters govern many of the design decisions in any engineering task. When two materials are in contact and microscopically small, relative motions (either vibratory or creeping) occur, and fretting fatigue can result. Fretting fatigue is a material response influenced by the materials in contact as well as by such variables as loading and vibratory conditions. Fretting produces fresh, clean interacting surfaces and induces adhesion, galling, and wear in the contact zone. Time, money, and materials are unnecessarily wasted when galling and wear result in excessive fretting fatigue that leads to poorly performing, unreliable mechanical systems. Fretting fatigue is a complex problem of significant interest to aircraft engine manufacturers. It can occur in a variety of engine components. Numerous approaches, depending on the component and the operating conditions, have been taken to address the fretting problems. The components of interest in this investigation were the low-pressure turbine blades and disks. The blades in this case were titanium aluminide, Ti-48Al-2Cr- 2Nb, and the disk was a nickel-base superalloy, Inconel 718 (IN 718). A concern for these airfoils is the fretting in fitted interfaces at the dovetail where the blade and disk are connected. Careful design can reduce fretting in most cases, but not completely eliminate it, because the airfoils frequently have a skewed (angled) blade-disk dovetail attachment, which leads to a complex stress state. Furthermore, the local stress state becomes more complex when the influence of the metal-metal contact and the edge of contact are considered.

Miyoshi, Kazuhisa↗

A hierarchical approach to reliability modeling of fault-tolerant systems

A methodology for performing fault tolerant system reliability analysis is presented. The method decomposes a system into its subsystems, evaluates vent rates derived from the subsystem's conditional state probability vector and incorporates those results into a hierarchical Markov model of the system. This is done in a manner that addresses failure sequence dependence associated with the system's redundancy management strategy. The method is derived for application to a specific system definition. Results are presented that compare the hierarchical model's unreliability prediction to that of a more complicated tandard Markov model of the system. The results for the example given indicate that the hierarchical method predicts system unreliability to a desirable level of accuracy while achieving significant computational savings relative to component level Markov model of the system.

Gossman, W. E.↗

Optimizing Air Transportation Service to Metroplex Airports: Analysis of Historical Data - Part 1

The air transportation system is a significant driver of the U.S. economy, providing safe, affordable, and rapid transportation. During the past three decades airspace and airport capacity has not grown in step with demand for air transportation (+4% annual growth), resulting in unreliable service and systemic delays. Estimates of the impact of delays and unreliable air transportation service on the economy range from $32B to $41B per year. This report describes the results of an analysis of airline strategic decision-making with regards to: (1) geographic access, (2) economic access, and (3) airline finances. This analysis evaluated markets-served, scheduled flights, aircraft size, airfares, and profit from 2005-2009. During this period, airlines experienced changes in costs of operation (due to fluctuations in hedged fuel prices), changes in travel demand (due to changes in the economy), and changes in infrastructure capacity (due to the capacity limits at EWR, JFK, and LGA). This analysis captures the impact of the implementation of capacity limits at airports, as well as the effect of increased costs of operation (i.e. hedged fuel prices). The increases in costs of operation serve as a proxy for increased costs per flight that might occur if auctions or congestion pricing are imposed.

Donohue, George↗

Semi-Markov Unreliability-Range Evaluator

Reconfigurable, fault-tolerant systems modeled. Semi-Markov unreliability-range evaluator (SURE) computer program is software tool for analysis of reliability of reconfigurable, fault-tolerant systems. Based on new method for computing death-state probabilities of semi-Markov model. Computes accurate upper and lower bounds on probability of failure of system. Written in PASCAL.

Butler, Ricky W.↗

Coverage modeling for dependability analysis of fault-tolerant systems

Several different models for predicting coverage in a fault-tolerant system, including models for permanent, intermittent, and transient errors, are discussed. Markov, semi-Markov, nonhomogeneous Markov, and extended stochastic Petri net models for computing coverage are developed. Two types of events that interfere with recovery are examined; and methods for modeling such events, whether they are deterministic or random, are given. The sensitivity of system reliability/availability to the coverage parameter and the sensitivity of the coverage parameter to various error-handling strategies are investigated. It is found that a policy of attempting transient recovery upon detection of an error can actually increase the unreliability of the system. This result is true if the error detectability is not nearly perfect, so that the risk of producing an undetectable error is greater than the benefit gained by not discarding the component.

Dugan, Joanne Bechta↗

On reliable control system designs with and without feedback reconfigurations

This paper contains an overview of a theoretical framework for the design of reliable multivariable control systems, with special emphasis on actuator failures and necessary actuator redundancy levels. Using a linear model of the system, with Markovian failure probabilities and quadratic performance index, an optimal stochastic control problem is posed and solved. The solution requires the iteration of a set of highly coupled Riccati-like matrix difference equations; if these converge one has a reliable design; if they diverge, the design is unreliable, and the system design cannot be stabilized. In addition, it is shown that the existence of a stabilizing constant feedback gain and the reliability of its implementation is equivalent to the convergence properties of a set of coupled Riccati-like matrix difference equations. In summary, these results can be used for offline studies relating the open loop dynamics, required performance, actuator mean time to failure, and functional or identical actuator redundancy, with and without feedback gain reconfiguration strategies.

Birdwell, J. D.↗