Search NASA⌕ Search

SEARCH · Search NASA

Results for “fault tree analysis”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 109 records · Page 6

The role of reliability graph models in assuring dependable operation of complex hardware/software systems

The complexity of computer systems currently being designed for critical applications in the scientific, commercial, and military arenas requires the development of new techniques for utilizing models of system behavior in order to assure 'ultra-dependability'. The complexity of these systems, such as Space Station Freedom and the Air Traffic Control System, stems from their highly integrated designs containing both hardware and software as critical components. Reliability graph models, such as fault trees and digraphs, are used frequently to model hardware systems. Their applicability for software systems has also been demonstrated for software safety analysis and the analysis of software fault tolerance. This paper discusses further uses of graph models in the design and implementation of fault management systems for safety critical applications.

Patterson-Hine, F. A.↗

Medical Resource Set Bulky Item Trade Space Analysis for Spaceflight Medical Risk

The NASA engineering community utilizes event-driven and fault-tree probabilistic techniques to classify risks in the space environment by taking advantage of the inherent knowledge of complex spaceflight system design and testing to quantify failure risk. In harmonizing the risk of human space flight, answering the question of ‘How do we balance health, performance and resource risks with other engineering risks on long duration space missions?’ remains a deeply challenging and largely qualitative practice. The Medical Extensible Dynamic Probabilistic Risk Assessment Tool (MEDPRAT) is one aspect of the efforts by NASA’s Human Research Program (HRP) to quantitatively assess the impact of health and performance risk. One of MEDPRAT’s key features is its high degree of computational efficiency. Coupled with the HRP High Performance Compute cluster located at NASA’s Glenn Research Center, MEDPRAT runs millions of simulated missions in a matter of minutes. This degree of computational efficiency provides the novel opportunity to explore the relationship between medical set mass, volume, and medical resource size. Of particular interest for future human spaceflight missions are ‘bulky’ items, medical resources like devices, which occupy a large portion of the small, allocated mass and volume for the medical set leaving less room for other resources. This talk will present results showing the quantitative impact of forced inclusion of several bulky items across a variety of medical kit constraints, and the effect that a potential research investment into reducing the bulky item mass and volume may have on risk.

Lauren Mcintyre↗

Simplified Phased-Mission System Analysis for Systems with Independent Component Repairs

Accurate analysis of reliability of system requires that it accounts for all major variations in system's operation. Most reliability analyses assume that the system configuration, success criteria, and component behavior remain the same. However, multiple phases are natural. We present a new computationally efficient technique for analysis of phased-mission systems where the operational states of a system can be described by combinations of components states (such as fault trees or assertions). Moreover, individual components may be repaired, if failed, as part of system operation but repairs are independent of the system state. For repairable systems Markov analysis techniques are used but they suffer from state space explosion. That limits the size of system that can be analyzed and it is expensive in computation. We avoid the state space explosion. The phase algebra is used to account for the effects of variable configurations, repairs, and success criteria from phase to phase. Our technique yields exact (as opposed to approximate) results. We demonstrate our technique by means of several examples and present numerical results to show the effects of phases and repairs on the system reliability/availability.

Somani, Arun K.↗

Comparative Analysis of Static and Dynamic Probabilistic Risk Assessment

This study examines three different methodologies for producing loss-of-mission (LOM) and loss-of-crew (LOC) risks estimates for probabilistic risk assessments (PRA) of crewed spacecraft. The three bottom-up, component-based PRA approaches examined are a traditional static fault tree, a dynamic Monte Carlo simulation, and a fault tree hybrid that incorporates some dynamic elements. These approaches were used to model the reaction control system thruster pod of a generic crewed spacecraft and mission, and a comparative analysis of the methods is presented. The methodologies are assessed in terms of the process of modeling a system, the actionable information produced for the design team, and the overall fidelity of the quantitative risk evaluation generated. The system modeling process is compared in terms of the effort required to generate the initial model, update the model in response to design changes, and support mass-versus-risk trade studies. The results are compared by examining the top-level LOM/LOC estimates and the relative risk driver rankings at the failure mode level. The fidelity of each modeling methodology is discussed in terms of its capability to handle real-world system dynamics such as cold-sparing, changes in mission operations due to loss of redundancy, and common cause failure modes. The paper also discusses the applicability of each methodology to different phases of system development and shows that a single methodology may not be suitable for all of the many purposes of a spacecraft PRA. The fault tree hybrid approach is shown to be best suited to the needs of early assessments during conceptual design phases. As the design begins to mature, the level of detail represented in the risk model must go beyond redundancy and nominal mission operations to include dynamic, time- and state-dependent system responses as well as diverse system capabilities. This is best accomplished using the dynamic simulation approach, since these phenomena are not easily captured by static methods. Ultimately, once the design has been finalized and the goal of the PRA is to provide design validation and requirement verification, more traditional, static fault tree approaches may become as appropriate as the simulation method.

Mattenberger, Christopher J.↗

Health Management Applications for International Space Station

Traditional mission and vehicle management involves teams of highly trained specialists monitoring vehicle status and crew activities, responding rapidly to any anomalies encountered during operations. These teams work from the Mission Control Center and have access to engineering support teams with specialized expertise in International Space Station (ISS) subsystems. Integrated System Health Management (ISHM) applications can significantly augment these capabilities by providing enhanced monitoring, prognostic and diagnostic tools for critical decision support and mission management. The Intelligent Systems Division of NASA Ames Research Center is developing many prototype applications using model-based reasoning, data mining and simulation, working with Mission Control through the ISHM Testbed and Prototypes Project. This paper will briefly describe information technology that supports current mission management practice, and will extend this to a vision for future mission control workflow incorporating new ISHM applications. It will describe ISHM applications currently under development at NASA and will define technical approaches for implementing our vision of future human exploration mission management incorporating artificial intelligence and distributed web service architectures using specific examples. Several prototypes are under development, each highlighting a different computational approach. The ISStrider application allows in-depth analysis of Caution and Warning (C&W) events by correlating real-time telemetry with the logical fault trees used to define off-nominal events. The application uses live telemetry data and the Livingstone diagnostic inference engine to display the specific parameters and fault trees that generated the C&W event, allowing a flight controller to identify the root cause of the event from thousands of possibilities by simply navigating animated fault tree models on their workstation. SimStation models the functional power flow for the ISS Electrical Power System and can predict power balance for nominal and off-nominal conditions. SimStation uses realtime telemetry data to keep detailed computational physics models synchronized with actual ISS power system state. In the event of failure, the application can then rapidly diagnose root cause, predict future resource levels and even correlate technical documents relevant to the specific failure. These advanced computational models will allow better insight and more precise control of ISS subsystems, increasing safety margins by speeding up anomaly resolution and reducing,engineering team effort and cost. This technology will make operating ISS more efficient and is directly applicable to next-generation exploration missions and Crew Exploration Vehicles.

Alena, Richard↗

Comparative Analysis of Static and Dynamic Probabilistic Risk Assessment

Implementation of risk-informed design allows the design team to thoroughly explore the risks of a system while iterating the operations concept, design, and requirements until the system meets mission objects and is achievable within constraints. To arrive at a space system design that is likely to meet all constraints placed upon mass, cost, performance and risk, the system requirements must be understood and traded against each other as early as the conceptual design phase. Depending on the project phase and the goals of the risk analysis, various PRA methodologies could be used to produce quantitative risk estimates to enable such a process. In order to better understand the applicability, advantages, and limitations of various PRA methodologies, a comparative analysis of three bottom-up, component-based PRA approaches was performed. The three methods examined are a traditional static fault tree, a fault tree hybrid, and a dynamic Monte Carlo simulation. Each approach was used to assess a generic reaction control system (RCS) thruster pod and mission. The methods are assessed in terms of the process of modeling a system, the actionable information produced for the design team, and the overall fidelity of the quantitative risk evaluation generated. The paper also discusses the applicability of each methodology to the different phases of system development.

Probablistic↗

Learning from examples - Generation and evaluation of decision trees for software resource analysis

A general solution method for the automatic generation of decision (or classification) trees is investigated. The approach is to provide insights through in-depth empirical characterization and evaluation of decision trees for software resource data analysis. The trees identify classes of objects (software modules) that had high development effort. Sixteen software systems ranging from 3,000 to 112,000 source lines were selected for analysis from a NASA production environment. The collection and analysis of 74 attributes (or metrics), for over 4,700 objects, captured information about the development effort, faults, changes, design style, and implementation style. A total of 9,600 decision trees were automatically generated and evaluated. The trees correctly identified 79.3 percent of the software modules that had high development effort or faults, and the trees generated from the best parameter combinations correctly identified 88.4 percent of the modules on the average.

Selby, Richard W.↗

Development and Validation of a Probabilistic Risk Assessment Model for a Generic Modular High Temperature Gas-Cooled Reactor

This study looks to develop and validate a probabilistic risk assessment (PRA) model for the modular high temperature gas-cooled reactor (MHTGR) using INL’s Systems Analysis Programs for Hands-on Integrated Reliability Evaluations (SAPHIRE) software. Validation against a General Atomics design involves matching event and fault trees to historical frequencies, with a goal of under 15% difference. The research aims to deliver a reliable PRA model to assess the safety of MHTGRs for use within high-temperature industrial applications.

22 GENERAL STUDIES OF NUCLEAR REACTORS↗

Reliability/safety analysis of a fly-by-wire system

An analysis technique has been developed to estimate the reliability of a very complex, safety-critical system by constructing a diagram of the reliability equations for the total system. This diagram has many of the characteristics of a fault-tree or success-path diagram, but is much easier to construct for complex redundant systems. The diagram provides insight into system failure characteristics and identifies the most likely failure modes. A computer program aids in the construction of the diagram and the computation of reliability. Analysis of the NASA F-8 Digital Fly-by-Wire Flight Control System is used to illustrate the technique.

Brock, L. D.↗

Communications and tracking expert systems study

The original objectives of the study consisted of five broad areas of investigation: criteria and issues for explanation of communication and tracking system anomaly detection, isolation, and recovery; data storage simplification issues for fault detection expert systems; data selection procedures for decision tree pruning and optimization to enhance the abstraction of pertinent information for clear explanation; criteria for establishing levels of explanation suited to needs; and analysis of expert system interaction and modularization. Progress was made in all areas, but to a lesser extent in the criteria for establishing levels of explanation suited to needs. Among the types of expert systems studied were those related to anomaly or fault detection, isolation, and recovery.

Leibfried, T. F.↗

Program Finds Minimal Cut Sets

CUTSETS computer program identifies all minimal cut sets for given node. Software package contains subprograms that solve for minimal cut sets of fault trees and digraphs by use of object-oriented programming techniques. Cut-set codes used to solve graph models for reliability analysis and identify potential single-point failures in modeled system. Includes utility subprogram that converts popular COD-format diagraph-model-description files into text input files suitable for use with other CUT-SETS subprograms. FEAT (MSC-21873) and FIRM (MSC-21860). Written in C language.

Iverson, D. L.↗

Method and system for dynamic probabilistic risk assessment

The DEFT methodology, system and computer readable medium extends the applicability of the PRA (Probabilistic Risk Assessment) methodology to computer-based systems, by allowing DFT (Dynamic Fault Tree) nodes as pivot nodes in the Event Tree (ET) model. DEFT includes a mathematical model and solution algorithm, supports all common PRA analysis functions and cutsets. Additional capabilities enabled by the DFT include modularization, phased mission analysis, sequence dependencies, and imperfect coverage.

Dugan, Joanne Bechta↗

Accounting for Point Estimate Uncertainty in Space Systems Reliability and Risk Analysis

Understanding and accounting for uncertainty in risk analysis is a critical step in the management and communication of risk in engineered systems. The component and system-level analysis to determine the probability of a negative outcome and its consequence is often quantified by a point estimate. Many Program and Enterprise decisions involving technical concerns and issues rely on reliability engineering activities to produce quantified risk analysis to inform the decision making process. At NASA, it is common to use a Probabilistic Risk Analysis (PRA) to inform the overall risk to Loss of Mission or Loss of Crew that involves integration across all spacecraft subsystem fault trees to produce an overall probability of mission failure. The point estimate is an estimate of this overall probability and is an immediate result of a fault tree model. It is the result of a model where the probability of each event is taken to be equal to its mean. The value provides an approximation of the overall mean without running any uncertainty calculations (e.g., no sampling). Using only the point estimate can lead to a false sense of precision and the point estimate may not match the resulting mean when uncertainty is taken into consideration. This paper will explore five conditions that can cause the PRA model mean to diverge from the point estimate and will provide engineers and managers insight into the importance of understanding uncertainty in the elements of PRA models.

Paul J Collier↗

Constellation Probabilistic Risk Assessment (PRA): Design Consideration for the Crew Exploration Vehicle

Managed by NASA's Office of Safety and Mission Assurance, a pilot probabilistic risk analysis (PRA) of the NASA Crew Exploration Vehicle (CEV) was performed in early 2006. The PRA methods used follow the general guidance provided in the NASA PRA Procedures Guide for NASA Managers and Practitioners'. Phased-mission based event trees and fault trees are used to model a lunar sortie mission of the CEV - involving the following phases: launch of a cargo vessel and a crew vessel; rendezvous of these two vessels in low Earth orbit; transit to th$: moon; lunar surface activities; ascension &om the lunar surface; and return to Earth. The analysis is based upon assumptions, preliminary system diagrams, and failure data that may involve large uncertainties or may lack formal validation. Furthermore, some of the data used were based upon expert judgment or extrapolated from similar components~systemsT. his paper includes a discussion of the system-level models and provides an overview of the analysis results used to identify insights into CEV risk drivers, and trade and sensitivity studies. Lastly, the PRA model was used to determine changes in risk as the system configurations or key parameters are modified.

Prassinos, Peter G.↗

Columbia Accident Investigation Board Report. Volume Two

Volume II of the Report contains appendices that were cited in Volume I. The Columbia Accident Investigation Board produced many of these appendices as working papers during the investigation into the February 1, 2003 destruction of the Space Shuttle Columbia. Other appendices were produced by other organizations (mainly NASA) in support of the Board investigation. In the case of documents that have been published by others, they are included here in the interest of establishing a complete record, but often at less than full page size. Contents include: CAIB Technical Documents Cited in the Report: Reader's Guide to Volume II; Appendix D. a Supplement to the Report; Appendix D.b Corrections to Volume I of the Report; Appendix D.1 STS-107 Training Investigation; Appendix D.2 Payload Operations Checklist 3; Appendix D.3 Fault Tree Closure Summary; Appendix D.4 Fault Tree Elements - Not Closed; Appendix D.5 Space Weather Conditions; Appendix D.6 Payload and Payload Integration; Appendix D.7 Working Scenario; Appendix D.8 Debris Transport Analysis; Appendix D.9 Data Review and Timeline Reconstruction Report; Appendix D.10 Debris Recovery; Appendix D.11 STS-107 Columbia Reconstruction Report; Appendix D.12 Impact Modeling; Appendix D.13 STS-107 In-Flight Options Assessment; Appendix D.14 Orbiter Major Modification (OMM) Review; Appendix D.15 Maintenance, Material, and Management Inputs; Appendix D.16 Public Safety Analysis; Appendix D.17 MER Manager's Tiger Team Checklist; Appendix D.18 Past Reports Review; Appendix D.19 Qualification and Interpretation of Sensor Data from STS-107; Appendix D.20 Bolt Catcher Debris Analysis.

CAIB (COLUMBIA ACCIDENT INVESTIGATION BOARD)↗

Reliability and Probabilistic Risk Assessment - How They Play Together

Since the Space Shuttle Challenger accident in 1986, NASA has extensively used probabilistic analysis methods to assess, understand, and communicate the risk of space launch vehicles. Probabilistic Risk Assessment (PRA), used in the nuclear industry, is one of the probabilistic analysis methods NASA utilizes to assess Loss of Mission (LOM) and Loss of Crew (LOC) risk for launch vehicles. PRA is a system scenario based risk assessment that uses a combination of fault trees, event trees, event sequence diagrams, and probability distributions to analyze the risk of a system, a process, or an activity. It is a process designed to answer three basic questions: 1) what can go wrong that would lead to loss or degraded performance (i.e., scenarios involving undesired consequences of interest), 2) how likely is it (probabilities), and 3) what is the severity of the degradation (consequences). Since the Challenger accident, PRA has been used in supporting decisions regarding safety upgrades for launch vehicles. Another area that was given a lot of emphasis at NASA after the Challenger accident is reliability engineering. Reliability engineering has been a critical design function at NASA since the early Apollo days. However, after the Challenger accident, quantitative reliability analysis and reliability predictions were given more scrutiny because of their importance in understanding failure mechanism and quantifying the probability of failure, which are key elements in resolving technical issues, performing design trades, and implementing design improvements. Although PRA and reliability are both probabilistic in nature and, in some cases, use the same tools, they are two different activities. Specifically, reliability engineering is a broad design discipline that deals with loss of function and helps understand failure mechanism and improve component and system design. PRA is a system scenario based risk assessment process intended to assess the risk scenarios that could lead to a major/top undesirable system event, and to identify those scenarios that are high-risk drivers. PRA output is critical to support risk informed decisions concerning system design. This paper describes the PRA process and the reliability engineering discipline in detail. It discusses their differences and similarities and how they work together as complementary analyses to support the design and risk assessment processes. Lessons learned, applications, and case studies in both areas are also discussed in the paper to demonstrate and explain these differences and similarities.

Safie, Fayssal↗

Reliability and Probabilistic Risk Assessment - How They Play Together

PRA methodology is one of the probabilistic analysis methods that NASA brought from the nuclear industry to assess the risk of LOM, LOV and LOC for launch vehicles. PRA is a system scenario based risk assessment that uses a combination of fault trees, event trees, event sequence diagrams, and probability and statistical data to analyze the risk of a system, a process, or an activity. It is a process designed to answer three basic questions: What can go wrong? How likely is it? What is the severity of the degradation? Since 1986, NASA, along with industry partners, has conducted a number of PRA studies to predict the overall launch vehicles risks. Planning Research Corporation conducted the first of these studies in 1988. In 1995, Science Applications International Corporation (SAIC) conducted a comprehensive PRA study. In July 1996, NASA conducted a two-year study (October 1996 - September 1998) to develop a model that provided the overall Space Shuttle risk and estimates of risk changes due to proposed Space Shuttle upgrades. After the Columbia accident, NASA conducted a PRA on the Shuttle External Tank (ET) foam. This study was the most focused and extensive risk assessment that NASA has conducted in recent years. It used a dynamic, physics-based, integrated system analysis approach to understand the integrated system risk due to ET foam loss in flight. Most recently, a PRA for Ares I launch vehicle has been performed in support of the Constellation program. Reliability, on the other hand, addresses the loss of functions. In a broader sense, reliability engineering is a discipline that involves the application of engineering principles to the design and processing of products, both hardware and software, for meeting product reliability requirements or goals. It is a very broad design-support discipline. It has important interfaces with many other engineering disciplines. Reliability as a figure of merit (i.e. the metric) is the probability that an item will perform its intended function(s) for a specified mission profile. In general, the reliability metric can be calculated through the analyses using reliability demonstration and reliability prediction methodologies. Reliability analysis is very critical for understanding component failure mechanisms and in identifying reliability critical design and process drivers. The following sections discuss the PRA process and reliability engineering in detail and provide an application where reliability analysis and PRA were jointly used in a complementary manner to support a Space Shuttle flight risk assessment.

Safie, Fayssal M.↗

The Containment Assurance Risk Framework of the Mars Sample Return Program

The Mars Sample Return campaign aims at bringing rock and atmospheric samples from Mars to Earth through a series of robotic missions. These missions would collect the samples being cached and deposited on Martian soil by the Perseverance rover, place them in a container, and launch them into Martian orbit for subsequent capture by an orbiter that would bring them back. Given there exists a non-zero probability that the samples contain biological material, precautions are being taken to design systems that would break the chain of contact between Mars and Earth. These include techniques such as sterilization of Martian particles, redundant containment vessels, and a robust reentry capsule capable of accurate landings without a parachute. Requirements exist that the probability of containment not assured of Martian-contaminated material into Earth’s biosphere be less than one in a million. To demonstrate compliance with this strict requirement, a statistical framework was developed to assess the likelihood of containment loss during each sample return phase and make a statement about the total combined mission probability of containment not assured. The work presented here describes this framework, which considers failure modes or fault conditions that can initiate failure sequences ultimately leading to containment not assured. Reliability estimates are generated from databases, design heritage, component specifications, or expert opinion in the form of probability density functions or point estimates and provided as inputs to the mathematical models that simulate the different failure sequences. The probabilistic outputs are then combined following the logic of several fault trees to compute the ultimate probability of containment not assured. Given the multidisciplinary nature of the problem and the different types of mathematical models used, the statistical tools needed for analysis are required to be computationally efficient. While standard Monte Carlo approaches are used for fast models, a multi-fidelity approach to rare event probabilities is proposed for expensive models. In this paradigm, inexpensive low-fidelity models are developed for computational acceleration purposes while the expensive high-fidelity model is kept in the loop to retain accuracy in the results. This work presents an example of end-to-end application of this framework highlighting the computational benefits of a multi-fidelity approach.

Giuseppe Cataldo↗