Search NASA⌕ Search

SEARCH · Search NASA

Results for “fault tree”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 181 records · Page 10

System Validation on the Europa Clipper mission in Early Implementation Phase

NASA’s next flag-ship mission - Europa Clipper, will embark on a journey to Jupiter’s icy moon Europa in 2024 to assess its environment and habitability with a highly capable spacecraft. Post Jupiter-Orbit-Insertion, the spacecraft will be commanded to perform intricate, yet meticulously planned Europa flybys to perform science investigations using a suite of instruments, while withstanding Jupiter’s harsh radiation environment. The success of this mission is dependent on a well-coordinated project and its elements such as the flight hardware and software, the ground support and mission operations teams, procedures and other cross-cutting elements. The Europa Clipper project needs to ensure that these elements are realized at a reasonable confidence level prior to launch and other mission critical events. System Validation test and analysis activities exercise and confirm the integrity of the system of all project elements in the expected flight environment with reasonable stressing conditions. These activities go beyond system design requirements verification and are driven by validation objectives that describe the end-to-end functional and operational capabilities required during nominal and off-nominal flight-like scenarios and critical events. The challenges associated with validating that the Europa Clipper project as a whole can function and perform correctly to meet the intended mission objectives with the as-delivered capabilities of all of its elements are daunting. This paper discusses the systematic methodology established in the early implementation phase of the Europa Clipper project for developing System Validation activities and their validation objectives, and addressing any validation-related challenges on the project. Approaches include decomposition of mission objectives using activity timelines in the Mission Design plan for developing nominal scenarios, use of fault trees for exploring off-nominal cases and system boundaries, and use of Model-based Systems Engineering (MBSE) tools for planning and prioritizing these activities.

Wang, Xu↗

System Safety Analysis of Complex NASA Systems with Model Based Engineering

The emergence of model-based engineering is transforming design and analysis methodologies [5]. A recognized benefit of model-based engineering is the existence of a “single source of truth” about the system that becomes the authoritative source of data and information for designers, analysts, and developers. This promotes consistency and efficiency as the design emerges and can be used to further optimize the design. Integrating System Safety Engineers to the “single source of truth” will ensure that the outputs of their assessments and analyses are relevant to the design as it evolves. Use of an integrated system model enables near immediate evaluation of a design change as well as development of operational processes for risk assessment and communication. Such models can enable efficient and timely analysis of system hazards (e.g., hazard fault tree analysis and procedure simulations) and produce complete, accurate, and more consistent products (e.g., hazard reports and safety requirement evaluations). Therefore, an agency-sponsored team at Goddard Space Flight Center (GSFC) recently completed a System Safety Study of modeling and testing capabilities as part of a Model-Based Safety and Mission Assurance Initiative (MBSMAI). Using an existing model developed for reliability analyses [1], GSFC modeling and system safety experts performed system safety analysis/modeling and produced safety products. The team evaluated model-based feasibility to support System Safety Engineering, developed safety analysis modeling processes, and identified tool capability advancement/development needs. These study results indicate model-based engineering is valid and useable for System

Model Based Engineering↗

The Containment Assurance Risk Framework of the Mars Sample Return Program

The Mars Sample Return campaign aims at bringing rock and atmospheric samples from Mars to Earth through a series of robotic missions. These missions would collect the samples being cached and deposited on Martian soil by the Perseverance rover, place them in a container, and launch them into Martian orbit for subsequent capture by an orbiter that would bring them back. Given there exists a non-zero probability that the samples contain biological material, precautions are being taken to design systems that would break the chain of contact between Mars and Earth. These include techniques such as sterilization of Martian particles, redundant containment vessels, and a robust reentry capsule capable of accurate landings without a parachute. Requirements exist that the probability of containment not assured of Martian-contaminated material into Earth’s biosphere be less than one in a million. To demonstrate compliance with this strict requirement, a statistical framework was developed to assess the likelihood of containment loss during each sample return phase and make a statement about the total combined mission probability of containment not assured. The work presented here describes this framework, which considers failure modes or fault conditions that can initiate failure sequences ultimately leading to containment not assured. Reliability estimates are generated from databases, design heritage, component specifications, or expert opinion in the form of probability density functions or point estimates and provided as inputs to the mathematical models that simulate the different failure sequences. The probabilistic outputs are then combined following the logic of several fault trees to compute the ultimate probability of containment not assured. Given the multidisciplinary nature of the problem and the different types of mathematical models used, the statistical tools needed for analysis are required to be computationally efficient. While standard Monte Carlo approaches are used for fast models, a multi-fidelity approach to rare event probabilities is proposed for expensive models. In this paradigm, inexpensive low-fidelity models are developed for computational acceleration purposes while the expensive high-fidelity model is kept in the loop to retain accuracy in the results. This work presents an example of end-to-end application of this framework highlighting the computational benefits of a multi-fidelity approach.

Giuseppe Cataldo↗

Medical Resource Set Bulky Item Trade Space Analysis for Spaceflight Medical Risk

The NASA engineering community utilizes event-driven and fault-tree probabilistic techniques to classify risks in the space environment by taking advantage of the inherent knowledge of complex spaceflight system design and testing to quantify failure risk. In harmonizing the risk of human space flight, answering the question of ‘How do we balance health, performance and resource risks with other engineering risks on long duration space missions?’ remains a deeply challenging and largely qualitative practice. The Medical Extensible Dynamic Probabilistic Risk Assessment Tool (MEDPRAT) is one aspect of the efforts by NASA’s Human Research Program (HRP) to quantitatively assess the impact of health and performance risk. One of MEDPRAT’s key features is its high degree of computational efficiency. Coupled with the HRP High Performance Compute cluster located at NASA’s Glenn Research Center, MEDPRAT runs millions of simulated missions in a matter of minutes. This degree of computational efficiency provides the novel opportunity to explore the relationship between medical set mass, volume, and medical resource size. Of particular interest for future human spaceflight missions are ‘bulky’ items, medical resources like devices, which occupy a large portion of the small, allocated mass and volume for the medical set leaving less room for other resources. This talk will present results showing the quantitative impact of forced inclusion of several bulky items across a variety of medical kit constraints, and the effect that a potential research investment into reducing the bulky item mass and volume may have on risk.

Lauren Mcintyre↗

Quantifying the Sensitivity of Condition Incidence Parameters in the Evidence Library

One approach to quantifying spaceflight risk at NASA makes use event driven probabilistic techniques. The Medical Extensible Dynamic Probabilistic Risk Assessment Tool (MEDPRAT) is such a tool that estimates medical risk metrics via simulation and enables optimization of medical resources subject to mission constraints [1]. Previous analyses have informed medical set composition, exercise countermeasures, and water intake, where each analysis quantifies the risk associated with proposed variations in system design. As future mission profiles extend beyond Low-Earth Orbit (LEO) and lengthen in duration, understanding these risks and contributing factors is critical. MEDPRAT employs Monte Carlo sampling techniques to simulate missions and track the occurrence of medical events. These events follow fault-tree-like progressions through levels of severity and mitigation via medical treatment to many possible outcomes and these are reported throughout the mission. Making this possible, are the medical databases that contain evidence gathered by the Human Research Program (HRP). Quantifying the impact of uncertainty or variability in the input data is an important step in evaluating the credibility of modeling and simulation results. In this work, we investigate the sensitivity of medical risk metrics with respect to the condition incidence parameters within the Evidence Library (EL) [2] as the medical database input for MEDPRAT. The medical conditions, contained in the EL, are equipped with incidence rates that describe the likelihood that the condition will occur. These incidence rates reflect historical spaceflight data or when appropriate, terrestrial data. In this presentation, we will explore how uncertainty in these rates propagate to the medical risk described by MEDPRAT. These results identify the conditions and parameters with the largest contribution to medical risks.

Ian Lim↗

The Future of Integrated Performance Modeling in the Crew Health and Performance – Probabilistic Risk Assessment Project

The NASA engineering community utilizes event-driven and fault-tree probabilistic techniques to classify risks in the space environment by taking advantage of the inherent knowledge of complex spaceflight system design and testing to quantify failure risk. In harmonizing the risk of human space flight, answering the question of ‘How do we balance health, performance and resource risks with other engineering risks on long duration space missions?’ remains a deeply challenging and largely qualitative practice. The Human Research Program’s Medical Extensible Dynamic Probabilistic Risk Assessment Tool (MEDPRAT) was a significant step forward in efforts to robustly quantify the risk to crew health for exploration missions. However, there remains a significant gap in the ability to comprehensively assess and characterize risk across the disparate functionalities and capabilities which comprise the Crew Health and Performance (CHP) system. The Crew Health and Performance – Probabilistic Risk Assessment (CHP-PRA) project seeks to characterize CHP risks by expanding beyond the foundation established by its PRA predecessors like IMM and MEDPRAT, that simulate medical risk metrics like loss of crew life and evacuations. One of the new risk measures in the CHP-PRA system is embodied in our Performance Risk Model (PRisM). PRisM provides a novel way of assessing crew performance on mission tasks, using a generalized framework which relates back to NASA-STD-3001. This approach allows PRisM to capture and integrate data from a variety of different domains into a single, unified, reproducible representation of astronaut performance. In this presentation, we discuss the motivation for the CHP-PRA work and give a high level overview of the goals of the project, outline the forward work for PRisM, and discuss collaboration opportunities for the community who might explore if their domain knowledge and data could be represented, integrated, and quantified with these tools, whose outcomes are metrics useful for supporting operational mission planning and decision making.

Lauren McIntyre↗

Development and Validation of a Probabilistic Risk Assessment Model for a Generic Modular High Temperature Gas-Cooled Reactor

This study looks to develop and validate a probabilistic risk assessment (PRA) model for the modular high temperature gas-cooled reactor (MHTGR) using INL’s Systems Analysis Programs for Hands-on Integrated Reliability Evaluations (SAPHIRE) software. Validation against a General Atomics design involves matching event and fault trees to historical frequencies, with a goal of under 15% difference. The research aims to deliver a reliable PRA model to assess the safety of MHTGRs for use within high-temperature industrial applications.

22 GENERAL STUDIES OF NUCLEAR REACTORS↗

An Evaluation of The Dynamic Physical Security Risk Assessment Methodology for Fleet-Wide Applications

The requirements for U.S. nuclear power plants to maintain a large onsite physical security force contribute to their high operational costs. The cost of maintaining the current physical security posture is approximately 10% of the overall operation and maintenance budget for commercial nuclear power plants. The goal of the Light Water Reactor Sustainability (LWRS) program’s physical security pathway is to develop tools, methods, and technologies and provide the technical basis for an optimized physical security posture. The conservatisms built into current security postures may be analyzed and minimized to reduce security costs while still ensuring adequate security and operational safety. The research performed at Idaho National Laboratory within LWRS program’s physical security pathway has successfully developed a dynamic force-on-force modeling framework using various computer simulation tools and integrating them with the dynamic assessment Event Modeling Risk Assessment using Linked Diagrams (EMRALD) tool. This integrated process for physical security analysis is named Modeling and Analysis for Safety Security using Dynamic EMRALD Framework (MASS-DEF). This document provides an update on the progress in applying the MASS-DEF process to an operating commercial nuclear power plant as well as additional industry feedback regarding use of the tool for other physical security risk-informed topics. This report is only a summary of the progress and does not contain specific modeling results as those contain sensitive security information. Previous reports described how a user could integrate their plant-specific force-on-force models with the dynamic simulation tool EMRALD, model operator actions, and integrate with probabilistic risk assessment tools, such as CAFTA (Computer Aided Fault Tree Analysis System) or SAPHIRE (Systems Analysis Programs for Hands-on Integrated Reliability Evaluations), and with thermal-hydraulic tools, such as RELAP-5 or MAAP. Previous reports applied various combinations of available simulations codes with EMRALD using generic plant models to demonstrate how to perform the analysis. This report is an update the progress of applying the dynamic computational framework to an actual nuclear facility using their security scenarios and timelines. This report also provides an update to the procedural guidance for the MASS-DEF process and an overview of the generic models available for use by utilities. This report does not contain any plant’s sensitive information and/or safeguards information. This study’s purpose was to verify that the results achieved using generic models are similar to actual plant results and refine our guidance on the use of the framework. This assessment enables further analysis, such as what-if scenarios and staff-reduction evaluation, thereby optimizing physical security at plants.

22 GENERAL STUDIES OF NUCLEAR REACTORS↗

Integration of Condition-Based, Diagnostic, Prognostic, And Anomaly Detection Data into Reliability Models to Support a Predictive Maintenance Context

Reliability data employed in plant reliability models are an approximated integral representation of the past industrywide operational experience, and they neglect the present asset health status (available, for example, from online monitoring data and diagnostic assessments) and forecasted health projection (when available from prognostic models). Ideally, in a predictive maintenance context, system reliability models should support decision making by propagating actual health information from the asset to the system level in order to provide a quantitative snapshot of system health and identify the most critical assets. Asset health should be informed solely by that specific asset’s current and historical performance data and should not be an approximated integral representation of the past industrywide operational experience (as currently performed by system reliability models through Bayesian updating processes). This paper proposes a reliability modeling approach that relies on asset diagnostic and prognostic assessments, along with monitoring data to measure asset health. We show how state-of-the art condition-based, diagnostic, prognostic, and anomaly detection models can be linked to system reliability models not in probability terms, but in terms of margin where margin is defined as the “distance” between the present status and an undesired event (e.g., failure or unacceptable performance). Then, we show how the propagation of margin data from the asset to the system level is performed through classical reliability models such as fault trees or reliability block diagrams. The described method is in fact able to propagate heterogenous health data from the asset to the system level in order to analytically assess system health.

97 MATHEMATICS AND COMPUTING↗

ON THE LANGUAGE OF RELIABILITY: A SYSTEM ENGINEER PERSPECTIVE

In its classical definition, risk is defined by three elements: what can go wrong, what are its consequences and how likely is it to occur. While this definition makes sense in a regulatory based framework to estimate risk associated to power plants (in terms of core damage frequency and large early release frequency), this approach does not provide a useful snapshot of the health of the plant. A possible alternate path can start by redefining the word “risk” to a broader meaning that better reflects the needs of a system health and asset management decision making process. Rather than asking how likely an event can occur (in probabilistic terms), we can ask how far this event is from occurring. We will show how, given the data available from plant equipment reliability and monitoring/diagnostic/prognostic centers, a margin can be described and determined for all type of maintenance approaches (e.g., corrective or predictive maintenance). We will show how to link SSC margin-based reliability models to system reliability models (i.e., fault trees) in order to assess system/plant health and how to perform margin-based system calculations. These calculations are not solved using classical probabilistic calculations applied to sets (as performed by any PRA code) but, instead, through metric spaces operations (i.e., distance/margin based approach).

97 - MATHEMATICS AND COMPUTING↗

Advancing Multi-Hazard Risk and Safety Considerations for Aging Nuclear Facilities

While probabilistic risk assessment (PRA) of nuclear facilities is expected to include internal and external hazards for a risk-informed and performance-based design, the current state of practice treats each hazard independently. However, such an independent treatment of hazards may not account for the correlations between different hazards and their response of and damage to the structures, systems, and components (SSCs) in a plant resulting in underestimating the overall risk. This project proposes to advance the multi-hazard PRA of nuclear facilities to more adequately evaluate concurrent hazards and contribute to an increased safety of nuclear plants. A framework for multi-hazard PRA will be developed by identifying concurrent hazard events (both internal and external) and event sequences that include interdependencies through the response of SSCs. An example application of the multi-hazard PRA framework will be demonstrated by considering a generic pressurized water reactor (PWR) subjected to seismic and internal flooding hazards. Computational models for the response of components will be developed to generated multi-hazard fragility surfaces under seismic and flooding loads. A PRA model consisting of event and fault trees will also be developed to quantify the multi-hazard risk profile and compare it with the independent hazard risk profile. Overall, by advancing the multi-hazard PRA of nuclear facilities, this project enhances nuclear safety and reduces costs by mitigating unforeseen consequences caused by correlations between concurrent hazards.

97 - MATHEMATICS AND COMPUTING↗

Reconfigurable tree architectures using subtree oriented fault tolerance

An approach to the design of reconfigurable tree architecture is presented in which spare processors are allocated at the leaves. The approach is unique in that spares are associated with subtrees and sharing of spares between these subtrees can occur. The Subtree Oriented Fault Tolerance (SOFT) approach is more reliable than previous approaches capable of tolerating link and switch failures for both single chip and multichip tree implementations while reducing redundancy in terms of both spare processors and links. VLSI layout is 0(n) for binary trees and is directly extensible to N-ary trees and fault tolerance through performance degradation.

Lowrie, Matthew B.↗

Reconfigurable tree architectures using subtree oriented fault tolerance

An approach to the design of reconfigurable tree architectures is presented in which spare processors are allocated at the leaves. The approach is unique in that spares are associated with subtrees, and sharing of spares between these subtrees can occur. The subtree-oriented fault-tolerance approach is more reliable than previous approaches capable of tolerating link and switch failures for both single-chip and multichip tree implementations while reducing redundancy in terms of both spare processors and links. VLSI layout is O(n) for binary trees and is directly extensible to N-ary trees and fault tolerance through performance degradation.

Lowrie, Matthew B.↗

Chemical hazards database and detection system for Microgravity and Materials Processing Facility (MMPF)

The ability to identify contaminants associated with experiments and facilities is directly related to the safety of the Space Station. A means of identifying these contaminants has been developed through this contracting effort. The delivered system provides a listing of the materials and/or chemicals associated with each facility, information as to the contaminant's physical state, a list of the quantity and/or volume of each suspected contaminant, a database of the toxicological hazards associated with each contaminant, a recommended means of rapid identification of the contaminants under operational conditions, a method of identifying possible failure modes and effects analysis associated with each facility, and a fault tree-type analysis that will provide a means of identifying potential hazardous conditions related to future planned missions.

Steele, Jimmy↗

Overcoming obstacles to the exchange of information between risk tools

Our work to date in connecting risk tools hs had successes, but also has revealed there to be significant impediments to information exchange between them. These impediments stem from the well-known phenomenon of 'semantic dissonance' - mismatch between conceptual assumptions made by the separately developed tools. This issue represents a fundamental challenge that arises regardless of the mechanism of information exchange. This paper explains the issue and illustrates it with reference to our experiences to date connecting several risk tools. We motivate this work, present and discuss the solutions we have adopted to surmount these impediments, and the implications this work has for future efforts to integrate risk tools.

semantic dissonance↗

Modeling uncertainty in requirements engineering decision support

One inherent characteristic of requrements engineering is a lack of certainty during this early phase of a project. Nevertheless, decisions about requirements must be made in spite of this uncertainty. Here we describe the context in which we are exploring this, and some initial work to support elicitation of uncertain requirements, and to deal with the combination of such information from multiple stakeholders.

risk analysis↗

Probabilistic Risk Assessment for Decision Making During Spacecraft Operations

Decisions made during the operational phase of a space mission often have significant and immediate consequences. Without the explicit consideration of the risks involved and their representation in a solid model, it is very likely that these risks are not considered systematically in trade studies. Wrong decisions during the operational phase of a space mission can lead to immediate system failure whereas correct decisions can help recover the system even from faulty conditions. A problem of special interest is the determination of the system fault protection strategies upon the occurrence of faults within the system. Decisions regarding the fault protection strategy also heavily rely on a correct understanding of the state of the system and an integrated risk model that represents the various possible scenarios and their respective likelihoods. Probabilistic Risk Assessment (PRA) modeling is applicable to the full lifecycle of a space mission project, from concept development to preliminary design, detailed design, development and operations. The benefits and utilities of the model, however, depend on the phase of the mission for which it is used. This is because of the difference in the key strategic decisions that support each mission phase. The focus of this paper is on describing the particular methods used for PRA modeling during the operational phase of a spacecraft by gleaning insight from recently conducted case studies on two operational Mars orbiters. During operations, the key decisions relate to the commands sent to the spacecraft for any kind of diagnostics, anomaly resolution, trajectory changes, or planning. Often, faults and failures occur in the parts of the spacecraft but are contained or mitigated before they can cause serious damage. The failure behavior of the system during operations provides valuable data for updating and adjusting the related PRA models that are built primarily based on historical failure data. The PRA models, in turn, provide insight into the effect of various faults or failures on the risk and failure drivers of the system and the likelihood of possible end case scenarios, thereby facilitating the decision making process during operations. This paper describes the process of adjusting PRA models based on observed spacecraft data, on one hand, and utilizing the models for insight into the future system behavior on the other hand. While PRA models are typically used as a decision aid during the design phase of a space mission, we advocate adjusting them based on the observed behavior of the spacecraft and utilizing them for decision support during the operations phase.

dynamic fault trees↗

Architectural Modeling and Analysis for Safety Engineering

Model-based development tools are increasingly being used for system-level development of safety-critical systems. Architectural and behavioral models provide important information that can be leveraged to improve the system safety analysis process. Model-based design artifacts produced in early stage development activities can be used to perform system safety analysis, reducing costs and providing accurate results throughout the system life-cycle. In this report we describe an extension to the Architecture Analysis and Design Language (AADL) that supports modeling of system behavior under failure conditions. This Safety Annex enables the independent modeling of component failures and allows safety engineers to weave various types of fault behavior into the nominal system model. The accompanying tool support uses model checking to propagate errors from their source to their effect on safety properties without the need to add separate propagation specifications. The tool also captures all minimal set of fault combinations that can cause violation of the safety properties, that can be compared to qualitative and quantitative objectives as part of the safety assessment process. We describe the Safety Annex, illustrate its use with a representative example, and discuss and demonstrate the tool support enabling an analyst to investigate the system behavior under failure conditions.

FTA↗