Search NASASearch

SEARCH · Search NASA

Results for “fault tree”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

AUTOMATIC GENERATION OF EVENT TREES AND FAULT TREES: A MODEL-BASED APPROACH

In the past few decades, increasing complexity in modern engineering systems has been driven by the integration of a large number of components and by the fact that the system operations involve many disciplines (e.g., thermal-hydraulics, plant operations, cyber-security). Current safety/reliability modeling approaches to such systems are labor intensive, difficult to learn, and rely heavily on simplistic Boolean logic to depict failure propagation and accident progression. While these methods serve well for simple systems (i.e., linear causal systems with limited small inter- and intra-system interactions), their results are difficult to verify when modeling complex systems (typically performed through the extensive use of modeling assumptions). The development of new methods is addressed to meet these challenges through a model-based system engineering (MBSE) lens. Under MBSE philosophy, every aspect of the system (form or function) is represented by a model that completely characterizes its architecture or behavior. MBSE approach greatly improves the management of design, analysis and verification of complex systems. An integration of Dynamic Probabilistic Risk Assessment (DPRA) methods with MBSE models is proposed to perform safety/reliability analyses of engineering systems. In particular, MBSE representation of the system (performed using Systems Modeling Language [SysML]) is coupled with DPRA methods to automatically generate event trees and fault trees.

97 - MATHEMATICS AND COMPUTING

Eucalyptus – An Analysis Suite for Fault Trees with Uncertainty Quantification

Eucalyptus is a novel code developed at Lawrence Livermore National Laboratory to incorporate uncertainty quantification into Fault Tree Analysis (FTA). This tool addresses the challenge of imperfect knowledge in “grey-box” systems by allowing analysts to incorporate and propagate uncertainty from component-level assessments to system-level effects. Eucalyptus facilitates a consistent evaluation of the impact of subject matter expert judgment and knowledge gaps on overall system response by Monte Carlo generation of possible system fault trees, sampling probabilities of the existence of subsystems and components. Here, the code supports the specification of fault trees through text and allows export to various formats, including auto-generated images, easing analysis and reducing errors. It has undergone extensive verification testing, demonstrating its reliability and readiness for deployment, and leverages on-node parallelism for rapid analysis. Example analyses are shown that include the identification of system failure paths and quantification of the value of further information about system components.

Fault Tree Analysis

Redefining Design for Remanufacturing: A Practical Methodology for Prioritizing Remanufacturing Design Rules

Products are often discarded when they fail or no longer meet user needs. These outcomes are frequently shaped by early design decisions. While remanufacturing offers a sustainable alternative by restoring products to like‐new condition, its potential is often limited by designs that do not consider remanufacturing from the outset. This research addresses that challenge by introducing a structured Design for Remanufacturing (DfRem) methodology and a CAD‐integrated tool to support real‐time design decisions. The DfRem framework introduces a new primary design function focused on preserving product functionality across its life cycle. It is supported by a fault tree that identifies failure modes that limit remanufacturing potential and a hierarchy of design principles including Prevent, Minimize, Relocate, Restore, and others. Each principle is linked to actionable design rules that help engineers reduce the need for remanufacturing or improve its efficiency when necessary. To operationalize this framework, we developed CAD plugins for Autodesk Inventor and PTC Creo. These tools use a state machine model to present prioritized design rules based on selected failure modes and user input. By embedding DfRem logic directly into widely used CAD environments, the tool enables engineers to make sustainability‐informed decisions without disrupting existing workflows. Furthermore, this approach highlights the critical role of design in enabling circular and resource‐efficient product development, making remanufacturing a more practical and accessible strategy during the early stages of product design.

CAD

Development and Demonstration of a Prototype Molten Salt Sampling System

Molten salt reactors (MSRs) offer potential operability and safety advantages when compared to commercial light water reactors (LWRs). However, operating experience with MSRs is sparse in comparison to what exists for LWRs. Further, the chemical and isotopic composition of the fuel and/or coolant salt is dynamic and difficult to characterize continuously, posing potential safety, operability, and safeguards unknowns that need to be addressed. A molten salt sampling system (MSSS) is regarded as a necessary subsystem within first generation MSRs used to obtain samples of salt for chemical and isotopic analysis in support of the need to monitor and control salt composition during operation. The MSSS is being developed using the Safety-in-Design (SiD) methodology, which incorporates incremental integration of safety analysis into the design process. The MSSS conceptual design emerging from the application of the early stages of the SiD methodology consists of a sample collection system and its housing, a freeze port, and inert gas control and delivery systems. This article describes the prototypes developed to test the functions of these MSSS subsystems, presents the results of testing in both dry and molten salt environments (including reliability data collection performed in accordance with the principles of SiD and the development of a semiquantitative fault tree model), and summarizes the opportunities for future design and testing enhancements based on the results of prototype testing.

molten salt reactor

Success Path Method: Introduction to the Success Path Method Software Tool©

As part of its commitment to advancing safety and reliability assessment methodologies, Argonne National Laboratory pioneered the use of an evaluation method called the Success Path Method (SPM) to improve risk management for offshore oil and gas operations. The development of the SPM at Argonne has been driven by the need to improve existing risk assessment methodologies by focusing on the steps necessary for success rather than failure modes alone. This is particularly important for industrial environments like offshore facilities that perform multiple functions under a continuously evolving set of operational conditions – such as water depth and temperature, currents, and weather conditions. In these dynamic environments, the traditional Probabilistic Risk Assessment (PRA) approach is far too complex as it focuses on what can go wrong – which comprises an infinite failure space that must be fully explored and understood. By shifting the focus to a finite space of success paths, the SPM enables operators and decision makers to prioritize a manageable number of steps that must go right to ensure success. Building on its five decades of experience in safety assessments for the nuclear industry, Argonne made major adaptations to existing risk assessment methods utilizing features similar to fault trees that are traditionally used in PRA to map all pathways in which the system can malfunction. In contrast, SPM identifies the components and processes that must function correctly to achieve specific outcomes – such as preventing the uncontrolled release of hydrocarbons during drilling operations. The SPM framework integrates equipment, procedures, software, processes, and human actions to ensure that physical barriers meet critical safety functions in dynamic operational conditions. This approach helps identify failure modes and improve operational risk management by narrowing the focus to key success elements, which in turn reduces uncertainty and helps users understand, manage, and respond to failures.

97 MATHEMATICS AND COMPUTING

Quantitative Risk Assessment for Fuel Cell Electric Bus Hydrogen Storage and Refueling Facility

It is necessary to understand the safety implications and risk mitigation options for fuel cell electric bus fleet deployment, especially for related facilities responsible for operations such as production, storage, compression, and dispensing of hydrogen for use by the buses. In this report, we present a quantitative risk assessment for a potential fuel cell electric bus fleet that was motivated by efforts to improve resilience at the Portland International Airport but can be applicable to a range of hydrogen case studies and use cases. We estimated risk for a facility that produces, stores, compresses, and dispenses hydrogen for the fleet of buses, with a focus on individual risk to people in terms of annual frequency of fatality. We considered the frequency of hydrogen leaks that could result in harmful physical outcomes like jet fires or explosions, and the consequences of those outcomes for people. We created customized fault trees to calculate the frequencies of different sizes of leaks and event sequence diagrams to calculate ignition probabilities for the various leak sizes. We also leveraged the HyRAM+ toolkit to use these inputs to calculate overall risk for the facility, which we separated into one section responsible for producing, storing, and compressing hydrogen, and one section responsible for dispensing the hydrogen to the buses. We found that the dispensing area seemed to have a higher risk than the production/storage/compression area of the facility, largely because of the inclusion of a component with a high leak frequency (the heat exchanger used to cool the hydrogen before entering the vehicle, to prevent overheating and expansion of hydrogen in the onboard tank). For the example production and refueling facility we evaluated and the data we used for the analysis, the leak frequency had a larger impact on the risk differences between the two sections on the facility, compared to the physical outcome consequence, which was slightly different due to the varying fuel conditions, but not substantially different. Actions can be taken to prevent these hazards (e.g., lowering leak frequencies in system components) or to mitigate the consequences if they do occur (e.g., installing barriers to protect people if ignition events occur). The choice of which actions to take depends not only on safety considerations but also on space, time, staffing, feasibility, and financial constraints. Therefore, the quantitative risk assessment approach can help understand relative risk contributions from different components, leak sizes, consequences, and human actions, to prioritize risk reduction strategies and balance these parameters. The outcomes of this report may be useful for a variety of stakeholders working in the hydrogen, transportation, vehicle, and aviation sector, including those responsible for aspects like facility design, operations, and regulations. There is not a single value of risk that determines whether a hypothetical system is “safe” or not. The insights about risk mitigations may be leveraged, and the quantitative risk assessment approach can be applied to other case studies to understand risk priorities and contributions specific to different FCEB and hydrogen facility uses.

08 HYDROGEN

Probability of Hydrogen Ignition: A Landscape Review and Gaps Assessment

The primary hazard of a leak from a hydrogen system is due to the immediate or delayed ignition of the fuel leading to a jet flame or explosion. Therefore, understanding the hydrogen ignition probability is critical for analyzing the risk of hydrogen systems. This report reviews the current understanding of hydrogen ignition mechanisms and methods for modeling their probability. The stoichiometry, ignition strength, and ignition source temperature are all important characteristics that can affect both the probability of ignition and the outcome of the subsequent combustion event. A brief review of diffusion ignition demonstrates that ignition probability models must account for seemingly spontaneous ignition of hydrogen in addition to scenarios where the ignition source is readily identified. State-of-the art models for both immediate and delayed ignition probabilities are presented, including different physical aspects of the scenarios (e.g., flow rate, ignition source characteristics) that are considered in the different modeling approaches. Current models often fail to account for the unique properties of hydrogen compared to other fuels, and most lack rigorous validation with hydrogen as a fuel. A fault tree framework is proposed to systematically evaluate the probability of ignition by integrating various ignition mechanisms and their uncertainties. Furthermore, this type of framework could enable additional insights into the most important mechanisms and would enable uncertainty quantification in risk assessment modeling. Recommendations for future research include the need for experimental validation of ignition models and the development of comprehensive methodologies that incorporate the specifics of hydrogen behavior in real-world scenarios.

hydrogen

Development and Validation of a Probabilistic Risk Assessment Model for a Generic Modular High Temperature Gas-Cooled Reactor

This study looks to develop and validate a probabilistic risk assessment (PRA) model for the modular high temperature gas-cooled reactor (MHTGR) using INL’s Systems Analysis Programs for Hands-on Integrated Reliability Evaluations (SAPHIRE) software. Validation against a General Atomics design involves matching event and fault trees to historical frequencies, with a goal of under 15% difference. The research aims to deliver a reliable PRA model to assess the safety of MHTGRs for use within high-temperature industrial applications.

22 GENERAL STUDIES OF NUCLEAR REACTORS

An Evaluation of The Dynamic Physical Security Risk Assessment Methodology for Fleet-Wide Applications

The requirements for U.S. nuclear power plants to maintain a large onsite physical security force contribute to their high operational costs. The cost of maintaining the current physical security posture is approximately 10% of the overall operation and maintenance budget for commercial nuclear power plants. The goal of the Light Water Reactor Sustainability (LWRS) program’s physical security pathway is to develop tools, methods, and technologies and provide the technical basis for an optimized physical security posture. The conservatisms built into current security postures may be analyzed and minimized to reduce security costs while still ensuring adequate security and operational safety. The research performed at Idaho National Laboratory within LWRS program’s physical security pathway has successfully developed a dynamic force-on-force modeling framework using various computer simulation tools and integrating them with the dynamic assessment Event Modeling Risk Assessment using Linked Diagrams (EMRALD) tool. This integrated process for physical security analysis is named Modeling and Analysis for Safety Security using Dynamic EMRALD Framework (MASS-DEF). This document provides an update on the progress in applying the MASS-DEF process to an operating commercial nuclear power plant as well as additional industry feedback regarding use of the tool for other physical security risk-informed topics. This report is only a summary of the progress and does not contain specific modeling results as those contain sensitive security information. Previous reports described how a user could integrate their plant-specific force-on-force models with the dynamic simulation tool EMRALD, model operator actions, and integrate with probabilistic risk assessment tools, such as CAFTA (Computer Aided Fault Tree Analysis System) or SAPHIRE (Systems Analysis Programs for Hands-on Integrated Reliability Evaluations), and with thermal-hydraulic tools, such as RELAP-5 or MAAP. Previous reports applied various combinations of available simulations codes with EMRALD using generic plant models to demonstrate how to perform the analysis. This report is an update the progress of applying the dynamic computational framework to an actual nuclear facility using their security scenarios and timelines. This report also provides an update to the procedural guidance for the MASS-DEF process and an overview of the generic models available for use by utilities. This report does not contain any plant’s sensitive information and/or safeguards information. This study’s purpose was to verify that the results achieved using generic models are similar to actual plant results and refine our guidance on the use of the framework. This assessment enables further analysis, such as what-if scenarios and staff-reduction evaluation, thereby optimizing physical security at plants.

22 GENERAL STUDIES OF NUCLEAR REACTORS

Integration of Condition-Based, Diagnostic, Prognostic, And Anomaly Detection Data into Reliability Models to Support a Predictive Maintenance Context

Reliability data employed in plant reliability models are an approximated integral representation of the past industrywide operational experience, and they neglect the present asset health status (available, for example, from online monitoring data and diagnostic assessments) and forecasted health projection (when available from prognostic models). Ideally, in a predictive maintenance context, system reliability models should support decision making by propagating actual health information from the asset to the system level in order to provide a quantitative snapshot of system health and identify the most critical assets. Asset health should be informed solely by that specific asset’s current and historical performance data and should not be an approximated integral representation of the past industrywide operational experience (as currently performed by system reliability models through Bayesian updating processes). This paper proposes a reliability modeling approach that relies on asset diagnostic and prognostic assessments, along with monitoring data to measure asset health. We show how state-of-the art condition-based, diagnostic, prognostic, and anomaly detection models can be linked to system reliability models not in probability terms, but in terms of margin where margin is defined as the “distance” between the present status and an undesired event (e.g., failure or unacceptable performance). Then, we show how the propagation of margin data from the asset to the system level is performed through classical reliability models such as fault trees or reliability block diagrams. The described method is in fact able to propagate heterogenous health data from the asset to the system level in order to analytically assess system health.

97 MATHEMATICS AND COMPUTING

ON THE LANGUAGE OF RELIABILITY: A SYSTEM ENGINEER PERSPECTIVE

In its classical definition, risk is defined by three elements: what can go wrong, what are its consequences and how likely is it to occur. While this definition makes sense in a regulatory based framework to estimate risk associated to power plants (in terms of core damage frequency and large early release frequency), this approach does not provide a useful snapshot of the health of the plant. A possible alternate path can start by redefining the word “risk” to a broader meaning that better reflects the needs of a system health and asset management decision making process. Rather than asking how likely an event can occur (in probabilistic terms), we can ask how far this event is from occurring. We will show how, given the data available from plant equipment reliability and monitoring/diagnostic/prognostic centers, a margin can be described and determined for all type of maintenance approaches (e.g., corrective or predictive maintenance). We will show how to link SSC margin-based reliability models to system reliability models (i.e., fault trees) in order to assess system/plant health and how to perform margin-based system calculations. These calculations are not solved using classical probabilistic calculations applied to sets (as performed by any PRA code) but, instead, through metric spaces operations (i.e., distance/margin based approach).

97 - MATHEMATICS AND COMPUTING

Advancing Multi-Hazard Risk and Safety Considerations for Aging Nuclear Facilities

While probabilistic risk assessment (PRA) of nuclear facilities is expected to include internal and external hazards for a risk-informed and performance-based design, the current state of practice treats each hazard independently. However, such an independent treatment of hazards may not account for the correlations between different hazards and their response of and damage to the structures, systems, and components (SSCs) in a plant resulting in underestimating the overall risk. This project proposes to advance the multi-hazard PRA of nuclear facilities to more adequately evaluate concurrent hazards and contribute to an increased safety of nuclear plants. A framework for multi-hazard PRA will be developed by identifying concurrent hazard events (both internal and external) and event sequences that include interdependencies through the response of SSCs. An example application of the multi-hazard PRA framework will be demonstrated by considering a generic pressurized water reactor (PWR) subjected to seismic and internal flooding hazards. Computational models for the response of components will be developed to generated multi-hazard fragility surfaces under seismic and flooding loads. A PRA model consisting of event and fault trees will also be developed to quantify the multi-hazard risk profile and compare it with the independent hazard risk profile. Overall, by advancing the multi-hazard PRA of nuclear facilities, this project enhances nuclear safety and reduces costs by mitigating unforeseen consequences caused by correlations between concurrent hazards.

97 - MATHEMATICS AND COMPUTING

Understanding drivers of oil and gas well integrity issues in the greater wattenberg area of Colorado

Well integrity is critically important to maintain to minimize the environmental impacts of oil and gas development and other subsurface energy operations. The Wattenberg Field of Colorado—a top producing field with >40,000 wells—has one of the most robust publicly reported well integrity programs in the country. Here, in this study, we analyzed annular pressure and annular-fluid geochemical test results collected from Wattenberg wells through the end of 2019 to characterize the frequency and spatial variability of integrity issues in the field and understand their drivers. Estimated frequencies of integrity issues among tested wells were 8.2-17.1% between 1955 and 2019 and 6.1-11.4% in 2019 alone. The frequency of integrity issues was nearly four times greater in wells located above the Longmont Wrench Fault Zone. Potential drivers of integrity issues were identified using ensemble decision tree models trained with a broad set of relevant information. Models show that well integrity issues are spatially clustered on regional and sub-regional scales and suggest the relatively high frequency of integrity issues observed is likely attributed to geologic factors. These findings are valuable for regulatory agencies and operators seeking to inform well integrity monitoring, plugging, and emissions reduction efforts and design future subsurface energy projects.

03 NATURAL GAS

Field-based AFDD for refrigerant undercharge in residential HVAC systems: enhancing reliability through false alarm mitigation

This study evaluated rule-based and machine learning (ML) based automated fault detection and diagnostics (AFDD) algorithms for detecting refrigerant undercharge faults in residential heating, ventilation, and air conditioning (HVAC) systems, using actual building data and a minimal set of features. The ML-based algorithms included Decision Tree (DT) and K-Nearest Neighbors (KNN). Both the rule-based and ML-based algorithms demonstrated the capability to detect refrigerant undercharge faults of -30% or more. Both types of algorithms exhibited false alarms before the implementation of a false alarm mitigation algorithm, which motivated the development of such a mitigation strategy. After applying the mitigation, false alarms were substantially reduced, with the rule-based algorithm decreasing to 0.6% and the ML-based algorithms reaching 0%, while maintaining strong detection performance. Although the rule-based algorithm initially showed lower performance compared to the ML-based algorithms, its detection accuracy improved after mitigation to a level comparable to the ML-based algorithms. These results confirm that combining false alarm mitigation with both rule-based and ML-based AFDD algorithms significantly enhances practical reliability while preserving robust fault detection capabilities. Furthermore, the findings demonstrate the potential for field deployment of these algorithms in residential HVAC systems and highlight the importance of minimizing false alarms.

False Alarm

Seismic Features Predict Ground Motions During Repeating Caldera Collapse Sequence

Abstract Applying machine learning to continuous acoustic emissions, signals previously deemed noise, from laboratory faults and slowly slipping subduction‐zone faults, demonstrates hidden signatures are emitted that describe physical details, including fault displacement and friction. However, no evidence currently exists to demonstrate that similar hidden signals occur during seismogenic stick‐slip on earthquake faults—the damaging earthquakes of most societal interest. We show that continuous seismic emissions emitted during the 2018 multi‐month caldera collapse sequence at the Kı̄lauea volcano in Hawai'i contain hidden signatures characterizing the earthquake cycle. Multi‐spectral data features extracted from 30 s intervals of the continuous seismic emission are used to train a gradient boosted tree regression model to predict the GNSS‐derived contemporaneous surface displacement and time‐to‐failure of the upcoming collapse event. This striking result suggests that at least some faults emit such signals and provide a potential path to characterizing the instantaneous and future behavior of earthquake faults.

58 GEOSCIENCES

Noisy quantum trees: infinite protection without correction

We study quantum networks with tree structures, in which information propagates from a root to leaves. At each node in the network, the received qubit unitarily interacts with fresh ancilla qubits, after which each qubit is sent through a noisy channel to a different node in the next level. Therefore, as the tree depth grows, there is a competition between the irreversible effect of noise and the protection against such noise achieved by the delocalization of information. In the classical setting, where each node simply copies the input bit into multiple output bits, this model has been studied as the broadcasting or reconstruction problem on trees, which has broad applications. In this work, we study the quantum version of this problem. We consider a Clifford encoder at each node that encodes the input qubit in a stabilizer code, along with a single qubit Pauli noise channel at each edge. Such noisy quantum trees describe a scenario in which one has access to a stream of fresh (low-entropy) ancilla qubits, but cannot perform error correction. Therefore, they provide a different perspective on quantum fault tolerance. Furthermore, they provide a useful model for describing the effect of noise within the encoders of concatenated codes. We prove that above certain noise thresholds, which depend on the properties of the code such as its distance, as well as the properties of the encoder, information decays exponentially with the depth of the tree. On the other hand, by studying certain efficient decoders, we prove that for codes with distance d ≥ 2 and for sufficiently small (but non-zero) noise, classical information and entanglement propagate over a noisy tree with infinite depth. Indeed, we find that this remains true even for binary trees with certain 2-qubit encoders at each node, which encodes the received qubit in the binary repetition code with distance d = 1.

Quantum information

Event-Based Energy Impact Tracking and Forecasting with Limited Measurements for Rooftop Units

Packaged air conditioning units and heat pumps, also known as rooftop units (RTUs), are responsible for almost 133 billion kWh of electricity usage annually on site for space cooling U.S. commercial buildings. In addition, the use of heat pumps is a trend we expect to accelerate as buildings transition from fossil fuel-based heating to electricity as a key step for decarbonizing the U.S. commercial buildings sector. However, the operation conditions and energy use of RTUs and heat pumps are usually not well monitored as they are not commonly integrated with building automation systems and lack exposed sensing and control points. To fill this gap, this paper proposes a framework for tracking and forecasting energy impacts resulting from degradation of performance and improved performance for unit servicing using limited data. The proposed framework makes use of a constrained dataset, specifically measurements of the outdoor air temperature and the power demand of individual RTUs, to track and forecast changes in energy use associated with changes in performance over various temporal horizons ranging from days to weeks. Following the detection of an RTU fault, performance degradation, or performance improvement, the framework employs a prediction model to assess the cumulative energy impact. We demonstrate the effectiveness of the method with field-collected data for servicing and degradation examples and compare the predicting accuracy of Gradient Boosting Decision Tree (GBDT) Regression models to Support Vector Regression and Linear Regression models. The results show that GBDT achieved the best accuracy for time-series validation datasets for the servicing and degradation cases, and the prediction model was able to track the cumulative energy impacts of events. The proposed framework can inform building owners of the cumulative change in energy usage of RTUs associated with performance degradation, performance improvement, or a fault.

packaged air conditioners, packaged heat pumps, ro

Automating Anomaly Detection for Target systems at Spallation Neutron Source

The Spallation Neutron Source (SNS) at Oak Ridge National Laboratory, produces the world’s most intense pulse neutrons beams. An accelerated proton beam is directed into a mercury target to generate neutrons via spallation. The target system accounted for over 40% of the overall downtime of the facility in 2022. Thus, early detection in anomalies in the target systems can enable taking corrective actions to avoid failures and reduce downtime. Fault prognostics and anomaly detection in accelerators, both at SNS and outside, has largely focused on the beam side. This paper presents one the first studies exploring leveraging machine learning to automate the detection of anomalies in the target system. The target system consists of over 30 different interconnected subsystems, and the present work focuses on the mercury process system as a use case. Analyzing data from 28 process variables from 2022 and 2023, tree-based and reconstruction-based algorithms are employed to detect anomalies in archived data. The algorithms detected previously unreported anomalies, several of which were deemed alert worthy by human experts, particularly those found by reconstruction-based algorithms. Using data from each production run in the accelerator increased the generalizability of the models in time. Efforts are now underway to implement a workflow for incorporating human feedback to update the models and evaluating performance on unseen data. The models will eventually be integrated into the existing System Tracking and Reliability system with a web interface for automated anomaly detection and reporting along with a pathway for incorporating human feedback for model updates.

Raj, Anant [ORNL] (ORCID:0000000306711244)