Search NASA⌕ Search

SEARCH · Search NASA

Results for “fault tree”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 109 records · Page 6

Using minimal spanning trees to compare the reliability of network topologies

Graph theoretic methods are applied to compute the reliability for several types of networks of moderate size. The graph theory methods used are minimal spanning trees for networks with bi-directional links and the related concept of strongly connected directed graphs for networks with uni-directional links. A comparison is conducted of ring networks and braided networks. The case is covered where just the links fail and the case where both links and nodes fail. Two different failure modes for the links are considered. For one failure mode, the link no longer carries messages. For the other failure mode, the link delivers incorrect messages. There is a description and comparison of link-redundancy versus path-redundancy as methods to achieve reliability. All the computations are carried out by means of a fault tree program.

Leister, Karen J.↗

Redefining Design for Remanufacturing: A Practical Methodology for Prioritizing Remanufacturing Design Rules

Products are often discarded when they fail or no longer meet user needs. These outcomes are frequently shaped by early design decisions. While remanufacturing offers a sustainable alternative by restoring products to like‐new condition, its potential is often limited by designs that do not consider remanufacturing from the outset. This research addresses that challenge by introducing a structured Design for Remanufacturing (DfRem) methodology and a CAD‐integrated tool to support real‐time design decisions. The DfRem framework introduces a new primary design function focused on preserving product functionality across its life cycle. It is supported by a fault tree that identifies failure modes that limit remanufacturing potential and a hierarchy of design principles including Prevent, Minimize, Relocate, Restore, and others. Each principle is linked to actionable design rules that help engineers reduce the need for remanufacturing or improve its efficiency when necessary. To operationalize this framework, we developed CAD plugins for Autodesk Inventor and PTC Creo. These tools use a state machine model to present prioritized design rules based on selected failure modes and user input. By embedding DfRem logic directly into widely used CAD environments, the tool enables engineers to make sustainability‐informed decisions without disrupting existing workflows. Furthermore, this approach highlights the critical role of design in enabling circular and resource‐efficient product development, making remanufacturing a more practical and accessible strategy during the early stages of product design.

CAD↗

Development and Demonstration of a Prototype Molten Salt Sampling System

Molten salt reactors (MSRs) offer potential operability and safety advantages when compared to commercial light water reactors (LWRs). However, operating experience with MSRs is sparse in comparison to what exists for LWRs. Further, the chemical and isotopic composition of the fuel and/or coolant salt is dynamic and difficult to characterize continuously, posing potential safety, operability, and safeguards unknowns that need to be addressed. A molten salt sampling system (MSSS) is regarded as a necessary subsystem within first generation MSRs used to obtain samples of salt for chemical and isotopic analysis in support of the need to monitor and control salt composition during operation. The MSSS is being developed using the Safety-in-Design (SiD) methodology, which incorporates incremental integration of safety analysis into the design process. The MSSS conceptual design emerging from the application of the early stages of the SiD methodology consists of a sample collection system and its housing, a freeze port, and inert gas control and delivery systems. This article describes the prototypes developed to test the functions of these MSSS subsystems, presents the results of testing in both dry and molten salt environments (including reliability data collection performed in accordance with the principles of SiD and the development of a semiquantitative fault tree model), and summarizes the opportunities for future design and testing enhancements based on the results of prototype testing.

molten salt reactor↗

Success Path Method: Introduction to the Success Path Method Software Tool©

As part of its commitment to advancing safety and reliability assessment methodologies, Argonne National Laboratory pioneered the use of an evaluation method called the Success Path Method (SPM) to improve risk management for offshore oil and gas operations. The development of the SPM at Argonne has been driven by the need to improve existing risk assessment methodologies by focusing on the steps necessary for success rather than failure modes alone. This is particularly important for industrial environments like offshore facilities that perform multiple functions under a continuously evolving set of operational conditions – such as water depth and temperature, currents, and weather conditions. In these dynamic environments, the traditional Probabilistic Risk Assessment (PRA) approach is far too complex as it focuses on what can go wrong – which comprises an infinite failure space that must be fully explored and understood. By shifting the focus to a finite space of success paths, the SPM enables operators and decision makers to prioritize a manageable number of steps that must go right to ensure success. Building on its five decades of experience in safety assessments for the nuclear industry, Argonne made major adaptations to existing risk assessment methods utilizing features similar to fault trees that are traditionally used in PRA to map all pathways in which the system can malfunction. In contrast, SPM identifies the components and processes that must function correctly to achieve specific outcomes – such as preventing the uncontrolled release of hydrocarbons during drilling operations. The SPM framework integrates equipment, procedures, software, processes, and human actions to ensure that physical barriers meet critical safety functions in dynamic operational conditions. This approach helps identify failure modes and improve operational risk management by narrowing the focus to key success elements, which in turn reduces uncertainty and helps users understand, manage, and respond to failures.

97 MATHEMATICS AND COMPUTING↗

Quantitative Risk Assessment for Fuel Cell Electric Bus Hydrogen Storage and Refueling Facility

It is necessary to understand the safety implications and risk mitigation options for fuel cell electric bus fleet deployment, especially for related facilities responsible for operations such as production, storage, compression, and dispensing of hydrogen for use by the buses. In this report, we present a quantitative risk assessment for a potential fuel cell electric bus fleet that was motivated by efforts to improve resilience at the Portland International Airport but can be applicable to a range of hydrogen case studies and use cases. We estimated risk for a facility that produces, stores, compresses, and dispenses hydrogen for the fleet of buses, with a focus on individual risk to people in terms of annual frequency of fatality. We considered the frequency of hydrogen leaks that could result in harmful physical outcomes like jet fires or explosions, and the consequences of those outcomes for people. We created customized fault trees to calculate the frequencies of different sizes of leaks and event sequence diagrams to calculate ignition probabilities for the various leak sizes. We also leveraged the HyRAM+ toolkit to use these inputs to calculate overall risk for the facility, which we separated into one section responsible for producing, storing, and compressing hydrogen, and one section responsible for dispensing the hydrogen to the buses. We found that the dispensing area seemed to have a higher risk than the production/storage/compression area of the facility, largely because of the inclusion of a component with a high leak frequency (the heat exchanger used to cool the hydrogen before entering the vehicle, to prevent overheating and expansion of hydrogen in the onboard tank). For the example production and refueling facility we evaluated and the data we used for the analysis, the leak frequency had a larger impact on the risk differences between the two sections on the facility, compared to the physical outcome consequence, which was slightly different due to the varying fuel conditions, but not substantially different. Actions can be taken to prevent these hazards (e.g., lowering leak frequencies in system components) or to mitigate the consequences if they do occur (e.g., installing barriers to protect people if ignition events occur). The choice of which actions to take depends not only on safety considerations but also on space, time, staffing, feasibility, and financial constraints. Therefore, the quantitative risk assessment approach can help understand relative risk contributions from different components, leak sizes, consequences, and human actions, to prioritize risk reduction strategies and balance these parameters. The outcomes of this report may be useful for a variety of stakeholders working in the hydrogen, transportation, vehicle, and aviation sector, including those responsible for aspects like facility design, operations, and regulations. There is not a single value of risk that determines whether a hypothetical system is “safe” or not. The insights about risk mitigations may be leveraged, and the quantitative risk assessment approach can be applied to other case studies to understand risk priorities and contributions specific to different FCEB and hydrogen facility uses.

08 HYDROGEN↗

Probability of Hydrogen Ignition: A Landscape Review and Gaps Assessment

The primary hazard of a leak from a hydrogen system is due to the immediate or delayed ignition of the fuel leading to a jet flame or explosion. Therefore, understanding the hydrogen ignition probability is critical for analyzing the risk of hydrogen systems. This report reviews the current understanding of hydrogen ignition mechanisms and methods for modeling their probability. The stoichiometry, ignition strength, and ignition source temperature are all important characteristics that can affect both the probability of ignition and the outcome of the subsequent combustion event. A brief review of diffusion ignition demonstrates that ignition probability models must account for seemingly spontaneous ignition of hydrogen in addition to scenarios where the ignition source is readily identified. State-of-the art models for both immediate and delayed ignition probabilities are presented, including different physical aspects of the scenarios (e.g., flow rate, ignition source characteristics) that are considered in the different modeling approaches. Current models often fail to account for the unique properties of hydrogen compared to other fuels, and most lack rigorous validation with hydrogen as a fuel. A fault tree framework is proposed to systematically evaluate the probability of ignition by integrating various ignition mechanisms and their uncertainties. Furthermore, this type of framework could enable additional insights into the most important mechanisms and would enable uncertainty quantification in risk assessment modeling. Recommendations for future research include the need for experimental validation of ignition models and the development of comprehensive methodologies that incorporate the specifics of hydrogen behavior in real-world scenarios.

hydrogen↗

Operations analysis (study 2.1). Contingency analysis

Future operational concepts for the space transportation system were studied in terms of space shuttle upper stage failure contingencies possible during deployment, retrieval, or space servicing of automated satellite programs. Problems anticipated during mission planning were isolated using a modified 'fault tree' technique, normally used in safety analyses. A comprehensive space servicing hazard analysis is presented which classifies possible failure modes under the catagories of catastrophic collision, failure to rendezvous and dock, servicing failure, and failure to undock. The failure contingencies defined are to be taken into account during design of the upper stage.

Source record↗

Reliability/safety analysis of a fly-by-wire system

An analysis technique has been developed to estimate the reliability of a very complex, safety-critical system by constructing a diagram of the reliability equations for the total system. This diagram has many of the characteristics of a fault-tree or success-path diagram, but is much easier to construct for complex redundant systems. The diagram provides insight into system failure characteristics and identifies the most likely failure modes. A computer program aids in the construction of the diagram and the computation of reliability. Analysis of the NASA F-8 Digital Fly-by-Wire Flight Control System is used to illustrate the technique.

Brock, L. D.↗

System safety in Stirling engine development

The DOE/NASA Stirling Engine Project Office has required that contractors make safety considerations an integral part of all phases of the Stirling engine development program. As an integral part of each engine design subtask, analyses are evolved to determine possible modes of failure. The accepted system safety analysis techniques (Fault Tree, FMEA, Hazards Analysis, etc.) are applied in various degrees of extent at the system, subsystem and component levels. The primary objectives are to identify critical failure areas, to enable removal of susceptibility to such failures or their effects from the system and to minimize risk.

Bankaitis, H.↗

Models for evaluating the performability of degradable computing systems

Recent advances in multiprocessor technology established the need for unified methods to evaluate computing systems performance and reliability. In response to this modeling need, a general modeling framework that permits the modeling, analysis and evaluation of degradable computing systems is considered. Within this framework, several user oriented performance variables are identified and shown to be proper generalizations of the traditional notions of system performance and reliability. Furthermore, a time varying version of the model is developed to generalize the traditional fault tree reliability evaluation methods of phased missions.

Wu, L. T.↗

Space reliability technology - A historical perspective

The progressive improvements in reliability of launch vehicles is traced from the Vanguard rocket to the STS. The Vanguard, built with minimal redundancy and a high mass ratio, was used as an operational vehicle midway through its test program in an attempt to meet the perceived challenge represented by the Sputnik. The fourth Vanguard failed due to inadequate contamination prevention and lack of inspection ports. Automatic firing sequences were adopted for the Titan rockets, which were an order of magnitude larger than the Vanguard and therefore had room for interior inspections. Qualification testing and reporting were introduced for components, along with X ray inspection of fuel tank welds. Dual systems were added for flight critical components when the Titan became man-rated for the Gemini program. Designs incorporated full failure mode effects and criticality analyses for the Apollo program, which exposed the limits of applicability of numerical reliability models. Fault tree analyses and program milestone reviews were initiated. The worth of man-in-the-loop in space activities for reliability was demonstrated with the rescue of Skylab after solar panel and meteoroid shield failures. It is now the reliability of the payload, rather than the vehicle, that is questioned for Shuttle launches.

Cohen, H.↗

Diagnostic reasoning in digital systems

Described is an efficient method for fault diagnosis in digital systems based on the technique of reasoning. The methodology operates on the observed erroneous behavior and the structure of the system. The behavior consists of the error(s) observed on the circuit's output lines and specific values on the circuit's input lines. The techniques described improve on previously published research on diagnostic reasoning in two ways. Previous work has stressed system independent techniques which could be used to diagnose any faulty system whose structure can be represented. By concentrating on the specific case of diagnosing faulty digital circuits, it is possible to simplify the representation of the structure of the system. This representation, in the form of an AND/OR fault tree, efficiently abstracts the structure of a faulty digital system. More importantly, a method for partitioning the digital system is introduced which can considerably reduce the runtime complexity of a diagnosis.

Thearling, Kurt Henry↗

STS-27R OV-104 Orbiter TPS damage review team, volume 1

Following the return to earth on December 2, 1988, of Orbiter OV-104, Atlantis, it was observed that there was substantial Thermal Protection System (TPS) tile damage present on the lower right fuselage and wing. Damage sites were more numerous than on previous flights and conversely, there was almost no damage present on Atlantis' left side. A review team investigated the cause beginning with a detailed inspection of the Atlantis TPS damage, and a review of related inspection reports to establish an indepth anomaly definition. An exhaustive data review followed. A fault tree and several failure scenarios were developed. Finally, the failure scenarios were categorized as either not possible, possible but not probable, or probable. This and other information gained during the review formed the basis for the team's findings and recommendations. The team concluded that the most probable cause of the severe STS-27R Orbiter tile damage is that the ablative insulating material covering the RH SRB Nose Cap dislodged and struck the Orbiter tile near 85 seconds into flight and possibly that debris from other sources, including repaired insulation and missing joint cork, caused minor tile damage. Findings are presented, and recommendations that are believed pertinent to minimizing the potential for inflight debris are described.

Thomas, John W.↗

An approach to solving large reliability models

This paper describes a unified approach to the problem of solving large realistic reliability models. The methodology integrates behavioral decomposition, state trunction, and efficient sparse matrix-based numerical methods. The use of fault trees, together with ancillary information regarding dependencies to automatically generate the underlying Markov model state space is proposed. The effectiveness of this approach is illustrated by modeling a state-of-the-art flight control system and a multiprocessor system. Nonexponential distributions for times to failure of components are assumed in the latter example. The modeling tool used for most of this analysis is HARP (the Hybrid Automated Reliability Predictor).

Boyd, Mark A.↗

Neural computing for numeric-to-symbolic conversion in control systems

A type of neural network, the multilayer perceptron, is used to classify numeric data and assign appropriate symbols to various classes. This numeric-to-symbolic conversion results in a type of information extraction, which is similar to what is called data reduction in pattern recognition. The use of the neural network as a numeric-to-symbolic converter is introduced, its application in autonomous control is discussed, and several applications are studied. The perceptron is used as a numeric-to-symbolic converter for a discrete-event system controller supervising a continuous variable dynamic system. It is also shown how the perceptron can implement fault trees, which provide useful information (alarms) in a biological system and information for failure diagnosis and control purposes in an aircraft example.

Passino, Kevin M.↗

Flammability and sensitivity of materials in oxygen-enriched atmospheres; Proceedings of the Fourth International Symposium, Las Cruces, NM, Apr. 11-13, 1989. Volume 4

The present volume discusses the ignition of nonmetallic materials by the impact of high-pressure oxygen, the promoted combustion of nine structural metals in high-pressure gaseous oxygen, the oxygen sensitivity/compatibility ranking of several materials by different test methods, the ignition behavior of silicon greases in oxygen atmospheres, fire spread rates along cylindrical metal rods in high-pressure oxygen, and the design of an ignition-resistant, high pressure/temperature oxygen valve. Also discussed are the promoted ignition of oxygen regulators, the ignition of PTFE-lined flexible hoses by rapid pressurization with oxygen, evolving nonswelling elastomers for high-pressure oxygen environments, the evaluation of systems for oxygen service through the use of the quantitative fault-tree analysis, and oxygen-enriched fires during surgery of the head and neck.

Stoltzfus, Joel M.↗

Program Models Propagation Of Failures

FIRM is software tool for identification of failure and management of risk based on directed-graph ("digraph") approach. Three core algorithms optimized for processing singletons and doubletons and also handle tripletons. FIRM identifies loops in digraphs and displays direct failure paths between any two nodes. Solves for reachability for given node without computing reachability for entire digraph. Represents hybrid between schematic-diagram and fault-tree approaches. Written in C.

Hackler, Donald B.↗

Hubble Space Telescope: SRM/QA observations and lessons learned

The Hubble Space Telescope (HST) Optical Systems Board of Investigation was established on July 2, 1990 to review, analyze, and evaluate the facts and circumstances regarding the manufacture, development, and testing of the HST Optical Telescope Assembly (OTA). Specifically, the board was tasked to ascertain what caused the spherical aberration and how it escaped notice until on-orbit operation. The error that caused the on-orbit spherical aberration in the primary mirror was traced to the assembly process of the Reflective Null Corrector, one of the three Null Correctors developed as special test equipment (STE) to measure and test the primary mirror. Therefore, the safety, reliability, maintainability, and quality assurance (SRM&QA) investigation covers the events and the overall product assurance environment during the manufacturing phase of the primary mirror and Null Correctors (from 1978 through 1981). The SRM&QA issues that were identified during the HST investigation are summarized. The crucial product assurance requirements (including nonconformance processing) for the HST are examined. The history of Quality Assurance (QA) practices at Perkin-Elmer (P-E) for the period under investigation are reviewed. The importance of the information management function is discussed relative to data retention/control issues. Metrology and other critical technical issues also are discussed. The SRM&QA lessons learned from the investigation are presented along with specific recommendations. Appendix A provides the MSFC SRM&QA report. Appendix B provides supplemental reference materials. Appendix C presents the findings of the independent optical consultants, Optical Research Associates (ORA). Appendix D provides further details of the fault-tree analysis portion of the investigation process.

Rodney, George A.↗