Search NASA⌕ Search

SEARCH · Search NASA

Results for “fault tree”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 91 records · Page 5

System Safety Analysis of Complex NASA Systems with Model-Based Engineering (REV B)

The emergence of model-based engineering is transforming design and analysis methodologies. A recognized benefit of model-based engineering is the existence of a “single source of truth” about the system that becomes the authoritative source of data and information for designers, analysts, and developers. This promotes consistency and efficiency as the design emerges and can be used to further optimize the design. Integrating System Safety Engineers to the “single source of truth” will ensure that the outputs of their assessments and analyses are relevant to the design as it evolves. Use of an integrated system model enables near immediate evaluation of a design change as well as development of operational processes for risk assessment and communication. Such models can enable efficient and timely analysis of system hazards (e.g., hazard fault tree analysis and procedure simulations) and produce complete, accurate, and more consistent products (e.g., hazard reports and safety requirement evaluations). Therefore, an agency-sponsored team at Goddard Space Flight Center (GSFC) recently completed a System Safety Study of modeling and testing capabilities as part of a Model-Based Safety and Mission Assurance Initiative (MBSMAI). Using an existing model developed for reliability analyses, GSFC modeling and system safety experts performed system safety analysis/modeling and produced safety products. The team evaluated model-based feasibility to support System Safety Engineering, developed safety analysis modeling processes, and identified tool capability advancement/development needs. These study results indicate model-based engineering is valid and useable for System Safety Engineering for NASA if adequate modeling processes and environment are established.

NASA↗

Failure Assessment

Three questions to which software developers want accurate, precise answers are "How can the software system fail?", "mat bad things will happen if the software fails?t', and "How many failures will the software experience?". Numerous techniques have been devised to answer these questions; three of the best known are: 1) Software Fault Tree Analysis (SFTA) 2) Software Failure Modes, Effects, and Criticality Analysis (SFMECA 3) Software Fault/Failure Modeling. SFTA and SFMECA have been successfully used to analyze the flight software for a number of robotic planetary exploration missions, including Galileo, Cassini, and Deep Space 1. Given the increasing interest in reusing software components from mission to mission, one of us has developed techniques for reusing the corresponding portions of the SFTA and SFMECA, reducing the effort required to conduct these analyses. SFTA has also been shown to be effective in analyzing the security aspects of software systems; intrusion mechanisms and effects can easily be modeled using these techniques. The Bi- Directional Safety Analysis (BDSA) method combines a forward search (similar to SFMECA) from potential failure modes to their effects, with a backward search (similar to SFTA) from feasible hazards to the contributing causes of each hazard. BDSA offers an efficient way to identify latent failures. Recent work has extended BDSA to product-line applications such as flight-instrumentation displays and developed tool support for the reuse of the failure-analysis artifacts within a product line. BDSA has also been streamlined to support those projects having tight cost and/or schedule constraints for their failure analysis efforts. We discuss lessons learned from practice, describe available tools, and identi@ some future directions for the topic. A substantial amount of research has been devoted to estimating the number of failures that a software system will experience during test and operations, as well as the number of faults that have been inserted into that system during its development. One of us has found that the amount of structural change to a system during its development is strongly related to the number of faults inserted into it. Using techniques requiring no additional effort on the part of the development organization, the required measurements of structural evolution can be easily obtained from a development effort's configuration management system and readily transformed into an estimate of fault content. So far, structure-fault relationships have been identified for source code; current work seeks to examine artifacts available earlier in the lifecycle to determine if similar relationships between structure and fault content can be found. In particular, relationships between requirements change requests and the number of faults inserted into the implemented system would provide a significant improvement in our ability to control software quality during the early development phases.

fault tree↗

Hybrid automated reliability predictor integrated work station (HiREL)

The Hybrid Automated Reliability Predictor (HARP) integrated reliability (HiREL) workstation tool system marks another step toward the goal of producing a totally integrated computer aided design (CAD) workstation design capability. Since a reliability engineer must generally graphically represent a reliability model before he can solve it, the use of a graphical input description language increases productivity and decreases the incidence of error. The captured image displayed on a cathode ray tube (CRT) screen serves as a documented copy of the model and provides the data for automatic input to the HARP reliability model solver. The introduction of dependency gates to a fault tree notation allows the modeling of very large fault tolerant system models using a concise and visually recognizable and familiar graphical language. In addition to aiding in the validation of the reliability model, the concise graphical representation presents company management, regulatory agencies, and company customers a means of expressing a complex model that is readily understandable. The graphical postprocessor computer program HARPO (HARP Output) makes it possible for reliability engineers to quickly analyze huge amounts of reliability/availability data to observe trends due to exploratory design changes.

Bavuso, Salvatore J.↗

Practical Application of PRA as an Integrated Design Tool for Space Systems

This paper presents the application of the first comprehensive Probabilistic Risk Assessment (PRA) during the design phase of a joint NASA/NOAA weather satellite program, Geostationary Operational Environmental Satellite Series R (GOES-R). GOES-R is the next generation weather satellite primarily to help understand the weather and help save human lives. PRA has been used at NASA for Human Space Flight for many years. PRA was initially adopted and implemented in the operational phase of manned space flight programs and more recently for the next generation human space systems. Since its first use at NASA, PRA has become recognized throughout the Agency as a method of assessing complex mission risks as part of an overall approach to assuring safety and mission success throughout project lifecycles. PRA is now included as a requirement during the design phase of both NASA next generation manned space vehicles as well as for high priority robotic missions. The influence of PRA on GOES-R design and operation concepts are discussed in detail. The GOES-R PRA is unique at NASA for its early implementation. It also represents a pioneering effort to integrate risks from both Spacecraft (SC) and Ground Segment (GS) to fully assess the probability of achieving mission objectives. PRA analysts were actively involved in system engineering and design engineering to ensure that a comprehensive set of technical risks were correctly identified and properly understood from a design and operations perspective. The analysis included an assessment of SC hardware and software, SC fault management system, GS hardware and software, common cause failures, human error, natural hazards, solar weather and infrastructure (such as network and telecommunications failures, fire). PRA findings directly resulted in design changes to reduce SC risk from micro-meteoroids. PRA results also led to design changes in several SC subsystems, e.g. propulsion, guidance, navigation and control (GNC), communications, mechanisms, and command and data handling (C&DH). The fault tree approach assisted in the development of the fault management system design. Human error analysis, which examined human response to failure, indicated areas where automation could reduce the overall probability of gaps in operation by half. In addition, the PRA brought to light many potential root causes of system disruptions, including earthquakes, inclement weather, solar storms, blackouts and other extreme conditions not considered in the typical reliability and availability analyses. Ultimately the PRA served to identify potential failures that, when mitigated, resulted in a more robust design, as well as to influence the program's concept of operations. The early and active integration of PRA with system and design engineering provided a well-managed approach for risk assessment that increased reliability and availability, optimized lifecyc1e costs, and unified the SC and GS developments.

Kalia, Prince↗

NASA Space Nuclear Propulsion (SNP) MBSE Initiatives

NASA’s Space Nuclear Propulsion (SNP) program is developing several MagicDraw SysML models to support the development of high performance Nuclear Thermal Rocket Engines (NTRE). Currently, the Demonstration Rocket for Agile Cislunar Operations (DRACO) project is aiming to perform the first ever flight demonstration of an NTRE, and NASA is developing a DRACO Insight Project Model Based Systems Engineering (MBSE) model to capture, define, analyze, and report on the flight and ground test system architecture, functional behavior, requirements, risks, and lessons learned. Additional models are in work for engine component trade trees, fault detection sensor coverage analysis using a Goal Function Tree (GFT) plugin, stakeholder engagement, and technology maturation projects. The GFT plugin is the Galois, Inc. Failure Recovery Instruction Generation using Automata derived from Traditional Engineering models (FRIGATE) tool. A new capability for Jira to MagicDraw data sharing using the OpenPDM collaboration platform is under development with partner Victory Solutions, Inc. to enhance risk impact analysis.

Space Nuclear Propulsion (SNP)↗

Methodology for Designing Fault-Protection Software

A document describes a methodology for designing fault-protection (FP) software for autonomous spacecraft. The methodology embodies and extends established engineering practices in the technical discipline of Fault Detection, Diagnosis, Mitigation, and Recovery; and has been successfully implemented in the Deep Impact Spacecraft, a NASA Discovery mission. Based on established concepts of Fault Monitors and Responses, this FP methodology extends the notion of Opinion, Symptom, Alarm (aka Fault), and Response with numerous new notions, sub-notions, software constructs, and logic and timing gates. For example, Monitor generates a RawOpinion, which graduates into Opinion, categorized into no-opinion, acceptable, or unacceptable opinion. RaiseSymptom, ForceSymptom, and ClearSymptom govern the establishment and then mapping to an Alarm (aka Fault). Local Response is distinguished from FP System Response. A 1-to-n and n-to- 1 mapping is established among Monitors, Symptoms, and Responses. Responses are categorized by device versus by function. Responses operate in tiers, where the early tiers attempt to resolve the Fault in a localized step-by-step fashion, relegating more system-level response to later tier(s). Recovery actions are gated by epoch recovery timing, enabling strategy, urgency, MaxRetry gate, hardware availability, hazardous versus ordinary fault, and many other priority gates. This methodology is systematic, logical, and uses multiple linked tables, parameter files, and recovery command sequences. The credibility of the FP design is proven via a fault-tree analysis "top-down" approach, and a functional fault-mode-effects-and-analysis via "bottoms-up" approach. Via this process, the mitigation and recovery strategy(s) per Fault Containment Region scope (width versus depth) the FP architecture.

Barltrop, Kevin↗

Analysis of a hardware and software fault tolerant processor for critical applications

Computer systems for critical applications must be designed to tolerate software faults as well as hardware faults. A unified approach to tolerating hardware and software faults is characterized by classifying faults in terms of duration (transient or permanent) rather than source (hardware or software). Errors arising from transient faults can be handled through masking or voting, but errors arising from permanent faults require system reconfiguration to bypass the failed component. Most errors which are caused by software faults can be considered transient, in that they are input-dependent. Software faults are triggered by a particular set of inputs. Quantitative dependability analysis of systems which exhibit a unified approach to fault tolerance can be performed by a hierarchical combination of fault tree and Markov models. A methodology for analyzing hardware and software fault tolerant systems is applied to the analysis of a hypothetical system, loosely based on the Fault Tolerant Parallel Processor. The models consider both transient and permanent faults, hardware and software faults, independent and related software faults, automatic recovery, and reconfiguration.

Dugan, Joanne B.↗

Automated Generation of Fault Management Artifacts from a Simple System Model

Our understanding of off-nominal behavior - failure modes and fault propagation - in complex systems is often based purely on engineering intuition; specific cases are assessed in an ad hoc fashion as a (fallible) fault management engineer sees fit. This work is an attempt to provide a more rigorous approach to this understanding and assessment by automating the creation of a fault management artifact, the Failure Modes and Effects Analysis (FMEA) through querying a representation of the system in a SysML model. This work builds off the previous development of an off-nominal behavior model for the upcoming Soil Moisture Active-Passive (SMAP) mission at the Jet Propulsion Laboratory. We further developed the previous system model to more fully incorporate the ideas of State Analysis, and it was restructured in an organizational hierarchy that models the system as layers of control systems while also incorporating the concept of "design authority". We present software that was developed to traverse the elements and relationships in this model to automatically construct an FMEA spreadsheet. We further discuss extending this model to automatically generate other typical fault management artifacts, such as Fault Trees, to efficiently portray system behavior, and depend less on the intuition of fault management engineers to ensure complete examination of off-nominal behavior.

Spinup and Orient↗

Toward a Model-Based Approach for Flight System Fault Protection

Use SysML/UML to describe the physical structure of the system This part of the model would be shared with other teams - FS Systems Engineering, Planning & Execution, V&V, Operations, etc., in an integrated model-based engineering environment Use the UML Profile mechanism, defining Stereotypes to precisely express the concepts of the FP domain This extends the UML/SysML languages to contain our FP concepts Use UML/SysML, along with our profile, to capture FP concepts and relationships in the model Generate typical FP engineering products (the FMECA, Fault Tree, MRD, V&V Matrices)

fault protection↗

Systems Engineering and Assurance Modeling (SEAM): A Web-Based Solution for Integrated Mission Assurance

We present an overview of the Systems Engineering and Assurance Modeling (SEAM) platform, a web-browser-based tool which is designed to help engineers evaluate the radiation vulnerabilities and develop an assurance approach for electronic parts in space systems. The SEAM framework consists of three interconnected modeling tools, a SysML compatible system description tool, a Goal Structuring Notation (GSN) visual argument tool, and Bayesian Net and Fault Tree extraction and export tools. The SysML and GSN sections also have a coverage check application that ensures that every radiation fault identified on the SysML side is also addressed in the assurance case in GSN. The SEAM platform works on space systems of any degree of radiation hardness but is especially helpful for assessing radiation performance in systems with commercial-off-the-shelf (COTS) electronic components.

SEAM↗

Evaluating Faulty State Occurrence in Wildfire UAS Missions Using Markov Chains

As autonomous technology advances, unmanned aircraft systems are increasingly integrated into emergency response missions, such as wildfire response. These systems must be be safe with less risk than non-autonomous counter parts, yet quantifying the risk associated with present-day and future systems conventionally relies solely on expert opinion and little data. Instead, combining narrative mishap reports with probabilistic analysis can provide a method for evolutionary and timely risk analysis. In this paper, we present a framework for a data-driven probabilistic risk assessment style analysis, where hazard events and rates originate from documented UAS mishaps. The framework is applied to a UAS mapping mission in wildfire response, including a fault tree analysis, event tree analysis, and probabilistic analysis using Markov Chains. The analysis provides an enumeration of hazards in the system, hazard events that can lead to faults, the probability of a mission experiencing any fault, the probability of experiencing a specific fault, and the expected time spent until faulty states occur in present-day operations.

risk analysis↗

Evaluating Faulty State Occurrence in Wildfire UAS Missions Using Markov Chains

As autonomous technology advances, unmanned aircraft systems are increasingly integrated into emergency response missions, such as wildfire response. These systems must be be safe with less risk than non-autonomous counter parts, yet quantifying the risk associated with present-day and future systems conventionally relies solely on expert opinion and little data. Instead, combining narrative mishap reports with probabilistic analysis can provide a method for evolutionary and timely risk analysis. In this paper, we present a framework for a data-driven probabilistic risk assessment style analysis, where hazard events and rates originate from documented UAS mishaps. The framework is applied to a UAS mapping mission in wildfire response, including a fault tree analysis, event tree analysis, and probabilistic analysis using Markov Chains. The analysis provides an enumeration of hazards in the system, hazard events that can lead to faults, the probability of a mission experiencing any fault, the probability of experiencing a specific fault, and the expected time spent until faulty states occur in present-day operations.

risk analysis↗

The role of reliability graph models in assuring dependable operation of complex hardware/software systems

The complexity of computer systems currently being designed for critical applications in the scientific, commercial, and military arenas requires the development of new techniques for utilizing models of system behavior in order to assure 'ultra-dependability'. The complexity of these systems, such as Space Station Freedom and the Air Traffic Control System, stems from their highly integrated designs containing both hardware and software as critical components. Reliability graph models, such as fault trees and digraphs, are used frequently to model hardware systems. Their applicability for software systems has also been demonstrated for software safety analysis and the analysis of software fault tolerance. This paper discusses further uses of graph models in the design and implementation of fault management systems for safety critical applications.

Patterson-Hine, F. A.↗

Model-Based Safety Analysis

System safety analysis techniques are well established and are used extensively during the design of safety-critical systems. Despite this, most of the techniques are highly subjective and dependent on the skill of the practitioner. Since these analyses are usually based on an informal system model, it is unlikely that they will be complete, consistent, and error free. In fact, the lack of precise models of the system architecture and its failure modes often forces the safety analysts to devote much of their effort to gathering architectural details about the system behavior from several sources and embedding this information in the safety artifacts such as the fault trees. This report describes Model-Based Safety Analysis, an approach in which the system and safety engineers share a common system model created using a model-based development process. By extending the system model with a fault model as well as relevant portions of the physical system to be controlled, automated support can be provided for much of the safety analysis. We believe that by using a common model for both system and safety engineering and automating parts of the safety analysis, we can both reduce the cost and improve the quality of the safety analysis. Here we present our vision of model-based safety analysis and discuss the advantages and challenges in making this approach practical.

Joshi, Anjali↗

Failure Analysis and Products in a Model-Based Environment

The work presented in this paper describes an approach, including a methodology and tools, which allows system engineers to capture failure-related information in a model and generate automatically key failure analysis products: the Failure Modes, Effects and Criticality Analysis (FMECA) and the Fault Tree Analysis (FTA). The work has been developed by Tietronix Software, Inc. and the NASA’s Jet Propulsion Laboratory (JPL), and the resulting auto-generated artifacts shown in this paper demonstrate the ability to obtain powerful reliability and fault management products in a model-based environment.

Castet, Jean-Francois↗

Fault Diagnosis of Power Components with Reliability Assessment in Extraterrestrial Microgrids

This research investigates the possible failures caused by aging and other environmental and external factors that could significantly impact the performance of extraterrestrial power systems. Additionally, it presents a reliability assessment model for the space microgrid based on fault tree analysis (FTA). The reliability assessment model developed in this paper represents a tool that can be used by engineers to harden the system design for operational and economic benefits. To improve the reliability of the system, this work provides a broad review of the different fault detection and diagnosis (FDD) algorithms used for power microgrids and space applications. Using data sets from the Habitat Simulator developed through the NASA-funded Resilient Extraterrestrial Habitat Institute, this paper compares the applicability and accuracy of the different FDD methods. The primary FDD approach proposed and assessed in this work is based on the Markov reliability model. It predicts and detects future faults in the space microgrids by using past data samples and categorizing them into different classes. Data-driven-based models such as artificial neural networks are also investigated, tested, and evaluated using simulation data sets. According to the simulation results and the broad FDD algorithm comparison, this study provides the crew or maintenance engineers with a clear methodology to detect and localize power system failures.

Leila Chebbo↗

Fault detection and fault tolerance in robotics

Robots are used in inaccessible or hazardous environments in order to alleviate some of the time, cost and risk involved in preparing men to endure these conditions. In order to perform their expected tasks, the robots are often quite complex, thus increasing their potential for failures. If men must be sent into these environments to repair each component failure in the robot, the advantages of using the robot are quickly lost. Fault tolerant robots are needed which can effectively cope with failures and continue their tasks until repairs can be realistically scheduled. Before fault tolerant capabilities can be created, methods of detecting and pinpointing failures must be perfected. This paper develops a basic fault tree analysis of a robot in order to obtain a better understanding of where failures can occur and how they contribute to other failures in the robot. The resulting failure flow chart can also be used to analyze the resiliency of the robot in the presence of specific faults. By simulating robot failures and fault detection schemes, the problems involved in detecting failures for robots are explored in more depth.

Visinsky, Monica↗