Search NASA⌕ Search

SEARCH · Search NASA

Results for “fault analysis”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 127 records · Page 7

Achieving Improved Reliability with Failure Analysis

Reliability is the ability of a product to properly function, within specified performance limits, for a specified period of time, under the life cycle application conditions. Failure analysis is a vital tool in the effort to ensure reliability of electronic products and systems throughout their product lifecycle. Today, organizations involved in activities within the electronics supply chain are facing new challenges, not just from complex assembly styles, harsher lifecycle environments, and sophisticated supply chains, but also from customers who are demanding a quicker turn-around. Unfortunately, root cause failure analysis is often performed incompletely, leading to a poor understanding of failure mechanisms and causes and, customer dissatisfaction due to recurring failures. The PDC (Professional Development Course) starts with an introduction to reliability concepts, physics of failure and an overview of failure mechanisms that affect PCBs (Printed Circuit Boards), PCBAs (Printed Circuit Board Assembly) and components. The PDC then dives into root cause hypothesizing techniques (Pareto, FMEA (Failure Modes and Effects Analysis), fishbone (Cause-And-Effect Diagram), FTA (Fault Tree Analysis)), non-destructive and destructive analysis and, materials characterization will be discussed. Numerous failure analysis case studies will be used to illustrate the techniques and analysis principles to arrive at the root cause(s) of field failures on printed circuit boards, active components, and assemblies. What Attendees will Learn: Topics include: Overview of Reliability Concepts Failure mechanisms of electronic products Root cause analysis Failure analysis techniques -Non-destructive techniques (optical, CSAM (Confocal Scanning Electron Microscopy) etc.) -Destructive analysis (DPA (Destructive Physical Analysis), Decap (Decapsulation), FIB (Focused Ion Beam) etc.) -Materials characterization (XRF (X-Ray Fluorescence) , EDS (Error Detection Sequential), TMA/DSC (Thermal Mechanical Analysis/Differential Scanning Calorimetry) etc.)

PCB quality↗

Model authoring system for fail safe analysis

The Model Authoring System is a prototype software application for generating fault tree analyses and failure mode and effects analyses for circuit designs. Utilizing established artificial intelligence and expert system techniques, the circuits are modeled as a frame-based knowledge base in an expert system shell, which allows the use of object oriented programming and an inference engine. The behavior of the circuit is then captured through IF-THEN rules, which then are searched to generate either a graphical fault tree analysis or failure modes and effects analysis. Sophisticated authoring techniques allow the circuit to be easily modeled, permit its behavior to be quickly defined, and provide abstraction features to deal with complexity.

Sikora, Scott E.↗

Functional Fault Model Development Process to Support Design Analysis and Operational Assessment

A functional fault model (FFM) is an abstract representation of the failure space of a given system. As such, it simulates the propagation of failure effects along paths between the origin of the system failure modes and points within the system capable of observing the failure effects. As a result, FFMs may be used to diagnose the presence of failures in the modeled system. FFMs necessarily contain a significant amount of information about the design, operations, and failure modes and effects. One of the important benefits of FFMs is that they may be qualitative, rather than quantitative and, as a result, may be implemented early in the design process when there is more potential to positively impact the system design. FFMs may therefore be developed and matured throughout the monitored system's design process and may subsequently be used to provide real-time diagnostic assessments that support system operations. This paper provides an overview of a generalized NASA process that is being used to develop and apply FFMs. FFM technology has been evolving for more than 25 years. The FFM development process presented in this paper was refined during NASA's Ares I, Space Launch System, and Ground Systems Development and Operations programs (i.e., from about 2007 to the present). Process refinement took place as new modeling, analysis, and verification tools were created to enhance FFM capabilities. In this paper, standard elements of a model development process (i.e., knowledge acquisition, conceptual design, implementation & verification, and application) are described within the context of FFMs. Further, newer tools and analytical capabilities that may benefit the broader systems engineering process are identified and briefly described. The discussion is intended as a high-level guide for future FFM modelers.

Verification↗

Abstractions for Fault-Tolerant Distributed System Verification

Four kinds of abstraction for the design and analysis of fault tolerant distributed systems are discussed. These abstractions concern system messages, faults, fault masking voting, and communication. The abstractions are formalized in higher order logic, and are intended to facilitate specifying and verifying such systems in higher order theorem provers.

Pike, Lee S.↗

Testing for Software Safety

This research focuses on testing whether or not the hazardous conditions identified by design-level fault tree analysis will occur in the target implementation. Part 1: Integrate fault tree models into functional specifications so as to identify testable interactions between intended behaviors and hazardous conditions. Part 2: Develop a test generator that produces not only functional tests but also safety tests for a target implementation in a cost-effective way. Part 3: Develop a testing environment for executing generated functional and safety tests and evaluating test results against expected behaviors or hazardous conditions. It includes a test harness as well as an environment simulation of external events and conditions.

Chen, Ken↗

Failure-Tolerant Avionics for Crewed Space Systems Recommended Best Practices

This paper provides an overview of some of the major steps needed to mature and justify the design of an avionics system for crewed spacecraft. It is organized as a collection of artifacts or pieces of evidence that NASA needs to assess the system at design reviews, including a functional failure modes and effects analysis (FFMEA), fault containment region (FCR) definitions, the failure hypothesis, and reliability analysis. This paper is intended as a reference for designers working on NASA crewed spaceflight projects, reliability engineers responsible for avionics system assessments, and program managers wanting to understand what evidence is required at design reviews to ensure crew safety and mission success.

Avionics↗

A study of mapping exogenous knowledge representations into CONFIG

Qualitative reasoning is reasoning with a small set of qualitative values that is an abstraction of a larger and perhaps infinite set of quantitative values. The use of qualitative and quantitative reasoning together holds great promise for performance improvement in applications that suffer from large and/or imprecise knowledge domains. Included among these applications are the modeling, simulation, analysis, and fault diagnosis of physical systems. Several research groups continue to discover and experiment with new qualitative representations and reasoning techniques. However, due to the diversity of these techniques, it is difficult for the programs produced to exchange system models easily. The availability of mappings to transform knowledge from the form used by one of these programs to that used by another would open the doors for comparative analysis of these programs in areas such as completeness, correctness, and performance. A group at the Johnson Space Center (JSC) is working to develop CONFIG, a prototype qualitative modeling, simulation, and analysis tool for fault diagnosis applications in the U.S. space program. The availability of knowledge mappings from the programs produced by other research groups to CONFIG may provide savings in CONFIG's development costs and time, and may improve CONFIG's performance. The study of such mappings is the purpose of the research described in this paper. Two other research groups that have worked with the JSC group in the past are the Northwest University Group and the University of Texas at Austin Group. The former has produced a qualitative reasoning tool named SIMGEN, and the latter has produced one named QSIM. Another program produced by the Austin group is CC, a preprocessor that permits users to develop input for eventual use by QSIM, but in a more natural format. CONFIG and CC are both based on a component-connection ontology, so a mapping from CC's knowledge representation to CONFIG's knowledge representation was chosen as the focus of this study. A mapping from CC to CONFIG was developed. Due to differences between the two programs, however, the mapping transforms some of the CC knowledge to CONFIG as documentation rather than as knowledge in a form useful to computation. The study suggests that it may be worthwhile to pursue the mappings further. By implementing the mapping as a program, actual comparisons of computational efficiency and quality of results can be made between the QSIM and CONFIG programs. A secondary study may reveal that the results of the two programs augment one another, contradict one another, or differ only slightly. If the latter, the qualitative reasoning techniques may be compared in other areas, such as computational efficiency.

Mayfield, Blayne E.↗

Functional Fault Model Development Process to Support Design Analysis and Operational Assessment

A functional fault model (FFM) is an abstract representation of the failure space of a givensystem. As such, it simulates the propagation of failure effects along paths between the origin ofthe system failure modes and points within the system capable of observing the failure effects. Asa result, FFMs may be used to diagnose the presence of failures in the modeled system. FFMsnecessarily contain a significant amount of information about the design, operations, and failuremodes and effects. One of the important benefits of FFMs is that they may be qualitative, ratherthan quantitative and, as a result, may be implemented early in the design process when there ismore potential to positively impact the system design. FFMs may therefore be developed andmatured throughout the monitored system's design process and may subsequently be used toprovide real-time diagnostic assessments that support system operations. This paper provides anoverview of a generalized NASA process that is being used to develop and apply FFMs. FFMtechnology has been evolving for more than 25 years. The FFM development process presented inthis paper was refined during NASA's Ares I, Space Launch System, and Ground SystemsDevelopment and Operations programs (i.e., from about 2007 to the present). Process refinementtook place as new modeling, analysis, and verification tools were created to enhance FFMcapabilities. In this paper, standard elements of a model development process (i.e., knowledgeacquisition, conceptual design, implementation & verification, and application) are describedwithin the context of FFMs. Further, newer tools and analytical capabilities that may benefit thebroader systems engineering process are identified and briefly described. The discussion isintended as a high-level guide for future FFM modelers.

Melcher, Kevin J.↗

From Informal Safety-Critical Requirements to Property-Driven Formal Validation

Most of the efforts in formal methods have historically been devoted to comparing a design against a set of requirements. The validation of the requirements themselves, however, has often been disregarded, and it can be considered a largely open problem, which poses several challenges. The first challenge is given by the fact that requirements are often written in natural language, and may thus contain a high degree of ambiguity. Despite the progresses in Natural Language Processing techniques, the task of understanding a set of requirements cannot be automatized, and must be carried out by domain experts, who are typically not familiar with formal languages. Furthermore, in order to retain a direct connection with the informal requirements, the formalization cannot follow standard model-based approaches. The second challenge lies in the formal validation of requirements. On one hand, it is not even clear which are the correctness criteria or the high-level properties that the requirements must fulfill. On the other hand, the expressivity of the language used in the formalization may go beyond the theoretical and/or practical capacity of state-of-the-art formal verification. In order to solve these issues, we propose a new methodology that comprises of a chain of steps, each supported by a specific tool. The main steps are the following. First, the informal requirements are split into basic fragments, which are classified into categories, and dependency and generalization relationships among them are identified. Second, the fragments are modeled using a visual language such as UML. The UML diagrams are both syntactically restricted (in order to guarantee a formal semantics), and enriched with a highly controlled natural language (to allow for modeling static and temporal constraints). Third, an automatic formal analysis phase iterates over the modeled requirements, by combining several, complementary techniques: checking consistency; verifying whether the requirements entail some desirable properties; verify whether the requirements are consistent with selected scenarios; diagnosing inconsistencies by identifying inconsistent cores; identifying vacuous requirements; constructing multiple explanations by enabling the fault-tree analysis related to particular fault models; verifying whether the specification is realizable.

Cimatti, Alessandro↗

Analysis of a hardware and software fault tolerant processor for critical applications

Computer systems for critical applications must be designed to tolerate software faults as well as hardware faults. A unified approach to tolerating hardware and software faults is characterized by classifying faults in terms of duration (transient or permanent) rather than source (hardware or software). Errors arising from transient faults can be handled through masking or voting, but errors arising from permanent faults require system reconfiguration to bypass the failed component. Most errors which are caused by software faults can be considered transient, in that they are input-dependent. Software faults are triggered by a particular set of inputs. Quantitative dependability analysis of systems which exhibit a unified approach to fault tolerance can be performed by a hierarchical combination of fault tree and Markov models. A methodology for analyzing hardware and software fault tolerant systems is applied to the analysis of a hypothetical system, loosely based on the Fault Tolerant Parallel Processor. The models consider both transient and permanent faults, hardware and software faults, independent and related software faults, automatic recovery, and reconfiguration.

Dugan, Joanne B.↗

Emulation applied to reliability analysis of reconfigurable, highly reliable, fault-tolerant computing systems

Emulation techniques applied to the analysis of the reliability of highly reliable computer systems for future commercial aircraft are described. The lack of credible precision in reliability estimates obtained by analytical modeling techniques is first established. The difficulty is shown to be an unavoidable consequence of: (1) a high reliability requirement so demanding as to make system evaluation by use testing infeasible; (2) a complex system design technique, fault tolerance; (3) system reliability dominated by errors due to flaws in the system definition; and (4) elaborate analytical modeling techniques whose precision outputs are quite sensitive to errors of approximation in their input data. Next, the technique of emulation is described, indicating how its input is a simple description of the logical structure of a system and its output is the consequent behavior. Use of emulation techniques is discussed for pseudo-testing systems to evaluate bounds on the parameter values needed for the analytical techniques. Finally an illustrative example is presented to demonstrate from actual use the promise of the proposed application of emulation.

Migneault, G. E.↗

Evaluating Faulty State Occurrence in Wildfire UAS Missions Using Markov Chains

As autonomous technology advances, unmanned aircraft systems are increasingly integrated into emergency response missions, such as wildfire response. These systems must be be safe with less risk than non-autonomous counter parts, yet quantifying the risk associated with present-day and future systems conventionally relies solely on expert opinion and little data. Instead, combining narrative mishap reports with probabilistic analysis can provide a method for evolutionary and timely risk analysis. In this paper, we present a framework for a data-driven probabilistic risk assessment style analysis, where hazard events and rates originate from documented UAS mishaps. The framework is applied to a UAS mapping mission in wildfire response, including a fault tree analysis, event tree analysis, and probabilistic analysis using Markov Chains. The analysis provides an enumeration of hazards in the system, hazard events that can lead to faults, the probability of a mission experiencing any fault, the probability of experiencing a specific fault, and the expected time spent until faulty states occur in present-day operations.

risk analysis↗

Evaluating Faulty State Occurrence in Wildfire UAS Missions Using Markov Chains

As autonomous technology advances, unmanned aircraft systems are increasingly integrated into emergency response missions, such as wildfire response. These systems must be be safe with less risk than non-autonomous counter parts, yet quantifying the risk associated with present-day and future systems conventionally relies solely on expert opinion and little data. Instead, combining narrative mishap reports with probabilistic analysis can provide a method for evolutionary and timely risk analysis. In this paper, we present a framework for a data-driven probabilistic risk assessment style analysis, where hazard events and rates originate from documented UAS mishaps. The framework is applied to a UAS mapping mission in wildfire response, including a fault tree analysis, event tree analysis, and probabilistic analysis using Markov Chains. The analysis provides an enumeration of hazards in the system, hazard events that can lead to faults, the probability of a mission experiencing any fault, the probability of experiencing a specific fault, and the expected time spent until faulty states occur in present-day operations.

risk analysis↗

Monte Carlo Simulation of Markov, Semi-Markov, and Generalized Semi- Markov Processes in Probabilistic Risk Assessment

A standard tool of reliability analysis used at NASA-JSC is the event tree. An event tree is simply a probability tree, with the probabilities determining the next step through the tree specified at each node. The nodal probabilities are determined by a reliability study of the physical system at work for a particular node. The reliability study performed at a node is typically referred to as a fault tree analysis, with the potential of a fault tree existing.for each node on the event tree. When examining an event tree it is obvious why the event tree/fault tree approach has been adopted. Typical event trees are quite complex in nature, and the event tree/fault tree approach provides a systematic and organized approach to reliability analysis. The purpose of this study was two fold. Firstly, we wanted to explore the possibility that a semi-Markov process can create dependencies between sojourn times (the times it takes to transition from one state to the next) that can decrease the uncertainty when estimating time to failures. Using a generalized semi-Markov model, we studied a four element reliability model and were able to demonstrate such sojourn time dependencies. Secondly, we wanted to study the use of semi-Markov processes to introduce a time variable into the event tree diagrams that are commonly developed in PRA (Probabilistic Risk Assessment) analyses. Event tree end states which change with time are more representative of failure scenarios than are the usual static probability-derived end states.

English, Thomas↗

Elevation changes near the San Gabriel Fault, Southern California

Analysis of repeated leveling observations in the vicinity of the San Gabriel Fault in Southern California indicate subsidence immediately south of the Fault relative to points to the north, south and east. These observations were previously interpreted as reflecting tectonic motions associated with either the 'Palmdale Bulge' or with preseismic effects of the San Fernando earthquake. Relative subsidence between 1953 and 1964 reaches approximately 9 cm and extends over a distance of more than 20 km. Subsidence occurs directly above the Saugus aquifer and shows a temporal correlation with the history of water level decline within the aquifer. The degree of subsidence of individual benchmarks is roughly proportional to the product of aquifer thickness and water level decline at the location of the benchmarks. Thses observations strongly suggest that movements of the surface near the San Gabriel Fault, previously inferred to be of tectonic origin, actually result from near surface sediment compaction within the Saugus basin.

Reilinger, R.↗

Crustal dynamics studies in China

Geodynamics of Mainland China and Taiwan are discussed. The following research was performed: (1) the tectonics along the Tanlu fault in eastern China; (2) tectonics in the Taiwan Strait behind the collision zone in Taiwan; and (3) analysis of faulting in the vicinity of the Altyn Tagn fault. It is found that the existence of the fault is traced back to at least Jurassic with the deposition of conglomerate sandstones in the troungh along the present Tanlu fault branches in the Shantung Province. Taiwan is the product of collision between the Phillipine plate and the Asian plate and Taiwan came into being because of a former island arc.

Wu, F. T.↗

Model Transformation for a System of Systems Dependability Safety Case

Software plays an increasingly larger role in all aspects of NASA's science missions. This has been extended to the identification, management and control of faults which affect safety-critical functions and by default, the overall success of the mission. Traditionally, the analysis of fault identification, management and control are hardware based. Due to the increasing complexity of system, there has been a corresponding increase in the complexity in fault management software. The NASA Independent Validation & Verification (IV&V) program is creating processes and procedures to identify, and incorporate safety-critical software requirements along with corresponding software faults so that potential hazards may be mitigated. This Specific to Generic ... A Case for Reuse paper describes the phases of a dependability and safety study which identifies a new, process to create a foundation for reusable assets. These assets support the identification and management of specific software faults and, their transformation from specific to generic software faults. This approach also has applications to other systems outside of the NASA environment. This paper addresses how a mission specific dependability and safety case is being transformed to a generic dependability and safety case which can be reused for any type of space mission with an emphasis on software fault conditions.

Murphy, Judy↗

A Review of Diagnostic Techniques for ISHM Applications

System diagnosis is an integral part of any Integrated System Health Management application. Diagnostic applications make use of system information from the design phase, such as safety and mission assurance analysis, failure modes and effects analysis, hazards analysis, functional models, fault propagation models, and testability analysis. In modern process control and equipment monitoring systems, topological and analytic , models of the nominal system, derived from design documents, are also employed for fault isolation and identification. Depending on the complexity of the monitored signals from the physical system, diagnostic applications may involve straightforward trending and feature extraction techniques to retrieve the parameters of importance from the sensor streams. They also may involve very complex analysis routines, such as signal processing, learning or classification methods to derive the parameters of importance to diagnosis. The process that is used to diagnose anomalous conditions from monitored system signals varies widely across the different approaches to system diagnosis. Rule-based expert systems, case-based reasoning systems, model-based reasoning systems, learning systems, and probabilistic reasoning systems are examples of the many diverse approaches ta diagnostic reasoning. Many engineering disciplines have specific approaches to modeling, monitoring and diagnosing anomalous conditions. Therefore, there is no "one-size-fits-all" approach to building diagnostic and health monitoring capabilities for a system. For instance, the conventional approaches to diagnosing failures in rotorcraft applications are very different from those used in communications systems. Further, online and offline automated diagnostic applications are integrated into an operations framework with flight crews, flight controllers and maintenance teams. While the emphasis of this paper is automation of health management functions, striking the correct balance between automated and human-performed tasks is a vital concern.

Patterson-Hine, Ann↗