Search NASA⌕ Search

SEARCH · Search NASA

Results for “fault analysis”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 109 records · Page 6

A Reference Model for Software and System Inspections. White Paper

Software Quality Assurance (SQA) is an important component of the software development process. SQA processes provide assurance that the software products and processes in the project life cycle conform to their specified requirements by planning, enacting, and performing a set of activities to provide adequate confidence that quality is being built into the software. Typical techniques include: (1) Testing (2) Simulation (3) Model checking (4) Symbolic execution (5) Management reviews (6) Technical reviews (7) Inspections (8) Walk-throughs (9) Audits (10) Analysis (complexity analysis, control flow analysis, algorithmic analysis) (11) Formal method Our work over the last few years has resulted in substantial knowledge about SQA techniques, especially the areas of technical reviews and inspections. But can we apply the same QA techniques to the system development process? If yes, what kind of tailoring do we need before applying them in the system engineering context? If not, what types of QA techniques are actually used at system level? And, is there any room for improvement.) After a brief examination of the system engineering literature (especially focused on NASA and DoD guidance) we found that: (1) System and software development process interact with each other at different phases through development life cycle (2) Reviews are emphasized in both system and software development. (Figl.3). For some reviews (e.g. SRR, PDR, CDR), there are both system versions and software versions. (3) Analysis techniques are emphasized (e.g. Fault Tree Analysis, Preliminary Hazard Analysis) and some details are given about how to apply them. (4) Reviews are expected to use the outputs of the analysis techniques. In other words, these particular analyses are usually conducted in preparation for (before) reviews. The goal of our work is to explore the interaction between the Quality Assurance (QA) techniques at the system level and the software level.

He, Lulu↗

Detection of faults and software reliability analysis

Specific topics briefly addressed include: the consistent comparison problem in N-version system; analytic models of comparison testing; fault tolerance through data diversity; and the relationship between failures caused by automatically seeded faults.

Knight, J. C.↗

Probabilistic evaluation of on-line checks in fault-tolerant multiprocessor systems

The analysis of fault-tolerant multiprocessor systems that use concurrent error detection (CED) schemes is much more difficult than the analysis of conventional fault-tolerant architectures. Various analytical techniques have been proposed to evaluate CED schemes deterministically. However, these approaches are based on worst-case assumptions related to the failure of system components. Often, the evaluation results do not reflect the actual fault tolerance capabilities of the system. A probabilistic approach to evaluate the fault detecting and locating capabilities of on-line checks in a system is developed. The various probabilities associated with the checking schemes are identified and used in the framework of the matrix-based model. Based on these probabilistic matrices, estimates for the fault tolerance capabilities of various systems are derived analytically.

Nair, V. S. S.↗

Methodology for Designing Fault-Protection Software

A document describes a methodology for designing fault-protection (FP) software for autonomous spacecraft. The methodology embodies and extends established engineering practices in the technical discipline of Fault Detection, Diagnosis, Mitigation, and Recovery; and has been successfully implemented in the Deep Impact Spacecraft, a NASA Discovery mission. Based on established concepts of Fault Monitors and Responses, this FP methodology extends the notion of Opinion, Symptom, Alarm (aka Fault), and Response with numerous new notions, sub-notions, software constructs, and logic and timing gates. For example, Monitor generates a RawOpinion, which graduates into Opinion, categorized into no-opinion, acceptable, or unacceptable opinion. RaiseSymptom, ForceSymptom, and ClearSymptom govern the establishment and then mapping to an Alarm (aka Fault). Local Response is distinguished from FP System Response. A 1-to-n and n-to- 1 mapping is established among Monitors, Symptoms, and Responses. Responses are categorized by device versus by function. Responses operate in tiers, where the early tiers attempt to resolve the Fault in a localized step-by-step fashion, relegating more system-level response to later tier(s). Recovery actions are gated by epoch recovery timing, enabling strategy, urgency, MaxRetry gate, hardware availability, hazardous versus ordinary fault, and many other priority gates. This methodology is systematic, logical, and uses multiple linked tables, parameter files, and recovery command sequences. The credibility of the FP design is proven via a fault-tree analysis "top-down" approach, and a functional fault-mode-effects-and-analysis via "bottoms-up" approach. Via this process, the mitigation and recovery strategy(s) per Fault Containment Region scope (width versus depth) the FP architecture.

Barltrop, Kevin↗

Assessing Seismic Risk for CO2 Geologic Storage: Comparative Analysis of the Delaware Basin and Basin and Range Province Projects

ABSTRACT: Effective management of induced seismicity is critical for safe and sustainable CO2 storage. This study evaluates fault slippage risks in the Delaware Basin (Texas) and Basin and Range Province (Utah), integrating geological, operational, and geomechanical parameters to assess fault stability and seismic hazard mitigation. In the Delaware Basin, two sites were analyzed under an injection rate of 20,000 bbl/day over 25 years. One site showed low fault slip risk, while the other exhibited higher reactivation potential due to proximity to critically stressed faults. Sensitivity analysis revealed that increased pore pressure significantly heightened slip potential, highlighting the necessity of precise pressure control and real-time monitoring. In the Basin and Range Province, fault stability was evaluated at Neck of the Desert, Escalante Desert, Parowan, and Beaver sites under injection rates of 8,750 bbl/day per site over 30 years. Minimal fault slip risk was observed at Neck of the Desert and Escalante Desert sites, whereas Parowan and Beaver sites exhibited elevated slip potential due to semi-critically stressed faults sensitive to modest pore pressure increases. The findings demonstrate that fault slippage analysis, combined with sensitivity analysis of pore pressure and friction coefficients, is essential for understanding seismic risks. Continuous monitoring, adaptive injection management, and rigorous geomechanical analysis are key strategies for minimizing induced seismicity in CO2 sequestration projects.

58 GEOSCIENCES↗

SIFT - Design and analysis of a fault-tolerant computer for aircraft control

SIFT (Software Implemented Fault Tolerance) is an ultrareliable computer for critical aircraft control applications that achieves fault tolerance by the replication of tasks among processing units. The main processing units are off-the-shelf minicomputers, with standard microcomputers serving as the interface to the I/O system. Fault isolation is achieved by using a specially designed redundant bus system to interconnect the processing units. Error detection and analysis and system reconfiguration are performed by software. Iterative tasks are redundantly executed, and the results of each iteration are voted upon before being used. Thus, any single failure in a processing unit or bus can be tolerated with triplication of tasks, and subsequent failures can be tolerated after reconfiguration. Independent execution by separate processors means that the processors need only be loosely synchronized, and a novel fault-tolerant synchronization method is described.

Wensley, J. H.↗

Fault Slip and Fluid Flow: Seismic Source Analysis to Assess Role of Multiple Slip Patches in Fault Permeability

The relationship between fault reactivation, microearthquakes (MEQs), and permeability evolution during fluid injection plays a critical role in energy harvesting and waste disposal. Recent studies have demonstrated the possibility of predicting fault permeability using cumulative seismic moments of MEQs quantitatively. To understand the underlying physical processes, we conduct fault reactivation experiments using Utah FORGE granitoid and analyze acoustic emission (AE) signals generated during stepwise increases in fluid injection pressure. Frequency analysis of thousands of calibrated AE signals reveals that fault reactivation produces multiple AE source patches with millimeter-scale radii—smaller than the sample fault radius. The cumulative area of the reactivated patches covers the fault multiple times over (∼10x–50x area) for each pressure step. These findings provide mechanistic insight that measured permeability enhancement is not driven by a single large slip event, but by the sequential and interacting activation of multiple slip patches that create a continuous flow pathway.

Nurshal, M. E. M. [Pennsylvania State University, ↗

A systematic risk management approach employed on the CloudSat project

The CloudSat Project has developed a simplified approach for fault tree analysis and probabilistic risk assessment. A system-level fault tree has been constructed to identify credible fault scenarios and failure modes leading up to a potential failure to meet the nominal mission success criteria.

risk management fault tree analysis probabilistic ↗

Faults Discovery By Using Mined Data

Fault discovery in the complex systems consist of model based reasoning, fault tree analysis, rule based inference methods, and other approaches. Model based reasoning builds models for the systems either by mathematic formulations or by experiment model. Fault Tree Analysis shows the possible causes of a system malfunction by enumerating the suspect components and their respective failure modes that may have induced the problem. The rule based inference build the model based on the expert knowledge. Those models and methods have one thing in common; they have presumed some prior-conditions. Complex systems often use fault trees to analyze the faults. Fault diagnosis, when error occurs, is performed by engineers and analysts performing extensive examination of all data gathered during the mission. International Space Station (ISS) control center operates on the data feedback from the system and decisions are made based on threshold values by using fault trees. Since those decision-making tasks are safety critical and must be done promptly, the engineers who manually analyze the data are facing time challenge. To automate this process, this paper present an approach that uses decision trees to discover fault from data in real-time and capture the contents of fault trees as the initial state of the trees.

Lee, Charles↗

Considerations for Introducing Artificial Intelligence into Nuclear Power Plants

Advanced computational tools and techniques such as artificial intelligence and machine learning (AI/ML) can transform the nuclear power industry. This is necessary given that the economic viability of the existing fleet is in jeopardy and its labor-centric approach to operations and maintenance. Currently, AI/ML research is being undertaken for reactor system design and analysis including fault and accident prognosis, nuclear risk analysis such as plant safety and security evaluation, and plant operations and maintenance including predictive maintenance. Applications include both existing and advanced reactor technologies with the aim of improving operational and business efficiencies. Most every aspect of the organization can benefit, from instrumentation and control, to work planning, to human-machine interactions and business management. AI/ML in nuclear can simplify complex problems and produce more effective decision-making. Nonetheless, careful consideration must be given to the implementation of an AI/ML initiative. The aims of this research are to 1) review barriers to AI/ML adoption within the nuclear power industry, and 2) suggest potential solutions. These barriers are organized along five distinct categories (Figure 1) that are interconnected. The first are historical barriers that track the industry’s development over the decades including worldwide nuclear events that shaped public perceptions. The resulting federal scrutiny and intense safety culture that emerged are discussed. Technical barriers to AI/ML adoption are considerable, and include data privacy concerns, data governance, and the current lack of AI/ML expert knowledge at the plants. The main business case barrier remains cost, but an absence of an industry-wide vision and wide-scale adoption also produces reluctance. Stakeholder readiness is reviewed with special attention given to regulatory readiness. The 5-year strategic plan for AI readiness recently published by the U.S. Nuclear Regulatory Commission is highlighted. Last, adoption barriers at the user level are addressed including the importance of user experience and explainable AI. The AI adoption barriers described here are inter-related and ideally should be addressed in a holistic fashion.

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND ↗

Model-Based Testability Assessment and Directed Troubleshooting of Shuttle Wiring Systems

We have recently completed a pilot study on the Space shuttle wiring system commissioned by the Wiring Integrity Research (WIRe) team at NASA Ames Research Center, As the space shuttle ages, it is experiencing wiring degradation problems including arcing, chaffing insulation breakdown and broken conductors. A systematic and comprehensive test process is required to thoroughly test and quality assure (QA) the wiring systems. The NASA WIRe team recognized the value of a formal model based analysis for risk-assessment and fault coverage analysis. However. wiring systems are complex and involve over 50,000 wire segments. Therefore, NASA commissioned this pilot study with Qualtech Systems. Inc. (QSI) to explore means of automatically extracting high fidelity multi-signal models from wiring information database for use with QSI's Testability Engineering and Maintenance System (TEAMS) tool.

Deb, Somnath↗

System Safety Analysis of Complex NASA Systems with Model Based Engineering

The emergence of model-based engineering is transforming design and analysis methodologies [5]. A recognized benefit of model-based engineering is the existence of a “single source of truth” about the system that becomes the authoritative source of data and information for designers, analysts, and developers. This promotes consistency and efficiency as the design emerges and can be used to further optimize the design. Integrating System Safety Engineers to the “single source of truth” will ensure that the outputs of their assessments and analyses are relevant to the design as it evolves. Use of an integrated system model enables near immediate evaluation of a design change as well as development of operational processes for risk assessment and communication. Such models can enable efficient and timely analysis of system hazards (e.g., hazard fault tree analysis and procedure simulations) and produce complete, accurate, and more consistent products (e.g., hazard reports and safety requirement evaluations). Therefore, an agency-sponsored team at Goddard Space Flight Center (GSFC) recently completed a System Safety Study of modeling and testing capabilities as part of a Model-Based Safety and Mission Assurance Initiative (MBSMAI). Using an existing model developed for reliability analyses [1], GSFC modeling and system safety experts performed system safety analysis/modeling and produced safety products. The team evaluated model-based feasibility to support System Safety Engineering, developed safety analysis modeling processes, and identified tool capability advancement/development needs. These study results indicate model-based engineering is valid and useable for System

Model Based Engineering↗

Reliability database development for use with an object-oriented fault tree evaluation program

A description is given of the development of a fault-tree analysis method using object-oriented programming. In addition, the authors discuss the programs that have been developed or are under development to connect a fault-tree analysis routine to a reliability database. To assess the performance of the routines, a relational database simulating one of the nuclear power industry databases has been constructed. For a realistic assessment of the results of this project, the use of one of existing nuclear power reliability databases is planned.

Heger, A. Sharif↗

Measurement and analysis of operating system fault tolerance

This paper demonstrates a methodology to model and evaluate the fault tolerance characteristics of operational software. The methodology is illustrated through case studies on three different operating systems: the Tandem GUARDIAN fault-tolerant system, the VAX/VMS distributed system, and the IBM/MVS system. Measurements are made on these systems for substantial periods to collect software error and recovery data. In addition to investigating basic dependability characteristics such as major software problems and error distributions, we develop two levels of models to describe error and recovery processes inside an operating system and on multiple instances of an operating system running in a distributed environment. Based on the models, reward analysis is conducted to evaluate the loss of service due to software errors and the effect of the fault-tolerance techniques implemented in the systems. Software error correlation in multicomputer systems is also investigated.

Lee, I.↗

Evaluate the application of modal test and analysis processes to structural fault detection in MSFC - STS project elements

The Space Transportation System (STS) is a complex and expensive flight system intended to carry unique payloads into low Earth orbit and return. A catastrophic failure, such as STS 51-L, resulted in the loss of both human life as well as expensive and unique hardware. The impact of this incident reaffirms the need to do everything possible to ensure the integrity and reliability of STS. One means of achieving this goal is to expand the number of inspection technologies available. Reported here is the evaluation of the use of modal analysis and test techniques for the purpose of assessing the structural integrity of STS components for which Marshall Space Flight Center has responsibility. This entailed reviewing existing literature and developing a low-level experimental program determine the feasibility of using this technology for structural fault detection.

Springer, William T.↗

Automatic translation of digraph to fault-tree models

The author presents a technique for converting digraph models, including those models containing cycles, to a fault-tree format. A computer program which automatically performs this translation using an object-oriented representation of the models has been developed. The fault-trees resulting from translations can be used for fault-tree analysis and diagnosis. Programs to calculate fault-tree and digraph cut sets and perform diagnosis with fault-tree models have also been developed. The digraph to fault-tree translation system has been successfully tested on several digraphs of varying size and complexity. Details of some representative translation problems are presented. Most of the computation performed by the program is dedicated to finding minimal cut sets for digraph nodes in order to break cycles in the digraph. Fault-trees produced by the translator have been successfully used with NASA's Fault-Tree Diagnosis System (FTDS) to produce automated diagnostic systems.

Iverson, David L.↗