Search NASASearch

SEARCH · Search NASA

Results for “software fault model”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 55 records · Page 3

Chip level simulation of fault tolerant computers

Chip level modeling techniques, functional fault simulation, simulation software development, a more efficient, high level version of GSP, and a parallel architecture for functional simulation are discussed.

Armstrong, J. R.

A theoretical basis for the analysis of redundant software subject to coincident errors

Fundamental to the development of redundant software techniques fault-tolerant software, is an understanding of the impact of multiple-joint occurrences of coincident errors. A theoretical basis for the study of redundant software is developed which provides a probabilistic framework for empirically evaluating the effectiveness of the general (N-Version) strategy when component versions are subject to coincident errors, and permits an analytical study of the effects of these errors. The basic assumptions of the model are: (1) independently designed software components are chosen in a random sample; and (2) in the user environment, the system is required to execute on a stationary input series. The intensity of coincident errors, has a central role in the model. This function describes the propensity to introduce design faults in such a way that software components fail together when executing in the user environment. The model is used to give conditions under which an N-Version system is a better strategy for reducing system failure probability than relying on a single version of software. A condition which limits the effectiveness of a fault-tolerant strategy is studied, and it is posted whether system failure probability varies monotonically with increasing N or whether an optimal choice of N exists.

Eckhardt, D. E., Jr.

Measurement and analysis of operating system fault tolerance

This paper demonstrates a methodology to model and evaluate the fault tolerance characteristics of operational software. The methodology is illustrated through case studies on three different operating systems: the Tandem GUARDIAN fault-tolerant system, the VAX/VMS distributed system, and the IBM/MVS system. Measurements are made on these systems for substantial periods to collect software error and recovery data. In addition to investigating basic dependability characteristics such as major software problems and error distributions, we develop two levels of models to describe error and recovery processes inside an operating system and on multiple instances of an operating system running in a distributed environment. Based on the models, reward analysis is conducted to evaluate the loss of service due to software errors and the effect of the fault-tolerance techniques implemented in the systems. Software error correlation in multicomputer systems is also investigated.

Lee, I.

A Sensitivity-driven Wide Area Protection (SWAP) Coordination Tool for High Penetration of Inverter-based Resources (IBR)

Traditionally, power system generation sources have been composed of synchronous generators, of which the fault current behavior is understood with minimal differences between generation size and types due to the physics of their construction. Present protection schemes and modeling methods are based upon these understood characteristics. Most renewable generation is composed of inverter-based resources (IBR), in which fault current is determined by switching control software and hardware limitations, each of which can vary between manufacturers and even between models of the same manufacturer. The resulting fault current is low in magnitude, low in negative-sequence current, unpredictable phase angles, and is a challenge to model. These characteristics also result in a challenge to traditional protection schemes and fault simulation software. To address several of these concerns, the project has the following goals: 1. Improve IBR models: Improve IBR models used in short circuit (SC) programs to accurately capture the response of IBRs at the bulk power system (BPS) level for fault and protection studies. 2. Develop automation tool: Develop an automation tool that allows engineers to identify protection coordination and sensitivity issues by performing SC and protection coordination studies in a high IBR-penetrated grid by applying variations to the IBR models, faults, contingencies, etc. 3. Develop schemes: Develop new protection mitigation solution schemes that complement the existing protection systems to ensure safe operation of the BPS with higher IBR penetration levels. The project team did not achieve this final goal, as the Department of Energy (DOE) stopped the project early due to changes in DOE funding priorities. The termination notice came at the beginning of the final project phase, while the team was identifying and beginning to investigate protection issues. It should be noted that the team discussed a 100% penetration scenario. However, this scenario would require the use of grid-forming IBR models that are not presently available. Since developing these models requires additional effort, the 100% penetration scenario was not pursued during this project. In the future, developing the methodology and models for the 100% scenario could benefit the industry.

14 SOLAR ENERGY

Measurement of fault latency in a digital avionic miniprocessor

The results of fault injection experiments utilizing a gate-level emulation of the central processor unit of the Bendix BDX-930 digital computer are presented. The failure detection coverage of comparison-monitoring and a typical avionics CPU self-test program was determined. The specific tasks and experiments included: (1) inject randomly selected gate-level and pin-level faults and emulate six software programs using comparison-monitoring to detect the faults; (2) based upon the derived empirical data develop and validate a model of fault latency that will forecast a software program's detecting ability; (3) given a typical avionics self-test program, inject randomly selected faults at both the gate-level and pin-level and determine the proportion of faults detected; (4) determine why faults were undetected; (5) recommend how the emulation can be extended to multiprocessor systems such as SIFT; and (6) determine the proportion of faults detected by a uniprocessor BIT (built-in-test) irrespective of self-test.

Mcgough, J. G.

The 13th Technology of Deep Space One

On October 24th, 1998, the Deep Space One (DS-1) spacecraft launched aboard a Delta II rocket as the first step towards the bold task of testing and validating 12 new technologies for future missions. This launch also represented yet another thrilling event; namely, the successful test and validation of a 13th heretofore undisclosed technology: model-base-code-generation of the spacecraft's system-level fault-protection (FP) software from behavioral state diagrams and structural models.

model-based-code-generation

The 13th Technology of Deep Space One - Abstract

On October 24th, 1998, the Deep Space One (DS-1) spacecraft launched aboard a Delta II rocket as the first step towards the bold task of testing and validating 12 new technologies for future missions. This launch also represented yet another thrilling event; namely, the successful test and validation of a 13th heretofore undisclosed technology: model-based code-generation of the spacecraft's system-level fault-protection (FP) software from behavioral state diagrams and structural models.In this paper, we describe the process we used to leverage model-based code generation from state diagrams and structural specifications to better respond to the evolving requirements and scope of DS- I's system-level fault-protection design, development, test and operation. The evolution of the high-level design and the low-level changes in the flight software architecture and interfaces contributed to multiplying the number and frequency of fault-protection software releases thereby creating a multitude of software integration issues. To address the resulting software integration issues, we broadened the scope of code -eneration to other forms of model- based analysis techniques more traditionally associated with first-principle's reasoning about physical models. Additionally, we describe our in-flight launch and initial acquisition experience.

Rouquette, Nicolas

Multi-version software reliability through fault-avoidance and fault-tolerance

A number of experimental and theoretical issues associated with the practical use of multi-version software to provide run-time tolerance to software faults were investigated. A specialized tool was developed and evaluated for measuring testing coverage for a variety of metrics. The tool was used to collect information on the relationships between software faults and coverage provided by the testing process as measured by different metrics (including data flow metrics). Considerable correlation was found between coverage provided by some higher metrics and the elimination of faults in the code. Back-to-back testing was continued as an efficient mechanism for removal of un-correlated faults, and common-cause faults of variable span. Software reliability estimation methods was also continued based on non-random sampling, and the relationship between software reliability and code coverage provided through testing. New fault tolerance models were formulated. Simulation studies of the Acceptance Voting and Multi-stage Voting algorithms were finished and it was found that these two schemes for software fault tolerance are superior in many respects to some commonly used schemes. Particularly encouraging are the safety properties of the Acceptance testing scheme.

Vouk, Mladen A.

On-Board Model Based Fault Diagnosis for CubeSat Attitude Control Subsystem: Flight Data Results

Self-sufficient, robotic spacecraft require estimates of their hardware health state in order to project future system state and plan actions toward achieving mission goals. In this paper, we report on integration of a Model-Based Fault Diagnosis (MBFD) model and reasoning engine into flight software leveraging the Arcsecond Space Telescope Enabling Research in Astrophysics (ASTERIA) mission, including test results against captured flight data using the ASTERIA system testbed. Our effort integrated the Model-based Off-Nominal State Identification and Detection (MONSID) model-based reasoning system, developed by Okean Solutions, into ASTERIA flight software using the F Prime software framework. The MONSID engine was supplied with a model of the Blue Canyon Technologies XACT attitude control system (ACS) and tested against flight data and seeded fault tests. While we were unable to conduct an on-board experiment due to the premature loss of ASTERIA, our effort proved the feasibility of on-board model-based fault management, demonstrating reliable and accurate diagnosis using captured data, and further supporting a closed-loop spacecraft autonomy demonstration including autonomous navigation in off-nominal conditions.

Prather, Maurice

Tutorial: Advanced fault tree applications using HARP

Reliability analysis of fault tolerant computer systems for critical applications is complicated by several factors. These modeling difficulties are discussed and dynamic fault tree modeling techniques for handling them are described and demonstrated. Several advanced fault tolerant computer systems are described, and fault tree models for their analysis are presented. HARP (Hybrid Automated Reliability Predictor) is a software package developed at Duke University and NASA Langley Research Center that is capable of solving the fault tree models presented.

Dugan, Joanne Bechta

Semi-Markov Unreliability-Range Evaluator

Reconfigurable, fault-tolerant systems modeled. Semi-Markov unreliability-range evaluator (SURE) computer program is software tool for analysis of reliability of reconfigurable, fault-tolerant systems. Based on new method for computing death-state probabilities of semi-Markov model. Computes accurate upper and lower bounds on probability of failure of system. Written in PASCAL.

Butler, Ricky W.

Effectiveness of back-to-back testing

Three models of back-to-back testing processes are described. Two models treat the case where there is no intercomponent failure dependence. The third model describes the more realistic case where there is correlation among the failure probabilities of the functionally equivalent components. The theory indicates that back-to-back testing can, under the right conditions, provide a considerable gain in software reliability. The models are used to analyze the data obtained in a fault-tolerant software experiment. It is shown that the expected gain is indeed achieved, and exceeded, provided the intercomponent failure dependence is sufficiently small. However, even with the relatively high correlation the use of several functionally equivalent components coupled with back-to-back testing may provide a considerable reliability gain. Implications of this finding are that the multiversion software development is a feasible and cost effective approach to providing highly reliable software components intended for fault-tolerant software systems, on condition that special attention is directed at early detection and elimination of correlated faults.

Vouk, Mladen A.

Graphical workstation capability for reliability modeling

In addition to computational capabilities, software tools for estimating the reliability of fault-tolerant digital computer systems must also provide a means of interfacing with the user. Described here is the new graphical interface capability of the hybrid automated reliability predictor (HARP), a software package that implements advanced reliability modeling techniques. The graphics oriented (GO) module provides the user with a graphical language for modeling system failure modes through the selection of various fault-tree gates, including sequence-dependency gates, or by a Markov chain. By using this graphical input language, a fault tree becomes a convenient notation for describing a system. In accounting for any sequence dependencies, HARP converts the fault-tree notation to a complex stochastic process that is reduced to a Markov chain, which it can then solve for system reliability. The graphics capability is available for use on an IBM-compatible PC, a Sun, and a VAX workstation. The GO module is written in the C programming language and uses the graphical kernal system (GKS) standard for graphics implementation. The PC, VAX, and Sun versions of the HARP GO module are currently in beta-testing stages.

Bavuso, Salvatore J.

Model-Based Fault Diagnosis: Performing Root Cause and Impact Analyses in Real Time

Generic, object-oriented fault models, built according to causal-directed graph theory, have been integrated into an overall software architecture dedicated to monitoring and predicting the health of mission- critical systems. Processing over the generic fault models is triggered by event detection logic that is defined according to the specific functional requirements of the system and its components. Once triggered, the fault models provide an automated way for performing both upstream root cause analysis (RCA), and for predicting downstream effects or impact analysis. The methodology has been applied to integrated system health management (ISHM) implementations at NASA SSC's Rocket Engine Test Stands (RETS).

Figueroa, Jorge F.

Study of fault-tolerant software technology

Presented is an overview of the current state of the art of fault-tolerant software and an analysis of quantitative techniques and models developed to assess its impact. It examines research efforts as well as experience gained from commercial application of these techniques. The paper also addresses the computer architecture and design implications on hardware, operating systems and programming languages (including Ada) of using fault-tolerant software in real-time aerospace applications. It concludes that fault-tolerant software has progressed beyond the pure research state. The paper also finds that, although not perfectly matched, newer architectural and language capabilities provide many of the notations and functions needed to effectively and efficiently implement software fault-tolerance.

Slivinski, T.

Repurposing Drilling Control Diagnostics for Subsurface Edge Detection and Boundary Advisement During Planetary Drilling

Informed decision-making during lunar drilling and sampling missions will require data monitoring tools and specialized ground data systems. Accurate and updated situational awareness, with ongoing data monitoring, is critical for timely responses by to incoming science data. Traverse plans and scheduled activities may need to be flexibly changed in order to react to unexpected data or situations. Unlike (for example) Mars missions, the relative lightspeed closeness of the Moon allows for near-real-time ground processing of incoming mission and instrument data. An Apollo-class lunar regolith drill will in a sense “travel” a meter or two vertically at a given subsurface characterization site. As the drill penetrates into lunar regolith, it is likely to encounter a range of material densities, orientations, fracture toughness, and (perhaps) ice percentages. Lunar drill telemetry can provide science teams with a valuable first look into the subsurface structure, the regolith bulk properties, and constituents at each drilled site. Real-time AI-based recognition and reaction to downhole situations has been developed for automated deeper drilling on Mars and beyond. We can leverage the same knowledge bases and pattern-matching as areal-time interpreter of the subsurface, a situational awareness tool during drilling operations. We recently (Sept. 2019) demonstrated this AI drilling monitoring and analysis capability, in control of in-situ drilling and sampling operations, mounted on a KREX-2 rover in Chile’s Atacama Desert. Terrestrial automated drilling log analyses in oil exploration have used similar machine learning techniques in classifying and identifying features in drilling logs –but these typically are designed assuming a drilling fluid influencing downhole measurements and data (permeability, resistivity). Drilling models and existing AI software designed to detect and respond to drilling faults and hard materials can be repurposed, for near-real-time (ground-based) interpretation of drilling telemetry –a potentially valuable advisory tool for strata boundaries and changes in drilling parameters. On the Moon, this approach could be used to study the structure and to some extent the composition of lunar regolith vs. borehole depth, based on recognizable variations in fracture hardness, drilling energy and penetration rates while actively drilling. Since the early 2000s, a series of increasingly-capable real-time drilling telemetry interpretation and characterization software tools have been developed. These subsurface models and software tools have monitored the real-time drilling data received, and automatically identified changes in drill behavior (e.g., encountering a harder target layer, bit inclusions, drill choking due to infall downhole, and others) correlating these with subsurface structures and features. We discuss the mappings between drill borehole parameters, faults or events detected, and modeled changes in rock layer boundaries, in examples drawn from field testing at analog sites in an Arctic impact crater, Rio Tinto, and Chile’s Atacama Desert. These demonstrate how subsurface structural boundaries led to fault detections and responses by the software.

drilling advisor

Model authoring system for fail safe analysis

The Model Authoring System is a prototype software application for generating fault tree analyses and failure mode and effects analyses for circuit designs. Utilizing established artificial intelligence and expert system techniques, the circuits are modeled as a frame-based knowledge base in an expert system shell, which allows the use of object oriented programming and an inference engine. The behavior of the circuit is then captured through IF-THEN rules, which then are searched to generate either a graphical fault tree analysis or failure modes and effects analysis. Sophisticated authoring techniques allow the circuit to be easily modeled, permit its behavior to be quickly defined, and provide abstraction features to deal with complexity.

Sikora, Scott E.

Lessons Learned on Implementing Fault Detection, Isolation, and Recovery (FDIR) in a Ground Launch Environment

This paper's main purpose is to detail issues and lessons learned regarding designing, integrating, and implementing Fault Detection Isolation and Recovery (FDIR) for Constellation Exploration Program (CxP) Ground Operations at Kennedy Space Center (KSC). Part of the0 overall implementation of National Aeronautics and Space Administration's (NASA's) CxP, FDIR is being implemented in three main components of the program (Ares, Orion, and Ground Operations/Processing). While not initially part of the design baseline for the CxP Ground Operations, NASA felt that FDIR is important enough to develop, that NASA's Exploration Systems Mission Directorate's (ESMD's) Exploration Technology Development Program (ETDP) initiated a task for it under their Integrated System Health Management (ISHM) research area. This task, referred to as the FDIIR project, is a multi-year multi-center effort. The primary purpose of the FDIR project is to develop a prototype and pathway upon which Fault Detection and Isolation (FDI) may be transitioned into the Ground Operations baseline. Currently, Qualtech Systems Inc (QSI) Commercial Off The Shelf (COTS) software products Testability Engineering and Maintenance System (TEAMS) Designer and TEAMS RDS/RT are being utilized in the implementation of FDI within the FDIR project. The TEAMS Designer COTS software product is being utilized to model the system with Functional Fault Models (FFMs). A limited set of systems in Ground Operations are being modeled by the FDIR project, and the entire Ares Launch Vehicle is being modeled under the Functional Fault Analysis (FFA) project at Marshall Space Flight Center (MSFC). Integration of the Ares FFMs and the Ground Processing FFMs is being done under the FDIR project also utilizing the TEAMS Designer COTS software product. One of the most significant challenges related to integration is to ensure that FFMs developed by different organizations can be integrated easily and without errors. Software Interface Control Documents (ICDs) for the FFMs and their usage will be addressed as the solution to this issue. In particular, the advantages and disadvantages of these ICDs across physically separate development groups will be delineated.

Ferrell, Bob A.