Search NASA⌕ Search

SEARCH · Search NASA

Results for “software fault model”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 271 records · Page 15

The SIFT computer and its development

Software Implemented Fault Tolerance (SIFT) is an aircraft control computer designed to allow failure probability of less than 10 to the -10th/hour. The system is based on advanced fault-tolerance computing and validation methodology. Since confirmation of reliability by observation is essentially impossible, system reliability is estimated by a Markov model. A mathematical proof is used to justify the validity of the Markov model. System design is represented by a hierarchy of abstract models, and the design proof comprises mathematical proofs that each model is, in fact, an elaboration of the next more abstract model.

Goldberg, J.↗

The Mars 2020 lander vision system field test

The Mars 2020 Lander Vision System estimates position relative to a map and provides this information to the spacecraft so that large hazards can be avoided during landing. The LVS is a new mission critical sensor and as such requires extensive validation. A field test conducted in May 2019 was the primary means to prove that the LVS will operate as designed. During this test over 600 independent real-time runs on engineering model LVS hardware and software were executed and clearly showed that it could meet a 40m position estimation requirement over a wide operational envelope. This paper will describe the test approach, operations and results. Specific examples as well as aggregate performance will be discussed along with off-nominal testing and fault recovery.

Montgomery, J.↗

The process group approach to reliable distributed computing

The difficulty of developing reliable distributed software is an impediment to applying distributed computing technology in many settings. Experience with the ISIS system suggests that a structured approach based on virtually synchronous process groups yields systems which are substantially easier to develop, fault-tolerance, and self-managing. Six years of research on ISIS are reviewed, describing the model, the types of applications to which ISIS was applied, and some of the reasoning that underlies a recent effort to redesign and reimplement ISIS as a much smaller, lightweight system.

Birman, Kenneth P.↗

Logic flowgraph methodology - A tool for modeling embedded systems

The logic flowgraph methodology (LFM), a method for modeling hardware in terms of its process parameters, has been extended to form an analytical tool for the analysis of integrated (hardware/software) embedded systems. In the software part of a given embedded system model, timing and the control flow among different software components are modeled by augmenting LFM with modified Petrinet structures. The objective of the use of such an augmented LFM model is to uncover possible errors and the potential for unanticipated software/hardware interactions. This is done by backtracking through the augmented LFM mode according to established procedures which allow the semiautomated construction of fault trees for any chosen state of the embedded system (top event). These fault trees, in turn, produce the possible combinations of lower-level states (events) that may lead to the top event.

Muthukumar, C. T.↗

The Mars 2020 Lander Vision System Field Test

The Mars 2020 Lander Vision System estimates position relative to a map and provides this information to the spacecraft so that large hazards can be avoided during landing. The LVS is a new mission critical sensor and as such requires extensive validation. A field test conducted in May 2019 was the primary means to prove that the LVS will operate as designed. During this test over 600 independent real-time runs on engineering model LVS hardware and software were executed and clearly showed that it could meet a 40m position estimation requirement over a wide operational envelope. This paper will describe the test approach, operations and results. Specific examples as well as aggregate per-formance will be discussed along with off-nominal testing and fault recovery.

Johnson, A.↗

Modernizing Accelerator Responsiveness and Controls in Operations

Accelerators increasingly use artificial intelligence (AI) and machine learning (ML) software and workflows for a variety of tasks, from optimization to fault detection and recovery. Efficient and sustainable application of these technologies necessitates specialized and facility-specific infrastructure commitments. Accelerator facilities also introduce unique radiation and security hazards, placing additional demands on operational infrastructure. These needs further escalate the prioritization of effective collaboration models and associated funding mechanisms and legal frameworks.

43 PARTICLE ACCELERATORS↗

Program Models Propagation Of Failures

FIRM is software tool for identification of failure and management of risk based on directed-graph ("digraph") approach. Three core algorithms optimized for processing singletons and doubletons and also handle tripletons. FIRM identifies loops in digraphs and displays direct failure paths between any two nodes. Solves for reachability for given node without computing reachability for entire digraph. Represents hybrid between schematic-diagram and fault-tree approaches. Written in C.

Hackler, Donald B.↗

Qualitative model-based diagnostics for rocket systems

A diagnostic software package is currently being developed at NASA LeRC that utilizes qualitative model-based reasoning techniques. These techniques can provide diagnostic information about the operational condition of the modeled rocket engine system or subsystem. The diagnostic package combines a qualitative model solver with a constraint suspension algorithm. The constraint suspension algorithm directs the solver's operation to provide valuable fault isolation information about the modeled system. A qualitative model of the Space Shuttle Main Engine's oxidizer supply components was generated. A diagnostic application based on this qualitative model was constructed to process four test cases: three numerical simulations and one actual test firing. The diagnostic tool's fault isolation output compared favorably with the input fault condition.

Maul, William↗

Qualitative model-based diagnostics for rocket systems

A diagnostic software package is currently being developed at NASA LeRC that utilizes qualitative model-based reasoning techniques. These techniques can provide diagnostic information about the operational condition of the modeled rocket engine system or subsystem. The diagnostic package combines a qualitative model solver with a constraint suspension algorithm. The constraint suspension algorithm directs the solver's operation to provide valuable fault isolation information about the modeled system. A qualitative model of the Space Shuttle Main Engine's oxidizer supply components was generated. A diagnostic application based on this qualitative model was constructed to process four test cases: three numerical simulations and one actual test firing. The diagnostic tool's fault isolation output compared favorably with the input fault condition.

Maul, William↗

Transparent Ada rendezvous in a fault tolerant distributed system

There are many problems associated with distributing an Ada program over a loosely coupled communication network. Some of these problems involve the various aspects of the distributed rendezvous. The problems addressed involve supporting the delay statement in a selective call and supporting the else clause in a selective call. Most of these difficulties are compounded by the need for an efficient communication system. The difficulties are compounded even more by considering the possibility of hardware faults occurring while the program is running. With a hardware fault tolerant computer system, it is possible to design a distribution scheme and communication software which is efficient and allows Ada semantics to be preserved. An Ada design for the communications software of one such system will be presented, including a description of the services provided in the seven layers of an International Standards Organization (ISO) Open System Interconnect (OSI) model communications system. The system capabilities (hardware and software) that allow this communication system will also be described.

Racine, Roger↗

Towards Real-time, On-board, Hardware-Supported Sensor and Software Health Management for Unmanned Aerial Systems

Unmanned aerial systems (UASs) can only be deployed if they can effectively complete their missions and respond to failures and uncertain environmental conditions while maintaining safety with respect to other aircraft as well as humans and property on the ground. In this paper, we design a real-time, on-board system health management (SHM) capability to continuously monitor sensors, software, and hardware components for detection and diagnosis of failures and violations of safety or performance rules during the flight of a UAS. Our approach to SHM is three-pronged, providing: (1) real-time monitoring of sensor and/or software signals; (2) signal analysis, preprocessing, and advanced on the- fly temporal and Bayesian probabilistic fault diagnosis; (3) an unobtrusive, lightweight, read-only, low-power realization using Field Programmable Gate Arrays (FPGAs) that avoids overburdening limited computing resources or costly re-certification of flight software due to instrumentation. Our implementation provides a novel approach of combining modular building blocks, integrating responsive runtime monitoring of temporal logic system safety requirements with model-based diagnosis and Bayesian network-based probabilistic analysis. We demonstrate this approach using actual data from the NASA Swift UAS, an experimental all-electric aircraft.

System & Software Health Management↗

Shielded-Twisted-Pair Cable Model for Chafe Fault Detection via Time-Domain Reflectometry

This report details the development, verification, and validation of an innovative physics-based model of electrical signal propagation through shielded-twisted-pair cable, which is commonly found on aircraft and offers an ideal proving ground for detection of small holes in a shield well before catastrophic damage occurs. The accuracy of this model is verified through numerical electromagnetic simulations using a commercially available software tool. The model is shown to be representative of more realistic (analytically intractable) cable configurations as well. A probabilistic framework is developed for validating the model accuracy with reflectometry data obtained from real aircraft-grade cables chafed in the laboratory.

Schuet, Stefan R.↗

The SSM/PMAD automated test bed project

The Space Station Module/Power Management and Distribution (SSM/PMAD) autonomous subsystem project was initiated in 1984. The project's goal has been to design and develop an autonomous, user-supportive PMAD test bed simulating the SSF Hab/Lab module(s). An eighteen kilowatt SSM/PMAD test bed model with a high degree of automated operation has been developed. This advanced automation test bed contains three expert/knowledge based systems that interact with one another and with other more conventional software residing in up to eight distributed 386-based microcomputers to perform the necessary tasks of real-time and near real-time load scheduling, dynamic load prioritizing, and fault detection, isolation, and recovery (FDIR).

Lollar, Louis F.↗

A Generic and Multifunctional Electromagnetic Transient Model for Grid-Following Inverters

This paper presents a generic and multifunctional electromagnetic transient (EMT) dynamic model of grid-following (GFL) inverter-based resources (IBRs) using the PSCAD software platform. The features of the model include flexibility in selecting various types and combinations of DC sources covering photovoltaic modules, battery modules, and ideal DC-source modules as well as flexibility in selecting either switching or averaged models of the inverter. This model also covers exhaustive lists of controller algorithm, including open-loop/closed-loop PQ dispatch control, DC voltage and AC terminal voltage control, and conventional current control designed in the dq-domain, the ..alpha....beta.. -domain, and the positive-/negative-sequence domain. Moreover, this model is equipped with flexibility in selecting various types of current-limiting schemes, including saturation-based and latching-based current limiters, and anti-windup protection. Also, the EMT model is agnostic to the MVA rating and is suitable for interfacing transmission systems by complying with IEEE Std. 2800. The generality in the power circuits and the multifunctional options in the operation and control of the developed EMT model make it suitable for both academia and industry to study various power system aspects, including, but not limited to, the fault behavior of GFL IBRs, the impacts on the protection system, and the transient stability of a system interfaced with large numbers of GFL IBRs.

integrated circuit modeling↗

Cassini Star Tracking and Identification Algorithms, Scene Simulation, and Testing

The Cassini mission will use autonomous star identification for initial attitude determination and a star tracking function for maintaining attitude. Because of the complexity of the StarID software, special software simulation tools were created to simulate the Stellar Reference Unit (SRU) output as a function of commands, spacecraft attitude, and star scene, and to allow the introduction of fault conditions. This paper gives the overview of the algorithm design and SRU simulation and a description of the simulation test results and a comparison with field test results obtained using the engineering model SRU.

Cassini↗

Fault Detection, Isolation and Recovery (FDIR) Portable Liquid Oxygen Hardware Demonstrator

The Fault Detection, Isolation and Recovery (FDIR) hardware demonstration will highlight the effort being conducted by Constellation's Ground Operations (GO) to provide the Launch Control System (LCS) with system-level health management during vehicle processing and countdown activities. A proof-of-concept demonstration of the FDIR prototype established the capability of the software to provide real-time fault detection and isolation using generated Liquid Hydrogen data. The FDIR portable testbed unit (presented here) aims to enhance FDIR by providing a dynamic simulation of Constellation subsystems that feed the FDIR software live data based on Liquid Oxygen system properties. The LO2 cryogenic ground system has key properties that are analogous to the properties of an electronic circuit. The LO2 system is modeled using electrical components and an equivalent circuit is designed on a printed circuit board to simulate the live data. The portable testbed is also be equipped with data acquisition and communication hardware to relay the measurements to the FDIR application running on a PC. This portable testbed is an ideal capability to perform FDIR software testing, troubleshooting, training among others.

Oostdyk, Rebecca L.↗

SWARM: Reimagining scientific workflow management systems in a distributed world

Modern scientific workflows process massive amounts of data from diverse instruments and sensors, leveraging geographically distributed, heterogeneous compute and storage resources—from leadership-class systems to edge devices—connected by high-performance networks. The diversity of resources introduces challenges in harnessing their full potential, with resilience issues arising across applications, system software, networks, storage, and hardware. Today, workflow management systems (WMS) coordinate the execution of computation and data management tasks across target resources. However, WMS’s centralized nature makes them vulnerable to faults and scalability issues that may result in failures of entire computational campaigns. In conclusion, this paper introduces a novel agentic framework for workflow management, fully distributing and decentralizing the WMS functions and modeling them as swarm intelligence agents infused with advanced artificial intelligence solutions and traditional distributed computing algorithms that can make coordinated decisions in the presence of failures of the underlying cyberinfrastructure.

Swarm intelligence↗

Hybrid Automated Diagnosis of Discrete/Continuous Systems

A recently conceived method of automated diagnosis of a complex electromechanical system affords a complete set of capabilities for hybrid diagnosis in the case in which the state of the electromechanical system is characterized by both continuous and discrete values (as represented by analog and digital signals, respectively). The method is an integration of two complementary diagnostic systems: (1) beacon-based exception analysis for multi-missions (BEAM), which is primarily useful in the continuous domain and easily performs diagnoses in the presence of transients; and (2) Livingstone, which is primarily useful in the discrete domain and is typically restricted to quasi-steady conditions. BEAM has been described in several prior NASA Tech Briefs articles: "Software for Autonomous Diagnosis of Complex Systems" (NPO-20803), Vol. 26, No. 3 (March 2002), page 33; "Beacon-Based Exception Analysis for Multimissions" (NPO-20827), Vol. 26, No. 9 (September 2002), page 32; "Wavelet-Based Real-Time Diagnosis of Complex Systems" (NPO-20830), Vol. 27, No. 1 (January 2003), page 67; and "Integrated Formulation of Beacon-Based Exception Analysis for Multimissions" (NPO-21126), Vol. 27, No. 3 (March 2003), page 74. Briefly, BEAM is a complete data-analysis method, implemented in software, for real-time or off-line detection and characterization of faults. The basic premise of BEAM is to characterize a system from all available observations and train the characterization with respect to normal phases of operation. The observations are primarily continuous in nature. BEAM isolates anomalies by analyzing the deviations from nominal for each phase of operation. Livingstone is a model-based reasoner that uses a model of a system, controller commands, and sensor observations to track the system s state, and detect and diagnose faults. Livingstone models a system within the discrete domain. Therefore, continuous sensor readings, as well as time, must be discretized. To reason about continuous systems, Livingstone uses monitors that discretize the sensor readings using trending and thresholding techniques. In development of the a hybrid method, BEAM results were sent to Livingstone to serve as an independent source of evidence that is in addition to the evidence gathered by Livingstone standard monitors. The figure depicts the flow of data in an early version of a hybrid system dedicated to diagnosing a simulated electromechanical system. In effect, BEAM served as a "smart" monitor for Livingstone. BEAM read the simulation data, processed the data to form observations, and stored the observations in a file. A monitor stub synchronized the events recorded by BEAM with the output of the Livingstone standard monitors according to time tags. This information was fed to a real-time interface, which buffered and fed the information to Livingstone, and requested diagnoses at the appropriate times. In a test, the hybrid system was found to correctly identify a failed component in an electromechanical system for which neither BEAM nor Livingstone alone yielded the correct diagnosis.

Park, Han↗