Search NASA⌕ Search

SEARCH · Search NASA

Results for “software failure”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 721 records · Page 40

Automatically Finding the Control Variables for Complex System Behavior

Testing large-scale systems is expensive in terms of both time and money. Running simulations early in the process is a proven method of finding the design faults likely to lead to critical system failures, but determining the exact cause of those errors is still time-consuming and requires access to a limited number of domain experts. It is desirable to find an automated method that explores the large number of combinations and is able to isolate likely fault points. Treatment learning is a subset of minimal contrast-set learning that, rather than classifying data into distinct categories, focuses on finding the unique factors that lead to a particular classification. That is, they find the smallest change to the data that causes the largest change in the class distribution. These treatments, when imposed, are able to identify the factors most likely to cause a mission-critical failure. The goal of this research is to comparatively assess treatment learning against state-of-the-art numerical optimization techniques. To achieve this, this paper benchmarks the TAR3 and TAR4.1 treatment learners against optimization techniques across three complex systems, including two projects from the Robust Software Engineering (RSE) group within the National Aeronautics and Space Administration (NASA) Ames Research Center. The results clearly show that treatment learning is both faster and more accurate than traditional optimization methods.

Gay, Gregory↗

Integrated restructurable flight control system demonstration results

The purpose of this study was to examine the complementary capabilities of several restructurable flight control system (RFCS) concepts through the integration of these technologies into a complete system. Performance issues were addressed through a re-examination of RFCS functional requirements, and through a qualitative analysis of the design issues that, if properly addressed during integration, will lead to the highest possible degree of fault-tolerant performance. Software developed under previous phases of this contract and under NAS1-18004 was modified and integrated into a complete RFCS subroutine for NASA's B-737 simulation. The integration of these modules involved the development of methods for dealing with the mismatch between the outputs of the failure detection module and the input requirements of the automatic control system redesign module. The performance of this demonstration system was examined through extensive simulation trials.

Weiss, Jerold L.↗

Analysis of SSEM Sensor Data Using BEAM

A report describes analysis of space shuttle main engine (SSME) sensor data using Beacon-based Exception Analysis for Multimissions (BEAM) [NASA Tech Briefs articles, the two most relevant being Beacon-Based Exception Analysis for Multimissions (NPO- 20827), Vol. 26, No.9 (September 2002), page 32 and Integrated Formulation of Beacon-Based Exception Analysis for Multimissions (NPO- 21126), Vol. 27, No. 3 (March 2003), page 74] for automated detection of anomalies. A specific implementation of BEAM, using the Dynamical Invariant Anomaly Detector (DIAD), is used to find anomalies commonly encountered during SSME ground test firings. The DIAD detects anomalies by computing coefficients of an autoregressive model and comparing them to expected values extracted from previous training data. The DIAD was trained using nominal SSME test-firing data. DIAD detected all the major anomalies including blade failures, frozen sense lines, and deactivated sensors. The DIAD was particularly sensitive to anomalies caused by faulty sensors and unexpected transients. The system offers a way to reduce SSME analysis time and cost by automatically indicating specific time periods, signals, and features contributing to each anomaly. The software described here executes on a standard workstation and delivers analyses in seconds, a computing time comparable to or faster than the test duration itself, offering potential for real-time analysis.

Zak, Michail↗

Trajectory Specification for Terminal Air Traffic: Pairwise Conflict Detection and Resolution

Trajectory Specification is the explicit bounding and control of aircraft trajectories such that the position at any point in time is constrained to a precisely defined volume of space. The bounding space is defined by cross-track, along-track, and vertical tolerances relative to a reference trajectory that specifies position as a function of time. The tolerances are dynamic and will be based on the aircraft navigation capabilities and the current traffic situation. Assuming conformance, Trajectory Specification can guarantee safe separation for an arbitrary period of time even in the event of an air traffic control (ATC) system or datalink failure; hence it can help to achieve the high level of safety and reliability needed for ATC automation. It can also reduce the reliance on tactical backup systems during normal operation. This paper applies it to the terminal area around a major airport and presents algorithms and software for detecting and resolving conflicts. A representative set of pairwise conflicts was generated, and a fast-time simulation was run on them. All conflicts were successfully resolved in real time, demonstrating the computational feasibility of the concept.

terminal area↗

A Vehicle Management End-to-End Testing and Analysis Platform for Validation of Mission and Fault Management Algorithms to Reduce Risk for NASA's Space Launch System

The engineering development of the new Space Launch System (SLS) launch vehicle requires cross discipline teams with extensive knowledge of launch vehicle subsystems, information theory, and autonomous algorithms dealing with all operations from pre-launch through on orbit operations. The characteristics of these spacecraft systems must be matched with the autonomous algorithm monitoring and mitigation capabilities for accurate control and response to abnormal conditions throughout all vehicle mission flight phases, including precipitating safing actions and crew aborts. This presents a large and complex system engineering challenge, which is being addressed in part by focusing on the specific subsystems involved in the handling of off-nominal mission and fault tolerance with response management. Using traditional model based system and software engineering design principles from the Unified Modeling Language (UML) and Systems Modeling Language (SysML), the Mission and Fault Management (M&FM) algorithms for the vehicle are crafted and vetted in specialized Integrated Development Teams (IDTs) composed of multiple development disciplines such as Systems Engineering (SE), Flight Software (FSW), Safety and Mission Assurance (S&MA) and the major subsystems and vehicle elements such as Main Propulsion Systems (MPS), boosters, avionics, Guidance, Navigation, and Control (GNC), Thrust Vector Control (TVC), and liquid engines. These model based algorithms and their development lifecycle from inception through Flight Software certification are an important focus of this development effort to further insure reliable detection and response to off-nominal vehicle states during all phases of vehicle operation from pre-launch through end of flight. NASA formed a dedicated M&FM team for addressing fault management early in the development lifecycle for the SLS initiative. As part of the development of the M&FM capabilities, this team has developed a dedicated testbed that integrates specific M&FM algorithms, specialized nominal and off-nominal test cases, and vendor-supplied physics-based launch vehicle subsystem models. Additionally, the team has developed processes for implementing and validating these algorithms for concept validation and risk reduction for the SLS program. The flexibility of the Vehicle Management End-to-end Testbed (VMET) enables thorough testing of the M&FM algorithms by providing configurable suites of both nominal and off-nominal test cases to validate the developed algorithms utilizing actual subsystem models such as MPS. The intent of VMET is to validate the M&FM algorithms and substantiate them with performance baselines for each of the target vehicle subsystems in an independent platform exterior to the flight software development infrastructure and its related testing entities. In any software development process there is inherent risk in the interpretation and implementation of concepts into software through requirements and test cases into flight software compounded with potential human errors throughout the development lifecycle. Risk reduction is addressed by the M&FM analysis group working with other organizations such as S&MA, Structures and Environments, GNC, Orion, the Crew Office, Flight Operations, and Ground Operations by assessing performance of the M&FM algorithms in terms of their ability to reduce Loss of Mission and Loss of Crew probabilities. In addition, through state machine and diagnostic modeling, analysis efforts investigate a broader suite of failure effects and associated detection and responses that can be tested in VMET to ensure that failures can be detected, and confirm that responses do not create additional risks or cause undesired states through interactive dynamic effects with other algorithms and systems. VMET further contributes to risk reduction by prototyping and exercising the M&FM algorithms early in their implementation and without any inherent hindrances such as meeting FSW processor scheduling constraints due to their target platform - ARINC 653 partitioned OS, resource limitations, and other factors related to integration with other subsystems not directly involved with M&FM such as telemetry packing and processing. The baseline plan for use of VMET encompasses testing the original M&FM algorithms coded in the same C++ language and state machine architectural concepts as that used by Flight Software. This enables the development of performance standards and test cases to characterize the M&FM algorithms and sets a benchmark from which to measure the effectiveness of M&FM algorithms performance in the FSW development and test processes.

Trevino, Luis↗

Hydrogen Plus Other Alternative Fuels Risk Assessment Models (HyRAM+) Technical Reference Manual (V.6.0)

The HyRAM+ software is an open-source toolkit that provides publicly available models and default input values to enable straightforward and consistent safety assessments for hydrogen and other alternative fuel systems, such as natural gas and propane. The HyRAM+ quantitative risk assessment calculation incorporates annual likelihood of leaks or failures for both compressed gaseous and liquefied flammable fuels, as well as probabilistic models for the effects of heat flux and overpressure. HyRAM

08 HYDROGEN↗

Advanced detection, isolation, and accommodation of sensor failures in turbofan engines: Real-time microcomputer implementation

The objective of the Advanced Detection, Isolation, and Accommodation Program is to improve the overall demonstrated reliability of digital electronic control systems for turbine engines. For this purpose, an algorithm was developed which detects, isolates, and accommodates sensor failures by using analytical redundancy. The performance of this algorithm was evaluated on a real time engine simulation and was demonstrated on a full scale F100 turbofan engine. The real time implementation of the algorithm is described. The implementation used state-of-the-art microprocessor hardware and software, including parallel processing and high order language programming.

Delaat, John C.↗

The Unparalleled Systems Engineering of MSL's Backup Entry, Descent, and Landing System: Second Chance

Second Chance (SECC) was a bare bones version of Mars Science Laboratory's (MSL) Entry Descent & Landing (EDL) flight software that ran on Curiosity's backup computer, which could have taken over swiftly in the event of a reset of Curiosity's prime computer, in order to land her safely on Mars. Without SECC, a reset of Curiosity's prime computer would have lead to catastrophic mission failure. Even though a reset of the prime computer never occurred, SECC had the important responsibility as EDL's guardian angel, and this responsibility would not have seen such success without unparalleled systems engineering. This paper will focus on the systems engineering behind SECC: Covering a brief overview of SECC's design, the intense schedule to use SECC as a backup system, the verification and validation of the system's "Do No Harm" mandate, the system's overall functional performance, and finally, its use on the fateful day of August 5th, 2012.

fault protection↗

The unparalleled systems engineering of MSL’s backup entry, descent, and landing system : second chance

Second Chance (SECC) was a bare bones version of Mars Science Laboratory’s (MSL) Entry Descent & Landing (EDL) flight software that ran on Curiosity’s backup computer, which could have taken over swiftly in the event of a reset of Curiosity’s prime computer, in order to land her safely on Mars. Without SECC, a reset of Curiosity’s prime computer would have lead to catastrophic mission failure. Even though a reset of the prime computer never occurred, SECC had the important responsibility as EDL’s guardian angel, and this responsibility would not have seen such success without unparalleled systems engineering. This paper will focus on the systems engineering behind SECC: Covering a brief overview of SECC’s design, the intense schedule to use SECC as a backup system, the verification and validation of the system’s “Do No Harm” mandate, the system’s overall functional performance, and finally, its use on the fateful day of August 5th, 2012.

Reeves, Glenn↗

Software Assists in Responding to Anomalous Conditions

Fault Induced Document Retrieval Officer (FIDO) is a computer program that reduces the need for a large and costly team of engineers and/or technicians to monitor the state of a spacecraft and associated ground systems and respond to anomalies. FIDO includes artificial-intelligence components that imitate the reasoning of human experts with reference to a knowledge base of rules that represent failure modes and to a database of engineering documentation. These components act together to give an unskilled operator instantaneous expert assistance and access to information that can enable resolution of most anomalies, without the need for highly paid experts. FIDO provides a system state summary (a configurable engineering summary) and documentation for diagnosis of a potentially failing component that might have caused a given error message or anomaly. FIDO also enables high-level browsing of documentation by use of an interface indexed to the particular error message. The collection of available documents includes information on operations and associated procedures, engineering problem reports, documentation of components, and engineering drawings. FIDO also affords a capability for combining information on the state of ground systems with detailed, hierarchically-organized, hypertext- enabled documentation.

James, Mark↗

Care 3, Phase 1, volume 1

A computer program to aid in accessing the reliability of fault tolerant avionics systems was developed. A simple mathematical expression was used to evaluate the reliability of any redundant configuration over any interval during which the failure rates and coverage parameters remained unaffected by configuration changes. Provision was made for convolving such expressions in order to evaluate the reliability of a dual mode system. A coverage model was also developed to determine the various relevant coverage coefficients as a function of the available hardware and software fault detector characteristics, and subsequent isolation and recovery delay statistics.

Stiffler, J. J.↗

Clementine Sensor Processing System

The design of the DSPSE Satellite Controller (DSC) is baselined as a single-string satellite controller. The DSC performs two main functions: health and maintenance of the spacecraft; and image capture, storage, and playback. The DSC contains two processors: a radiation-hardened Mil-Std-1750, and a commercial R3000. The Mil-Std-1750 processor performs all housekeeping operations, while the R3000 is mainly used to perform the image processing functions associated with the navigation functions, as well as performing various experiments. The DSC also contains a data handling unit (DHU) used to interface to various spacecraft imaging sensors and to capture, compress, and store selected images onto the solid-state data recorder. The development of the DSC evolved from several key requirements; the DSPSE satellite was to do the following: (1) have a radiation-hardened spacecraft control system and be immune to single-event upsets (SEU's); (2) use an R3000-based processor to run the star tracker software that was developed by SDIO (due to schedule and cost constraints, there was no time to port the software to a radiation-hardened processor); and (3) fly a commercial processor to verify its suitability for use in a space environment. In order to enhance the DSC reliability, the system was designed with multiple processing paths. These multiple processing paths provide for greater tolerance to various component failures. The DSC was designed so that all housekeeping processing functions are performed by either the Mil-Std-1750 processor or the R3000 processor. The image capture and storage is performed either by the DHU or the R3000 processor.

Feldstein, A. A.↗

System Risk Balancing Profiles: Software Component

The Software QA / V&V guide will be reviewed and updated based on feedback from NASA organizations and others with a vested interest in this area. Hardware, EEE Parts, Reliability, and Systems Safety are a sample of the future guides that will be developed. Cost Estimates, Lessons Learned, Probability of Failure and PACTS (Prevention, Avoidance, Control or Test) are needed to provide a more complete risk management strategy. This approach to risk management is designed to help balance the resources and program content for risk reduction for NASA's changing environment.

Kelly, John C.↗

High-Performance Acousto-Ultrasonic Scan System Being Developed

Acousto-ultrasonic (AU) interrogation is a single-sided nondestructive evaluation (NDE) technique employing separated sending and receiving transducers. It is used for assessing the microstructural condition and distributed damage state of the material between the transducers. AU is complementary to more traditional NDE methods, such as ultrasonic cscan, x-ray radiography, and thermographic inspection, which tend to be used primarily for discrete flaw detection. Throughout its history, AU has been used to inspect polymer matrix composites, metal matrix composites, ceramic matrix composites, and even monolithic metallic materials. The development of a high-performance automated AU scan system for characterizing within-sample microstructural and property homogeneity is currently in a prototype stage at NASA. This year, essential AU technology was reviewed. In addition, the basic hardware and software configuration for the scanner was developed, and preliminary results with the system were described. Mechanical and environmental loads applied to composite materials can cause distributed damage (as well as discrete defects) that plays a significant role in the degradation of physical properties. Such damage includes fiber/matrix debonding (interface failure), matrix microcracking, and fiber fracture and buckling. Investigations at the NASA Glenn Research Center have shown that traditional NDE scan inspection methods such as ultrasonic c-scan, x-ray imaging, and thermographic imaging tend to be more suited to discrete defect detection rather than the characterization of accumulated distributed micro-damage in composites. Since AU is focused on assessing the distributed micro-damage state of the material in between the sending and receiving transducers, it has proven to be quite suitable for assessing the relative composite material state. One major success story at Glenn with AU measurements has been the correlation between the ultrasonic decay rate obtained during AU inspection and the mechanical modulus (stiffness) seen during fatigue experiments with silicon carbide/silicon carbide (SiC/SiC) ceramic matrix composite samples. As shown in the figure, ultrasonic decay increased as the modulus decreased for the ceramic matrix composite tensile fatigue samples. The likely microstructural reason for the decrease in modulus (and increase in ultrasonic decay) is the matrix microcracking that commonly occurs during fatigue testing of these materials. Ultrasonic decay has shown the capability to track the pattern of transverse cracking and fiber breakage in these composites.

Roth, Don J.↗

High-Performance Acousto-Ultrasonic Scan System Being Developed

Acousto-ultrasonic (AU) interrogation is a single-sided nondestructive evaluation (NDE) technique employing separated sending and receiving transducers. It is used for assessing the microstructural condition and distributed damage state of the material between the transducers. AU is complementary to more traditional NDE methods, such as ultrasonic cscan, x-ray radiography, and thermographic inspection, which tend to be used primarily for discrete flaw detection. Throughout its history, AU has been used to inspect polymer matrix composites, metal matrix composites, ceramic matrix composites, and even monolithic metallic materials. The development of a high-performance automated AU scan system for characterizing within-sample microstructural and property homogeneity is currently in a prototype stage at NASA. This year, essential AU technology was reviewed. In addition, the basic hardware and software configuration for the scanner was developed, and preliminary results with the system were described. Mechanical and environmental loads applied to composite materials can cause distributed damage (as well as discrete defects) that plays a significant role in the degradation of physical properties. Such damage includes fiber/matrix debonding (interface failure), matrix microcracking, and fiber fracture and buckling. Investigations at the NASA Glenn Research Center have shown that traditional NDE scan inspection methods such as ultrasonic c-scan, x-ray imaging, and thermographic imaging tend to be more suited to discrete defect detection rather than the characterization of accumulated distributed microdamage in composites. Since AU is focused on assessing the distributed microdamage state of the material in between the sending and receiving transducers, it has proven to be quite suitable for assessing the relative composite material state. One major success story at Glenn with AU measurements has been the correlation between the ultrasonic decay rate obtained during AU inspection and the mechanical modulus (stiffness) seen during fatigue experiments with silicon carbide/silicon carbide (SiC/SiC) ceramic matrix composite samples. As shown in the figure, ultrasonic decay increased as the modulus decreased for the ceramic matrix composite tensile fatigue samples. The likely microstructural reason for the decrease in modulus (and increase in ultrasonic decay) is the matrix microcracking that commonly occurs during fatigue testing of these materials. Ultrasonic decay has shown the capability to track the pattern of transverse cracking and fiber breakage in these composites.

Roth, Don J.↗

More About Software for No-Loss Computing

A document presents some additional information on the subject matter of "Integrated Hardware and Software for No- Loss Computing" (NPO-42554), which appears elsewhere in this issue of NASA Tech Briefs. To recapitulate: The hardware and software designs of a developmental parallel computing system are integrated to effectuate a concept of no-loss computing (NLC). The system is designed to reconfigure an application program such that it can be monitored in real time and further reconfigured to continue a computation in the event of failure of one of the computers. The design provides for (1) a distributed class of NLC computation agents, denoted introspection agents, that effects hierarchical detection of anomalies; (2) enhancement of the compiler of the parallel computing system to cause generation of state vectors that can be used to continue a computation in the event of a failure; and (3) activation of a recovery component when an anomaly is detected.

Edmonds, Iarina↗

NASA Spacecraft Trade Modeling System (NSTRDMS)

A rapid mission analysis tool is developed to support the ongoing design of the Lunar Transit trajectory of the Power and Propulsion Element (PPE). A 50-kW class electric propulsion system is envisioned to transit a massive vehicle be-tween a Medium Earth Orbit (MEO) parking orbit and a lunar L2 southern Near Rectilinear Halo Orbit (NRHO). A parameterization is developed by which the Lunar Transit can be analyzed in the context of varying vehicle mass, solar elec-tric propulsion (SEP) configurations, and solar array power output. A rapid and novel mission analysis tool enables a wide array of these trade analyses to be completed without the need for extensive computing resources or time. This tool is shown to be useful in the analysis of a reference trajectory, where changes to the baseline vehicle architecture or off-nominal operational scenarios (such as electric thruster failures) can be rapidly assessed by the mission designer.

Low thrust↗

The predictive information obtained by testing multiple software versions

Multiversion programming is a redundancy approach to developing highly reliable software. In applications of this method, two or more versions of a program are developed independently by different programmers and the versions are combined to form a redundant system. One variation of this approach consists of developing a set of n program versions and testing the versions to predict the failure probability of a particular program or a system formed from a subset of the programs. The precision that might be obtained, and also the effect of programmer variability if predictions are made over repetitions of the process of generating different program versions, are examined.

Lee, Larry D.↗