Search NASA⌕ Search

SEARCH · Search NASA

Results for “Software Failures”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 469 records · Page 26

Validation of the Mars 2020 Fault Protection Design: Navigating the Infinity of the Off-Nominal

On July 30th 2020, the Mars 2020 mission successfully launched out of Cape Canaveral, Florida, passed through the Earth’s shadow, and began its short cruise to Mars. Less than seven months later, the Perseverance rover touched down safely in Jezero Crater to begin its ambitious mission that includes looking for signs of ancient life and collecting samples for future return to Earth. Getting to the successful landing, or “Tango Delta Nominal,” could not have been achieved without also considering the off-nominal. One of the teams supporting this ambitious mission is the fault protection (FP) team. This team is tasked with assessing the various failures, or faults, that could prevent mission success and with ensuring that the autonomous behaviors built into the software and hardware can detect faults and recover the vehicle to a safe state. As part of its charter, the FP team designed a test campaign to provide confidence in the system’s robustness to off-nominal scenarios across all of Mars 2020’s mission phases. The greatest challenge associated with designing such a validation campaign was reducing the infinite number of anomalous scenarios into a finite test suite. In addition, the tests needed to be executed efficiently in order to utilize the team’s limited test venue access, but still needed to maintain a level of rigor that guaranteed confidence in the test outcomes. Given that each test scenario generated massive amounts of data, the team also developed methods for quickly ascertaining whether the autonomous fault protection behaviors maintained vehicle safety in the presence of an anomaly. This paper summarizes the processes that the Mars 2020 fault protection team employed to execute its off-nominal validation campaign. It captures both the methods of generating a suite of off-nominal tests, as well as reducing it to a subset that can be realistically executed within schedule and resource constraints. It also describes the various processes and philosophies that the team utilized to execute the tests efficiently, including creating a standardized procedure template, keeping the test cases modular so that they could be easily interchanged, and capturing common fault injections in a change-controlled database. Finally, it will describe the tools and processes for assessing the test data, focusing in particular on a tool that evaluated vehicle state using “secondary” sources of data to validate that the software had truly configured the spacecraft to the expected safe state.

Morantz, Chaz↗

SpecFIDLER User Manual (Software V.2.6.0)

The Spectroscopic Field Instrument for Detection of Low Energy Radiation (SpecFIDLER) allows response teams to detect and quantify plutonium contamination on the ground. Notional scenarios include dispersion from a weapon accident, or the launch failure of a space probe containing a radioisotope thermoelectric generator. Unlike other instruments, the thin-window sodium iodide detector is sensitive to the low-energy gamma rays emitted by plutonium isotopes. The system supports both mobile survey as well as stationary sampling. This manual provides information about installing, maintaining, and troubleshooting the SpecFIDLER. The scope of this document includes the physical hardware, software for data acquisition, and algorithms for data analysis. Recent changes to the software and algorithms aim to streamline the operation of the system.

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND ↗

Emerging technologies for V&V of ISHM software for space exploration

Systems1,2 required to exhibit high operational reliability often rely on some form of fault protection to recognize and respond to faults, preventing faults' escalation to catastrophic failures. Integrated System Health Management (ISHM) extends the functionality of fault protection to both scale to more complex systems (and systems of systems), and to maintain capability rather than just avert catastrophe. Forms of ISHM have been utilized to good effect in the maintenance phase of systems' total lifecycles (often referred to as 'condition-based mainte-nance'), but less so in a 'fault protection' role during actual operations. One of the impediments to such use lies in the challenges of verification, validation and certification of ISHM systems themselves. This paper makes the case that state-of-the-practice V&V and certification techniques will not suffice for emerging forms of ISHM systems; however, a number of maturing software engineering assurance technologies show particular promise for addressing these ISHM V&V challenges.

fault detection, isolation, and recovery↗

Development, Implementation and Application of Micromechanical Analysis Tools for Advanced High Temperature Composites

This document contains the final report to the NASA Glenn Research Center (GRC) for the research project entitled Development, Implementation, and Application of Micromechanical Analysis Tools for Advanced High-Temperature Composites. The research supporting this initiative has been conducted by Dr. Brett A. Bednarcyk, a Senior Scientist at OM in Brookpark, Ohio from the period of August 1998 to March 2005. Most of the work summarized herein involved development, implementation, and application of enhancements and new capabilities for NASA GRC's Micromechanics Analysis Code with Generalized Method of Cells (MAC/GMC) software package. When the project began, this software was at a low TRL (3-4) and at release version 2.0. Due to this project, the TRL of MAC/GMC has been raised to 7 and two new versions (3.0 and 4.0) have been released. The most important accomplishments with respect to MAC/GMC are: (1) A multi-scale framework has been built around the software, enabling coupled design and analysis from the global structure scale down to the micro fiber-matrix scale; (2) The software has been expanded to analyze smart materials; (3) State-of-the-art micromechanics theories have been implemented and validated within the code; (4) The damage, failure, and lifing capabilities of the code have been expanded from a very limited state to a vast degree of functionality and utility; and (5) The user flexibility of the code has been significantly enhanced. MAC/GMC is now the premier code for design and analysis of advanced composite and smart materials. It is a candidate for the 2005 NASA Software of the Year Award. The work completed over the course of the project is summarized below on a year by year basis. All publications resulting from the project are listed at the end of this report.

Source record↗

Fault Detection and Diagnosis in Spacecraft Electrical Power Systems

The ability to accurately identify and isolate failures in the electrical power system (EPS) is critical to ensure the reliability of spacecraft. This paper proposes a novel solution to the problem of fault detection and diagnosis in direct current (DC) electric power systems for spacecraft. Autonomous operation becomes essential during deep space missions that lack the ability to monitor and control the spacecraft from ground locations. The current state of EPS fault supervision is insufficient to guarantee highly reliable operation. To solve this issue, a combination of model-based and knowledge-based techniques are used in a hierarchical framework to improve the diagnostic performance of the system. Noise, disturbances, and modeling errors are considered in the design of the fault detection system. Practical considerations related to spacecraft flight hardware and software are accounted for in the system design for flight applications. To assess the functionality of the design, a wide array of failures are simulated in a series of experiments. The experiments showed that the technique improved the capability of the autonomous system by increasing the number of fault types diagnosed. The significance of this study is to provide a framework capable of advanced diagnostics of an EPS with little to no interaction from human operators.

Autonomous Power Systems↗

Spacecraft Onboard Software Maintenance: An Effective Approach which Reduces Costs and Increases Science Return

Flight software (FSW) is a mission critical element of spacecraft functionality and performance. When ground operations personnel interface to a spacecraft, they are dealing almost entirely with onboard software. This software, even more than ground/flight communications systems, is expected to perform perfectly at all times during all phases of on-orbit mission life. Due to the fact that FSW can be reconfigured and reprogrammed to accommodate new spacecraft conditions, the on-orbit FSW maintenance team is usually significantly responsible for the long-term success of a science mission. Failure of FSW can result in very expensive operations work-around costs and lost science opportunities. There are three basic approaches to staffing on-orbit software maintenance, namely: (1) using the original developers, (2) using mission operations personnel, or (3) assembling a Center of Excellence for multi-spacecraft on-orbit FSW support. This paper explains a National Aeronautics and Space Administration, Goddard Space Flight Center (NASA/GSFC) experience related to the roles of on-orbit FSW maintenance personnel. It identifies the advantages and disadvantages of each of the three approaches to staffing the FSW roles, and demonstrates how a cost efficient on-orbit FSW Maintenance Center of Excellence can be established and maintained with significant return on the investment.

Shell, Elaine M.↗

Software For Fault-Tree Diagnosis Of A System

Fault Tree Diagnosis System (FTDS) computer program is automated-diagnostic-system program identifying likely causes of specified failure on basis of information represented in system-reliability mathematical models known as fault trees. Is modified implementation of failure-cause-identification phase of Narayanan's and Viswanadham's methodology for acquisition of knowledge and reasoning in analyzing failures of systems. Knowledge base of if/then rules replaced with object-oriented fault-tree representation. Enhancement yields more-efficient identification of causes of failures and enables dynamic updating of knowledge base. Written in C language, C++, and Common LISP.

Iverson, Dave↗

Dam Failure Inundation Map Project

At the end of the first year, we remain on schedule. Property owners were identified and contacted for land access purposes. A prototype software package has been completed and was demonstrated to the Division of Land and Natural Resources (DLNR), National Weather Service (NWS) and Pacific Disaster Center (PDC). A field crew gathered data and surveyed the areas surrounding two dams in Waimea. (A field report is included in the annual report.) Data sensitivity analysis was initiated and completed. A user's manual has been completed. Beta testing of the software was initiated, but not completed. The initial TNK and property owner data collection for the additional test sites on Oahu and Kauai have been initiated.

Johnson, Carl↗

Design of a modular digital computer system DRL 4 and 5

Design and development efforts for a spaceborne modular computer system are reported. An initial baseline description is followed by an interface design that includes definition of the overall system response to all classes of failure. Final versions for the register level designs for all module types were completed. Packaging, support and control executive software, including memory utilization estimates and design verification plan, were formalized to insure a soundly integrated design of the digital computer system.

Source record↗

Prevention of design flaws in multicomputer systems

Report summarizes research on failure mode analysis for multicomputer systems where two or more computers may serve as redundant set. Failure modes such as data bus monopolization, shutdown due to transients, loss of control system equalization, memory alteration, and software errors are discussed.

Romberg, J. M.↗

Reliability measurement during software development

Measurement of software reliability was carried out during the development of data base software for a multi-sensor tracking system. Every run made during this project was scored as success or failure, and supporting data were collected on forms for further analysis. The failure ratio (number of failures per calendar interval divided by total number of runs) and failure rate (number of failures divided by CPU time for the interval) were found to be consistent measures, on a month-to-month basis as well as from module to module, and therefore considered valid indicators of reliability in this environment. Trend lines could be established from these measurements that provide good visualization of the progress on the job as a whole as well as on individual modules. Over one-half of the observed failures were due to factors associated with the specific run submission rather than with the code proper.

Hecht, H.↗

NASA V/STOL Propulsion Control Analysis - Phase I and II program status

NASA-Lewis Research Center initiated the V/STOL Propulsion Control Analysis Program in 1979 in order to take advantage of advanced electronic control technology for the next generation of V/STOL aircraft. The rationale, methods, and criteria developed during Phase I and Phase II of the program are discussed first. The development of an integrated flight and propulsion control system is then described. Most V/STOL aircraft under consideration depend on reaction thrusters or separate lift engines for attitude control and on engine thrust variations for height control. S/CTOL aircraft vector-thrust for pitch control during low speed operation. The baseline propulsion control system provides a thrust level response of 10 rad/sec and a vectoring response of 22 rad/sec. The requirements of the control components include: (1) redundant electronic control computers with minimum software and 100% Fault Coverage; (2) prime reliable electromechanical actuator interfaces; (3) miniature, digitally compatible sensors with easily detected failure modes; and (4) dual element, fail operational fuel pumps.

Miller, R. J.↗

Design for validation, based on formal methods

Validation of ultra-reliable systems decomposes into two subproblems: (1) quantification of probability of system failure due to physical failure; (2) establishing that Design Errors are not present. Methods of design, testing, and analysis of ultra-reliable software are discussed. It is concluded that a design-for-validation based on formal methods is needed for the digital flight control systems problem, and also that formal methods will play a major role in the development of future high reliability digital systems.

Butler, Ricky W.↗

Safety

Software requirements, design, implementation, verification and validation, and especially management are affected by the need to produce safe software. This paper discusses the changes in the software life cycle that are necessary to ensure that software will execute without resulting in unacceptable risk. Software is being used increasingly to monitor and control safety-critical processes in which a run-time failure or error could result in unacceptable losses such as death, injury, loss of property, or environmental harm. Examples of such processes maybe found in transportation, energy, aerospace, basic industry, medicine, and defense systems.

Leveson, Nancy G.↗

Branch recovery with compiler-assisted multiple instruction retry

In processing systems where rapid recovery from transient faults is important, schemes for multiple instruction rollback recovery may be appropriate. Multiple instruction retry has been implemented in hardware by researchers and also in mainframe computers. This paper extends compiler-assisted instruction retry to a broad class of code execution failures. Five benchmarks were used to measure the performance penalty of hazard resolution. Results indicate that the enhanced pure software approach can produce performance penalties consistent with existing hardware techniques. A combined compiler/hardware resolution strategy is also described and evaluated. Experimental results indicate a lower performance penalty than with either a totally hardware or totally software approach.

Alewine, N. J.↗

Risk-Based Object Oriented Testing

Software testing is a well-defined phase of the software development life cycle. Functional ("black box") testing and structural ("white box") testing are two methods of test case design commonly used by software developers. A lesser known testing method is risk-based testing, which takes into account the probability of failure of a portion of code as determined by its complexity. For object oriented programs, a methodology is proposed for identification of risk-prone classes. Risk-based testing is a highly effective testing technique that can be used to find and fix the most important problems as quickly as possible.

Rosenberg, Linda H.↗

Modeling Code Is Helping Cleveland Develop New Products

Master Builders, Inc., is a 350-person company in Cleveland, Ohio, that develops and markets specialty chemicals for the construction industry. Developing new products involves creating many potential samples and running numerous tests to characterize the samples' performance. Company engineers enlisted NASA's help to replace cumbersome physical testing with computer modeling of the samples' behavior. Since the NASA Lewis Research Center's Structures Division develops mathematical models and associated computation tools to analyze the deformation and failure of composite materials, its researchers began a two-phase effort to modify Lewis' Integrated Composite Analyzer (ICAN) software for Master Builders' use. Phase I has been completed, and Master Builders is pleased with the results. The company is now working to begin implementation of Phase II.

Source record↗

Explosion/Blast Dynamics for Constellation Launch Vehicles Assessment

An assessment methodology is developed to guide quantitative predictions of adverse physical environments and the subsequent effects on the Ares-1 crew launch vehicle associated with the loss of containment of cryogenic liquid propellants from the upper stage during ascent. Development of the methodology is led by a team at Sandia National Laboratories (SNL) with guidance and support from a number of National Aeronautics and Space Administration (NASA) personnel. The methodology is based on the current Ares-1 design and feasible accident scenarios. These scenarios address containment failure from debris impact or structural response to pressure or blast loading from an external source. Once containment is breached, the envisioned assessment methodology includes predictions for the sequence of physical processes stemming from cryogenic tank failure. The investigative techniques, analysis paths, and numerical simulations that comprise the proposed methodology are summarized and appropriate simulation software is identified in this report.

Baer, Mel↗