Search NASASearch

SEARCH · Search NASA

Results for “software failure”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 181 records · Page 10

The implementation and use of Ada on distributed systems with high reliability requirements

The use and implementation of Ada in distributed environments in which reliability is the primary concern is investigated. Emphasis is placed on the possibility that a distributed system may be programmed entirely in ADA so that the individual tasks of the system are unconcerned with which processors they are executing on, and that failures may occur in the software or underlying hardware. The primary activities are: (1) Continued development and testing of our fault-tolerant Ada testbed; (2) consideration of desirable language changes to allow Ada to provide useful semantics for failure; (3) analysis of the inadequacies of existing software fault tolerance strategies.

Knight, J. C.

Risk-Significant Adverse Condition Awareness Strengthens Assurance of Fault Management Systems

As spaceflight systems increase in complexity, Fault Management (FM) systems are ranked high in risk-based assessment of software criticality, emphasizing the importance of establishing highly competent domain expertise to provide assurance. Adverse conditions (ACs) and specific vulnerabilities encountered by safety- and mission-critical software systems have been identified through efforts to reduce the risk posture of software-intensive NASA missions. Acknowledgement of potential off-nominal conditions and analysis to determine software system resiliency are important aspects of hazard analysis and FM. A key component of assuring FM is an assessment of how well software addresses susceptibility to failure through consideration of ACs. Focus on significant risk predicted through experienced analysis conducted at the NASA Independent Verification Validation (IVV) Program enables the scoping of effective assurance strategies with regard to overall asset protection of complex spaceflight as well as ground systems. Research efforts sponsored by NASA's Office of Safety and Mission Assurance defined terminology, categorized data fields, and designed a baseline repository that centralizes and compiles a comprehensive listing of ACs and correlated data relevant across many NASA missions. This prototype tool helps projects improve analysis by tracking ACs and allowing queries based on project, mission type, domaincomponent, causal fault, and other key characteristics. Vulnerability in off-nominal situations, architectural design weaknesses, and unexpected or undesirable system behaviors in reaction to faults are curtailed with the awareness of ACs and risk-significant scenarios modeled for analysts through this database. Integration within the Enterprise Architecture at NASA IVV enables interfacing with other tools and datasets, technical support, and accessibility across the Agency. This paper discusses the development of an improved workflow process utilizing this database for adaptive, risk-informed FM assurance that critical software systems will safely and securely protect against faults and respond to ACs in order to achieve successful missions.

Fault management

Risk-Significant Adverse Condition Awareness Strengthens Assurance of Fault Management Systems

As spaceflight systems increase in complexity, Fault Management (FM) systems are ranked high in risk-based assessment of software criticality, emphasizing the importance of establishing highly competent domain expertise to provide assurance. Adverse conditions (ACs) and specific vulnerabilities encountered by safety- and mission-critical software systems have been identified through efforts to reduce the risk posture of software-intensive NASA missions. Acknowledgement of potential off-nominal conditions and analysis to determine software system resiliency are important aspects of hazard analysis and FM. A key component of assuring FM is an assessment of how well software addresses susceptibility to failure through consideration of ACs. Focus on significant risk predicted through experienced analysis conducted at the NASA Independent Verification & Validation (IV&V) Program enables the scoping of effective assurance strategies with regard to overall asset protection of complex spaceflight as well as ground systems. Research efforts sponsored by NASAs Office of Safety and Mission Assurance (OSMA) defined terminology, categorized data fields, and designed a baseline repository that centralizes and compiles a comprehensive listing of ACs and correlated data relevant across many NASA missions. This prototype tool helps projects improve analysis by tracking ACs and allowing queries based on project, mission type, domain/component, causal fault, and other key characteristics. Vulnerability in off-nominal situations, architectural design weaknesses, and unexpected or undesirable system behaviors in reaction to faults are curtailed with the awareness of ACs and risk-significant scenarios modeled for analysts through this database. Integration within the Enterprise Architecture at NASA IV&V enables interfacing with other tools and datasets, technical support, and accessibility across the Agency. This paper discusses the development of an improved workflow process utilizing this database for adaptive, risk-informed FM assurance that critical software systems will safely and securely protect against faults and respond to ACs in order to achieve successful missions.

IV&V

NASA Tech Briefs, July 2013

Dielectrophoresis-Based Particle Sensor Using Nanoelectrode Arrays; Multi-Dimensional Damage Detection for Surfaces and Structures; ULTRA: Underwater Localization for Transit and Reconnaissance Autonomy; Autonomous Cryogenic Leak Detector for Improving Launch Site Operations; Submillimeter Planetary Atmospheric Chemistry Exploration Sounder; Method for Reduction of Silver Biocide Plating on Metal Surfaces; Silicon Micromachined Microlens Array for THz Antennas; Forward-Looking IED Detector Ground Penetrating Radar; Fully Printed, Flexible, Phased Array Antenna for Lunar Surface Communication, Battery Charge Equalizer with Transformer Array; An Efficient, Highly Flexible Multi-Channel Digital Downconverter Architecture; Dimmable Electronic Ballast for a Gas Discharge Lamp; Conductive Carbon Nanotube Inks for Use with Desktop Inkjet Printing Technology; Enhanced Schapery Theory Software Development for Modeling Failure of Fiber-Reinforced Laminates; High-Performance, Low-Temperature-Operating, Long-Lifetime Aerospace Lubricants; Carbon Nanotube Microarrays Grown on Nanoflake Substrates; Differential Muon Tomography to Continuously Monitor Changes in the Composition of Subsurface Fluids; Microgravity Drill and Anchor System; 20 Granular Media-Based Tunable Passive Vibration Suppressor; 21 Miga Aero Actuator and 2D Machined Mechanical Binary Latch; Micro-XRF for In Situ Geological Exploration of Other Planets; Hydrogen-Enhanced Lunar Oxygen Extraction and Storage Using Only Solar Power; Uplift of Ionospheric Oxygen Ions During Extreme Magnetic Storms; Miniaturized, High-Speed, Modulated X-Ray Source; Hollow-Fiber Spacesuit Water Membrane Evaporator 25 High-Power Single-Mode 2.65-micrometers InGaAsSb/AlInGaAsSb Diode Lasers; Optical Device for Converting a Laser Beam Into Two Co-aligned but Oppositely Directed Beams; A Hybrid Fiber/Solid-State Regenerative Amplifier with Tunable Pulse Widths for Satellite Laser Ranging; X-Ray Diffractive Optics; SynGenics Optimization System (SynOptSys); 29 CFD Script for Rapid TPS Damage Assessment; radEq Add-On Module for CFD Solver Loci-CHEM; Science Opportunity Analyzer (SOA) Version 8; 30 Autonomous Byte Stream Randomizer; Distributed Engine Control Empirical/Analytical Verification Tools; Dynamic Server-Based KML Code Generator Method for Level-of-Detail Traversal of Geospatial Data; Automated Planning of Science Products Based on Nadir Overflights and Alerts for Onboard and Ground Processing; Linked Autonomous Interplanetary Satellite Orbit Navigation; Risk-Constrained Dynamic Programming for Optimal Mars Entry, Descent, and Landing; Scheduling Operations for Massive Heterogeneous Clusters; Deepak Condenser Model (DeCoM); Flight Software Math Library; Recirculating 1-K-Pot for Pulse-Tube Cryostats; 35 Method for Processing Lunar Regolith Using Microwaves; Wells for In Situ Extraction of Volatiles from Regolith (WIEVR); and Estimating the Backup Reaction Wheel Orientation Using Reaction Wheel Spin Rates Flight Telemetry from a Spacecraft.

Source record

Machine learning for photovoltaic single axis tracker fault detection and classification

More than 81% of the annual capacity of utility-scale photovoltaic (PV) power plants in the U.S. use single-axis trackers (SATs) due to SATs delivering 4% in capacity factor on average over fixed-array systems. However, SATs are subject to faults, such as software misconfigurations and mechanical failures, resulting in suboptimal tracking. If left undetected, the overall power yield of the PV power plant is reduced significantly. Minimizing downtime and ensuring efficient operation of SATs requires robust detection and diagnosis mechanisms for SAT faults. We present a machine learning framework for implementing real-time SAT fault detection and classification. Our implementation of the proposed framework reliably identifies measurements taken from a test PV system undergoing emulated SAT faults relative to state-of-the-art algorithms and produces nearly zero false positives on our testing days. Code and data are available at https://pvpmc.sandia.gov/tools.

Fault classification

A guide to onboard checkout. Volume 1: Guidance, navigation and control

The results are presented of a study of onboard checkout techniques, as they relate to space station subsystems, as a guide to those who may need to implement onboard checkout in similar subsystems. Guidance, navigation, and control subsystems, and their reliability and failure analyses are presented. Software and testing procedures are also given.

Source record

Multiple IMU system development, volume 1

A redundant gimballed inertial system is described. System requirements and mechanization methods are defined and hardware and software development is described. Failure detection and isolation algorithms are presented and technology achievements described. Application of the system as a test tool for shuttle avionics concepts is outlined.

Landey, M.

Expert systems for adaptive control of large space structures

It is expected that space systems for the future will evolve to structures of unprecedented size with associated extreme control requirements. A method is necessary that is sufficiently general to initiate stable control of a vehicle and subsequently learn the true nature of the structure. It is suggested that a suitable constructed expert system (ES) would be capable of learning by appending observations to a knowledge base. To verify that an ES can control a large space structure, numerical simulations of a simple structure subjected to periodic vibrations and the performance of a classical controllers were performed. The ES was then exercised to show its ability to truthfully mimic nominal control and to demonstrate its superiority to the classical controller, given sensor failures. An ES generating software package, The Intelligent Machine Model, was employed. It uses the pattern matching technique. Results of this program are discussed.

Gartrell, Charles F.

A fault-tolerant intelligent robotic control system

This paper describes the concept, design, and features of a fault-tolerant intelligent robotic control system being developed for space and commercial applications that require high dependability. The comprehensive strategy integrates system level hardware/software fault tolerance with task level handling of uncertainties and unexpected events for robotic control. The underlying architecture for system level fault tolerance is the distributed recovery block which protects against application software, system software, hardware, and network failures. Task level fault tolerance provisions are implemented in a knowledge-based system which utilizes advanced automation techniques such as rule-based and model-based reasoning to monitor, diagnose, and recover from unexpected events. The two level design provides tolerance of two or more faults occurring serially at any level of command, control, sensing, or actuation. The potential benefits of such a fault tolerant robotic control system include: (1) a minimized potential for damage to humans, the work site, and the robot itself; (2) continuous operation with a minimum of uncommanded motion in the presence of failures; and (3) more reliable autonomous operation providing increased efficiency in the execution of robotic tasks and decreased demand on human operators for controlling and monitoring the robotic servicing routines.

Marzwell, Neville I.

Some Improvements in Utilization of Flash Memory Devices

Two developments improve the utilization of flash memory devices in the face of the following limitations: (1) a flash write element (page) differs in size from a flash erase element (block), (2) a block must be erased before its is rewritten, (3) lifetime of a flash memory is typically limited to about 1,000,000 erases, (4) as many as 2 percent of the blocks of a given device may fail before the expected end of its life, and (5) to ensure reliability of reading and writing, power must not be interrupted during minimum specified reading and writing times. The first development comprises interrelated software components that regulate reading, writing, and erasure operations to minimize migration of data and unevenness in wear; perform erasures during idle times; quickly make erased blocks available for writing; detect and report failed blocks; maintain the overall state of a flash memory to satisfy real-time performance requirements; and detect and initialize a new flash memory device. The second development is a combination of hardware and software that senses the failure of a main power supply and draws power from a capacitive storage circuit designed to hold enough energy to sustain operation until reading or writing is completed.

Gender, Thomas K.

Development of Methodology for Programming Autonomous Agents

A brief report discusses the rationale for, and the development of, a methodology for generating computer code for autonomous-agent-based systems. The methodology is characterized as enabling an increase in the reusability of the generated code among and within such systems, thereby making it possible to reduce the time and cost of development of the systems. The methodology is also characterized as enabling reduction of the incidence of those software errors that are attributable to the human failure to anticipate distributed behaviors caused by the software. A major conceptual problem said to be addressed in the development of the methodology was that of how to efficiently describe the interfaces between several layers of agent composition by use of a language that is both familiar to engineers and descriptive enough to describe such interfaces unambivalently

Erol, Kutluhan

Maintaining the Health of Software Monitors

Software health management (SWHM) techniques complement the rigorous verification and validation processes that are applied to safety-critical systems prior to their deployment. These techniques are used to monitor deployed software in its execution environment, serving as the last line of defense against the effects of a critical fault. SWHM monitors use information from the specification and implementation of the monitored software to detect violations, predict possible failures, and help the system recover from faults. Changes to the monitored software, such as adding new functionality or fixing defects, therefore, have the potential to impact the correctness of both the monitored software and the SWHM monitor. In this work, we describe how the results of a software change impact analysis technique, Directed Incremental Symbolic Execution (DiSE), can be applied to monitored software to identify the potential impact of the changes on the SWHM monitor software. The results of DiSE can then be used by other analysis techniques, e.g., testing, debugging, to help preserve and improve the integrity of the SWHM monitor as the monitored software evolves.

Runtime Monitor

Micrometeoroid and Orbital Debris Risk Assessment With Bumper 3

The Bumper 3 computer code is the primary tool used by NASA for micrometeoroid and orbital debris (MMOD) risk analysis. Bumper 3 (and its predecessors) have been used to analyze a variety of manned and unmanned spacecraft. The code uses NASA's latest micrometeoroid (MEM-R2) and orbital debris (ORDEM 3.0) environment definition models and is updated frequently with ballistic limit equations that describe the hypervelocity impact performance of spacecraft materials. The Bumper 3 program uses these inputs along with a finite element representation of spacecraft geometry to provide a deterministic calculation of the expected number of failures. The Bumper 3 software is configuration controlled by the NASA/JSC Hypervelocity Impact Technology (HVIT) Group. This paper will demonstrate MMOD risk assessment techniques with Bumper 3 used by NASA's HVIT Group. The Permanent Multipurpose Module (PMM) was added to the International Space Station in 2011. A Bumper 3 MMOD risk assessment of this module will show techniques used to create the input model and assign the property IDs. The methodology used to optimize the MMOD shielding for minimum mass while still meeting structural penetration requirements will also be demonstrated.

Hyde, J.

Software Error Incident Categorizations in Aerospace

Since the first use of computers in space and aircraft, software errors have occurred. These errors can manifest as loss-of-life or less catastrophically. As the demand for automation increases, software in safety-critical systems should be designed to be tolerant to the most likely software faults. This paper categorizes historic aerospace software errors to determine trends of how and where automation is most likely to fail. A distinction between software producing wrong (erroneous) output versus no output (fail-silent) is introduced. Of the historical incidents analyzed, 87% were from software acting unexpectedly rather than simply stopping. Rebooting was found to be ineffective to clear erroneous behavior, and only partially effective for silent software. Errors were traced back to the software logic itself in 62% of cases, 13% within configurable data, and 25% introduced through input. Thirty percent (30%) of unexpected software behavior was caused by the absence of software and 20% was due to “unknown-unknowns”. These findings indicate that to achieve fault tolerance in safety-critical systems, backup strategies must be employed to detect and respond to erroneous software behavior beyond only fail-silent cases, and robust off-nominal testing should be performed to uncover unanticipated situations.

Software

The implementation and use of Ada on distributed systems with high reliability requirements

The use and implementation of Ada in distributed environments in which reliability is the primary concern were investigated. In particular, the concept that a distributed system may be programmed entirely in Ada so that the individual tasks of the system are unconcerned with which processors they are executing on, and that failures may occur in the software or underlying hardware was examined. Progress is discussed for the following areas: continued development and testing of the fault-tolerant Ada testbed; development of suggested changes to Ada so that it might more easily cope with the failure of interest; and design of new approaches to fault-tolerant software in real-time systems, and integration of these ideas into Ada.

Knight, J. C.

Command and Control System Software Development

With the first launch of the National Aeronautics and Space Administration's Space Launch System heavy-lift expendable launch vehicle and Lockheed Martin's Orion Multi-Purpose Crew Vehicle scheduled for the year 2020, there exists a need to complete development of a new command and control system that will provide systems monitoring and launch control for NASA's Exploration Missions. One remaining task necessary for completion of this command and control system is to create and maintain comprehensive unit tests of the control system software packages. These tests should verify that the implementation of all required and desired functionality works as intended. This testing infrastructure is mostly in place, but the control system's open source automation server still reports software "bugs" (possible flaws or failures which may lead to unintended behavior) and intermittently failing unit tests. Since code correctness is of critical importance for human rated software systems, I was assigned to diagnose the root cause of failing unit tests, eliminate non-determinism in these tests, and fix bugs as reported by the automation server.

GUI

DECIDER

This software offers methods and functions for building failure detectors for deep image classification models with the aid of vision-language models and LLMs. It includes functionalities for training baseline image classifiers, debiasing classifiers using vision-language models and LLMs, evaluating failure between models along with baselines. Developed using PyTorch, this software is compatible with standard neural network architectures used for imaging data. Additionally, it provides capabilities to compute evaluation metrics for assessing the performance and quality of the detectors.

Narayanaswamy, Vivek Sivaraman

Model Based Engineering for Software Assurance

NASA's successful development of next generation space vehicles, habitats, and robotic systems will require reliable hardware and software systems. The aim of this initiative is to develop modeling methodology and tools to support Model-Based Systems Engineering (MBSE) for software assurance and reliability analysis. This effort expands the Unified Modeling Language (UML) software design models to include fault data for the extraction of Failure Modes and Effects Criticality Analysis (FMECA) and Fault Tree Analysis (FTA) for software. We explored different modeling approaches to integrate the UML software design models with the Systems Modeling Language (SysML) system models to generate an integrated model and reliability tools that take into account software and hardware interfaces.The benefits of this concept directly affect the safety community with quick turnarounds to produce software assurance and reliability analysis artifacts and the ability to visualize failure effects, both hardware and software. The result is enhanced system design integrity and early identification of system risks. This initiative will enable software assurance activities early in the system design lifecycle, facilitating the discovery of design weaknesses and enhancing the capability to produce safe, hazard-free systems

Wang, Lui