Search NASA⌕ Search

SEARCH · Search NASA

Results for “Software Reliability”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 685 records · Page 38

Integrated Software Health Management for Aircraft GN and C

Modern aircraft rely heavily on dependable operation of many safety-critical software components. Despite careful design, verification and validation (V&V), on-board software can fail with disastrous consequences if it encounters problematic software/hardware interaction or must operate in an unexpected environment. We are using a Bayesian approach to monitor the software and its behavior during operation and provide up-to-date information about the health of the software and its components. The powerful reasoning mechanism provided by our model-based Bayesian approach makes reliable diagnosis of the root causes possible and minimizes the number of false alarms. Compilation of the Bayesian model into compact arithmetic circuits makes SWHM feasible even on platforms with limited CPU power. We show initial results of SWHM on a small simulator of an embedded aircraft software system, where software and sensor faults can be injected.

Schumann, Johann↗

Automated Translation of Safety Critical Application Software Specifications into PLC Ladder Logic

The numerous benefits of automatic application code generation are widely accepted within the software engineering community. A few of these benefits include raising the abstraction level of application programming, shorter product development time, lower maintenance costs, and increased code quality and consistency. Surprisingly, code generation concepts have not yet found wide acceptance and use in the field of programmable logic controller (PLC) software development. Software engineers at the NASA Kennedy Space Center (KSC) recognized the need for PLC code generation while developing their new ground checkout and launch processing system. They developed a process and a prototype software tool that automatically translates a high-level representation or specification of safety critical application software into ladder logic that executes on a PLC. This process and tool are expected to increase the reliability of the PLC code over that which is written manually, and may even lower life-cycle costs and shorten the development schedule of the new control system at KSC. This paper examines the problem domain and discusses the process and software tool that were prototyped by the KSC software engineers.

Leucht, Kurt W.↗

Functional Near-Infrared Spectroscopy Signals Measure Neuronal Activity in the Cortex

Functional near infrared spectroscopy (fNIRS) is an emerging optical neuroimaging technology that indirectly measures neuronal activity in the cortex via neurovascular coupling. It quantifies hemoglobin concentration ([Hb]) and thus measures the same hemodynamic response as functional magnetic resonance imaging (fMRI), but is portable, non-confining, relatively inexpensive, and is appropriate for long-duration monitoring and use at the bedside. Like fMRI, it is noninvasive and safe for repeated measurements. Patterns of [Hb] changes are used to classify cognitive state. Thus, fNIRS technology offers much potential for application in operational contexts. For instance, the use of fNIRS to detect the mental state of commercial aircraft operators in near real time could allow intelligent flight decks of the future to optimally support human performance in the interest of safety by responding to hazardous mental states of the operator. However, many opportunities remain for improving robustness and reliability. It is desirable to reduce the impact of motion and poor optical coupling of probes to the skin. Such artifacts degrade signal quality and thus cognitive state classification accuracy. Field application calls for further development of algorithms and filters for the automation of bad channel detection and dynamic artifact removal. This work introduces a novel adaptive filter method for automated real-time fNIRS signal quality detection and improvement. The output signal (after filtering) will have had contributions from motion and poor coupling reduced or removed, thus leaving a signal more indicative of changes due to hemodynamic brain activations of interest. Cognitive state classifications based on these signals reflect brain activity more reliably. The filter has been tested successfully with both synthetic and real human subject data, and requires no auxiliary measurement. This method could be implemented as a real-time filtering option or bad channel rejection feature of software used with frequency domain fNIRS instruments for signal acquisition and processing. Use of this method could improve the reliability of any operational or real-world application of fNIRS in which motion is an inherent part of the functional task of interest. Other optical diagnostic techniques (e.g., for NIR medical diagnosis) also may benefit from the reduction of probe motion artifact during any use in which motion avoidance would be impractical or limit usability.

Harrivel, Angela↗

Functional Near-Infrared Spectroscopy Signals Measure Neuronal Activity in the Cortex

Functional near infrared spectroscopy (fNIRS) is an emerging optical neuroimaging technology that indirectly measures neuronal activity in the cortex via neurovascular coupling. It quantifies hemoglobin concentration ([Hb]) and thus measures the same hemodynamic response as functional magnetic resonance imaging (fMRI), but is portable, non-confining, relatively inexpensive, and is appropriate for long-duration monitoring and use at the bedside. Like fMRI, it is noninvasive and safe for repeated measurements. Patterns of [Hb] changes are used to classify cognitive state. Thus, fNIRS technology offers much potential for application in operational contexts. For instance, the use of fNIRS to detect the mental state of commercial aircraft operators in near real time could allow intelligent flight decks of the future to optimally support human performance in the interest of safety by responding to hazardous mental states of the operator. However, many opportunities remain for improving robustness and reliability. It is desirable to reduce the impact of motion and poor optical coupling of probes to the skin. Such artifacts degrade signal quality and thus cognitive state classification accuracy. Field application calls for further development of algorithms and filters for the automation of bad channel detection and dynamic artifact removal. This work introduces a novel adaptive filter method for automated real-time fNIRS signal quality detection and improvement. The output signal (after filtering) will have had contributions from motion and poor coupling reduced or removed, thus leaving a signal more indicative of changes due to hemodynamic brain activations of interest. Cognitive state classifications based on these signals reflect brain activity more reliably. The filter has been tested successfully with both synthetic and real human subject data, and requires no auxiliary measurement. This method could be implemented as a real-time filtering option or bad channel rejection feature of software used with frequency domain fNIRS instruments for signal acquisition and processing. Use of this method could improve the reliability of any operational or real-world application of fNIRS in which motion is an inherent part of the functional task of interest. Other optical diagnostic techniques (e.g., for NIR medical diagnosis) also may benefit from the reduction of probe motion artifact during any use in which motion avoidance would be impractical or limit usability.

Harrivel, Angela↗

Simulation-Based Verification of Autonomous Controllers via Livingstone PathFinder

AI software is often used as a means for providing greater autonomy to automated systems, capable of coping with harsh and unpredictable environments. Due in part to the enormous space of possible situations that they aim to addrs, autonomous systems pose a serious challenge to traditional test-based verification approaches. Efficient verification approaches need to be perfected before these systems can reliably control critical applications. This publication describes Livingstone PathFinder (LPF), a verification tool for autonomous control software. LPF applies state space exploration algorithms to an instrumented testbed, consisting of the controller embedded in a simulated operating environment. Although LPF has focused on NASA s Livingstone model-based diagnosis system applications, the architecture is modular and adaptable to other systems. This article presents different facets of LPF and experimental results from applying the software to a Livingstone model of the main propulsion feed subsystem for a prototype space vehicle.

Lindsey, A. E.↗

Operating executive for the DSIF tracking subsystem software

The advanced engineering model of the DSIF tracking subsystem (DTS) is being developed by the Deep Space Instrumentation Facility. The DTS will provide effective and reliable tracking and data acquisition support for the complex planetary and interplanetary space flight missions planned for the 1970's. The nucleus of the subsystem is a Honeywell H832 digital computer. The design and capabilities of the real-time operating executive software are described.

Poulson, P. L.↗

Data flow modeling techniques

There have been a number of simulation packages developed for the purpose of designing, testing and validating computer systems, digital systems and software systems. Complex analytical tools based on Markov and semi-Markov processes have been designed to estimate the reliability and performance of simulated systems. Petri nets have received wide acceptance for modeling complex and highly parallel computers. In this research data flow models for computer systems are investigated. Data flow models can be used to simulate both software and hardware in a uniform manner. Data flow simulation techniques provide the computer systems designer with a CAD environment which enables highly parallel complex systems to be defined, evaluated at all levels and finally implemented in either hardware or software. Inherent in data flow concept is the hierarchical handling of complex systems. In this paper we will describe how data flow can be used to model computer system.

Kavi, K. M.↗

The Software Correlator of the Chinese VLBI Network

The software correlator of the Chinese VLBI Network (CVN) has played an irreplaceable role in the CVN routine data processing, e.g., in the Chinese lunar exploration project. This correlator will be upgraded to process geodetic and astronomical observation data. In the future, with several new stations joining the network, CVN will carry out crustal movement observations, quick UT1 measurements, astrophysical observations, and deep space exploration activities. For the geodetic or astronomical observations, we need a wide-band 10-station correlator. For spacecraft tracking, a realtime and highly reliable correlator is essential. To meet the scientific and navigation requirements of CVN, two parallel software correlators in the multiprocessor environments are under development. A high speed, 10-station prototype correlator using the mixed Pthreads and MPI (Massage Passing Interface) parallel algorithm on a computer cluster platform is being developed. Another real-time software correlator for spacecraft tracking adopts the thread-parallel technology, and it runs on the SMP (Symmetric Multiple Processor) servers. Both correlators have the characteristic of flexible structure and scalability.

Zheng, Weimin↗

Adding Assurance to Automatically Generated Code

Code to estimate position and attitude of a spacecraft or aircraft belongs to the most safety-critical parts of flight software. The complex underlying mathematics and abundance of design details make it error-prone and reliable implementations costly. AutoFilter is a program synthesis tool for the automatic generation of state estimation code from compact specifications. It can automatically produce additional safety certificates which formally guarantee that each generated program individually satisfies a set of important safety policies. These safety policies (e.g.. array-bounds, variable initialization) form a core of properties which are essential for high-assurance software. Here we describe the AutoFilter system and its certificate generator and compare our approach to the static analysis tool PolySpace.

Denney, Ewen↗

Enabling Reliable, Fault-Tolerant Autonomous Lunar Habitats with High-Performance Spaceflight Computing

The lunar surface presents unfavorable constraints and harsh living conditions. To address these challenges, autonomous habitats will require complex integrated systems that combine advanced software, high-performance hardware, and cutting-edge sensors to ensure sustainability, safety, and operational efficiency. Consequently, maintaining a sustainable presence on the Moon requires reliable infrastructure and efficient development, precise monitoring, and utilization of resources within a lunar installation. These elements are essential not only to ensure that lunar settlement can be long-term, self-sustaining, and resource-efficient, but also to serve as a foundation for future missions and eventual human habitation on Mars. Humans are not native to the Moon; therefore, our survival and ability to thrive will depend on autonomous systems that can foster safety and resilience through high-availability architectures, graceful degradation, and highly fault-tolerant spaceflight hardware capable of continuing operation during failures. This requires advanced human-rated distributed systems architectures with specialized electronics, scalable capabilities, and an integrated design approach. Unlike current practices focused on short-term missions and regularly maintained components, permanent lunar compute systems must be designed for extended operations beyond mission durations. This paper explores the necessity of transitioning toward fault- tolerant, highly autonomous hardware systems designed for multi-year missions. It also identifies critical subsystems that require high levels of autonomy, supported by radiation-hardened processors and extreme thermal loads, which are essential to mitigate long-term degradation and ensure sustainable lunar habitation. Finally, the paper aligns with NASA’s identified Civil Space Shortfalls, particularly in high-performance onboard computing, advanced data acquisition, extreme-environment avionics, radiation monitoring and countermeasures, and autonomous health management. It proposes NASA’s new High-Performance Spaceflight Computing (HPSC) processor as a turnkey solution, delivering 100 times the performance-per-watt of legacy rad-hard CPUs and enabling onboard AI, edge computing, and fault-tolerant features essential for sustained lunar autonomy and beyond.

Sarkis S Mikaelian↗

Probability approach for strength calculations

The use of probabilistic structural analysis methods (PSAM) to predict structural reliability is the subject of an on-going NASA research program. The elements of the new technology developed to date is reported. Applications of the developed software to structural problems are demonstrated for simple validation problems and for large scale application problems. On-going research to support component and system reliability predictions suitable for analytical certification of aerospace structures is briefly reviewed.

Chamis, Christos C.↗

An Analysis of Failure Handling in Chameleon, A Framework for Supporting Cost-Effective Fault Tolerant Services

The desire for low-cost reliable computing is increasing. Most current fault tolerant computing solutions are not very flexible, i.e., they cannot adapt to reliability requirements of newly emerging applications in business, commerce, and manufacturing. It is important that users have a flexible, reliable platform to support both critical and noncritical applications. Chameleon, under development at the Center for Reliable and High-Performance Computing at the University of Illinois, is a software framework. for supporting cost-effective adaptable networked fault tolerant service. This thesis details a simulation of fault injection, detection, and recovery in Chameleon. The simulation was written in C++ using the DEPEND simulation library. The results obtained from the simulation included the amount of overhead incurred by the fault detection and recovery mechanisms supported by Chameleon. In addition, information about fault scenarios from which Chameleon cannot recover was gained. The results of the simulation showed that both critical and noncritical applications can be executed in the Chameleon environment with a fairly small amount of overhead. No single point of failure from which Chameleon could not recover was found. Chameleon was also found to be capable of recovering from several multiple failure scenarios.

Haakensen, Erik Edward↗

Advances in Distributed Operations and Mission Activity Planning for Mars Surface Exploration

A centralized mission activity planning system for any long-term mission, such as the Mars Exploration Rover Mission (MER), is completely infeasible due to budget and geographic constraints. A distributed operations system is key to addressing these constraints; therefore, future system and software engineers must focus on the problem of how to provide a secure, reliable, and distributed mission activity planning system. We will explain how Maestro, the next generation mission activity planning system, with its heavy emphasis on portability and distributed operations has been able to meet these design challenges. MER has been an excellent proving ground for Maestro's new approach to distributed operations. The backend that has been developed for Maestro could benefit many future missions by reducing the cost of centralized operations system architecture.

Mars Exploration Rover (MER)↗

Securing Ground Data System Applications for Space Operations

The increasing prevalence and sophistication of cyber attacks has prompted the Multimission Ground Systems and Services (MGSS) Program Office at Jet Propulsion Laboratory (JPL) to initiate the Common Access Manager (CAM) effort to protect software applications used in Ground Data Systems (GDSs) at JPL and other NASA Centers. The CAM software provides centralized services and software components used by GDS subsystems to meet access control requirements and ensure data integrity, confidentiality, and availability. In this paper we describe the CAM software; examples of its integration with spacecraft commanding software applications and an information management service; and measurements of its performance and reliability.

Security↗

Advances in distributed operations and mission activity planning for Mars surface exploration

A centralized mission activity planning system for any long-term mission, such as the Mars Exploration Rover Mission (MER), is completely infeasible due to budget and geographic constraints. A distributed operations system is key to addressing these constraints; therefore, future system and software engineers must focus on the problem of how to provide a secure, reliable, and distributed mission activity planning system. We will explain how Maestro, the next generation mission activity planning system, with its heavy emphasis on portability and distributed operations has been able to meet these design challenges. MER has been an excellent proving ground for Maestro's new approach to distributed operations. The backend that has been developed for Maestro could benefit many future missions by reducing the cost of centralized operations system architecture.

Shams, Khawaja↗

Risk considerations for autonomy software

This paper summarizes key findings related to methods for the risk considerations of autonomy software. Existing methods for Verification and Validation (V&V) of autonomy software are summarized and a method for the assurance of autonomy software is suggested and demonstrated on a command execution use case. Risk and reliability are defined in the context of autonomy and an approach for risk assessment of autonomy is presented using an example use case. Key insights regarding areas of uncertainty for autonomy are provided, along with a suggested architecture for the systematic consideration of reliability within the context of a given autonomous planner.

Meshkat, Leila↗

A Sustainable, Reliable Mission-Systems Architecture that Supports a System of Systems Approach to Space Exploration

A mission-systems architecture based on a highly modular "systems of systems" infrastructure utilizing open-standards hardware and software interfaces as the enabling technology is absolutely essential for an affordable and sustainable space exploration program. This architecture requires (a) robust communication between heterogeneous systems, (b) high reliability, (c) minimal mission-to-mission reconfiguration, (d) affordable development, system integration, and verification of systems, and (e) minimum sustaining engineering. This paper proposes such an architecture. Lessons learned from the space shuttle program are applied to help define and refine the model.

Watson, Steve↗

Historical Aerospace Software Errors Categorized to Influence Fault Tolerance

Since the first use of computers in space and aircraft, software errors have occurred. These errors can manifest as loss-of-life or less catastrophically. As the demand for automation increases, software in mission or safety-critical systems should be designed to be tolerant to the most likely software faults. This paper categorizes a set of 55 historic aerospace software error incidents from 1962 to 2023 to determine trends of how and where automation is most likely to fail, behaving unexpectedly. A distinction between software producing unexpected (erroneous) output versus no output (failsilent) is introduced. Of the historical incidents analyzed, 85% were from software producing wrong output rather than simply stopping. Rebooting was found to be ineffective to clear erroneous behavior, and not reliable to recover from silent failures. Error origin was within the code/logic itself in 58% of cases, 16% from configurable data, 15% from unexpected sensor input, and 11% from command/operator input. A substantial forty percent (40%) of unexpected software behavior was indicated by the absence of code, arising from unanticipated situations and missing requirements, and 16% of incidents were subjectively deemed “unknown-unknowns”. No incidents were found to be the result of programming language, compiler, tool, or operating system; and only sixteen percent (16%) of all incidents were considered errors traditional computer science/programming in nature. These findings indicate that for fault tolerance, erroneous automation behavior must be a primary consideration especially at critical moments, and reboot recoverability may not be viable. Special care should be taken to validate configurable data and commands prior to use. “Test-like-you-fly”, including hardware-in-the-loop combined with robust off-nominal testing should be used to uncover missing logic arising from unanticipated situations not covered by requirements alone. This study uniquely focuses on manifestations of unexpected flight software behavior, independent of ultimate root cause. We characterize software error behavior and origin to improve software design, test, and operations for resilience to the most common manifestations, and provide a rich dataset for further study.

Aerospace↗