Search NASA⌕ Search

SEARCH · Search NASA

Results for “fault tolerant computing”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 577 records · Page 32

Logical Shadow Tomography: Efficient Estimation of Error-mitigated Observables

In near-term quantum applications, reducing errors and improving device reliability is an essential task. Towards these ends, various techniques have been introduced in recent literature, collectively referred to as quantum error mitigation techniques, for reducing errors in pre-fault-tolerant devices. Here, we introduce logical shadow tomography as a versatile error mitigation method. Our technique uses a stabilizer code to encode information in a logical state. Instead of doing active error correction, quantum states will be measured at the end of computation via shadow tomography and non-logical errors are projected out in the classical post-processing. Relative to quantum subspace expansion which requires O(2(M-1)L) experiments to estimate an logical Pauli observable encoded by an [[M, L, d]] code, our technique only requires 2L experiments, an important practical reduction in resources.

Hong-Ye Hu↗

DEPEND: A simulation-based environment for system level dependability analysis

The design and evaluation of highly reliable computer systems is a complex issue. Designers mostly develop such systems based on prior knowledge and experience and occasionally from analytical evaluations of simplified designs. A simulation-based environment called DEPEND which is especially geared for the design and evaluation of fault-tolerant architectures is presented. DEPEND is unique in that it exploits the properties of object-oriented programming to provide a flexible framework with which a user can rapidly model and evaluate various fault-tolerant systems. The key features of the DEPEND environment are described, and its capabilities are illustrated with a detailed analysis of a real design. In particular, DEPEND is used to simulate the Unix based Tandem Integrity fault-tolerance and evaluate how well it handles near-coincident errors caused by correlated and latent faults. Issues such as memory scrubbing, re-integration policies, and workload dependent repair times which affect how the system handles near-coincident errors are also evaluated. Issues such as the method used by DEPEND to simulate error latency and the time acceleration technique that provides enormous simulation speed up are also discussed. Unlike any other simulation-based dependability studies, the use of these approaches and the accuracy of the simulation model are validated by comparing the results of the simulations, with measurements obtained from fault injection experiments conducted on a production Tandem Integrity machine.

Goswami, Kumar↗

Software Error Incident Categorizations in Aerospace

Since the first use of computers in space and aircraft, software errors have occurred. These errors can manifest as loss-of-life or less catastrophically. As the demand for automation increases, software in safety-critical systems should be designed to be tolerant to the most likely software faults. This paper categorizes historic aerospace software errors to determine trends of how and where automation is most likely to fail. A distinction between software producing wrong (erroneous) output versus no output (fail-silent) is introduced. Of the historical incidents analyzed, 87% were from software acting unexpectedly rather than simply stopping. Rebooting was found to be ineffective to clear erroneous behavior, and only partially effective for silent software. Errors were traced back to the software logic itself in 62% of cases, 13% within configurable data, and 25% introduced through input. Thirty percent (30%) of unexpected software behavior was caused by the absence of software and 20% was due to “unknown-unknowns”. These findings indicate that to achieve fault tolerance in safety-critical systems, backup strategies must be employed to detect and respond to erroneous software behavior beyond only fail-silent cases, and robust off-nominal testing should be performed to uncover unanticipated situations.

Software↗

Quantum error thresholds for gauge-redundant digitizations of lattice field theories

In the quantum simulation of lattice gauge theories, gauge symmetry can be either fixed or encoded as a redundancy of the Hilbert space. While gauge-fixing reduces the number of qubits, keeping the gauge redundancy can provide code space to mitigate and correct quantum errors by checking and restoring Gauss’s law. In this work, we consider the correctable errors for generic finite gauge groups and design the quantum circuits to detect and correct them. We calculate the error thresholds below which the gauge-redundant digitization with Gauss’s law error correction has better fidelity than the gauge-fixed digitization involving only gauge-invariant states. Our results provide guidance for fault-tolerant quantum simulations of lattice gauge theories. Published by the American Physical Society 2024

97 MATHEMATICS AND COMPUTING↗

The embedded operating system project

This progress report describes research towards the design and construction of embedded operating systems for real-time advanced aerospace applications. The applications concerned require reliable operating system support that must accommodate networks of computers. The report addresses the problems of constructing such operating systems, the communications media, reconfiguration, consistency and recovery in a distributed system, and the issues of realtime processing. A discussion is included on suitable theoretical foundations for the use of atomic actions to support fault tolerance and data consistency in real-time object-based systems. In particular, this report addresses: atomic actions, fault tolerance, operating system structure, program development, reliability and availability, and networking issues. This document reports the status of various experiments designed and conducted to investigate embedded operating system design issues.

Campbell, R. H.↗

The embedded operating system project

The design and construction of embedded operating systems for real-time advanced aerospace applications was investigated. The applications require reliable operating system support that must accommodate computer networks. Problems that arise in the construction of such operating systems, reconfiguration, consistency and recovery in a distributed system, and the issues of real-time processing are reported. A thesis that provides theoretical foundations for the use of atomic actions to support fault tolerance and data consistency in real-time object-based system is included. The following items are addressed: (1) atomic actions and fault-tolerance issues; (2) operating system structure; (3) program development; (4) a reliable compiler for path Pascal; and (5) mediators, a mechanism for scheduling distributed system processes.

Campbell, R. H.↗

Study on advanced information processing system

Issues related to the reliability of a redundant system with large main memory are addressed. In particular, the Fault-Tolerant Processor (FTP) for Advanced Launch System (ALS) is used as a basis for our presentation. When the system is free of latent faults, the probability of system crash due to nearly-coincident channel faults is shown to be insignificant even when the outputs of computing channels are infrequently voted on. In particular, using channel error maskers (CEMs) is shown to improve reliability more effectively than increasing the number of channels for applications with long mission times. Even without using a voter, most memory errors can be immediately corrected by CEMs implemented with conventional coding techniques. In addition to their ability to enhance system reliability, CEMs--with a low hardware overhead--can be used to reduce not only the need of memory realignment, but also the time required to realign channel memories in case, albeit rare, such a need arises. Using CEMs, we have developed two schemes, called Scheme 1 and Scheme 2, to solve the memory realignment problem. In both schemes, most errors are corrected by CEMs, and the remaining errors are masked by a voter.

Shin, Kang G.↗

Redundant Buses For Loosely Coupled Dual Computers

Digital electronic network contains two data processors loosely coupled to each other and to four remote terminals via three MIL-STD-1553B data buses. Three-bus configuration provides redundancy for protection against failure of one of the processors and against failures of one or two buses. Triple-bus architecture used for fault tolerance in similar terrestrial digital electronic systems.

Katz, Richard B.↗

Heavy Lift Vehicle (HLV) Avionics Flight Computing Architecture Study

A NASA multi-Center study team was assembled from LaRC, MSFC, KSC, JSC and WFF to examine potential flight computing architectures for a Heavy Lift Vehicle (HLV) to better understand avionics drivers. The study examined Design Reference Missions (DRMs) and vehicle requirements that could impact the vehicles avionics. The study considered multiple self-checking and voting architectural variants and examined reliability, fault-tolerance, mass, power, and redundancy management impacts. Furthermore, a goal of the study was to develop the skills and tools needed to rapidly assess additional architectures should requirements or assumptions change.

Hodson, Robert F.↗

SpaceVPX Interoperability Assessment

The existing VMEbus (VersaModular Eurocard bus) International Trade Association (VITA)-78 industry standard, also known as SpaceVPX, is an avionics board- and chassis-level standard derived from the OpenVPX standard as defined in VITA-65. While VITA-65 defines backplane and board-level profiles from COTS vendors to ensure interoperability of products used in developing systems and subsystems, the VITA-78 standard defines SpaceVPX to incorporate fault tolerance features that are required by many spaceflight systems. However, VITA-78 allows so much flexibility that interoperability between modules cannot be assured. This assessment provides guidelines on the use of, and extensions to, the VITA-78 standard to enable avionics interoperability for future NASA missions. The assessment team was comprised of subject matter experts (SMEs) from Goddard Space Flight Center (GSFC), the Jet Propulsion Laboratory (JPL), Johnson Space Center (JSC), and Langley Research Center (LaRC). The team included valuable external consulting support from a SME who was a key participant in the development of the VITA-78 standard. The team had extensive collaboration with the NASA Space Technology Mission Directorate (STMD) High Performance Spaceflight Computing (HPSC) project, specifically in the development of SpaceVPX interconnect findings, observations, and NESC recommendations. To provide an understanding of the breadth of implementations that SpaceVPX must accommodate, multiple NASA use cases were analyzed to assess the requirements for SpaceVPX implementations across a wide range of NASA missions (Appendix C). Applications included crewed missions, science missions, and orbital and surface robotic systems. Product surveys were conducted to assess the level of industry support for SpaceVPX, applications, and the variations in their implementations (Appendix D). In-depth analysis was conducted in the areas of: (a) power management and distribution, (b) form factors and daughtercards, (c) interconnect, and (d) fault tolerance. Leveraging the use cases, product surveys, and SMEs from multiple NASA Centers, these areas were analyzed to determine the range of implementations permitted by the VITA-78 standard and potential interoperability issues. Applicable findings and NESC recommendations were provided for each area. During this assessment, there were multiple opportunities to engage with other agencies to learn about their interest in SpaceVPX, their strategies for implementing SpaceVPX-based systems, and their internal development efforts. These engagements also generated findings and NESC recommendations. Based on this assessment analysis, NESC recommendations were made regarding the feature set and module profiles to support NASA SpaceVPX implementations. This feature set includes restrictions on features in VITA-78, and extensions to the standard. Key recommendations in this area include the use of 10 Gigabit Ethernet and Peripheral Component Interconnect Express (PCIe) as high bandwidth interconnect on the backplane, the retention of SpaceWire interconnect for control functions, and support for 3U (unit) and 6U, form factors for NASA systems. Restrictions were proposed on the usage of user-defined signals to promote interoperability, and specific power managements and distribution schemes for 3U systems. Beyond the technical implementation of SpaceVPX, recommendations were made on areas that warrant further investigation. Primary among these is the recommendation for NASA to collaborate with other space-going agencies and industry to incorporate recommendations into a future ‘dot spec’ of VITA-78. This would ensure wide adoption and availability of the modules that comply with the specification. The assessment includes appendices with candidate module profiles that can be considered as a starting point for this activity, and example systems based on the recommendations. Follow-on studies are recommended for architectures beyond SpaceVPX to address potential enhancements including condensed set of interconnect, software required to implement protocol layers on the interconnect (and other features), alternative power architectures, and system-level testability.

SpaceVPX↗

Study on fault-tolerant processors for advanced launch system

Issues related to the reliability of a redundant system with large main memory are addressed. The Fault-Tolerant Processor (FTP) for the Advanced Launch System (ALS) is used as a basis for the presentation. When the system is free of latent faults, the probability of system crash due to multiple channel faults is shown to be insignificant even when voting on the outputs of computing channels is infrequent. Using channel error maskers (CEMs) is shown to improve reliability more effectively than increasing redundancy or the number of channels for applications with long mission times. Even without using a voter, most memory errors can be immediately corrected by those CEMs implemented with conventional coding techniques. In addition to their ability to enhance system reliability, CEMs (with a very low hardware overhead) can be used to dramatically reduce not only the need of memory realignment, but also the time required to realign channel memories in case, albeit rare, such a need arises. Using CEMs, two different schemes were developed to solve the memory realignment problem. In both schemes, most errors are corrected by CEMs, and the remaining errors are masked by a voter.

Shin, Kang G.↗

Evaluation of Advanced Computing Techniques and Technologies: Reconfigurable Computing

The focus of this project was to survey the technology of reconfigurable computing determine its level of maturity and suitability for NASA applications. To better understand and assess the effectiveness of the reconfigurable design paradigm that is utilized within the HAL-15 reconfigurable computer system. This system was made available to NASA MSFC for this purpose, from Star Bridge Systems, Inc. To implement on at least one application that would benefit from the performance levels that are possible with reconfigurable hardware. It was originally proposed that experiments in fault tolerance and dynamically reconfigurability would be perform but time constraints mandated that these be pursued as future research.

Wells, B. Earl↗

Hierarchical specification of the SIFT fault tolerant flight control system

The specification and mechanical verification of the Software Implemented Fault Tolerance (SIFT) flight control system is described. The methodology employed in the verification effort is discussed, and a description of the hierarchical models of the SIFT system is given. To meet the objective of NASA for the reliability of safety critical flight control systems, the SIFT computer must achieve a reliability well beyond the levels at which reliability can be actually measured. The methodology employed to demonstrate rigorously that the SIFT computer meets as reliability requirements is described. The hierarchy of design specifications from very abstract descriptions of system function down to the actual implementation is explained. The most abstract design specifications can be used to verify that the system functions correctly and with the desired reliability since almost all details of the realization were abstracted out. A succession of lower level models refine these specifications to the level of the actual implementation, and can be used to demonstrate that the implementation has the properties claimed of the abstract design specifications.

Melliar-Smith, P. M.↗

Mechanical verification of a schematic Byzantine clock synchronization algorithm

Schneider generalizes a number of protocols for Byzantine fault tolerant clock synchronization and presents a uniform proof for their correctness. The authors present a machine checked proof of this schematic protocol that revises some of the details in Schneider's original analysis. The verification was carried out with the EHDM system developed at the SRI Computer Science Laboratory. The mechanically checked proofs include the verification that the egocentric mean function used in Lamport and Melliar-Smith's Interactive Convergence Algorithm satisfies the requirements of Schneider's protocol.

Shankar, Natarajan↗

Reliability Model Generator for fault-tolerant systems

This paper discusses an analysis tool, the Reliability Model Generator, that reasons from structural and functional system design specifications to generate a reliability model for the system under investigation. The resultant model defines a system state space sufficient to characterize the effects of single and multiple component failures. This model may then be examined by the analyst or used as input to an existing reliability evaluation tool, the Semi-Markov Unreliability Range Evaluator (SURE), to compute numeric bounds for system reliability (i.e., safety, mission success, availability, etc.). In defining the input to the Reliability Model Generator, a separation of the component function from structural specification is proposed to allow easy modification for analysis of alternative architectures. A hierarchical system description paradigm promotes multiple abstractions, thereby enabling analysis at all phases of the design process. The work described in this paper has been supported under NASA contract NAS1-10899, Integrated Airframe/Propulsion Control System Architecture (IAPSA II). IAPSA II promotes a system engineering methodology and supporting analytical tools to evaluate flight control architecture configuration alternatives during early phases of the design cycle when the cost and schedule impact of design revisions is minimal. Emphasis is placed on traceability to ensure that the requirements drive the resulting design.

Catherine M McCann↗

Intelligent neuroprocessors for in-situ launch vehicle propulsion systems health management

Efficacy of existing on-board propulsion systems health management systems (HMS) are severely impacted by computational limitations (e.g., low sampling rates); paradigmatic limitations (e.g., low-fidelity logic/parameter redlining only, false alarms due to noisy/corrupted sensor signatures, preprogrammed diagnostics only); and telemetry bandwidth limitations on space/ground interactions. Ultra-compact/light, adaptive neural networks with massively parallel, asynchronous, fast reconfigurable and fault-tolerant information processing properties have already demonstrated significant potential for inflight diagnostic analyses and resource allocation with reduced ground dependence. In particular, they can automatically exploit correlation effects across multiple sensor streams (plume analyzer, flow meters, vibration detectors, etc.) so as to detect anomaly signatures that cannot be determined from the exploitation of single sensor. Furthermore, neural networks have already demonstrated the potential for impacting real-time fault recovery in vehicle subsystems by adaptively regulating combustion mixture/power subsystems and optimizing resource utilization under degraded conditions. A class of high-performance neuroprocessors, developed at JPL, that have demonstrated potential for next-generation HMS for a family of space transportation vehicles envisioned for the next few decades, including HLLV, NLS, and space shuttle is presented. Of fundamental interest are intelligent neuroprocessors for real-time plume analysis, optimizing combustion mixture-ratio, and feedback to hydraulic, pneumatic control systems. This class includes concurrently asynchronous reprogrammable, nonvolatile, analog neural processors with high speed, high bandwidth electronic/optical I/O interfaced, with special emphasis on NASA's unique requirements in terms of performance, reliability, ultra-high density ultra-compactness, ultra-light weight devices, radiation hardened devices, power stringency, and long life terms.

Gulati, S.↗

Use of Field Programmable Gate Array Technology in Future Space Avionics

Fulfilling NASA's new vision for space exploration requires the development of sustainable, flexible and fault tolerant spacecraft control systems. The traditional development paradigm consists of the purchase or fabrication of hardware boards with fixed processor and/or Digital Signal Processing (DSP) components interconnected via a standardized bus system. This is followed by the purchase and/or development of software. This paradigm has several disadvantages for the development of systems to support NASA's new vision. Building a system to be fault tolerant increases the complexity and decreases the performance of included software. Standard bus design and conventional implementation produces natural bottlenecks. Configuring hardware components in systems containing common processors and DSPs is difficult initially and expensive or impossible to change later. The existence of Hardware Description Languages (HDLs), the recent increase in performance, density and radiation tolerance of Field Programmable Gate Arrays (FPGAs), and Intellectual Property (IP) Cores provides the technology for reprogrammable Systems on a Chip (SOC). This technology supports a paradigm better suited for NASA's vision. Hardware and software production are melded for more effective development; they can both evolve together over time. Designers incorporating this technology into future avionics can benefit from its flexibility. Systems can be designed with improved fault isolation and tolerance using hardware instead of software. Also, these designs can be protected from obsolescence problems where maintenance is compromised via component and vendor availability.To investigate the flexibility of this technology, the core of the Central Processing Unit and Input/Output Processor of the Space Shuttle AP101S Computer were prototyped in Verilog HDL and synthesized into an Altera Stratix FPGA.

Ferguson, Roscoe C.↗

SEE Test Results for SAMA5D3

ARM processors power a class of high-performance, lower power system on a chip devices. In the absence of radiation effects, these devices are highly desirable for space use. The processor core architecture for ARM devices is licensed to provide computing on multiple hardware platforms. The A5 processor is in a unique pioneering space for providing detailed radiation response data to explore the baseline performance of these devices. These data can help set options for ARM processors and possibly impact design choices for the next generation of ARM fault tolerance capabilities. The SAMA5D3 was tested to establish general SEE performance for a relatively simple implementation of the ARM A5 core. This testing observed SRAM sensitivity starting at an LET of about 3 MeV-cm2/mg, with a saturated cross section of about 2x10-8cm2/bit, and this was determined by both active write and read of the caches, in addition to the use of a debugger to provide test results. Crash/SEFI data was collected using both Linux and bare metal C-code. The onset LET for crashes was about LET 1.5 MeV-cm2/mg, with saturated cross sections of about 2x10-5 cm2 for bare metal (low utilization), and 2x10-4cm2 for Linux (high utilization) tests.

Daniel, Andrew C.↗