Search NASA⌕ Search

SEARCH · Search NASA

Results for “fault tolerant computing”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 253 records · Page 14

Enabling Reliable, Fault-Tolerant Autonomous Lunar Habitats with High-Performance Spaceflight Computing

The lunar surface presents unfavorable constraints and harsh living conditions. To address these challenges, autonomous habitats will require complex integrated systems that combine advanced software, high-performance hardware, and cutting-edge sensors to ensure sustainability, safety, and operational efficiency. Consequently, maintaining a sustainable presence on the Moon requires reliable infrastructure and efficient development, precise monitoring, and utilization of resources within a lunar installation. These elements are essential not only to ensure that lunar settlement can be long-term, self-sustaining, and resource-efficient, but also to serve as a foundation for future missions and eventual human habitation on Mars. Humans are not native to the Moon; therefore, our survival and ability to thrive will depend on autonomous systems that can foster safety and resilience through high-availability architectures, graceful degradation, and highly fault-tolerant spaceflight hardware capable of continuing operation during failures. This requires advanced human-rated distributed systems architectures with specialized electronics, scalable capabilities, and an integrated design approach. Unlike current practices focused on short-term missions and regularly maintained components, permanent lunar compute systems must be designed for extended operations beyond mission durations. This paper explores the necessity of transitioning toward fault- tolerant, highly autonomous hardware systems designed for multi-year missions. It also identifies critical subsystems that require high levels of autonomy, supported by radiation-hardened processors and extreme thermal loads, which are essential to mitigate long-term degradation and ensure sustainable lunar habitation. Finally, the paper aligns with NASA’s identified Civil Space Shortfalls, particularly in high-performance onboard computing, advanced data acquisition, extreme-environment avionics, radiation monitoring and countermeasures, and autonomous health management. It proposes NASA’s new High-Performance Spaceflight Computing (HPSC) processor as a turnkey solution, delivering 100 times the performance-per-watt of legacy rad-hard CPUs and enabling onboard AI, edge computing, and fault-tolerant features essential for sustained lunar autonomy and beyond.

Sarkis S Mikaelian↗

Computer Reliability

Using a NASA developed program, Dr. J. Walter Bond is creating a course in computer reliability modeling. The course will examine three different computer programs, one of them NASA's Care III, the others UCLA's Aries 78 and Aries 82. All three are designed to help estimate the reliability of complex, redundant, fault tolerant system. In computer design, software of this kind can predict or model the effects of various hardware or software failures, a process called reliability modeling.

Source record↗

Distributed computing for autonomous on board planning and sequence validations

We propose a new conceptual approach to system-level autonomy that exploits in a synergistic way recent breakthroughs in three specific areas: automatic generation of embeddable planning and validation software, integration of telecommunications forecaster and planning tools, and fault-tolerant assignment of computing tasks to multiple processors.

autonomy synergy software telecommunications compu↗

Design and verification of a multiple fault tolerant control system for STS applications using computer simulation

General Dynamics/Convair is under NASA contract to integrate the Centaur upper stage into the space transportation system for future planetary missions. This requires that control of all safety critical functions be two-failure tolerant. The control system developed consists of five asynchronous computers, each contributing at their outputs to a 3-out-of-5 voting plane. Subsystem control is based on an end function redundancy management scheme. Analysis of multiple component failures and worst-case time-phase asynchrony among the computers is performed by a real-time computer simulation. The simulation emulates the hardware and subsystem interfaces, wire by wire, providing assessibility to any component for the insertion of preprogrammed failures. Observability is provided via a graphics system and diagnostic software. The simulation provides an engineering tool where the integrity of control system hardware and imbedded software can be demonstrated.

Szatkowski, G. P.↗

Digital avionics design and reliability analyzer

The description and specifications for a digital avionics design and reliability analyzer are given. Its basic function is to provide for the simulation and emulation of the various fault-tolerant digital avionic computer designs that are developed. It has been established that hardware emulation at the gate-level will be utilized. The primary benefit of emulation to reliability analysis is the fact that it provides the capability to model a system at a very detailed level. Emulation allows the direct insertion of faults into the system, rather than waiting for actual hardware failures to occur. This allows for controlled and accelerated testing of system reaction to hardware failures. There is a trade study which leads to the decision to specify a two-machine system, including an emulation computer connected to a general-purpose computer. There is also an evaluation of potential computers to serve as the emulation computer.

Source record↗

The Application of Microtechnology to Spacecraft On-Board Computing(abstract)

In this report, we will survey recent advances in chip packaging and stacking techniques that allow miniature computers to be developed for space applications. Several orders of magnitude reduction in mass, volume, and power consumption are possible using these techniques. Moreover, performance improvements can be achieved by increasing the scale of multiprocessing. Most importantly, long-term survivability can potentially be improved by increasing the level of redundancy and fault tolerance.

microtechnology computers multiprocessing packagin↗

Care 3, phase 1, volume 2

A computer program was developed as a general purpose reliability tool for fault tolerant avionics systems. The computer program requirements, together with several appendices containing computer printouts are presented.

Stiffler, J. J.↗

FTMP (Fault Tolerant Multiprocessor) programmer's manual

The Fault Tolerant Multiprocessor (FTMP) computer system was constructed using the Rockwell/Collins CAPS-6 processor. It is installed in the Avionics Integration Research Laboratory (AIRLAB) of NASA Langley Research Center. It is hosted by AIRLAB's System 10, a VAX 11/750, for the loading of programs and experimentation. The FTMP support software includes a cross compiler for a high level language called Automated Engineering Design (AED) System, an assembler for the CAPS-6 processor assembly language, and a linker. Access to this support software is through an automated remote access facility on the VAX which relieves the user of the burden of learning how to use the IBM 4381. This manual is a compilation of information about the FTMP support environment. It explains the FTMP software and support environment along many of the finer points of running programs on FTMP. This will be helpful to the researcher trying to run an experiment on FTMP and even to the person probing FTMP with fault injections. Much of the information in this manual can be found in other sources; we are only attempting to bring together the basic points in a single source. If the reader should need points clarified, there is a list of support documentation in the back of this manual.

Feather, F. E.↗

Improving system reliability through formal analysis and use of checks in software

Software is playing increasingly important roles in avionics systems. It is widely used in navigation and, in some cases, in control loops that maintain aircraft stability. To guarantee the safety of flight systems, the FAA requires that critical components have a probability of failure no greater than 10(exp -9) per hour of flight. Software is being used to diagnose system components for failure. SIFT (Software Implemented Fault Tolerance) was a computer system developed to study the use of software to check for failure and manage processor reconfiguration. To guarantee that software satisfies its specifications, formal verification can be used. With this a program and its specification are viewed as mathematical objects, and a mathematical proof is used to show that the program and its specification are equivalent. In previous research, a theory of checking was developed to offer assistance in analyzing specifications and designing run-time checks. In the theory, checking is considered abstractly in terms of n-ary relations much like those of relational database theory. Within the theory check are categorized, checks on input and checks on results are considered, and formal attention is given to the minimization and logical combination of checks. The focus is upon input checks and the obstacles in checking input to critical systems. A central concern is with a property referred to as independence. The concern is with circumstances under which it is possible to apply isolated, independent checks to separate sensor inputs and be assure that all illegal input will be properly detected. Presently, independence is being investigated and checked in the context of the GCS (Guidance and Control System). The GCS simulator is intended for testing software that implements control laws for landing spacecraft. The large number of inputs and their complex interrelationships provide an exciting context in which to investigate independence and the difficulties of supplying input checks.

Staknis, Mark E.↗

Investigation of Air Transportation Technology at Princeton University, 1989-1990

The Air Transportation Technology Program at Princeton University proceeded along six avenues during the past year: microburst hazards to aircraft; machine-intelligent, fault tolerant flight control; computer aided heuristics for piloted flight; stochastic robustness for flight control systems; neural networks for flight control; and computer aided control system design. These topics are briefly discussed, and an annotated bibliography of publications that appeared between January 1989 and June 1990 is given.

Stengel, Robert F.↗

A Numerical Simulation and Statistical Modeling of High Intensity Radiated Fields Experiment Data

Tests are conducted on a quad-redundant fault tolerant flight control computer to establish upset characteristics of an avionics system in an electromagnetic field. A numerical simulation and statistical model are described in this work to analyze the open loop experiment data collected in the reverberation chamber at NASA LaRC as a part of an effort to examine the effects of electromagnetic interference on fly-by-wire aircraft control systems. By comparing thousands of simulation and model outputs, the models that best describe the data are first identified and then a systematic statistical analysis is performed on the data. All of these efforts are combined which culminate in an extrapolation of values that are in turn used to support previous efforts used in evaluating the data.

Smith, Laura J.↗

Method and apparatus for fault tolerance

A method and apparatus for achieving fault tolerance in a computer system having at least a first central processing unit and a second central processing unit. The method comprises the steps of first executing a first algorithm in the first central processing unit on input which produces a first output as well as a certification trail. Next, executing a second algorithm in the second central processing unit on the input and on at least a portion of the certification trail which produces a second output. The second algorithm has a faster execution time than the first algorithm for a given input. Then, comparing the first and second outputs such that an error result is produced if the first and second outputs are not the same. The step of executing a first algorithm and the step of executing a second algorithm preferably takes place over essentially the same time period.

Masson, Gerald M.↗

Improved Fermion Hamiltonians for Quantum Simulations

Constructing improved hamiltonians for gauge theories coupled to fermonic matter will be important for improving continuum limit extrapolations of quantum computations. In this talk we will present a formulation for simulating ASQTAD fermions for lattice computation and provide fault tolerant resource costs in terms of primitive group operations. We additionally show that the scaling of energies with respect to the lattice spacing are better than for the unimproved Hamiltonian for toy models.

Erik Joseph Gustafson↗

Improved Fermion Hamiltonians for Quantum Simulations

Constructing improved hamiltonians for gauge theories coupled to fermonic matter will be important for improving continuum limit extrapolations of quantum computations. In this talk we will present a formulation for simulating ASQTAD fermions for lattice computation and provide fault tolerant resource costs in terms of primitive group operations. We additionally show that the scaling of energies with respect to the lattice spacing are better than for the unimproved Hamiltonian for toy models.

Erik Gustafson↗

Improved Fermion Hamiltonians for Quantum Simulations

Constructing improved hamiltonians for gauge theories coupled to fermonic matter will be important for improving continuum limit extrapolations of quantum computations. In this talk we will present a formulation for simulating ASQTAD fermions for lattice computation and provide fault tolerant resource costs in terms of primitive group operations. We additionally show that the scaling of energies with respect to the lattice spacing are better than for the unimproved Hamiltonian for toy models.

Quantum Algorithms↗

Navigation Ground Data System Engineering for the Cassini/Huygens Mission

The launch of the Cassini/Huygens mission on October 15, 1997, began a seven year journey across the solar system that culminated in the entry of the spacecraft into Saturnian orbit on June 30, 2004. Cassini/Huygens Spacecraft Navigation is the result of a complex interplay between several teams within the Cassini Project, performed on the Ground Data System. The work of Spacecraft Navigation involves rigorous requirements for accuracy and completeness carried out often under uncompromising critical time pressures. To support the Navigation function, a fault-tolerant, high-reliability/high-availability computational environment was necessary to support data processing. Configuration Management (CM) was integrated with fault tolerant design and security engineering, according to the cornerstone principles of Confidentiality, Integrity, and Availability. Integrated with this approach are security benchmarks and validation to meet strict confidence levels. In addition, similar approaches to CM were applied in consideration of the staffing and training of the system administration team supporting this effort. As a result, the current configuration of this computational environment incorporates a secure, modular system, that provides for almost no downtime during tour operations.

Beswick, R. M.↗