Search NASA⌕ Search

SEARCH · Search NASA

Results for “fault tolerant computing”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 307 records · Page 17

Real-time computer simulation/emulation for verification of multi-fault-tolerant control of Centaur-in-Shuttle

NASA has contracted with General Dynamics to design and develop an advanced Centaur liquid upper stage for support of the Galileo and Solar Polar interplanetary missions in 1985-86. The control of the Centaur while it resides in the Shuttle cargo bay must meet the STS safety requirements to be dual failure tolerant in all mission critical functions. The demonstration of the integrity of this control system in the event of multiple component failures and worst-case time-phase asynchroniety among the system's computers is performed by a real-time computer simulation. The simulation emulates the control hardware, subsystem interfaces, and imbedded software processes, wire-by-wire, to provide accessibility for fault insertion. Observability is provided via graphics and diagnostic software. Verification is the product of Monte Carlo simulation analysis.

Szatkowski, G. P.↗

Software-implemented Fault Tolerance for Supercomputing in Space

The NASA Jet Propulsion Laboratory Remote Exploration and Experimentaion (REE) Project is a large multi-year technology demonstration project which will develop low-power, scalable, fault-tolerant, high- performance computing for use in space and will demonstrate that significant onboard processing capability enables a new class of science missions.

software fault tolerance supercomputing space remo↗

RAMP - A fault tolerant distributed microcomputer structure for aircraft navigation and control

Design methodologies for realizing future high authority autoflight control systems are being investigated, taking into account also the study of distributed microcomputer architectures. Attention is given to the redundant asynchronous microprocessor (RAMP) structure. RAMP comprises a connected network of microcomputers which has as input command and sensor information, and which generates servo information to drive actuators, and thrust linkages. Tolerance to hardware failures is achieved by static redundancy. Results of a failed microcomputer are simply rejected. This is done in lieu of dynamic redundancy wherein the distributed computer system performs real time fault detection and reconfiguration of the system. Attention is given to the RAMP network structure and operation, flight control with parallel asynchronous computers, and intermittent fault tolerance.

Dunn, W. R.↗

Achieving reliability - The evolution of redundancy in American manned spacecraft computers

The Shuttle is the first launch system deployed by NASA with full redundancy in the on-board computer systems. Fault-tolerance, i.e., restoring to a backup with less capabilities, was the method selected for Apollo. The Gemini capsule was the first to carry a computer, which also served as backup for Titan launch vehicle guidance. Failure of the Gemini computer resulted in manual control of the spacecraft. The Apollo system served vehicle flight control and navigation functions. The redundant computer on Skylab provided attitude control only in support of solar telescope pointing. The STS digital, fly-by-wire avionics system requires 100 percent reliability. The Orbiter carries five general purpose computers, four being fully-redundant and the fifth being soley an ascent-descent tool. The computers are synchronized at input and output points at a rate of about six times a second. The system is projected to cause a loss of an Orbiter only four times in a billion flights.

Tomayko, J. E.↗

Enabling Reliable, Fault-Tolerant Autonomous Lunar Habitats with High-Performance Spaceflight Computing

The lunar surface presents unfavorable constraints and harsh living conditions. To address these challenges, autonomous habitats will require complex integrated systems that combine advanced software, high-performance hardware, and cutting-edge sensors to ensure sustainability, safety, and operational efficiency. Consequently, maintaining a sustainable presence on the Moon requires reliable infrastructure and efficient development, precise monitoring, and utilization of resources within a lunar installation. These elements are essential not only to ensure that lunar settlement can be long-term, self-sustaining, and resource-efficient, but also to serve as a foundation for future missions and eventual human habitation on Mars. Humans are not native to the Moon; therefore, our survival and ability to thrive will depend on autonomous systems that can foster safety and resilience through high-availability architectures, graceful degradation, and highly fault-tolerant spaceflight hardware capable of continuing operation during failures. This requires advanced human-rated distributed systems architectures with specialized electronics, scalable capabilities, and an integrated design approach. Unlike current practices focused on short-term missions and regularly maintained components, permanent lunar compute systems must be designed for extended operations beyond mission durations. This paper explores the necessity of transitioning toward fault- tolerant, highly autonomous hardware systems designed for multi-year missions. It also identifies critical subsystems that require high levels of autonomy, supported by radiation-hardened processors and extreme thermal loads, which are essential to mitigate long-term degradation and ensure sustainable lunar habitation. Finally, the paper aligns with NASA’s identified Civil Space Shortfalls, particularly in high-performance onboard computing, advanced data acquisition, extreme-environment avionics, radiation monitoring and countermeasures, and autonomous health management. It proposes NASA’s new High-Performance Spaceflight Computing (HPSC) processor as a turnkey solution, delivering 100 times the performance-per-watt of legacy rad-hard CPUs and enabling onboard AI, edge computing, and fault-tolerant features essential for sustained lunar autonomy and beyond.

Sarkis S Mikaelian↗

Computer Reliability

Using a NASA developed program, Dr. J. Walter Bond is creating a course in computer reliability modeling. The course will examine three different computer programs, one of them NASA's Care III, the others UCLA's Aries 78 and Aries 82. All three are designed to help estimate the reliability of complex, redundant, fault tolerant system. In computer design, software of this kind can predict or model the effects of various hardware or software failures, a process called reliability modeling.

Source record↗

Distributed computing for autonomous on board planning and sequence validations

We propose a new conceptual approach to system-level autonomy that exploits in a synergistic way recent breakthroughs in three specific areas: automatic generation of embeddable planning and validation software, integration of telecommunications forecaster and planning tools, and fault-tolerant assignment of computing tasks to multiple processors.

autonomy synergy software telecommunications compu↗

Design and verification of a multiple fault tolerant control system for STS applications using computer simulation

General Dynamics/Convair is under NASA contract to integrate the Centaur upper stage into the space transportation system for future planetary missions. This requires that control of all safety critical functions be two-failure tolerant. The control system developed consists of five asynchronous computers, each contributing at their outputs to a 3-out-of-5 voting plane. Subsystem control is based on an end function redundancy management scheme. Analysis of multiple component failures and worst-case time-phase asynchrony among the computers is performed by a real-time computer simulation. The simulation emulates the hardware and subsystem interfaces, wire by wire, providing assessibility to any component for the insertion of preprogrammed failures. Observability is provided via a graphics system and diagnostic software. The simulation provides an engineering tool where the integrity of control system hardware and imbedded software can be demonstrated.

Szatkowski, G. P.↗

Digital avionics design and reliability analyzer

The description and specifications for a digital avionics design and reliability analyzer are given. Its basic function is to provide for the simulation and emulation of the various fault-tolerant digital avionic computer designs that are developed. It has been established that hardware emulation at the gate-level will be utilized. The primary benefit of emulation to reliability analysis is the fact that it provides the capability to model a system at a very detailed level. Emulation allows the direct insertion of faults into the system, rather than waiting for actual hardware failures to occur. This allows for controlled and accelerated testing of system reaction to hardware failures. There is a trade study which leads to the decision to specify a two-machine system, including an emulation computer connected to a general-purpose computer. There is also an evaluation of potential computers to serve as the emulation computer.

Source record↗

The Application of Microtechnology to Spacecraft On-Board Computing(abstract)

In this report, we will survey recent advances in chip packaging and stacking techniques that allow miniature computers to be developed for space applications. Several orders of magnitude reduction in mass, volume, and power consumption are possible using these techniques. Moreover, performance improvements can be achieved by increasing the scale of multiprocessing. Most importantly, long-term survivability can potentially be improved by increasing the level of redundancy and fault tolerance.

microtechnology computers multiprocessing packagin↗

Braiding of Majorana zero modes in vortex cores

Here, we demonstrate the successful simulation of $\sqrt{𝑍-}$, $\sqrt{𝑋-}$, and 𝑋-quantum gates using Majorana zero modes (MZMs) that emerge in magnetic vortices located in topological superconductors. We compute the transition probabilities and geometric phase differences accounting for the full many-body dynamics and show that qubit states can be read out by fusing the vortex core MZMs and measuring the resulting charge density. We visualize the gate processes using the time- and energy-dependent nonequilibrium local density of states. Our results demonstrate the feasibility of employing vortex core MZMs for the realization of fault-tolerant topological quantum computing.

Majorana bound states↗

Realizing string-net condensation: Fibonacci anyon braiding for universal gates and sampling chromatic polynomials

Abstract The remarkable complexity of a topologically ordered many-body quantum system is encoded in the characteristics of its anyons. Quintessential predictions emanating from this complexity employ the Fibonacci string net condensate (Fib SNC) and its anyons: sampling Fib-SNC would estimate chromatic polynomials while exchanging its anyons would implement universal quantum computation. However, physical realizations remained elusive. We introduce a scalable dynamical string net preparation (DSNP) that constructs Fib SNC and its anyons on reconfigurable graphs suitable for near-term superconducting processors. Coupling the DSNP approach with composite error-mitigation on deep circuits, we create, measure, and braids Fibonacci anyons; charge measurements show 94% accuracy, and exchanging the anyons yields the expected golden ratioϕwith 98% average accuracy. We then sample the Fib SNC to estimate chromatic polynomial atϕ + 2 for several graphs. Our results establish the proof of principle for using Fib-SNC and its anyons for fault-tolerant universal quantum computation and aim at a classically hard problem.

Science & Technology - Other Topics↗

Care 3, phase 1, volume 2

A computer program was developed as a general purpose reliability tool for fault tolerant avionics systems. The computer program requirements, together with several appendices containing computer printouts are presented.

Stiffler, J. J.↗

FTMP (Fault Tolerant Multiprocessor) programmer's manual

The Fault Tolerant Multiprocessor (FTMP) computer system was constructed using the Rockwell/Collins CAPS-6 processor. It is installed in the Avionics Integration Research Laboratory (AIRLAB) of NASA Langley Research Center. It is hosted by AIRLAB's System 10, a VAX 11/750, for the loading of programs and experimentation. The FTMP support software includes a cross compiler for a high level language called Automated Engineering Design (AED) System, an assembler for the CAPS-6 processor assembly language, and a linker. Access to this support software is through an automated remote access facility on the VAX which relieves the user of the burden of learning how to use the IBM 4381. This manual is a compilation of information about the FTMP support environment. It explains the FTMP software and support environment along many of the finer points of running programs on FTMP. This will be helpful to the researcher trying to run an experiment on FTMP and even to the person probing FTMP with fault injections. Much of the information in this manual can be found in other sources; we are only attempting to bring together the basic points in a single source. If the reader should need points clarified, there is a list of support documentation in the back of this manual.

Feather, F. E.↗

Robustness of Vacancy-Bound Non-Abelian Anyons in the Kitaev Model in a Magnetic Field

Non-Abelian anyons in quantum spin liquids (QSLs) provide a promising route to fault-tolerant topological quantum computation. In the exactly solvable Kitaev honeycomb model, such anyons of the QSL state can be bound to nonmagnetic spin vacancies and endowed with non-Abelian statistics by an infinitesimal magnetic field. Here, we investigate how this approach for stabilizing non-Abelian anyons extends to a finite magnetic field represented by a proper Zeeman term. Through large-scale density-matrix renormalization group simulations, we compute the vacancy-anyon binding energy as a function of magnetic field for both the ferromagnetic and antiferromagnetic Kitaev models. Here, we find that anyon binding remains robust within the entire QSL phase for the ferromagnetic Kitaev model but breaks down already inside this phase for the antiferromagnetic Kitaev model. To compute a binding energy several orders of magnitude below the magnetic energy scale, we introduce both a refined definition and an extrapolation scheme based on carefully tailored perturbations.

Xiao, Bo [Oak Ridge National Laboratory (ORNL), Oa↗

Improving system reliability through formal analysis and use of checks in software

Software is playing increasingly important roles in avionics systems. It is widely used in navigation and, in some cases, in control loops that maintain aircraft stability. To guarantee the safety of flight systems, the FAA requires that critical components have a probability of failure no greater than 10(exp -9) per hour of flight. Software is being used to diagnose system components for failure. SIFT (Software Implemented Fault Tolerance) was a computer system developed to study the use of software to check for failure and manage processor reconfiguration. To guarantee that software satisfies its specifications, formal verification can be used. With this a program and its specification are viewed as mathematical objects, and a mathematical proof is used to show that the program and its specification are equivalent. In previous research, a theory of checking was developed to offer assistance in analyzing specifications and designing run-time checks. In the theory, checking is considered abstractly in terms of n-ary relations much like those of relational database theory. Within the theory check are categorized, checks on input and checks on results are considered, and formal attention is given to the minimization and logical combination of checks. The focus is upon input checks and the obstacles in checking input to critical systems. A central concern is with a property referred to as independence. The concern is with circumstances under which it is possible to apply isolated, independent checks to separate sensor inputs and be assure that all illegal input will be properly detected. Presently, independence is being investigated and checked in the context of the GCS (Guidance and Control System). The GCS simulator is intended for testing software that implements control laws for landing spacecraft. The large number of inputs and their complex interrelationships provide an exciting context in which to investigate independence and the difficulties of supplying input checks.

Staknis, Mark E.↗