Search NASA⌕ Search

SEARCH · Search NASA

Results for “Fault-tolerance”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 235 records · Page 13

Test experience on an ultrareliable computer communication network

The dispersed sensor processing mesh (DSPM) is an experimental, ultrareliable, fault-tolerant computer communications network that exhibits an organic-like ability to regenerate itself after suffering damage. The regeneration is accomplished by two routines - grow and repair. This paper discusses the DSPM concept for achieving fault tolerance and provides a brief description of the mechanization of both the experiment and the six-node experimental network. The main topic of this paper is the system performance of the growth algorithm contained in the grow routine. The characteristics imbued to DSPM by the growth algorithm are also discussed. Data from an experimental DSPM network and software simulation of larger DSPM-type networks are used to examine the inherent limitation on growth time by the growth algorithm and the relationship of growth time to network size and topology.

Abbott, L. W.↗

Autonomous spacecraft design methodology

A methodology for autonomous spacecraft design blends autonomy requirements with traditional mission requirements and assesses the impact of autonomy upon the total system resources available to support fault-tolerance and automation. A baseline functional design can be examined for autonomy implementation impacts, and the costs, risk, and benefits of various options can be assessed. The result of the process is a baseline design that includes autonomous control functions.

Divita, E. L.↗

Validation of fault-free behavior of a reliable multiprocessor system - FTMP: A case study

A program of experiments has been conducted at NASA-Langley to test the fault-free performance of a Fault-Tolerant Multiprocessor (FTMP) avionics system for next-generation aircraft. Baseline measurements of an operating FTMP system were obtained with respect to the following parameters: instruction execution time, frame size, and the variation of clock ticks. The mechanisms of frame stretching were also investigated. The experimental results are summarized in a table. Areas of interest for future tests are identified, with emphasis given to the implementation of a synthetic workload generation mechanism on FTMP.

Clune, E.↗

A user's view of CARE III

The present computerized reliability predictor for digital fault-tolerant systems whose sizes are of the order of one million Markovian equivalent states, employs advanced stochastic modeling techniques and implements a mixed Markov model that enables it to drastically reduce the state size of hitherto computationally unobtainable models. Attention is given to the concepts of failure, fault, and error, in the context of the novel system's fault/error-handling models. Examples are drawn from the system's user-friendly interface dialog.

Bavuso, S. J.↗

Automated ultrareliability models - A review

Analytic models are required to assess the reliability of systems designed to ultrareliability requirements. This paper reviews the capabilities and limitations of five currently available automated reliability models which are applicable to fault-tolerant flight control systems. 'System' includes sensors, computers, and actuators. A set of review criteria including validation, configuration adaptability, and resource requirements for model evaluation are described. Five models, ARIES, CARE II, CARE III, CARSRA, and CAST, are assessed against the criteria, thereby characterizing their capabilities and limitations. This review should be helpful to potential users of the models.

Bridgman, M. S.↗

Fault-free behavior of reliable multiprocessor systems: FTMP experiments in AIRLAB

This report describes a set of experiments which were implemented on the Fault tolerant Multi-Processor (FTMP) at NASA/Langley's AIRLAB facility. These experiments are part of an effort to formulate and evaluate validation methodologies for fault-tolerant computers. This report deals with the measurement of single parameters (baselines) of a fault free system. The initial set of baseline experiments lead to the following conclusions: (1) The system clock is constant and independent of workload in the tested cases; (2) the instruction execution times are constant; (3) the R4 frame size is 40mS with some variation; (4) the frame stretching mechanism has some flaws in its implementation that allow the possibility of an infinite stretching of frame duration. Future experiments are planned. Some will broaden the results of these initial experiments. Others will measure the system more dynamically. The implementation of a synthetic workload generation mechanism for FTMP is planned to enhance the experimental environment of the system.

Clune, E.↗

The embedded operating system project

The design and construction of embedded operating systems for real-time advanced aerospace applications was investigated. The applications require reliable operating system support that must accommodate computer networks. Problems that arise in the construction of such operating systems, reconfiguration, consistency and recovery in a distributed system, and the issues of real-time processing are reported. A thesis that provides theoretical foundations for the use of atomic actions to support fault tolerance and data consistency in real-time object-based system is included. The following items are addressed: (1) atomic actions and fault-tolerance issues; (2) operating system structure; (3) program development; (4) a reliable compiler for path Pascal; and (5) mediators, a mechanism for scheduling distributed system processes.

Campbell, R. H.↗

Integrated analysis of error detection and recovery

An integrated modeling and analysis of error detection and recovery is presented. When fault latency and/or error latency exist, the system may suffer from multiple faults or error propagations which seriously deteriorate the fault-tolerant capability. Several detection models that enable analysis of the effect of detection mechanisms on the subsequent error handling operations and the overall system reliability were developed. Following detection of the faulty unit and reconfiguration of the system, the contaminated processes or tasks have to be recovered. The strategies of error recovery employed depend on the detection mechanisms and the available redundancy. Several recovery methods including the rollback recovery are considered. The recovery overhead is evaluated as an index of the capabilities of the detection and reconfiguration mechanisms.

Shin, K. G.↗

Achieving reliability - The evolution of redundancy in American manned spacecraft computers

The Shuttle is the first launch system deployed by NASA with full redundancy in the on-board computer systems. Fault-tolerance, i.e., restoring to a backup with less capabilities, was the method selected for Apollo. The Gemini capsule was the first to carry a computer, which also served as backup for Titan launch vehicle guidance. Failure of the Gemini computer resulted in manual control of the spacecraft. The Apollo system served vehicle flight control and navigation functions. The redundant computer on Skylab provided attitude control only in support of solar telescope pointing. The STS digital, fly-by-wire avionics system requires 100 percent reliability. The Orbiter carries five general purpose computers, four being fully-redundant and the fifth being soley an ascent-descent tool. The computers are synchronized at input and output points at a rate of about six times a second. The system is projected to cause a loss of an Orbiter only four times in a billion flights.

Tomayko, J. E.↗

A theoretical basis for the analysis of multiversion software subject to coincident errors

Fundamental to the development of redundant software techniques (known as fault-tolerant software) is an understanding of the impact of multiple joint occurrences of errors, referred to here as coincident errors. A theoretical basis for the study of redundant software is developed which: (1) provides a probabilistic framework for empirically evaluating the effectiveness of a general multiversion strategy when component versions are subject to coincident errors, and (2) permits an analytical study of the effects of these errors. An intensity function, called the intensity of coincident errors, has a central role in this analysis. This function describes the propensity of programmers to introduce design faults in such a way that software components fail together when executing in the application environment. A condition under which a multiversion system is a better strategy than relying on a single version is given.

Eckhardt, D. E., Jr.↗

Digital avionics design and reliability analyzer

The description and specifications for a digital avionics design and reliability analyzer are given. Its basic function is to provide for the simulation and emulation of the various fault-tolerant digital avionic computer designs that are developed. It has been established that hardware emulation at the gate-level will be utilized. The primary benefit of emulation to reliability analysis is the fact that it provides the capability to model a system at a very detailed level. Emulation allows the direct insertion of faults into the system, rather than waiting for actual hardware failures to occur. This allows for controlled and accelerated testing of system reaction to hardware failures. There is a trade study which leads to the decision to specify a two-machine system, including an emulation computer connected to a general-purpose computer. There is also an evaluation of potential computers to serve as the emulation computer.

Source record↗

Primitive Quantum Gates for an $SU(3)$ Discrete Subgroup: $Σ(72\times3)$

We construct a primitive gate set for the digital quantum simulation of a discrete subgroup of $SU(3)$: the 216-element $Σ(72\times3)$. The necessary primitives are the inversion gate, the group multiplication gate, the trace gate, and the group Fourier transform, for which we provide qubit decompositions. The resulting fault-tolerant T gate costs for a fiducial calculation of shear viscosity would require about $10^{12}$ T gates which compares favorably to other modern estimates.

Perez, Sebastian Osorio [Fermilab; Maryland U.]↗

Power and Limitations of Linear Programming Decoder for Quantum LDPC Codes

Decoding quantum error-correcting codes is a key challenge in enabling fault-tolerant quantum computation. In the classical setting, linear programming (LP) decoders offer provable performance guarantees and can leverage fast practical optimization algorithms. Although LP decoders have been proposed for quantum codes, their performance and limitations remain relatively underexplored. In this work, we uncover a key limitation of LP decoding for quantum low-density parity-check (LDPC) codes: certain constant-weight error patterns lead to ambiguous fractional solutions that cannot be resolved through independent rounding. To address this issue, we incorporate a post-processing technique known as ordered statistics decoding (OSD), which significantly enhances LP decoding performance in practice. Our results show that LP decoding, when augmented with OSD, can outperform belief propagation with the same post-processing for intermediate code sizes of up to hundreds of qubits. These findings suggest that LP-based decoders, equipped with effective post-processing, offer a promising approach for decoding near-term quantum LDPC codes.

Gu, Shouzhen [Yale U.]↗

Oxide-nitride heteroepitaxy for low-loss dielectrics in superconducting quantum circuits

Superconducting qubits show great promise for the realization of fault-tolerant quantum computing, but lossy, amorphous dielectrics limit current technology. Identifying highly crystalline and stoichiometric dielectrics with intrinsically low microwave loss is therefore a central materials challenge, yet experimentally validated platforms remain scarce. In this work, we integrate a crystalline dielectric into a heteroepitaxial TiN/$γ$-Al$_2$O$_3$/TiN trilayer grown via pulsed laser deposition. Correlative high-resolution imaging, diffraction, and spectroscopy measurements confirm the single-crystal quality and chemical integrity of all layers, with minimal defects and limited anion interdiffusion across the oxide-nitride interfaces. Using microwave lumped-element resonators with parallel-plate capacitors, we report the first direct measurement of the dielectric loss of epitaxial $γ$-Al$_2$O$_3$, for which we find a low intrinsic two-level system loss, $δ_{\text{TLS}}^0 = (2.8 \pm 0.1) \times 10^{-5}$. These results establish heteroepitaxial oxides on transition metal nitrides as an attractive materials platform for superconducting quantum circuits, particularly for integration into compact device architectures such as merged-element transmons and microwave kinetic inductance detectors.

Garcia-Wetten, David A. [Northwestern U.]↗

Preparing Fermions via Classical Sampling and Linear Combinations of Unitaries

We present an extension of the Evolving density matrices on Qubits (E$ρ$OQ) framework that enables efficient fault-tolerant preparation of fermionic quantum states. The original method circumvents state preparation by stochastic sampling, but faces a sign problem in fermionic systems leading to a large number of circuits necessary. We resolve this by combining classical stochastic sampling with a linear combination of unitaries method that avoids the exponential circuit scaling that plagued naïve implementations. The resulting algorithm requires $\mathcal{O}(M^2)$$R_Z$ rotations for circuit preparation, where $M$ is the number of retained basis states. We validate the method for ground and excited states in the Thirring model, including by computing two-point correlation functions relevant to scattering. In this model for fixed accuracy $\varepsilon$, $M$ is found to scale empirically as $M \propto \frac{1}{mg}\log(1/g)\log(1/m)$.

Gustafson, Erik J. [RIACS, Mtn. View] (ORCID:00000↗

Power and Limitations of Linear Programming Decoder for Quantum LDPC Codes

Decoding quantum error-correcting codes is a key challenge in enabling fault-tolerant quantum computation. In the classical setting, linear programming (LP) decoders offer provable performance guarantees and can leverage fast practical optimization algorithms. Although LP decoders have been proposed for quantum codes, their performance and limitations remain relatively underexplored. In this work, we uncover a key limitation of LP decoding for quantum low-density parity-check (LDPC) codes: certain constant-weight error patterns lead to ambiguous fractional solutions that cannot be resolved through independent rounding. To address this issue, we incorporate a post-processing technique known as ordered statistics decoding (OSD), which significantly enhances LP decoding performance in practice. Our results show that LP decoding, when augmented with OSD, can outperform belief propagation with the same post-processing for intermediate code sizes of up to hundreds of qubits. These findings suggest that LP-based decoders, equipped with effective post-processing, offer a promising approach for decoding near-term quantum LDPC codes.

Gu, Shouzhen [Yale U.]↗

Measuring quasiparticle dynamics for particle impact reconstruction in a superconducting qubit chip

Quasiparticle poisoning following particle impacts poses a significant challenge to the development of fault-tolerant superconducting quantum computers, as a sudden excess of quasiparticles can simultaneously degrade the coherence of multiple qubits across large device arrays. In this work, we present a statistical analysis that models the time evolution of radiation-induced qubit energy relaxation through quasiparticle density dynamics. This study provides insight into quasiparticle loss processes by distinguishing between recombination and trapping decay channels and assessing their respective impact on qubit performance. We precisely measure quasiparticle recombination in multiple transmon qubits and uncover an unexpected dependence of qubit relaxation dynamics on deposited energy. By linking correlated relaxation events across qubits to ballistic phonon propagation, we introduce a statistical localization approach to extract the energy deposited in the substrate, which is in good agreement with Monte Carlo simulation. This work establishes the quantitative framework for using an arbitrary subset of superconducting transmon qubits in a QPU as energy-resolving witness particle detectors.

Celi, E. [Northwestern U.]↗

Real-Time Detection of Charge Jumps in Superconducting Qubits with a Convolutional Neural Network

Ionizing radiation from cosmic rays and gammas can induce discontinuous jumps in the environmental charge of superconducting qubits (charge jumps), causing correlated errors that challenge fault-tolerant quantum computing while simultaneously providing a detection signature for quantum sensing applications. Current detection methods operate offline, introducing latency incompatible with in-the-loop qubit control. In this paper, an online detector of charge jumps for superconducting qubits, based on a dilated causal convolutional neural network (DCCNN) designed for in-the-loop deployment on the Quantum Instrumentation Control Kit (QICK) platform, is presented. The network is trained on synthetic Ramsey tomography scans generated from qubit templates measured at the Northwestern Experimental Underground Site (NEXUS) at Fermilab, and translated to FPGA firmware via hls4ml with ap_fixed$\langle 16,6 \rangle$ quantization, reaching a per-inference latency of $6.19 μ$s on the Zynq UltraScale+ RFSoC ZCU216. At this operating point the DCCNN matches the detection efficiency of the established offline $χ^2$ algorithm ($0.843 \pm 0.022$ vs. $0.866 \pm 0.020$ on $|Δq| \in [0.1, 0.5] e$ at matched false-positive rate), while requiring no per-qubit hyperparameter tuning. This shifts charge-jump detection from a post-hoc diagnostic to a control-loop primitive, enabling adaptive protocols that respond to radiation-induced events in situ, with applications to quantum-computing error mitigation and to the use of superconducting qubits as particle detectors.

Gaytan-Villarreal, Daniel [Carnegie Mellon U.]↗