Search NASA⌕ Search

SEARCH · Search NASA

Results for “Fault-tolerance”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 145 records · Page 8

A unified method for analyzing mission reliability for fault tolerant computer systems.

For fault-tolerant computer systems consisting of multiple classes of modules, a unified method for analyzing mission reliability is proposed and evaluated. The analysis proceeds by generalizing the notions of standby and N modular redundancy into a concept called hybrid-degraded redundancy. The probabilistic evaluation of the unified redundancy concept is then developed to yield, for a given modular class, the joint distribution of success and the number of nonfailed modules from that class, at special times. With this information, a Markov chain analysis gives the reliability of an entire sequence of phases (mission profile).

Bricker, J. L.↗

Error-correcting codes for high-speed digital computers

Published document discusses method for correcting errors. According to this method, computer operation becomes fault-tolerant, i.e., its operation is error-free in spite of single hardware element malfunction. Also, method provides for detection and correction of repetitive and spurious processing and transmission errors.

Campbell, R. D.↗

On-line diagnosis of sequential systems, 2

The theory and techniques applicable to the on-line diagnosis of sequential systems, were investigated. A complete model for the study of on-line diagnosis is developed. First an appropriate class of system models is formulated which can serve as a basis for a theoretical study of on-line diagnosis. Then notions of realization, fault, fault-tolerance and diagnosability are formalized which have meaningful interpretations in the the context of on-line diagnosis. The diagnosis of systems which are structurally decomposed and are represented as a network of smaller systems is studied. The fault set considered is the set of faults which only affect one component system is the network. A characterization of those networks which can be diagnosed using a purely combinational detector is achieved. A technique is given which can be used to realize any network by a network which is diagnosable in the above sense. Limits are found on the amount of redundancy involved in any such technique.

Sundstrom, R. J.↗

Management and design of long-life systems; Proceedings of the Symposium, Denver, Colo., April 24-26, 1973

The long life of Pioneer interplanetary spacecraft is considered along with a general accelerated methodology for long-life mechanical components, dependable long-lived household appliances, and the design and development philosophy to achieve reliability and long life in large turbine generators. Other topics discussed include an integrated management approach to long life in space, artificial heart reliability factors, and architectural concepts and redundancy techniques in fault-tolerant computers. Individual items are announced in this issue.

Schurmeier, H. M.↗

A forward view on reliable computers for flight control

The requirements for fault-tolerant computers for flight control of commercial aircraft are examined; it is concluded that the reliability requirements far exceed those typically quoted for space missions. Examination of circuit technology and alternative computer architectures indicates that the desired reliability can be achieved with several different computer structures, though there are obvious advantages to those that are more economic, more reliable, and, very importantly, more certifiable as to fault tolerance. Progress in this field is expected to bring about better computer systems that are more rigorously designed and analyzed even though computational requirements are expected to increase significantly.

Goldberg, J.↗

Flight test results of the Strapdown hexad Inertial Reference Unit (SIRU). Volume 1: Flight test summary

Flight test results of the strapdown inertial reference unit (SIRU) navigation system are presented. The fault-tolerant SIRU navigation system features a redundant inertial sensor unit and dual computers. System software provides for detection and isolation of inertial sensor failures and continued operation in the event of failures. Flight test results include assessments of the system's navigational performance and fault tolerance.

Hruby, R. J.↗

Fault tolerant computing: A preamble for assuring viability of large computer systems

The need for fault-tolerant computing is addressed from the viewpoints of (1) why it is needed, (2) how to apply it in the current state of technology, and (3) what it means in the context of the Phoenix computer system and other related systems. To this end, the value of concurrent error detection and correction is described. User protection, program retry, and repair are among the factors considered. The technology of algebraic codes to protect memory systems and arithmetic codes to protect memory systems and arithmetic codes to protect arithmetic operations is discussed.

Lim, R. S.↗

The unified data system - A distributed processing network for control and data handling on a spacecraft

This paper presents the results obtained in a continuing investigation of real-time distributed processing systems which is being conducted at the Jet Propulsion Laboratory. A distributed processor architecture has been developed for control and data handling on a planetary spacecraft. This system, designated the Unified Data System, has been implemented in a feasibility breadboard. The following aspects of the Unified Data System are described: (1) standardized building block elements and their configuration into microprocessor modules, (2) a highly redundant bus structure which connects the various modules and facilitates intercommunications with minimal software support, (3) the structure of software within the individual modules and its coordination between modules, and (4) the mechanisms by which fault-tolerance can be implemented within the network. Through the attributes of multilevel standardization, simplicity, and flexibility, this system is expected to result in significant cost savings to future spacecraft missions.

Rennels, D. A.↗

A maintenance model for k-out-of-n subsystems aboard a fleet of advanced commercial aircraft

Proposed highly reliable fault-tolerant reconfigurable digital control systems for a future generation of commercial aircraft consist of several k-out-of-n subsystems. Each of these flight-critical subsystems will consist of n identical components, k of which must be functioning properly in order for the aircraft to be dispatched. Failed components are recoverable; they are repaired in a shop. Spares are inventoried at a main base where they may be substituted for failed components on planes during layovers. Penalties are assessed when failure of a k-out-of-n subsystem causes a dispatch cancellation or delay. A maintenance model for a fleet of aircraft with such control systems is presented. The goals are to demonstrate economic feasibility and to optimize.

Miller, D. R.↗

Computer-aided reliability estimation

Computer-aided reliability estimation (CARE) programs are developed to improve the tools available for estimating the reliability of fault-tolerant systems. A description is presented of a program, called CARE II, which was developed after the first program reported by Mathur (1971). Attention is given to the CARE II reliability model, the CARE II coverage model, and CARE II limitations which are to be rectified in CARE III. It is pointed out that the present coverage model in CARE II is extremely versatile. The major limitation is related to the burden placed on the user to determine the basic parameters from which the coverage calculations are made.

Stiffler, J. J.↗

Conference on Advanced Technology for Future Space Systems, Hampton, Va., May 8-10, 1979, Technical Papers

Propulsion systems for spacecraft, satellite communications technology, the design of large light-weight erectable structures for assembly in space, electronics and information processing for spacecraft, and self-diagnostic, fault-tolerant controls based on high memory and processing capabilities are discussed. Topics of the papers include the design of large delta wings for earth-to-orbit transports, dual-fuel propulsion units, magnetoplasmadynamic thrusters, heating rates on blunt-nosed bodies at various angles of attack, remote manipulators for space assembly tasks, solar electric propulsion for planetary missions, deployable space platforms with multiple payloads, the design of large offset-fed antennas, a nonlinear stress-strain relationship for metallic meshes, and adaptive sensors for spacecraft.

Source record↗

Primitive Quantum Gates for an $SU(3)$ Discrete Subgroup: $Σ(72\times3)$

We construct a primitive gate set for the digital quantum simulation of a discrete subgroup of $SU(3)$: the 216-element $Σ(72\times3)$. The necessary primitives are the inversion gate, the group multiplication gate, the trace gate, and the group Fourier transform, for which we provide qubit decompositions. The resulting fault-tolerant T gate costs for a fiducial calculation of shear viscosity would require about $10^{12}$ T gates which compares favorably to other modern estimates.

Perez, Sebastian Osorio [Fermilab; Maryland U.]↗

Power and Limitations of Linear Programming Decoder for Quantum LDPC Codes

Decoding quantum error-correcting codes is a key challenge in enabling fault-tolerant quantum computation. In the classical setting, linear programming (LP) decoders offer provable performance guarantees and can leverage fast practical optimization algorithms. Although LP decoders have been proposed for quantum codes, their performance and limitations remain relatively underexplored. In this work, we uncover a key limitation of LP decoding for quantum low-density parity-check (LDPC) codes: certain constant-weight error patterns lead to ambiguous fractional solutions that cannot be resolved through independent rounding. To address this issue, we incorporate a post-processing technique known as ordered statistics decoding (OSD), which significantly enhances LP decoding performance in practice. Our results show that LP decoding, when augmented with OSD, can outperform belief propagation with the same post-processing for intermediate code sizes of up to hundreds of qubits. These findings suggest that LP-based decoders, equipped with effective post-processing, offer a promising approach for decoding near-term quantum LDPC codes.

Gu, Shouzhen [Yale U.]↗

Oxide-nitride heteroepitaxy for low-loss dielectrics in superconducting quantum circuits

Superconducting qubits show great promise for the realization of fault-tolerant quantum computing, but lossy, amorphous dielectrics limit current technology. Identifying highly crystalline and stoichiometric dielectrics with intrinsically low microwave loss is therefore a central materials challenge, yet experimentally validated platforms remain scarce. In this work, we integrate a crystalline dielectric into a heteroepitaxial TiN/$γ$-Al$_2$O$_3$/TiN trilayer grown via pulsed laser deposition. Correlative high-resolution imaging, diffraction, and spectroscopy measurements confirm the single-crystal quality and chemical integrity of all layers, with minimal defects and limited anion interdiffusion across the oxide-nitride interfaces. Using microwave lumped-element resonators with parallel-plate capacitors, we report the first direct measurement of the dielectric loss of epitaxial $γ$-Al$_2$O$_3$, for which we find a low intrinsic two-level system loss, $δ_{\text{TLS}}^0 = (2.8 \pm 0.1) \times 10^{-5}$. These results establish heteroepitaxial oxides on transition metal nitrides as an attractive materials platform for superconducting quantum circuits, particularly for integration into compact device architectures such as merged-element transmons and microwave kinetic inductance detectors.

Garcia-Wetten, David A. [Northwestern U.]↗

Preparing Fermions via Classical Sampling and Linear Combinations of Unitaries

We present an extension of the Evolving density matrices on Qubits (E$ρ$OQ) framework that enables efficient fault-tolerant preparation of fermionic quantum states. The original method circumvents state preparation by stochastic sampling, but faces a sign problem in fermionic systems leading to a large number of circuits necessary. We resolve this by combining classical stochastic sampling with a linear combination of unitaries method that avoids the exponential circuit scaling that plagued naïve implementations. The resulting algorithm requires $\mathcal{O}(M^2)$$R_Z$ rotations for circuit preparation, where $M$ is the number of retained basis states. We validate the method for ground and excited states in the Thirring model, including by computing two-point correlation functions relevant to scattering. In this model for fixed accuracy $\varepsilon$, $M$ is found to scale empirically as $M \propto \frac{1}{mg}\log(1/g)\log(1/m)$.

Gustafson, Erik J. [RIACS, Mtn. View] (ORCID:00000↗

Power and Limitations of Linear Programming Decoder for Quantum LDPC Codes

Decoding quantum error-correcting codes is a key challenge in enabling fault-tolerant quantum computation. In the classical setting, linear programming (LP) decoders offer provable performance guarantees and can leverage fast practical optimization algorithms. Although LP decoders have been proposed for quantum codes, their performance and limitations remain relatively underexplored. In this work, we uncover a key limitation of LP decoding for quantum low-density parity-check (LDPC) codes: certain constant-weight error patterns lead to ambiguous fractional solutions that cannot be resolved through independent rounding. To address this issue, we incorporate a post-processing technique known as ordered statistics decoding (OSD), which significantly enhances LP decoding performance in practice. Our results show that LP decoding, when augmented with OSD, can outperform belief propagation with the same post-processing for intermediate code sizes of up to hundreds of qubits. These findings suggest that LP-based decoders, equipped with effective post-processing, offer a promising approach for decoding near-term quantum LDPC codes.

Gu, Shouzhen [Yale U.]↗

Measuring quasiparticle dynamics for particle impact reconstruction in a superconducting qubit chip

Quasiparticle poisoning following particle impacts poses a significant challenge to the development of fault-tolerant superconducting quantum computers, as a sudden excess of quasiparticles can simultaneously degrade the coherence of multiple qubits across large device arrays. In this work, we present a statistical analysis that models the time evolution of radiation-induced qubit energy relaxation through quasiparticle density dynamics. This study provides insight into quasiparticle loss processes by distinguishing between recombination and trapping decay channels and assessing their respective impact on qubit performance. We precisely measure quasiparticle recombination in multiple transmon qubits and uncover an unexpected dependence of qubit relaxation dynamics on deposited energy. By linking correlated relaxation events across qubits to ballistic phonon propagation, we introduce a statistical localization approach to extract the energy deposited in the substrate, which is in good agreement with Monte Carlo simulation. This work establishes the quantitative framework for using an arbitrary subset of superconducting transmon qubits in a QPU as energy-resolving witness particle detectors.

Celi, E. [Northwestern U.]↗

Real-Time Detection of Charge Jumps in Superconducting Qubits with a Convolutional Neural Network

Ionizing radiation from cosmic rays and gammas can induce discontinuous jumps in the environmental charge of superconducting qubits (charge jumps), causing correlated errors that challenge fault-tolerant quantum computing while simultaneously providing a detection signature for quantum sensing applications. Current detection methods operate offline, introducing latency incompatible with in-the-loop qubit control. In this paper, an online detector of charge jumps for superconducting qubits, based on a dilated causal convolutional neural network (DCCNN) designed for in-the-loop deployment on the Quantum Instrumentation Control Kit (QICK) platform, is presented. The network is trained on synthetic Ramsey tomography scans generated from qubit templates measured at the Northwestern Experimental Underground Site (NEXUS) at Fermilab, and translated to FPGA firmware via hls4ml with ap_fixed$\langle 16,6 \rangle$ quantization, reaching a per-inference latency of $6.19 μ$s on the Zynq UltraScale+ RFSoC ZCU216. At this operating point the DCCNN matches the detection efficiency of the established offline $χ^2$ algorithm ($0.843 \pm 0.022$ vs. $0.866 \pm 0.020$ on $|Δq| \in [0.1, 0.5] e$ at matched false-positive rate), while requiring no per-qubit hyperparameter tuning. This shifts charge-jump detection from a post-hoc diagnostic to a control-loop primitive, enabling adaptive protocols that respond to radiation-induced events in situ, with applications to quantum-computing error mitigation and to the use of superconducting qubits as particle detectors.

Gaytan-Villarreal, Daniel [Carnegie Mellon U.]↗