Search NASA⌕ Search

SEARCH · Search NASA

Results for “Fault-tolerance”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 217 records · Page 12

Operational characteristics of the dispersed sensor processor mesh

The dispersed sensor processing mesh (DSPM) is part of an experimental system used to pursue a proof of concept for an advanced fault-tolerant communication strategy. The DSPM experiments incorporate sensors and effectors into an ultra-reliable nodal network in which a central bus controller performs algorithms to grow and maintain the network. The experiments demonstrate that DSPM is an achievable communication network that can sustain failures and continue to function properly. The results of the experiment give insights into the unique characteristics of the DSPM network and the network's capacity to limit physical damage propagation, limit damaged-data propagation, and to survive cyclic communication paths. This paper describes the concept, mechanization, and performance of the DSPM system in a real-time flight-critical environment.

Abbott, L. W.↗

Fault detection, isolation and reconfiguration in FTMP Methods and experimental results

The Fault-Tolerant Multiprocessor (FTMP) is a highly reliable computer designed to meet a goal of 10 to the -10th failures per hour and built with the objective of flying an active-control transport aircraft. Fault detection, identification, and recovery software is described, and experimental results obtained by injecting faults in the pin level in the FTMP are presented. Over 21,000 faults were injected in the CPU, memory, bus interface circuits, and error detection, masking, and error reporting circuits of one LRU of the multiprocessor. Detection, isolation, and reconfiguration times were recorded for each fault, and the results were found to agree well with earlier assumptions made in reliability modeling.

Lala, J. H.↗

Space station - Technology development

The NASA manned space station program's systems technology effort involves the development of novel techniques that will reduce the scope of tasks neeeded for design, development, testing and evaluation of the hardware. Operations technology efforts encompass analyses that will define those techniques best able to improve the efficiency and reduce the costs of space station functions. The technology objective for data management calls for a fault-tolerant, distributed, expandable and adaptable, as well as repairable and user-friendly, flight data management system that employs state-of-the-art hardware and software. The space station's power system includes the largest element, a 'solar blanket', and the heaviest component, the batteries, of all the subsystems. A thermal management system for the power system is of paramount importance. Attention is also given to the exacting demands of attitude control and stabilization and a regenerative life support system of the requisite capacity and reliability.

Carlisle, R. F.↗

Test experience on an ultrareliable computer communication network

The dispersed sensor processing mesh (DSPM) is an experimental, ultra-reliable, fault-tolerant computer communications network that exhibits an organic-like ability to regenerate itself after suffering damage. The regeneration is accomplished by two routines - grow and repair. This paper discusses the DSPM concept for achieving fault tolerance and provides a brief description of the mechanization of both the experiment and the six-node experimental network. The main topic of this paper is the system performance of the growth algorithm contained in the grow routine. The characteristics imbued to DSPM by the growth algorithm are also discussed. Data from an experimental DSPM network and software simulation of larger DSPM-type networks are used to examine the inherent limitation on growth time by the growth algorithm and the relationship of growth time to network size and topology.

Abbott, L. W.↗

Data and results of a laboratory investigation of microprocessor upset caused by simulated lightning-induced analog transients

Advanced composite aircraft designs include fault-tolerant computer-based digital control systems with thigh reliability requirements for adverse as well as optimum operating environments. Since aircraft penetrate intense electromagnetic fields during thunderstorms, onboard computer systems maya be subjected to field-induced transient voltages and currents resulting in functional error modes which are collectively referred to as digital system upset. A methodology was developed for assessing the upset susceptibility of a computer system onboard an aircraft flying through a lightning environment. Upset error modes in a general-purpose microprocessor were studied via tests which involved the random input of analog transients which model lightning-induced signals onto interface lines of an 8080-based microcomputer from which upset error data were recorded. The application of Markov modeling to upset susceptibility estimation is discussed and a stochastic model development.

Belcastro, C. M.↗

Peer Review of a Formal Verification/Design Proof Methodology

The role of formal verification techniques in system validation was examined. The value and the state of the art of performance proving for fault-tolerant compuers were assessed. The investigation, development, and evaluation of performance proving tools were reviewed. The technical issues related to proof methodologies are examined. The technical issues discussed are summarized.

Source record↗

Some methods of estimating a coverage parameter

To demonstrate the high reliability of fault-tolerant computing systems, experimental testing can be performed by injecting faults into their hardware and observing the times needed to detect and isolate failed components and reconfigure the good components. Coverage, a parameter measuring continued system success, is the probability that recovery occurs before other, perhaps catastrophic, component failures occur. Methods of estimating coverage are given for the case of randomly selecting a subset of the pins or chips for testing. Pin-level samples of times to recovery are treated by the Type I censoring model often used in life testing.

Lee, L. D.↗

Test experience on an ultrareliable computer communication network

The dispersed sensor processing mesh (DSPM) is an experimental, ultrareliable, fault-tolerant computer communications network that exhibits an organic-like ability to regenerate itself after suffering damage. The regeneration is accomplished by two routines - grow and repair. This paper discusses the DSPM concept for achieving fault tolerance and provides a brief description of the mechanization of both the experiment and the six-node experimental network. The main topic of this paper is the system performance of the growth algorithm contained in the grow routine. The characteristics imbued to DSPM by the growth algorithm are also discussed. Data from an experimental DSPM network and software simulation of larger DSPM-type networks are used to examine the inherent limitation on growth time by the growth algorithm and the relationship of growth time to network size and topology.

Abbott, L. W.↗

Autonomous spacecraft design methodology

A methodology for autonomous spacecraft design blends autonomy requirements with traditional mission requirements and assesses the impact of autonomy upon the total system resources available to support fault-tolerance and automation. A baseline functional design can be examined for autonomy implementation impacts, and the costs, risk, and benefits of various options can be assessed. The result of the process is a baseline design that includes autonomous control functions.

Divita, E. L.↗

Validation of fault-free behavior of a reliable multiprocessor system - FTMP: A case study

A program of experiments has been conducted at NASA-Langley to test the fault-free performance of a Fault-Tolerant Multiprocessor (FTMP) avionics system for next-generation aircraft. Baseline measurements of an operating FTMP system were obtained with respect to the following parameters: instruction execution time, frame size, and the variation of clock ticks. The mechanisms of frame stretching were also investigated. The experimental results are summarized in a table. Areas of interest for future tests are identified, with emphasis given to the implementation of a synthetic workload generation mechanism on FTMP.

Clune, E.↗

A user's view of CARE III

The present computerized reliability predictor for digital fault-tolerant systems whose sizes are of the order of one million Markovian equivalent states, employs advanced stochastic modeling techniques and implements a mixed Markov model that enables it to drastically reduce the state size of hitherto computationally unobtainable models. Attention is given to the concepts of failure, fault, and error, in the context of the novel system's fault/error-handling models. Examples are drawn from the system's user-friendly interface dialog.

Bavuso, S. J.↗

Automated ultrareliability models - A review

Analytic models are required to assess the reliability of systems designed to ultrareliability requirements. This paper reviews the capabilities and limitations of five currently available automated reliability models which are applicable to fault-tolerant flight control systems. 'System' includes sensors, computers, and actuators. A set of review criteria including validation, configuration adaptability, and resource requirements for model evaluation are described. Five models, ARIES, CARE II, CARE III, CARSRA, and CAST, are assessed against the criteria, thereby characterizing their capabilities and limitations. This review should be helpful to potential users of the models.

Bridgman, M. S.↗

Digital avionics design and reliability analyzer

The description and specifications for a digital avionics design and reliability analyzer are given. Its basic function is to provide for the simulation and emulation of the various fault-tolerant digital avionic computer designs that are developed. It has been established that hardware emulation at the gate-level will be utilized. The primary benefit of emulation to reliability analysis is the fact that it provides the capability to model a system at a very detailed level. Emulation allows the direct insertion of faults into the system, rather than waiting for actual hardware failures to occur. This allows for controlled and accelerated testing of system reaction to hardware failures. There is a trade study which leads to the decision to specify a two-machine system, including an emulation computer connected to a general-purpose computer. There is also an evaluation of potential computers to serve as the emulation computer.

Source record↗

Primitive Quantum Gates for an $SU(3)$ Discrete Subgroup: $Σ(72\times3)$

We construct a primitive gate set for the digital quantum simulation of a discrete subgroup of $SU(3)$: the 216-element $Σ(72\times3)$. The necessary primitives are the inversion gate, the group multiplication gate, the trace gate, and the group Fourier transform, for which we provide qubit decompositions. The resulting fault-tolerant T gate costs for a fiducial calculation of shear viscosity would require about $10^{12}$ T gates which compares favorably to other modern estimates.

Perez, Sebastian Osorio [Fermilab; Maryland U.]↗

Power and Limitations of Linear Programming Decoder for Quantum LDPC Codes

Decoding quantum error-correcting codes is a key challenge in enabling fault-tolerant quantum computation. In the classical setting, linear programming (LP) decoders offer provable performance guarantees and can leverage fast practical optimization algorithms. Although LP decoders have been proposed for quantum codes, their performance and limitations remain relatively underexplored. In this work, we uncover a key limitation of LP decoding for quantum low-density parity-check (LDPC) codes: certain constant-weight error patterns lead to ambiguous fractional solutions that cannot be resolved through independent rounding. To address this issue, we incorporate a post-processing technique known as ordered statistics decoding (OSD), which significantly enhances LP decoding performance in practice. Our results show that LP decoding, when augmented with OSD, can outperform belief propagation with the same post-processing for intermediate code sizes of up to hundreds of qubits. These findings suggest that LP-based decoders, equipped with effective post-processing, offer a promising approach for decoding near-term quantum LDPC codes.

Gu, Shouzhen [Yale U.]↗

Oxide-nitride heteroepitaxy for low-loss dielectrics in superconducting quantum circuits

Superconducting qubits show great promise for the realization of fault-tolerant quantum computing, but lossy, amorphous dielectrics limit current technology. Identifying highly crystalline and stoichiometric dielectrics with intrinsically low microwave loss is therefore a central materials challenge, yet experimentally validated platforms remain scarce. In this work, we integrate a crystalline dielectric into a heteroepitaxial TiN/$γ$-Al$_2$O$_3$/TiN trilayer grown via pulsed laser deposition. Correlative high-resolution imaging, diffraction, and spectroscopy measurements confirm the single-crystal quality and chemical integrity of all layers, with minimal defects and limited anion interdiffusion across the oxide-nitride interfaces. Using microwave lumped-element resonators with parallel-plate capacitors, we report the first direct measurement of the dielectric loss of epitaxial $γ$-Al$_2$O$_3$, for which we find a low intrinsic two-level system loss, $δ_{\text{TLS}}^0 = (2.8 \pm 0.1) \times 10^{-5}$. These results establish heteroepitaxial oxides on transition metal nitrides as an attractive materials platform for superconducting quantum circuits, particularly for integration into compact device architectures such as merged-element transmons and microwave kinetic inductance detectors.

Garcia-Wetten, David A. [Northwestern U.]↗

Preparing Fermions via Classical Sampling and Linear Combinations of Unitaries

We present an extension of the Evolving density matrices on Qubits (E$ρ$OQ) framework that enables efficient fault-tolerant preparation of fermionic quantum states. The original method circumvents state preparation by stochastic sampling, but faces a sign problem in fermionic systems leading to a large number of circuits necessary. We resolve this by combining classical stochastic sampling with a linear combination of unitaries method that avoids the exponential circuit scaling that plagued naïve implementations. The resulting algorithm requires $\mathcal{O}(M^2)$$R_Z$ rotations for circuit preparation, where $M$ is the number of retained basis states. We validate the method for ground and excited states in the Thirring model, including by computing two-point correlation functions relevant to scattering. In this model for fixed accuracy $\varepsilon$, $M$ is found to scale empirically as $M \propto \frac{1}{mg}\log(1/g)\log(1/m)$.

Gustafson, Erik J. [RIACS, Mtn. View] (ORCID:00000↗

Power and Limitations of Linear Programming Decoder for Quantum LDPC Codes

Decoding quantum error-correcting codes is a key challenge in enabling fault-tolerant quantum computation. In the classical setting, linear programming (LP) decoders offer provable performance guarantees and can leverage fast practical optimization algorithms. Although LP decoders have been proposed for quantum codes, their performance and limitations remain relatively underexplored. In this work, we uncover a key limitation of LP decoding for quantum low-density parity-check (LDPC) codes: certain constant-weight error patterns lead to ambiguous fractional solutions that cannot be resolved through independent rounding. To address this issue, we incorporate a post-processing technique known as ordered statistics decoding (OSD), which significantly enhances LP decoding performance in practice. Our results show that LP decoding, when augmented with OSD, can outperform belief propagation with the same post-processing for intermediate code sizes of up to hundreds of qubits. These findings suggest that LP-based decoders, equipped with effective post-processing, offer a promising approach for decoding near-term quantum LDPC codes.

Gu, Shouzhen [Yale U.]↗