Search NASA⌕ Search

SEARCH · Search NASA

Results for “fault tolerant computing”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 505 records · Page 28

The Dangers of Failure Masking in Fault-Tolerant Software: Aspects of a Recent In-Flight Upset Event

On 1 August 2005, a Boeing Company 777-200 aircraft, operating on an international passenger flight from Australia to Malaysia, was involved in a significant upset event while flying on autopilot. The Australian Transport Safety Bureau's investigation into the event discovered that an anomaly existed in the component software hierarchy that allowed inputs from a known faulty accelerometer to be processed by the air data inertial reference unit (ADIRU) and used by the primary flight computer, autopilot and other aircraft systems. This anomaly had existed in original ADIRU software, and had not been detected in the testing and certification process for the unit. This paper describes the software aspects of the incident in detail, and suggests possible implications concerning complex, safety-critical, fault-tolerant software.

Johnson, C. W.↗

Quantum-classical embedding via ghost Gutzwiller approximation for enhanced simulations of correlated electron systems

Simulating correlated materials on present-day quantum hardware remains challenging due to limited quantum resources. Quantum embedding methods offer a promising route by reducing computational complexity through the mapping of bulk systems onto effective impurity models, allowing more feasible simulations on pre- and early-fault-tolerant quantum devices. Here, this work develops a quantum-classical embedding framework based on the ghost Gutzwiller approximation to enable quantum-enhanced simulations of ground-state properties and spectral functions of correlated electron systems. Circuit complexity is analyzed using an adaptive variational quantum algorithm on a statevector simulator, applied to the infinite-dimensional Hubbard model with increasing ghost mode numbers from 3 to 5, resulting in circuit depths growing from 16 to 104. Noise effects are examined using a realistic error model, revealing significant impact on the spectral weight of the Hubbard bands. To mitigate these effects, the Iceberg quantum error detection code is employed, achieving up to 40% error reduction in simulations. Finally, the accuracy of the density matrix estimation and the derived spectral function is benchmarked on IBM and Quantinuum quantum hardware, featuring distinct qubit-connectivity and employing multiple levels of error mitigation techniques.

Chen, I-Chi [Ames Laboratory (AMES), Ames, IA (Uni↗

Fault-Tolerant Decentralized Control for Large-Scale Inverter-Based Resources for Active Power Tracking

Integration of inverter-based resources (IBRs) which lack the intrinsic characteristics such as the inertial response of the traditional synchronous-generator (SG)-based sources presents a new challenge in the form of analyzing the grid stability under their presence. While the dynamic composition of IBRs differs from that of the SGs, the control objective remains similar in terms of tracking the desired active power. This letter presents a decentralized primal-dual-based fault-tolerant control framework for the power allocation in IBRs. Overall, a hierarchical control algorithm is developed with a lower level addressing the current control and the parameter estimation for the IBRs and the higher level acting as the reference power generator to the low level based on the desired active power profile. The decentralized network-based algorithm adaptively splits the desired power between the IBRs taking into consideration the health of the IBRs transmission lines. The proposed framework is tested through a simulation on the network of IBRs and the high-level controller performance is compared against the existing framework in the literature. The proposed algorithm shows significant performance improvement in the magnitude of power deviation and settling time to the nominal value under faulty conditions as compared to the algorithm in the literature.

24 POWER TRANSMISSION AND DISTRIBUTION↗

Enhancing quantum utility: Simulating large-scale quantum spin chains on superconducting quantum computers

We present the quantum simulation of the frustrated quantum spin- 1 2 antiferromagnetic Heisenberg spin chain with competing nearest-neighbor ( J 1 ) and next-nearest-neighbor ( J 2 ) exchange interactions in the real superconducting quantum computer with qubits ranging up to 100. In particular, we implement the Hamiltonian with the next-nearest neighbor exchange interaction in conjunction with the nearest-neighbor interaction on IBM's superconducting quantum computer and carry out the time evolution of the spin chain by employing the first-order Trotterization. Furthermore, our implementation of the second-order Trotterization for the isotropic Heisenberg spin chain, involving only nearest-neighbor exchange interaction, enables precise measurement of the expectation values of staggered magnetization observable across a range of up to 100 qubits. Notably, in both cases, our approach results in a constant circuit depth in each Trotter step, independent of the number of qubits. Our demonstration of the accurate measurement of expectation values for the large-scale quantum system using superconducting quantum computers designates the quantum utility of these devices for investigating various properties of many-body quantum systems. This will be a stepping stone to achieving the quantum advantage over classical ones in simulating quantum systems before the fault tolerance quantum era. Published by the American Physical Society 2024

97 MATHEMATICS AND COMPUTING↗

Realistic Cost to Execute Practical Quantum Circuits using Direct Clifford+T Lattice Surgery Compilation

We report a resource estimation pipeline that explicitly compiles quantum circuits expressed using the Clifford+T gate set into a surface code lattice surgery instruction set. The cadence of magic state requests from the compiled circuit enables the optimization of magic state distillation and storage requirements in a post-hoc analysis. To compile logical circuits into lattice surgery operations, we build upon the open-source Lattice Surgery Compiler. The revised compiler operates in two stages: the first translates logical gates into an abstract, layout-independent instruction set; the second compiles these into local lattice surgery instructions that are allocated to hardware tiles according to a specified resource layout. The second stage retains logical parallelism while avoiding resource contention in the fault-tolerant layer, aiding realism. Additionally, users can specify dedicated tiles at which magic states are replenished, enabling resource costs from the logical computation to be considered independently from magic state distillation and storage. We demonstrate the applicability of our pipeline to large practical quantum circuits by providing resource estimates for the ground state estimation of molecules. Finally, we find that variable magic state consumption rates in real circuits can cause the resource costs of magic state storage to dominate unless production is varied to suit.

97 MATHEMATICS AND COMPUTING↗

HiRel: Hybrid Automated Reliability Predictor (HARP) integrated reliability tool system, (version 7.0). Volume 1: HARP introduction and user's guide

The Hybrid Automated Reliability Predictor (HARP) integrated Reliability (HiRel) tool system for reliability/availability prediction offers a toolbox of integrated reliability/availability programs that can be used to customize the user's application in a workstation or nonworkstation environment. HiRel consists of interactive graphical input/output programs and four reliability/availability modeling engines that provide analytical and simulative solutions to a wide host of reliable fault-tolerant system architectures and is also applicable to electronic systems in general. The tool system was designed to be compatible with most computing platforms and operating systems, and some programs have been beta tested, within the aerospace community for over 8 years. Volume 1 provides an introduction to the HARP program. Comprehensive information on HARP mathematical models can be found in the references.

Bavuso, Salvatore J.↗

RAID Unbound: Storage Fault Tolerance in a Distributed Environment

Mirroring, data replication, backup, and more recently, redundant arrays of independent disks (RAID) are all technologies used to protect and ensure access to critical company data. A new set of problems has arisen as data becomes more and more geographically distributed. Each of the technologies listed above provides important benefits; but each has failed to adapt fully to the realities of distributed computing. The key to data high availability and protection is to take the technologies' strengths and 'virtualize' them across a distributed network. RAID and mirroring offer high data availability, which data replication and backup provide strong data protection. If we take these concepts at a very granular level (defining user, record, block, file, or directory types) and them liberate them from the physical subsystems with which they have traditionally been associated, we have the opportunity to create a highly scalable network wide storage fault tolerance. The network becomes the virtual storage space in which the traditional concepts of data high availability and protection are implemented without their corresponding physical constraints.

Ritchie, Brian↗

Advanced Launch System Multi-Path Redundant Avionics Architecture Analysis and Characterization

The objective of the Multi-Path Redundant Avionics Suite (MPRAS) program is the development of a set of avionic architectural modules which will be applicable to the family of launch vehicles required to support the Advanced Launch System (ALS). To enable ALS cost/performance requirements to be met, the MPRAS must support autonomy, maintenance, and testability capabilities which exceed those present in conventional launch vehicles. The multi-path redundant or fault tolerance characteristics of the MPRAS are necessary to offset a reduction in avionics reliability due to the increased complexity needed to support these new cost reduction and performance capabilities and to meet avionics reliability requirements which will provide cost-effective reductions in overall ALS recurring costs. A complex, real-time distributed computing system is needed to meet the ALS avionics system requirements. General Dynamics, Boeing Aerospace, and C.S. Draper Laboratory have proposed system architectures as candidates for the ALS MPRAS. The purpose of this document is to report the results of independent performance and reliability characterization and assessment analyses of each proposed candidate architecture and qualitative assessments of testability, maintainability, and fault tolerance mechanisms. These independent analyses were conducted as part of the MPRAS Part 2 program and were carried under NASA Langley Research Contract NAS1-17964, Task Assignment 28.

Baker, Robert L.↗

HOME - An application of fault-tolerant techniques and system self-testing

Hard Over Monitoring Equipment (HOME) has been designed to complement and enhance the flight safety of a flight research helicopter. HOME is an independent, highly reliable, and fail-safe special purpose computer that monitors the flight control commands issued by the flight control computer of the helicopter. In particular, HOME detects the issuance of a hazardous hard-over command for any of the four flight control axes and transfers the control of the helicopter to the flight safety pilot. The design of HOME incorporates certain reliability and fail-safe enhancement design features, such as triple modular redundancy, majority logic voting, fail-safe dual circuits, independent status monitors, in-flight self-test, and a built-in preflight exerciser. The HOME design and operation is described with special emphasis on the reliability and fail-safe aspects of the design.

Holden, D. G.↗

The inclusion of semi-Markov reconfiguration transitions into the computer-aided Markov evaluator (CAME) program

The modifications to the rule-based CAME program which allow it to more accurately model the fault-handling processes of fault-tolerant systems are described. This new capability is added to the CAME program by modeling the fault-handling processes of fault-tolerant systems. The integrated airframe/propulsion control system architecture (IAPSA II) reference configuration currently under development is detailed.

Rosch, Gene↗

Exploiting Kubernetes to Simplify the Deployment and Management of the Multi-purpose CMS Pilot Job Factory

GlideinWMS, a widely utilized workload management system in high-energy physics (HEP) research, serves as the backbone for efficient job provisioning across distributed computing resources. It is utilized by various experiments and organizations, including CMS, OSG, Dune, and FIFE, to create HTCondor pools as large as 600k cores. In particular, a shared factory service historically deployed at UCSD has been configured to interface with more than 500 routes to compute clusters. As part of our team’s initiative to modernize infrastructure and enhance scalability, we undertook the migration of the GlideinWMS factory service into the Kubernetes environment. Leveraging the flexibility and orchestration capabilities of Kubernetes, we successfully deployed the factory service within the OSG Tiger Kubernetes cluster. The major benefits Kubernetes gives us is it streamlines the management and monitoring of the factory infrastructure, and improves fault tolerance through its resilient deployment strategies. Through this case study, we aim to share insights, challenges, and best practices encountered during the migration process. Our experience underscores the benefits of embracing containerization and Kubernetes orchestration for HEP computing infrastructure, paving the way for scalability and resilience in distributed computing environments.

Dost, Jeffrey Michael [UC, San Diego (main)]↗

Harnessing Quantum Computing for Energy Materials: Opportunities and Challenges

Developing high-performance materials is critical for diverse energy applications to increase efficiency, improve sustainability and reduce costs. Classical computational methods have enabled important breakthroughs in energy materials development, but they face scaling and time-complexity limitations, particularly for high-dimensional or strongly correlated material systems. Quantum computing (QC) promises to offer a paradigm shift by exploiting quantum bits with their superposition and entanglement to address challenging problems intractable for classical approaches. This Perspective discusses the opportunities in leveraging QC to advance energy materials research and the challenges QC faces in solving complex and high-dimensional problems. We present cases on how QC, when combined with classical computing methods, can be used for the design and simulation of practical energy materials. We also outline the outlook for error-corrected, fault-tolerant QC capable of achieving predictive accuracy and quantum advantage for complex material systems.

Algorithms↗

A distributed programming environment for Ada

Despite considerable commercial exploitation of fault tolerance systems, significant and difficult research problems remain in such areas as fault detection and correction. A research project is described which constructs a distributed computing test bed for loosely coupled computers. The project is constructing a tool kit to support research into distributed control algorithms, including a distributed Ada compiler, distributed debugger, test harnesses, and environment monitors. The Ada compiler is being written in Ada and will implement distributed computing at the subsystem level. The design goal is to provide a variety of control mechanics for distributed programming while retaining total transparency at the code level.

Brennan, Peter↗

Laboratory test methodology for evaluating the effects of electromagnetic disturbances on fault-tolerant control systems

Control systems for advanced aircraft, especially those with relaxed static stability, will be critical to flight and will, therefore, have very high reliability specifications which must be met for adverse as well as nominal operating conditions. Adverse conditions can result from electromagnetic disturbances caused by lightning, high energy radio frequency transmitters, and nuclear electromagnetic pulses. Tools and techniques must be developed to verify the integrity of the control system in adverse operating conditions. The most difficult and illusive perturbations to computer based control systems caused by an electromagnetic environment (EME) are functional error modes that involve no component damage. These error modes are collectively known as upset, can occur simultaneously in all of the channels of a redundant control system, and are software dependent. A methodology is presented for performing upset tests on a multichannel control system and considerations are discussed for the design of upset tests to be conducted in the lab on fault tolerant control systems operating in a closed loop with a simulated plant.

Belcastro, Celeste M.↗

Modeling and measurement of error propagation in a multimodule computing system

An error propagation model has been developed for multimodule computing systems in which the main parameters are the distribution functions of error propagation times. A digraph model is used to represent a multimodule computing system, and error propagation in the system is modeled by general distributions of error propagation times between all pairs of modules. Two algorithms are developed to compute systematically and efficiently the distributions of error propagation times. Experiments are also conducted to measure the distributions of error propagation times with the fault-tolerant microprocessor (FTMP). Statistical analysis of experimental data shows that the error propagation times in FTMP do not follow a well-known distribution, thus justifying the use of general distributions in the present model.

Shin, Kang G.↗

Coverage modeling for dependability analysis of fault-tolerant systems

Several different models for predicting coverage in a fault-tolerant system, including models for permanent, intermittent, and transient errors, are discussed. Markov, semi-Markov, nonhomogeneous Markov, and extended stochastic Petri net models for computing coverage are developed. Two types of events that interfere with recovery are examined; and methods for modeling such events, whether they are deterministic or random, are given. The sensitivity of system reliability/availability to the coverage parameter and the sensitivity of the coverage parameter to various error-handling strategies are investigated. It is found that a policy of attempting transient recovery upon detection of an error can actually increase the unreliability of the system. This result is true if the error detectability is not nearly perfect, so that the risk of producing an undetectable error is greater than the benefit gained by not discarding the component.

Dugan, Joanne Bechta↗

Periodic Application of Concurrent Error Detection in Processor Array Architectures

Processor arrays can provide an attractive architecture for some applications. Featuring modularity, regular interconnection and high parallelism, such arrays are well-suited for VLSI/WSI implementations, and applications with high computational requirements, such as real-time signal processing. Preserving the integrity of results can be of paramount importance for certain applications. In these cases, fault tolerance should be used to ensure reliable delivery of a system's service. One aspect of fault tolerance is the detection of errors caused by faults. Concurrent error detection (CED) techniques offer the advantage that transient and intermittent faults may be detected with greater probability than with off-line diagnostic tests. Applying time-redundant CED techniques can reduce hardware redundancy costs. However, most time-redundant CED techniques degrade a system's performance.

Chen, Paul Peichuan↗

Dexterous end effector flight demonstration

The Dexterous End Effector Flight Experiment is a flight demonstration of newly developed equipment and methods which make for more dexterous manipulation of robotic arms. The following concepts are to be demonstrated: The Force Torque Sensor is a six axis load cell located at the end of the RMS which displays load data to the operator on the orbiter CCTV monitor. TRAC is a target system which provides six axis positional information to the operator. It has the characteristic of having high sensitivity to attitude misalignment while being flat. AUTO-TRAC is a variation of TRAC in which a computer analyzes a target, displays translational and attitude misalignment information, and provides cues to the operator for corrective inputs. The Magnetic End Effector is a fault tolerant end effector which grapples payloads using magnetic attraction. The Carrier Latch Assembly is a fault tolerant payload carrier, which uses mechanical latches and/or magnetic attraction to hold small payloads during launch/landing and to release payloads as desired. The flight experiment goals and objectives are explained. The experiment equipment is described, and the tasks to be performed during the demonstration are discussed.

Carter, Edward L.↗