Search NASA⌕ Search

SEARCH · Search NASA

Results for “fault tolerance challenges”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 37 records · Page 2

LLRF System Analysis for the Fermilab PIP-II LINAC

Developing long-lived quantum processing units (QPUs) capable of supporting high-fidelity quantum operations is a crucial challenge on the path toward fault-tolerant quantum computing. TESLA-shaped superconducting RF (SRF) cavities, known for photon relaxation times on the order of seconds, provide an excellent foundation for 3D QPUs and quantum memory. This talk presents a novel design that leverages TESLA cavity modes coupled to ancillary transmon qubits, optimized to preserve coherence and control. By carefully engineering the package geometry, optimizing Hamiltonian parameters, and minimizing lossy participation ratios, we achieve photon relaxation times of over 16 ms and 20 ms for the two cavity modes, representing a significant improvement over previous multimode quantum memories. Despite the reduced coupling between the qubit and cavity modes, which is necessary to preserve long lifetimes, the platform supports robust and universal control schemes that are not limited by low coupling strength. We will also discuss how this architecture can lead to scalable, modular quantum computing systems.

Varghese, P. [Fermilab]↗

Architecture for Survivable System Processing (ASSP)

The Architecture for Survivable System Processing (ASSP) Program is a multi-phase effort to implement Department of Defense (DOD) and commercially developed high-tech hardware, software, and architectures for reliable space avionics and ground based systems. System configuration options provide processing capabilities to address Time Dependent Processing (TDP), Object Dependent Processing (ODP), and Mission Dependent Processing (MDP) requirements through Open System Architecture (OSA) alternatives that allow for the enhancement, incorporation, and capitalization of a broad range of development assets. High technology developments in hardware, software, and networking models, address technology challenges of long processor life times, fault tolerance, reliability, throughput, memories, radiation hardening, size, weight, power (SWAP) and security. Hardware and software design, development, and implementation focus on the interconnectivity/interoperability of an open system architecture and is being developed to apply new technology into practical OSA components. To insure for widely acceptable architecture capable of interfacing with various commercial and military components, this program provides for regular interactions with standardization working groups (e.g.) the International Standards Organization (ISO), American National Standards Institute (ANSI), Society of Automotive Engineers (SAE), and Institute of Electrical and Electronic Engineers (IEEE). Selection of a viable open architecture is based on the widely accepted standards that implement the ISO/OSI Reference Model.

Wood, Richard J.↗

Technology Challenges for Deep-Throttle Cryogenic Engines for Space Exploration

Historically, cryogenic rocket engines have not been used for in-space applications due to their additional complexity, the mission need for high reliability, and the challenges of propellant boil-off. While the mission and vehicle architectures are not yet defined for the lunar and Martian robotic and human exploration objectives, cryogenic rocket engines offer the potential for higher performance and greater architecture/mission flexibility. In-situ cryogenic propellant production could enable a more robust exploration program by significantly reducing the propellant mass delivered to low earth orbit, thus warranting the evaluation of cryogenic rocket engines versus the hypergolic bi-propellant engines used in the Apollo program. A multi-use engine. one which can provide the functionality that separate engines provided in the Apollo mission architecture, is desirable for lunar and Mars exploration missions because it increases overall architecture effectiveness through commonality and modularity. The engine requirement derivation process must address each unique mission application and each unique phase within each mission. The resulting requirements, such as thrust level, performance, packaging, bum duration, number of operations; required impulses for each trajectory phase; operation after extended space or surface exposure; availability for inspection and maintenance; throttle range for planetary descent, ascent, acceleration limits and many more must be addressed. Within engine system studies, the system and component technology, capability, and risks must be evaluated and a balance between the appropriate amount of technology-push and technology-pull must be addressed. This paper will summarize many of the key technology challenges associated with using high-performance cryogenic liquid propellant rocket engine systems and components in the exploration program architectures. The paper is divided into two areas. The first area describes how the mission requirements affect the engine system requirements and create system level technology challenges. An engine system architecture for multiple applications or a family of engines based upon a set of core technologies, design, and fabrication approaches may reduce overall programmatic cost and risk. The engine system discussion will also address the characterization of engine cycle figures of merit, configurations, and design approaches for some in-space vehicle alternatives under consideration. The second area evaluates the component-level technology challenges induced from the system requirements. Component technology issues are discussed addressing injector, thrust chamber, ignition system, turbopump assembly, and valve design for the challenging requirements of high reliability, robustness, fault tolerance, deep throttling, reasonable performance (with respect to weight and specific impulse).

Brown, Kendall K.↗

A High-Reliability Photoelectric Detection System for Mars Sample Return’s Orbiting Sample

The Mars Sample Return campaign is an endeavor of unprecedented technological complexity and coordination that attempts to answer fundamental questions about the habitability of Mars by returning the first samples of Martian material to Earth for analysis. The third mission in the campaign consists of the NASA-provided Capture, Containment, and Return System (CCRS) onboard the European Space Agency’s Earth Return Orbiter, which will retrieve the Orbiting Sample (OS) container from its orbit around Mars. Retrieving a passive sample container from a planetary orbit has never been attempted by any spacecraft and requires the development of new technology to succeed in this ambitious task. This paper introduces the high-reliability Capture Sensor Suite (CSS), a novel optical detection system that provides CCRS with the capability to autonomously detect the OS as it is captured. This article will discuss the challenges and requirements for the fault-tolerant design of the CSS.

planetary sampling↗

Measuring quasiparticle dynamics for particle impact reconstruction in a superconducting qubit chip

Quasiparticle poisoning following particle impacts poses a significant challenge to the development of fault-tolerant superconducting quantum computers, as a sudden excess of quasiparticles can simultaneously degrade the coherence of multiple qubits across large device arrays. In this work, we present a statistical analysis that models the time evolution of radiation-induced qubit energy relaxation through quasiparticle density dynamics. This study provides insight into quasiparticle loss processes by distinguishing between recombination and trapping decay channels and assessing their respective impact on qubit performance. We precisely measure quasiparticle recombination in multiple transmon qubits and uncover an unexpected dependence of qubit relaxation dynamics on deposited energy. By linking correlated relaxation events across qubits to ballistic phonon propagation, we introduce a statistical localization approach to extract the energy deposited in the substrate, which is in good agreement with Monte Carlo simulation. This work establishes the quantitative framework for using an arbitrary subset of superconducting transmon qubits in a QPU as energy-resolving witness particle detectors.

Celi, E. [Northwestern U.]↗

A fault-tolerant neutral-atom architecture for universal quantum computation

Quantum error correction (QEC) is essential for the realization of large-scale quantum computers. However, owing to the complexity of operating on the encoded ‘logical’ qubits, understanding the physical principles for building fault-tolerant quantum devices and combining them into efficient architectures is an outstanding scientific challenge. Here we use reconfigurable arrays of up to 448 neutral atoms to implement the key elements of a universal, fault-tolerant quantum processing architecture and experimentally explore their underlying working mechanisms. We first use surface codes to study how repeated QEC suppresses errors, demonstrating 2.14(13)x below-threshold performance in a four-round characterization circuit by leveraging atom loss detection and machine learning decoding. We then investigate logical entanglement using transversal gates and lattice surgery and extend it to universal logic through transversal teleportation with three-dimensional [[15,1,3]] codes, enabling arbitrary-angle synthesis with polylogarithmic overhead. Finally, we develop mid-circuit qubit reuse16, increasing experimental cycle rates by two orders of magnitude and enabling deep-circuit protocols with dozens of logical qubits and hundreds of logical teleportations with [[7,1,3]] and high-rate [[16,6,4]] codes while maintaining constant internal entropy. Our experiments show key principles for efficient architecture design, involving the interplay between quantum logic and entropy removal, judiciously using physical entanglement in logic gates and magic state generation, and leveraging teleportations for universality and physical qubit reset. These results establish foundations for scalable, universal error-corrected processing and its practical implementation in neutral atom systems.

atomic and molecular physics↗

A Multi-Mission Testbed for Advanced Technologies

The mission of the Center for Space Integrated Microsystem (CSIM) at the Jet Propulsion Laboratory is to develop advanced avionics systems for future deep space missions. The Advanced Micro Spacecraft (AMS) task is building a multi-mission testbed facility to enable the infusion of CSIM technologies into future missions. The testbed facility will also perform experimentation for advanced avionics technologies and architectures to meet challenging power, performance, mass, volume, reliability, and fault tolerance of future missions. The testbed facility has two levels of testbeds: (1) a Proof-of-Concept (POC) Testbed and (2) an Engineering Model Testbed. The methodology of the testbed development and the process of technology infusion are presented in a separate paper in this conference. This paper focuses only on the design, implementation, and application of the POC testbed. Additional information is contained in the original extended abstract.

Chau, S. N.↗

Fault-Tolerant Deep Learning Cache with Hash Ring for Load Balancing in HPC Systems

Large-scale DL on HPC systems like Frontier and Summit uses distributed node-local caching to address scalability and performance challenges. However, as these systems grow more complex, the risk of node failures increases, and current caching approaches lack fault tolerance, jeopardizing large-scale training jobs. We analyzed six months of SLURM job logs from Frontier and found that over 30% of jobs failed after an average of 75 minutes. To address this, we propose fault-tolerance strategies that recache data lost from failed nodes using a hash ring technique for balanced data recaching in the distributed node-local caching, reducing reliance on the PFS. Our extensive evaluations on Frontier showed that the hash ring-based recaching approach reduced training time by approximately 25% compared to the approach that redirects I/O to the PFS after node failures and demonstrated effective load balancing of training data across nodes.

Lee, Seoyeong↗

QECC-Synth: A Layout Synthesizer for Quantum Error Correction Codes on Sparse Architectures

Quantum Error Correction (QEC) codes are essential for achieving fault-tolerant quantum computing (FTQC). However, their implementation faces significant challenges due to disparity between required dense qubit connectivity and sparse hardware architectures. Current approaches often either underutilize QEC circuit features or focus on manual designs tailored to specific codes and architectures, limiting their capability and generality. In response, we introduce QECC-Synth, an automated compiler for QEC code implementation that addresses these challenges. We leverage the ancilla bridge technique tailored to the requirements of QEC circuits and introduces a systematic classification of its design space flexibilities. We then formalize this problem using the MaxSAT framework to optimize these flexibilities. Evaluation shows that our method significantly outperforms existing methods while demonstrating broader applicability across diverse QEC codes and hardware architectures.

Yin, Keyi [University of California, San Diego]↗

Mars Science Laboratory Sample Acquisition, Sample Processing and Handling: Subsystem Design and Test Challenges

The Sample Acquisition/Sample Processing and Handling subsystem for the Mars Science Laboratory is a highly-mechanized, Rover-based sampling system that acquires powdered rock and regolith samples from the Martian surface, sorts the samples into fine particles through sieving, and delivers small portions of the powder into two science instruments inside the Rover. SA/SPaH utilizes 17 actuated degrees-of-freedom to perform the functions needed to produce 5 sample pathways in support of the scientific investigation on Mars. Both hardware redundancy and functional redundancy are employed in configuring this sampling system so some functionality is retained even with the loss of a degree-of-freedom. Intentional dynamic environments are created to move sample while vibration isolators attenuate this environment at the sensitive instruments located near the dynamic sources. In addition to the typical flight hardware qualification test program, two additional types of testing are essential for this kind of sampling system: characterization of the intentionally-created dynamic environment and testing of the sample acquisition and processing hardware functions using Mars analog materials in a low pressure environment. The overall subsystem design and configuration are discussed along with some of the challenges, tradeoffs, and lessons learned in the areas of fault tolerance, intentional dynamic environments, and special testing

Jandura, Louise↗

Adaptive Fault Tolerance for Many-Core Based Space-Borne Computing

This paper describes an approach to providing software fault tolerance for future deep-space robotic NASA missions, which will require a high degree of autonomy supported by an enhanced on-board computational capability. Such systems have become possible as a result of the emerging many-core technology, which is expected to offer 1024-core chips by 2015. We discuss the challenges and opportunities of this new technology, focusing on introspection-based adaptive fault tolerance that takes into account the specific requirements of applications, guided by a fault model. Introspection supports runtime monitoring of the program execution with the goal of identifying, locating, and analyzing errors. Fault tolerance assertions for the introspection system can be provided by the user, domain-specific knowledge, or via the results of static or dynamic program analysis. This work is part of an on-going project at the Jet Propulsion Laboratory in Pasadena, California.

fault tolerance↗

Federated Learning for Efficient Condition Monitoring and Anomaly Detection in Industrial Cyber-Physical Systems

Detecting and localizing anomalies in cyber-physical systems (CPS) has become increasingly challenging as systems grow in complexity, particularly due to varying sensor reliability and node failures in distributed environments. While federated learning (FL) offers a foundation for distributed model training, existing approaches lack mechanisms to handle these CPS-specific challenges. This paper presents an enhanced FL framework that introduces three key innovations: adaptive model aggregation based on sensor reliability, dynamic node selection for resource optimization, and Weibull-based checkpointing for fault tolerance. Our framework enables reliable condition monitoring while addressing the computational and reliability challenges of industrial CPS deployments. Experiments on NASA Bearing and Hydraulic System Datasets demonstrate superior performance over state-of-the-art FL methods, achieving 99.5% AUC-ROC in anomaly detection and maintaining accuracy under node failures. Statistical validation using Mann-Whitney (U) test confirms significant improvements (p < 0.05) in both detection accuracy and computational efficiency across diverse operational scenarios.1

Marfo, William [University of Texas at El Paso,Dep↗

First-Principles Assessment of ZnTe and CdSe as Prospective Tunnel Barriers at the InAs/Al Interface

Majorana zero modes are predicted to emerge in semiconductor/ superconductor interfaces, such as InAs/Al. Majorana modes could be utilized for fault tolerant topological qubits. However, their realization is hindered by materials challenges. The coupling between the superconductor and the semiconductor may be too strong for Majorana modes to emerge, due to effective doping of the semiconductor by the metallic contact. This could be mediated by adding a tunnel barrier of controlled thickness. We use density functional theory (DFT) with Hubbard U corrections, whose values are machine-learned via Bayesian optimization (BO), to assess ZnTe and CdSe as prospective tunnel barriers for the InAs/Al interface. The results of DFT +U(BO) for ZnTe are validated by comparison to angle resolved photoemission spectroscopy (ARPES). We then study bilayer interfaces of the three semiconductors with each other and with Al, as well as trilayer interfaces with a varying number of ZnTe or CdSe layers inserted between InAs and Al. We find that 16 atomic layers of either material completely insulate the InAs from metal induced gap states (MIGS). However, ZnTe and CdSe differ significantly in their band alignment, such that ZnTe forms an effective barrier for electrons, whereas CdSe forms a barrier for holes. Because of Fermi level pinning in the conduction band at the interface, only electron transport is relevant for InAs-based Majorana devices. Therefore, ZnTe is the better choice. Based on the results of our simulations, we suggest conducting experiments with ZnTe barriers in the thickness range of 6–18 atomic layers.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Scalable Circuit Cutting and Scheduling in a Resource-constrained and Distributed Quantum System

Despite quantum computing's rapid development, current systems remain limited in practical applications due to their limited qubit count and quality. Various technologies, such as superconducting, trapped ions, and neutral atom quantum computing technologies are progressing towards a fault tolerant era, however they all face a diverse set of challenges in scalability and control. Recent efforts have focused on multi-node quantum systems that connect multiple smaller quantum devices to execute larger circuits. Future demonstrations hope to use quantum channels to couple systems, however current demonstrations can leverage classical communication with circuit cutting techniques. This involves cutting large circuits into smaller subcircuits and reconstructing them post-execution. However, existing cutting methods are hindered by lengthy search times as the number of qubits and gates increases. Additionally, they often fail to effectively utilize the resources of various worker configurations in a multi-node system. To address these challenges, we introduce FitCut, a novel approach that transforms quantum circuits into weighted graphs and utilizes a community-based, bottom-up approach to cut circuits according to resource constraints, e.g., qubit counts, on each worker. FitCut also includes a scheduling algorithm that optimizes resource utilization across workers. Implemented with Qiskit and evaluated extensively, FitCut significantly outperforms the Qiskit Circuit Knitting Toolbox, reducing time costs by factors ranging from 3 to 2000 and improving resource utilization rates by up to 3.88 times on the worker side, achieving a system-wide improvement of 2.86 times.

Kan, Shuwen [Fordham University]↗

Integrated Application of Active Controls (IAAC) technology to an advanced subsonic transport project: Test act system validation

The primary objective of the Test Active Control Technology (ACT) System laboratory tests was to verify and validate the system concept, hardware, and software. The initial lab tests were open loop hardware tests of the Test ACT System as designed and built. During the course of the testing, minor problems were uncovered and corrected. Major software tests were run. The initial software testing was also open loop. These tests examined pitch control laws, wing load alleviation, signal selection/fault detection (SSFD), and output management. The Test ACT System was modified to interface with the direct drive valve (DDV) modules. The initial testing identified problem areas with DDV nonlinearities, valve friction induced limit cycling, DDV control loop instability, and channel command mismatch. The other DDV issue investigated was the ability to detect and isolate failures. Some simple schemes for failure detection were tested but were not completely satisfactory. The Test ACT System architecture continues to appear promising for ACT/FBW applications in systems that must be immune to worst case generic digital faults, and be able to tolerate two sequential nongeneric faults with no reduction in performance. The challenge in such an implementation would be to keep the analog element sufficiently simple to achieve the necessary reliability.

Source record↗

Enabling Reliable, Fault-Tolerant Autonomous Lunar Habitats with High-Performance Spaceflight Computing

The lunar surface presents unfavorable constraints and harsh living conditions. To address these challenges, autonomous habitats will require complex integrated systems that combine advanced software, high-performance hardware, and cutting-edge sensors to ensure sustainability, safety, and operational efficiency. Consequently, maintaining a sustainable presence on the Moon requires reliable infrastructure and efficient development, precise monitoring, and utilization of resources within a lunar installation. These elements are essential not only to ensure that lunar settlement can be long-term, self-sustaining, and resource-efficient, but also to serve as a foundation for future missions and eventual human habitation on Mars. Humans are not native to the Moon; therefore, our survival and ability to thrive will depend on autonomous systems that can foster safety and resilience through high-availability architectures, graceful degradation, and highly fault-tolerant spaceflight hardware capable of continuing operation during failures. This requires advanced human-rated distributed systems architectures with specialized electronics, scalable capabilities, and an integrated design approach. Unlike current practices focused on short-term missions and regularly maintained components, permanent lunar compute systems must be designed for extended operations beyond mission durations. This paper explores the necessity of transitioning toward fault- tolerant, highly autonomous hardware systems designed for multi-year missions. It also identifies critical subsystems that require high levels of autonomy, supported by radiation-hardened processors and extreme thermal loads, which are essential to mitigate long-term degradation and ensure sustainable lunar habitation. Finally, the paper aligns with NASA’s identified Civil Space Shortfalls, particularly in high-performance onboard computing, advanced data acquisition, extreme-environment avionics, radiation monitoring and countermeasures, and autonomous health management. It proposes NASA’s new High-Performance Spaceflight Computing (HPSC) processor as a turnkey solution, delivering 100 times the performance-per-watt of legacy rad-hard CPUs and enabling onboard AI, edge computing, and fault-tolerant features essential for sustained lunar autonomy and beyond.

Sarkis S Mikaelian↗

Fault-Tolerant, Radiation-Hard DSP

Commercial digital signal processors (DSPs) for use in high-speed satellite computers are challenged by the damaging effects of space radiation, mainly single event upsets (SEUs) and single event functional interrupts (SEFIs). Innovations have been developed for mitigating the effects of SEUs and SEFIs, enabling the use of very-highspeed commercial DSPs with improved SEU tolerances. Time-triple modular redundancy (TTMR) is a method of applying traditional triple modular redundancy on a single processor, exploiting the VLIW (very long instruction word) class of parallel processors. TTMR improves SEU rates substantially. SEFIs are solved by a SEFI-hardened core circuit, external to the microprocessor. It monitors the health of the processor, and if a SEFI occurs, forces the processor to return to performance through a series of escalating events. TTMR and hardened-core solutions were developed for both DSPs and reconfigurable field-programmable gate arrays (FPGAs). This includes advancement of TTMR algorithms for DSPs and reconfigurable FPGAs, plus a rad-hard, hardened-core integrated circuit that services both the DSP and FPGA. Additionally, a combined DSP and FPGA board architecture was fully developed into a rad-hard engineering product. This technology enables use of commercial off-the-shelf (COTS) DSPs in computers for satellite and other space applications, allowing rapid deployment at a much lower cost. Traditional rad-hard space computers are very expensive and typically have long lead times. These computers are either based on traditional rad-hard processors, which have extremely low computational performance, or triple modular redundant (TMR) FPGA arrays, which suffer from power and complexity issues. Even more frustrating is that the TMR arrays of FPGAs require a fixed, external rad-hard voting element, thereby causing them to lose much of their reconfiguration capability and in some cases significant speed reduction. The benefits of COTS high-performance signal processing include significant increase in onboard science data processing, enabling orders of magnitude reduction in required communication bandwidth for science data return, orders of magnitude improvement in onboard mission planning and critical decision making, and the ability to rapidly respond to changing mission environments, thus enabling opportunistic science and orders of magnitude reduction in the cost of mission operations through reduction of required staff. Additional benefits of COTS-based, high-performance signal processing include the ability to leverage considerable commercial and academic investments in advanced computing tools, techniques, and infra structure, and the familiarity of the science and IT community with these computing environments.

Czajkowski, David↗

Phoenix Mars Scout UHF Relay-Only Operations

The Phoenix Mars Scout Lander will launch in August 2007 and land on the northern plains of Mars in May of 2008. In a departure from traditional planetary surface mission operations, it will have no direct-to-Earth communications capability and will rely entirely on Mars-orbiting relays in order to facilitate command and control as well as the return of science and engineering data. The Mars Exploration Rover missions have demonstrated the robust data-return capability using this architecture, and also have demonstrated the capability of using this method for command and control. The Phoenix mission will take the next step and incorporate this as the sole communications link. Operations for 90 Sols will need to work within the constraints of Odyssey and Mars Reconnaissance Orbiter communications availability, anomalies must be diagnosed and responded to through an intermediary and on-board fault responses must be tolerant to loss of a relay. These and other issues pose interesting challenges and changes in paradigm for traditional space operations and spacecraft architecture, and the approach proposed for the Phoenix mission is detailed herein.

fault protection↗