Search NASA⌕ Search

SEARCH · Search NASA

Results for “Hardware fault mitigations”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 37 records · Page 2

Quantum utility-scale error mitigation for quantum quench dynamics in Heisenberg spin chains

Here, we implement a quantum error mitigation method termed self-mitigation, which is comparable to zero-noise extrapolation, at large scales to achieve quantum utility on near-term, noisy quantum computers. We investigate the effectiveness of several quantum error mitigation strategies, including self-mitigation, by simulating quantum quench dynamics for Heisenberg spin chains with system sizes up to 104 qubits using IBM quantum processors. In particular, we discuss the limitations of zero-noise extrapolation and the advantages offered by self-mitigation at large scales. The self-mitigation method demonstrates stable accuracy with large systems of 104 qubits comprising more than 3,000 CNOT gates. Also, we combine the discussed quantum error mitigation methods with practical entanglement entropy measuring methods, and it shows a good agreement with the theoretical estimation. Our study illustrates the usefulness of near-term noisy quantum hardware in examining the quantum quench dynamics of many-body systems at large scales and lays the groundwork for surpassing classical simulations with quantum methods prior to the development of fault-tolerant quantum computers.

97 MATHEMATICS AND COMPUTING↗

Driving Curiosity: Mars Rover Mobility Trends During the First Seven Years

NASA’s Mars Science Laboratory (MSL) mission landed the Curiosity rover on Mars on August 6, 2012. As of August 6, 2019 (sol 2488), Curiosity has driven 21,318.5 meters over a variety of terrain types and slopes, employing multiple drive modes with varying amounts of onboard autonomy. Curiosity’s drive distances each sol have ranged from its shortest drive of 2.6 centimeters to its longest drive of 142.5 meters, with an average drive distance of 28.9 meters. Real-time human intervention during Curiosity drives on Mars is not possible due to the latency in uplinking commands and downlinking telemetry, so the operations team relies on the rover’s flight software to prevent an unsafe state during driving. Over the first seven years of the mission, Curiosity has attempted 738 drives. While 622 drives have completed successfully, 116 drives were prevented or stopped early by the rover’s fault protection software. The primary risks to mobility success have been wheel wear, wheel entrapment, progressive wheel sinkage (which can lead to rover embedding), and terrain interactions or hardware or cabling failures that result in an inability to command one or more steer or drive actuators. In this paper, we describe mobility trends over the first 21.3km of the mission, operational aspects of the mobility fault protection, and risk mitigation strategies that will support continued mobility success for the remainder of the mission.

Rankin, Arturo↗

Reconfigurable Processing Module

To accommodate a wide spectrum of applications and technologies, NASA s Exploration System's Missions Directorate has called for reconfigurable and modular technologies to support future missions to the moon and Mars. In response, Langley Research Center is leading a program entitled Reconfigurable Scaleable Computing (RSC) that is centered on the development of FPGA-based computing resources in a stackable form factor. This paper details the architecture and implementation of the Reconfigurable Processing Module (RPM), which is the key element of the RSC system. The RPM is an FPGA-based, space-qualified printed circuit assembly leveraging terrestrial/commercial design standards into the space applications domain. The form factor is similar to, and backwards compatible with, the PCI-104 standard utilizing only the PCI interface. The size is expanded to accommodate the required functionality while still better than 30% smaller than a 3U CompactPCI(TradeMark)card and without the overhead of the backplane. The architecture is built around two FPGA devices, one hosting PCI and memory interfaces, and another hosting mission application resources; both of which are connected with a high-speed data bus. The PCI interface FPGA provides access via the PCI bus to onboard SDRAM, flash PROM, and the application resources; both configuration management as well as runtime interaction. The reconfigurable FPGA, referred to as the Application FPGA - or simply "the application" - is a radiation-tolerant Xilinx Virtex-4 FX60 hosting custom application specific logic or soft microprocessor IP. The RPM implements various SEE mitigation techniques including TMR, EDAC, and configuration scrubbing of the reconfigurable FPGA. Prototype hardware and formal modeling techniques are used to explore the performability trade space. These models provide a novel way to calculate quality-of-service performance measures while simultaneously considering fault-related behavior due to SEE soft errors.

Somervill, Kevin↗

Methodology for Designing Fault-Protection Software

A document describes a methodology for designing fault-protection (FP) software for autonomous spacecraft. The methodology embodies and extends established engineering practices in the technical discipline of Fault Detection, Diagnosis, Mitigation, and Recovery; and has been successfully implemented in the Deep Impact Spacecraft, a NASA Discovery mission. Based on established concepts of Fault Monitors and Responses, this FP methodology extends the notion of Opinion, Symptom, Alarm (aka Fault), and Response with numerous new notions, sub-notions, software constructs, and logic and timing gates. For example, Monitor generates a RawOpinion, which graduates into Opinion, categorized into no-opinion, acceptable, or unacceptable opinion. RaiseSymptom, ForceSymptom, and ClearSymptom govern the establishment and then mapping to an Alarm (aka Fault). Local Response is distinguished from FP System Response. A 1-to-n and n-to- 1 mapping is established among Monitors, Symptoms, and Responses. Responses are categorized by device versus by function. Responses operate in tiers, where the early tiers attempt to resolve the Fault in a localized step-by-step fashion, relegating more system-level response to later tier(s). Recovery actions are gated by epoch recovery timing, enabling strategy, urgency, MaxRetry gate, hardware availability, hazardous versus ordinary fault, and many other priority gates. This methodology is systematic, logical, and uses multiple linked tables, parameter files, and recovery command sequences. The credibility of the FP design is proven via a fault-tree analysis "top-down" approach, and a functional fault-mode-effects-and-analysis via "bottoms-up" approach. Via this process, the mitigation and recovery strategy(s) per Fault Containment Region scope (width versus depth) the FP architecture.

Barltrop, Kevin↗

A Cryogenic Muon Tagging System Integrated with a Superconducting Qubit Device for Radiation-Induced Error Mitigation

Superconducting qubits are highly sensitive to ionizing radiation, which can induce correlated errors and limit scalable fault-tolerant quantum computing. In particular, cosmic-ray muons can deposit energy in the substrate, generating phonon bursts that break Cooper pairs and produce quasiparticles, leading to correlated decoherence events across multiple qubits. We present the development of a cryogenic muon tagging system based on Kinetic Inductance Detectors (KIDs) and its integration with superconducting quantum hardware. Originally developed within the ACE-SuperQ project and validated as a standalone detector, the system demonstrated a muon tagging efficiency of approximately 90% and excellent agreement with Monte Carlo simulations. Building on this validation, the tagging system has been integrated with a multi-qubit superconducting chip operated in a dilution refrigerator. The detector configuration consists of a multi-layer KID stack arranged above and below the quantum device, enabling time-coincident identification of muon-induced events within the same cryogenic environment. The integrated setup has been successfully commissioned, enabling simultaneous operation of the qubit chip and the muon tagging system. A first measurement campaign has been carried out, and preliminary data show time-correlated events between the muon tagging detectors and the qubit readout. A quantitative analysis of radiation-induced effects on qubit performance is currently ongoing. This work represents a step toward the implementation of event-level radiation tagging as a tool for characterizing and potentially mitigating correlated errors in superconducting quantum processors, while establishing a modular platform for future studies at the interface between particle physics and quantum information science.

Roy, Tanay [Fermilab] (ORCID:000000019442862X)↗

Digital Twin Framework for PIP-II Linac: AI-Driven Multi-Scale Modeling from Ion Source to 800 MeV

The PIP-II linac will enable >1.2 MW beam power for DUNE, requiring unprecedented operational reliability across its warm front-end (RFQ, MEBT) and five distinct SRF sections operating at 162.5/325/650 MHz. We present a comprehensive digital twin framework uniquely combining a fully differentiable fast beam transport code with neural network surrogates trained on high-fidelity PIC simulations, capturing space charge and nonlinear dynamics beyond traditional envelope codes while achieving 10⁴ speedup at <1% accuracy. End-to-end differentiability enables gradient-based optimization across 500+ parameters simultaneously previously impossible with conventional tools while the model incorporates static/dynamic errors and serves as a virtual commissioning platform for diverse hardware integration. The framework facilitates reinforcement learning for pulsed/CW mode transitions, predictive maintenance through anomaly detection, and autonomous tuning algorithm development with real-time execution capability. Validation against physics simulations shows excellent agreement for the front-end, with initial results demonstrating potential for 30% commissioning time reduction and proactive fault mitigation, providing a scalable blueprint for operating next-generation high-intensity accelerators.

Pathak, Abhishek [Fermilab] (ORCID:000000021704208↗

Digital Twin Framework for PIP-II Linac: AI-Driven Multi-Scale Modeling from Ion Source to 800 MeV

The PIP-II linac will enable >1.2 MW beam power for DUNE, requiring unprecedented operational reliability across its warm front-end (RFQ, MEBT) and five distinct SRF sections operating at 162.5/325/650 MHz. We present a comprehensive digital twin framework uniquely combining a fully differentiable fast beam transport code with neural network surrogates trained on high-fidelity PIC simulations, capturing space charge and nonlinear dynamics beyond traditional envelope codes while achieving 10⁴× speedup at <1% accuracy. End-to-end differentiability enables gradient-based optimization across 500+ parameters simultaneously—previously impossible with conventional tools—while the model incorporates static/dynamic errors and serves as a virtual commissioning platform for diverse hardware integration. The framework facilitates reinforcement learning for pulsed/CW mode transitions, predictive maintenance through anomaly detection, and autonomous tuning algorithm development with real-time execution capability. Validation against physics simulations shows excellent agreement for the front-end, with initial results demonstrating potential for 30% commissioning time reduction and proactive fault mitigation, providing a scalable blueprint for operating next-generation high-intensity accelerators.

Pathak, Abhishek [Fermilab] (ORCID:000000021704208↗

Model Transformation for a System of Systems Dependability Safety Case

Software plays an increasingly larger role in all aspects of NASA's science missions. This has been extended to the identification, management and control of faults which affect safety-critical functions and by default, the overall success of the mission. Traditionally, the analysis of fault identification, management and control are hardware based. Due to the increasing complexity of system, there has been a corresponding increase in the complexity in fault management software. The NASA Independent Validation & Verification (IV&V) program is creating processes and procedures to identify, and incorporate safety-critical software requirements along with corresponding software faults so that potential hazards may be mitigated. This Specific to Generic ... A Case for Reuse paper describes the phases of a dependability and safety study which identifies a new, process to create a foundation for reusable assets. These assets support the identification and management of specific software faults and, their transformation from specific to generic software faults. This approach also has applications to other systems outside of the NASA environment. This paper addresses how a mission specific dependability and safety case is being transformed to a generic dependability and safety case which can be reused for any type of space mission with an emphasis on software fault conditions.

Murphy, Judy↗

Study the Protection Improvements for a Weak Grid Area With High IBRs

NLR is collaborating with Florida Power & Light (FPL) and GE to investigate power system stability and protection reliability challenges in a weak-grid region with high penetration of inverter-based resources (IBRs). This presentation will primarily focus on the protection aspects of the study. We will share key insights from this real-world project, including best practices for developing high-fidelity fault study models, establishing a controller-hardware-in-the-loop (CHIL) platform for testing physical relays, identifying system-level protection challenges, and designing enhanced protection schemes to address those issues. Through this discussion, the audience will gain practical understanding of protection studies in IBR-dominated systems, the emerging challenges associated with reduced fault current and altered transient behavior, and effective mitigation strategies. In particular, we will highlight the critical importance of IBR compliance with IEEE 2800-2022 to ensure dependable and secure protection relay operation in modern transmission systems.

24 POWER TRANSMISSION AND DISTRIBUTION↗

Voyager Interstellar Mission: Challenges of Flying a Very Old Spacecraft on a Very Long Mission

Two Voyager spacecraft were launched in 1977. After the successful flybys of Jupiter and Saturn by both Voyagers and Uranus and Neptune by Voyager 2, the mission has been extended for another 30 years in search of the transition region between the dominance of the solar energy and interstellar energy. The Voyager Interstellar Mission (VIM) started on January 1, 1990. It can be characterized by several factors including extremely long communication distances, aging hardware, reduced staffing levels and difficulty in obtaining Deep Space Network (DSN) resources necessitated by the increasing distance between the spacecraft and Earth. The mission was redesigned to compensate for such factors while maximizing the science return. After 25 years of VIM and several significant science discoveries, both Voyager spacecraft are still functioning well and the Voyager flight team is preparing for an even longer mission - until the year 2025 and beyond. In order to work around the challenges and to continue the mission even further, the team has been implementing numerous changes, mainly through flight software modifications and hardware reconfiguration. The major drivers for the changes are two-fold: resource constraints (such as decreasing power output and difficulty in obtaining the necessary DSN coverage) and anomalies due to the aging hardware. The majority of changes occur through flight software modifications so the state of the on-board responses is appropriate for the changing space environment and mission phase, and the flight software is compatible in allowing the maximum data gathering. The on-board flight software routines such as baseline sequence, fault protection routines, the High Gain Antenna POINTing to Earth (HPOINT) table, and long-term events table need to be maintained through flight software updates. The changes also occur through hardware reconfiguration such as selecting the backup Hybrid Buffer Interface Circuits (HYBIC) or attitude propulsion thrusters. This paper will describe the challenges of VIM and what has been done to overcome or mitigate those challenges. The primary focus will be the major flight software changes made during VIM and the changes that are in store for the near future in preparation for continuing the extended mission, from the originally projected year of 2020 out to the year 2025 and possibly beyond.

Matsumoto, Sun Kang↗

Quantum-classical embedding via ghost Gutzwiller approximation for enhanced simulations of correlated electron systems

Simulating correlated materials on present-day quantum hardware remains challenging due to limited quantum resources. Quantum embedding methods offer a promising route by reducing computational complexity through the mapping of bulk systems onto effective impurity models, allowing more feasible simulations on pre- and early-fault-tolerant quantum devices. Here, this work develops a quantum-classical embedding framework based on the ghost Gutzwiller approximation to enable quantum-enhanced simulations of ground-state properties and spectral functions of correlated electron systems. Circuit complexity is analyzed using an adaptive variational quantum algorithm on a statevector simulator, applied to the infinite-dimensional Hubbard model with increasing ghost mode numbers from 3 to 5, resulting in circuit depths growing from 16 to 104. Noise effects are examined using a realistic error model, revealing significant impact on the spectral weight of the Hubbard bands. To mitigate these effects, the Iceberg quantum error detection code is employed, achieving up to 40% error reduction in simulations. Finally, the accuracy of the density matrix estimation and the derived spectral function is benchmarked on IBM and Quantinuum quantum hardware, featuring distinct qubit-connectivity and employing multiple levels of error mitigation techniques.

Chen, I-Chi [Ames Laboratory (AMES), Ames, IA (Uni↗

Neural Networks for Flight Control

Neural networks are being developed at NASA Ames Research Center to permit real-time adaptive control of time varying nonlinear systems, enhance the fault-tolerance of mission hardware, and permit online system reconfiguration. In general, the problem of controlling time varying nonlinear systems with unknown structures has not been solved. Adaptive neural control techniques show considerable promise and are being applied to technical challenges including automated docking of spacecraft, dynamic balancing of the space station centrifuge, online reconfiguration of damaged aircraft, and reducing cost of new air and spacecraft designs. Our experiences have shown that neural network algorithms solved certain problems that conventional control methods have been unable to effectively address. These include damage mitigation in nonlinear reconfiguration flight control, early performance estimation of new aircraft designs, compensation for damaged planetary mission hardware by using redundant manipulator capability, and space sensor platform stabilization. This presentation explored these developments in the context of neural network control theory. The discussion began with an overview of why neural control has proven attractive for NASA application domains. The more important issues in control system development were then discussed with references to significant technical advances in the literature. Examples of how these methods have been applied were given, followed by projections of emerging application needs and directions.

Jorgensen, Charles C.↗

Portable Health Algorithms Test System

A document discusses the Portable Health Algorithms Test (PHALT) System, which has been designed as a means for evolving the maturity and credibility of algorithms developed to assess the health of aerospace systems. Comprising an integrated hardware-software environment, the PHALT system allows systems health management algorithms to be developed in a graphical programming environment, to be tested and refined using system simulation or test data playback, and to be evaluated in a real-time hardware-in-the-loop mode with a live test article. The integrated hardware and software development environment provides a seamless transition from algorithm development to real-time implementation. The portability of the hardware makes it quick and easy to transport between test facilities. This hard ware/software architecture is flexible enough to support a variety of diagnostic applications and test hardware, and the GUI-based rapid prototyping capability is sufficient to support development execution, and testing of custom diagnostic algorithms. The PHALT operating system supports execution of diagnostic algorithms under real-time constraints. PHALT can perform real-time capture and playback of test rig data with the ability to augment/ modify the data stream (e.g. inject simulated faults). It performs algorithm testing using a variety of data input sources, including real-time data acquisition, test data playback, and system simulations, and also provides system feedback to evaluate closed-loop diagnostic response and mitigation control.

Melcher, Kevin J.↗

Use of Soft Computing Technologies For Rocket Engine Control

The problem to be addressed in this paper is to explore how the use of Soft Computing Technologies (SCT) could be employed to further improve overall engine system reliability and performance. Specifically, this will be presented by enhancing rocket engine control and engine health management (EHM) using SCT coupled with conventional control technologies, and sound software engineering practices used in Marshall s Flight Software Group. The principle goals are to improve software management, software development time and maintenance, processor execution, fault tolerance and mitigation, and nonlinear control in power level transitions. The intent is not to discuss any shortcomings of existing engine control and EHM methodologies, but to provide alternative design choices for control, EHM, implementation, performance, and sustaining engineering. The approaches outlined in this paper will require knowledge in the fields of rocket engine propulsion, software engineering for embedded systems, and soft computing technologies (i.e., neural networks, fuzzy logic, and Bayesian belief networks), much of which is presented in this paper. The first targeted demonstration rocket engine platform is the MC-1 (formerly FASTRAC Engine) which is simulated with hardware and software in the Marshall Avionics & Software Testbed laboratory that

Trevino, Luis C.↗

NASA Space Flight Vehicle Fault Isolation Challenges

The Space Launch System (SLS) is the new NASA heavy lift launch vehicle and is scheduled for its first mission in 2017. The goal of the first mission, which will be uncrewed, is to demonstrate the integrated system performance of the SLS rocket and spacecraft before a crewed flight in 2021. SLS has many of the same logistics challenges as any other large scale program. Common logistics concerns for SLS include integration of discrete programs geographically separated, multiple prime contractors with distinct and different goals, schedule pressures and funding constraints. However, SLS also faces unique challenges. The new program is a confluence of new hardware and heritage, with heritage hardware constituting seventy-five percent of the program. This unique approach to design makes logistics concerns such as testability of the integrated flight vehicle especially problematic. The cost of fully automated diagnostics can be completely justified for a large fleet, but not so for a single flight vehicle. Fault detection is mandatory to assure the vehicle is capable of a safe launch, but fault isolation is another issue. SLS has considered various methods for fault isolation which can provide a reasonable balance between adequacy, timeliness and cost. This paper will address the analyses and decisions the NASA Logistics engineers are making to mitigate risk while providing a reasonable testability solution for fault isolation.

Bramon, Christopher↗

ByzSec — A Multi-layered Byzantine Resilient Architecture for Bulk Power System Protective Relays

Reliability, selectivity, and sensitivity are the fundamental attributes of any protection system, acting as the main drivers in the selection of schemes, and equipment. In high-voltage systems, microprocessor-based relays represent the industry’s preferred solution, providing engineers with a vast array of benefits. However, they remain vulnerable to cybersecurity events that may compromise their functionality. To help mitigate against potential cybersecurity risks, this paper presents a fault-tolerant, Byzantine Resilient (BR) architecture that significantly increases the cybersecurity attributes of a protection system while minimizing the amount of performance impacts and integration overheads introduced. The solution relies on an array of independent relays that utilize robust consensus methods (based on Spire [1], [2]) to ensure correct system behavior is achieved even when a relay has been compromised. Furthermore, the solution has been complemented with a custom-built Situational Awareness engine that can be used to detect and identify potential threats. The implemented solution has been developed in consultation with three hardware vendors and has been tested to comply with the performance requirements of a 345kV differential protection scheme (87T). The results indicate that the proposed architecture is a comprehensive solution that: supports the strict correctness and performance requirements of the bulk power grid while providing a cost-effective alternative that offers a seamless, long-term solution.

byzantine security, Fault Tolerant Application Sof↗

Use of Soft Computing Technologies for a Qualitative and Reliable Engine Control System for Propulsion Systems

The problem to be addressed in this paper is to explore how the use of Soft Computing Technologies (SCT) could be employed to improve overall vehicle system safety, reliability, and rocket engine performance by development of a qualitative and reliable engine control system (QRECS). Specifically, this will be addressed by enhancing rocket engine control using SCT, innovative data mining tools, and sound software engineering practices used in Marshall's Flight Software Group (FSG). The principle goals for addressing the issue of quality are to improve software management, software development time, software maintenance, processor execution, fault tolerance and mitigation, and nonlinear control in power level transitions. The intent is not to discuss any shortcomings of existing engine control methodologies, but to provide alternative design choices for control, implementation, performance, and sustaining engineering, all relative to addressing the issue of reliability. The approaches outlined in this paper will require knowledge in the fields of rocket engine propulsion (system level), software engineering for embedded flight software systems, and soft computing technologies (i.e., neural networks, fuzzy logic, data mining, and Bayesian belief networks); some of which are briefed in this paper. For this effort, the targeted demonstration rocket engine testbed is the MC-1 engine (formerly FASTRAC) which is simulated with hardware and software in the Marshall Avionics & Software Testbed (MAST) laboratory that currently resides at NASA's Marshall Space Flight Center, building 4476, and is managed by the Avionics Department. A brief plan of action for design, development, implementation, and testing a Phase One effort for QRECS is given, along with expected results. Phase One will focus on development of a Smart Start Engine Module and a Mainstage Engine Module for proper engine start and mainstage engine operations. The overall intent is to demonstrate that by employing soft computing technologies, the quality and reliability of the overall scheme to engine controller development is further improved and vehicle safety is further insured. The final product that this paper proposes is an approach to development of an alternative low cost engine controller that would be capable of performing in unique vision spacecraft vehicles requiring low cost advanced avionics architectures for autonomous operations from engine pre-start to engine shutdown.

Trevino, Luis↗

Flight Evaluation of the Army/NASA Variable Stability Fly-by-Wire Rotorcraft Aircrew Systems Concept Airborne Laboratory (RASCAL) JUH-60A

NASA Ames Research Center and the U.S. Army Aeroflight dynamics Directorate (AFDD) have performed initial flight evaluations of the Research Flight Control System (RFCS) integrated into the Army/NASA Rotorcraft Aircrew Systems Concepts Airborne Laboratory (RASCAL) JUH-GOA. The highly modified JUH-GOA Black Hawk helicopter is a full authority, high bandwidth, variable stability, in-flight simulator designed to support development of advanced flight control, sensor, and integrated display and control technologies in a fail safe environment. Preparation for flight test required an extensive hazard analysis and ground testing to ensure proper system operation. A hardware in the loop development facility was utilized to evaluate control law stability following software changes, assess servo hardover upset conditions during manual and monitor disengagements and provide pilot familiarization of test techniques and software changes prior to flight. First engagement of the RFCS was conducted on 31 Aug 2001. RFCS transfer system operation, envelope expansion and a limited rate monitor evaluation have been completed with low bandwidth and model following control laws. The presentation will discuss the following - System overview including aircraft modifications and integrated development facilities used with the RASCAL facility. - Preliminary hazard identification and mitigation prior to flight test. - Ground testing used to qualify the RFCS transfer system and verify fault monitor operation. - Flight test results of low-bandwidth and model following control law evaluations including maneuver agility, control limitations, fault monitor reliability, and recovery from manual and monitor disengagement. - Lessons learned including test techniques using a passive three-axis sidearm controller, the value of the development facility in reducing risk and crew coordination issues related to the operation of a full authority, variable stability platform. - Future research and modifications planned for the RASCAL aircraft.

Dave Arterburn↗