Search NASASearch

SEARCH · Search NASA

Results for “Hardware fault mitigations”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

Signal and Power Integrity Design Methodology for High-Performance Flight Computing Systems

Computing capabilities of space systems have in-creased onboard performance by orders of magnitude with the use of radiation-tolerant field-programmable gate arrays (FPGA)and processors. The incorporation of signal and power integrity analysis with printed circuit board (PCB) design in reliable computing architectures for space systems has become critical to enable future mission capabilities. Developers launch high-performance processors into a breadth of orbits and missions, running varying applications that create challenges for designing reliable computing hardware. Specifically, for these designs, academic and industry research has focused on component radiation performance, fault mitigation, and reliable architectures. How-ever, other design parameters including electromagnetic interference (EMI), PCB stackup, signal integrity (SI), voltage regulator module (VRM) design, and power distribution network (PDN)are often deprioritized or disregarded as the design matures. Since these characteristics are becoming more significant in high-performance processor designs, this research presents a hardware design and analysis methodology for high-performance, space-computing systems that focuses on a holistic design approach and PDN reliability. While these challenges exist across all space hardware, the reduced PCB dimensions imposed by SmallSats and CubeSats introduce additional hurdles, specifically to VRM and decoupling design. By examining the relationship between the PDN and radiation performance, an analytical relationship is developed that incorporates Total Ionizing Dose and Single-Event Transients to ensure reliability throughout the mission duration. The presented design methodology is applied to the SpaceCube v3.0 Mini, an FPGA-based on-board science data processing system developed at NASA Goddard Space Flight Center.

Advanced avionics

A Testbed for Evaluating Lunar Habitat Autonomy Architectures

A lunar outpost will involve a habitat with an integrated set of hardware and software that will maintain a safe environment for human activities. There is a desire for a paradigm shift whereby crew will be the primary mission operators, not ground controllers. There will also be significant periods when the outpost is uncrewed. This will require that significant automation software be resident in the habitat to maintain all system functions and respond to faults. JSC is developing a testbed to allow for early testing and evaluation of different autonomy architectures. This will allow evaluation of different software configurations in order to: 1) understand different operational concepts; 2) assess the impact of failures and perturbations on the system; and 3) mitigate software and hardware integration risks. The testbed will provide an environment in which habitat hardware simulations can interact with autonomous control software. Faults can be injected into the simulations and different mission scenarios can be scripted. The testbed allows for logging, replaying and re-initializing mission scenarios. An initial testbed configuration has been developed by combining an existing life support simulation and an existing simulation of the space station power distribution system. Results from this initial configuration will be presented along with suggested requirements and designs for the incremental development of a more sophisticated lunar habitat testbed.

Lawler, Dennis G.

Fault Management Architectures and the Challenges of Providing Software Assurance

Fault Management (FM) is focused on safety, the preservation of assets, and maintaining the desired functionality of the system. How FM is implemented varies among missions. Common to most missions is system complexity due to a need to establish a multi-dimensional structure across hardware, software and spacecraft operations. FM is necessary to identify and respond to system faults, mitigate technical risks and ensure operational continuity. Generally, FM architecture, implementation, and software assurance efforts increase with mission complexity. Because FM is a systems engineering discipline with a distributed implementation, providing efficient and effective verification and validation (V&V) is challenging. A breakout session at the 2012 NASA Independent Verification & Validation (IV&V) Annual Workshop titled "V&V of Fault Management: Challenges and Successes" exposed this issue in terms of V&V for a representative set of architectures. NASA's Software Assurance Research Program (SARP) has provided funds to NASA IV&V to extend the work performed at the Workshop session in partnership with NASA's Jet Propulsion Laboratory (JPL). NASA IV&V will extract FM architectures across the IV&V portfolio and evaluate the data set, assess visibility for validation and test, and define software assurance methods that could be applied to the various architectures and designs. This SARP initiative focuses efforts on FM architectures from critical and complex projects within NASA. The identification of particular FM architectures and associated V&V/IV&V techniques provides a data set that can enable improved assurance that a system will adequately detect and respond to adverse conditions. Ultimately, results from this activity will be incorporated into the NASA Fault Management Handbook providing dissemination across NASA, other agencies and the space community. This paper discusses the approach taken to perform the evaluations and preliminary findings from the research.

Fault Management

Fault Management Architectures and the Challenges of Providing Software Assurance

The satellite systems Fault Management (FM) is focused on safety, the preservation of assets, and maintaining the desired functionality of the system. How FM is implemented varies among missions. Common to most is system complexity due to a need to establish a multi-dimensional structure across hardware, software and operations. This structure is necessary to identify and respond to system faults, mitigate technical risks and ensure operational continuity. These architecture, implementation and software assurance efforts increase with mission complexity. Because FM is a systems engineering discipline with a distributed implementation, providing efficient and effective verification and validation (VV) is challenging. A breakout session at the 2012 NASA Independent Verification Validation (IVV) Annual Workshop titled VV of Fault Management: Challenges and Successes exposed these issues in terms of VV for a representative set of architectures. NASA's IVV is funded by NASA's Software Assurance Research Program (SARP) in partnership with NASA's Jet Propulsion Laboratory (JPL) to extend the work performed at the Workshop session. NASA IVV will extract FM architectures across the IVV portfolio and evaluate the data set for robustness, assess visibility for validation and test, and define software assurance methods that could be applied to the various architectures and designs. This work focuses efforts on FM architectures from critical and complex projects within NASA. The identification of particular FM architectures, visibility, and associated VVIVV techniques provides a data set that can enable higher assurance that a satellite system will adequately detect and respond to adverse conditions. Ultimately, results from this activity will be incorporated into the NASA Fault Management Handbook providing dissemination across NASA, other agencies and the satellite community. This paper discusses the approach taken to perform the evaluations and preliminary findings from the research including identification of FM architectures, visibility observations, and methods utilized for VVIVV.

Fault Management

NASA Spacecraft Fault Management Workshop Results

Fault Management is a critical aspect of deep-space missions. For the purposes of this paper, fault management is defined as the ability of a system to detect, isolate, and mitigate events that impact, or have the potential to impact, nominal mission operations. The fault management capabilities are commonly distributed across flight and ground subsystems, impacting hardware, software, and mission operations designs. The National Aeronautics and Space Administration (NASA) Discovery & New Frontiers (D&NF) Program Office at Marshall Space Flight Center (MSFC) recently studied cost overruns and schedule delays for 5 missions. The goal was to identify the underlying causes for the overruns and delays, and to develop practical mitigations to assist the D&NF projects in identifying potential risks and controlling the associated impacts to proposed mission costs and schedules. The study found that 4 out of the 5 missions studied had significant overruns due to underestimating the complexity and support requirements for fault management. As a result of this and other recent experiences, the NASA Science Mission Directorate (SMD) Planetary Science Division (PSD) commissioned a workshop to bring together invited participants across government, industry, academia to assess the state of the art in fault management practice and research, identify current and potential issues, and make recommendations for addressing these issues. The workshop was held in New Orleans in April of 2008. The workshop concluded that fault management is not being limited by technology, but rather by a lack of emphasis and discipline in both the engineering and programmatic dimensions. Some of the areas cited in the findings include different, conflicting, and changing institutional goals and risk postures; unclear ownership of end-to-end fault management engineering; inadequate understanding of the impact of mission-level requirements on fault management complexity; and practices, processes, and tools that have not kept pace with the increasing complexity of mission requirements and spacecraft systems. This paper summarizes the findings and recommendations from that workshop, as well as opportunities identified for future investment in tools, processes, and products to facilitate the development of space flight fault management capabilities.

Newhouse, Marilyn

Error Mitigation of Point-to-Point Communication for Fault-Tolerant Computing

Fault tolerant systems require the ability to detect and recover from physical damage caused by the hardware s environment, faulty connectors, and system degradation over time. This ability applies to military, space, and industrial computing applications. The integrity of Point-to-Point (P2P) communication, between two microcontrollers for example, is an essential part of fault tolerant computing systems. In this paper, different methods of fault detection and recovery are presented and analyzed.

Akamine, Robert L.

Software Health Management: A Short Review of Challenges and Existing Techniques

Modern spacecraft (as well as most other complex mechanisms like aircraft, automobiles, and chemical plants) rely more and more on software, to a point where software failures have caused severe accidents and loss of missions. Software failures during a manned mission can cause loss of life, so there are severe requirements to make the software as safe and reliable as possible. Typically, verification and validation (V&V) has the task of making sure that all software errors are found before the software is deployed and that it always conforms to the requirements. Experience, however, shows that this gold standard of error-free software cannot be reached in practice. Even if the software alone is free of glitches, its interoperation with the hardware (e.g., with sensors or actuators) can cause problems. Unexpected operational conditions or changes in the environment may ultimately cause a software system to fail. Is there a way to surmount this problem? In most modern aircraft and many automobiles, hardware such as central electrical, mechanical, and hydraulic components are monitored by IVHM (Integrated Vehicle Health Management) systems. These systems can recognize, isolate, and identify faults and failures, both those that already occurred as well as imminent ones. With the help of diagnostics and prognostics, appropriate mitigation strategies can be selected (replacement or repair, switch to redundant systems, etc.). In this short paper, we discuss some challenges and promising techniques for software health management (SWHM). In particular, we identify unique challenges for preventing software failure in systems which involve both software and hardware components. We then present our classifications of techniques related to SWHM. These classifications are performed based on dimensions of interest to both developers and users of the techniques, and hopefully provide a map for dealing with software faults and failures.

Pipatsrisawat, Knot

SDN-Based Dynamic Cybersecurity Framework of IEC-61850 Communications in Smart Grid

In recent years, critical infrastructure and power grids have experienced a series of cyber-attacks, leading to temporary, widespread blackouts of considerable magnitude. Since most substations are unmanned and have limited physical security protection, cyber breaches into power grid substations present a risk. Nowadays, the susceptibility of SDN architecture to cyber-attacks has exhibited a notable increase in recent years, as indicated by research findings. This suggests a growing concern regarding the potential for cybersecurity breaches within the SDN framework. In this paper, we propose a hybrid intrusion detection system (IDS)-integrated SDN architecture for detecting and preventing the injection of malicious IEC 61850-based generic object-oriented system event (GOOSE) messages in a digital substation. Additionally, this program locates the fault’s location and, as a form of mitigation, disables a certain port. Furthermore, implementation examples are demonstrated and verified using a hardware-in-the-loop (HIL) testbed that mimics the functioning of a digital substation.

Liu, Chen-Ching [Virginia Tech] (ORCID:00000002894

FPGA-Based, Self-Checking, Fault-Tolerant Computers

A proposed computer architecture would exploit the capabilities of commercially available field-programmable gate arrays (FPGAs) to enable computers to detect and recover from bit errors. The main purpose of the proposed architecture is to enable fault-tolerant computing in the presence of single-event upsets (SEUs). [An SEU is a spurious bit flip (also called a soft error) caused by a single impact of ionizing radiation.] The architecture would also enable recovery from some soft errors caused by electrical transients and, to some extent, from intermittent and permanent (hard) errors caused by aging of electronic components. A typical FPGA of the current generation contains one or more complete processor cores, memories, and highspeed serial input/output (I/O) channels, making it possible to shrink a board-level processor node to a single integrated-circuit chip. Custom, highly efficient microcontrollers, general-purpose computers, custom I/O processors, and signal processors can be rapidly and efficiently implemented by use of FPGAs. Unfortunately, FPGAs are susceptible to SEUs. Prior efforts to mitigate the effects of SEUs have yielded solutions that degrade performance of the system and require support from external hardware and software. In comparison with other fault-tolerant- computing architectures (e.g., triple modular redundancy), the proposed architecture could be implemented with less circuitry and lower power demand. Moreover, the fault-tolerant computing functions would require only minimal support from circuitry outside the central processing units (CPUs) of computers, would not require any software support, and would be largely transparent to software and to other computer hardware. There would be two types of modules: a self-checking processor module and a memory system (see figure). The self-checking processor module would be implemented on a single FPGA and would be capable of detecting its own internal errors. It would contain two CPUs executing identical programs in lock step, with comparison of their outputs to detect errors. It would also contain various cache local memory circuits, communication circuits, and configurable special-purpose processors that would use self-checking checkers. (The basic principle of the self-checking checker method is to utilize logic circuitry that generates error signals whenever there is an error in either the checker or the circuit being checked.) The memory system would comprise a main memory and a hardware-controlled check-pointing system (CPS) based on a buffer memory denoted the recovery cache. The main memory would contain random-access memory (RAM) chips and FPGAs that would, in addition to everything else, implement double-error-detecting and single-error-correcting memory functions to enable recovery from single-bit errors.

Some, Raphael

Lessons Learned from Astrobee Operations on the International Space Station

Since its launch in 2019, NASA has been operating three Astrobee free-flying robots providing an autonomous and adaptable research platform aboard the International Space Station (ISS). These robots have not only facilitated a myriad of national and international research endeavors in microgravity but have also served as a STEM outreach platform for student competitions aboard the ISS. Amidst its extensive operational tenure, spanning over five years and exceeding 1200 hours of cumulative free-flyer operation as of April 2024, the Astrobee robots have encountered software and hardware anomalies. Despite its inherent design for on-orbit repair or replacement, certain anomalies have proven to be complex, necessitating remote resolution via software and firmware updates or, in extreme cases, hardware replacements or the return of faulty units to NASA's ground facilities for repair. Such challenges underscore the delicate balance between the autonomous functionality of Astrobee and the occasional need for human intervention to maintain optimal performance. One recurring point of failure identified during Astrobee's operational lifespan has been the SD card, a critical component utilized by the different Astrobee processors and the Dock Station. The occurrence of SD card anomalies, both on orbit and within ground units, has provided invaluable insights into the improvement of Astrobee's systems and mitigation to future faults. This presentation will focus on four key areas: 1. Overview of Faults and Anomalies: A comprehensive examination of the diverse array of faults and anomalies encountered by Astrobee and its associated systems both in orbit and on the ground. From software glitches to hardware malfunctions, this section provides insights into the challenges faced during Astrobee's operational tenure. 2. Resolution Processes and Procedures: An in-depth discussion of the methodologies and procedures implemented to resolve the encountered anomalies. This includes remote troubleshooting, software patches, firmware updates, and, when necessary, the logistics involved in hardware replacements or down-massing for repair. 3. Implementation of Software Updates and Hardware Upgrades: A detailed exploration of the strategies employed to mitigate the risk of recurring anomalies through the implementation of software updates and hardware upgrades. This section highlights the iterative nature of Astrobee's development, emphasizing the continuous pursuit of robustness and reliability. 4. Lessons Learned and Future Directions: Reflecting on the insights gained from addressing anomalies, this section examines the lessons learned and outlines future directions for enhancing Astrobee's robustness and resilience. It underscores the iterative nature of space exploration and the importance of adaptability and continuous improvement in the pursuit of scientific discovery. Through a nuanced examination of Astrobee's operational challenges and the strategies employed to overcome them, this presentation sheds light on the complexities of operating autonomous robotic systems in the ISS environment. It underscores NASA's commitment to pushing the boundaries of exploration and innovation while navigating the inherent challenges of space exploration.

Astrobee

Lessons Learned from Astrobee Operations on the International Space Station

Since its launch in 2019, NASA has been operating three Astrobee free-flying robots providing an autonomous and adaptable research platform aboard the International Space Station (ISS). These robots have not only facilitated a myriad of national and international research endeavors in microgravity but have also served as a STEM outreach platform for student competitions aboard the ISS. Amidst its extensive operational tenure, spanning over five years and exceeding 1200 hours of cumulative free-flyer operation as of April 2024, the Astrobee robots have encountered software and hardware anomalies. Despite its inherent design for on-orbit repair or replacement, certain anomalies have proven to be complex, necessitating remote resolution via software and firmware updates or, in extreme cases, hardware replacements or the return of faulty units to NASA's ground facilities for repair. Such challenges underscore the delicate balance between the autonomous functionality of Astrobee and the occasional need for human intervention to maintain optimal performance. One recurring point of failure identified during Astrobee's operational lifespan has been the SD card, a critical component utilized by the different Astrobee processors and the Dock Station. The occurrence of SD card anomalies, both on orbit and within ground units, has provided invaluable insights into the improvement of Astrobee's systems and mitigation to future faults. This presentation will focus on four key areas: 1. Overview of Faults and Anomalies: A comprehensive examination of the diverse array of faults and anomalies encountered by Astrobee and its associated systems both in orbit and on the ground. From software glitches to hardware malfunctions, this section provides insights into the challenges faced during Astrobee's operational tenure. 2. Resolution Processes and Procedures: An in-depth discussion of the methodologies and procedures implemented to resolve the encountered anomalies. This includes remote troubleshooting, software patches, firmware updates, and, when necessary, the logistics involved in hardware replacements or down-massing for repair. 3. Implementation of Software Updates and Hardware Upgrades: A detailed exploration of the strategies employed to mitigate the risk of recurring anomalies through the implementation of software updates and hardware upgrades. This section highlights the iterative nature of Astrobee's development, emphasizing the continuous pursuit of robustness and reliability. 4. Lessons Learned and Future Directions: Reflecting on the insights gained from addressing anomalies, this section examines the lessons learned and outlines future directions for enhancing Astrobee's robustness and resilience. It underscores the iterative nature of space exploration and the importance of adaptability and continuous improvement in the pursuit of scientific discovery. Through a nuanced examination of Astrobee's operational challenges and the strategies employed to overcome them, this presentation sheds light on the complexities of operating autonomous robotic systems in the ISS environment. It underscores NASA's commitment to pushing the boundaries of exploration and innovation while navigating the inherent challenges of space exploration.

Astrobee

Novel concept for detection of a fluid flow fault in a pumped fluid heat rejection system

A pumped fluid heat rejection system (HRS) requires continuous flow of the working fluid to ensure that the components controlled by the HRS stay within their allowable temperature limits. An interruption of flow could result in violations of hardware qualification limits and mission failure. Some of these violations can happen within a few hours. Hence, quick detection of a flow fault to invoke mitigation measures is very critical. In typical pumped fluid HRS, dual or triple pumps are employed to switch the backup units in case the primary unit were to fail. The key metric for the selection of a fault detection system is that it should detect the fault much before the fault’s impact on the thermal health of the HRS controlled components is realized. A trade study was conducted to select such a system for the Europa Clipper Mission to Europa, a moon of Jupiter that is planned for a launch in 2023. Out of the several concepts studied, the most attractive one was a novel and simple concept that uses a low power film heater attached to a section of the HRS tubing. While the fluid flows at its nominal rate, the high thermal coupling of the flowing fluid leads to the tube being close to the fluid’s temperature. However, when the flow stops, the heater warms the small thermal mass of the tubing to a high temperature in a short span of time (~15 minutes). This large temperature rise would then imply that the flow must have stopped. This paper will describe the various concepts considered, the chosen concept, its implementation, and the results of developments tests to validate its performance.

Schmidt, Tyler

Electrified Aircraft Propulsion Systems: Potential Failure Modes and Failure Mitigation Strategies

Electrified aircraft propulsion (EAP) systems hold great potential for the reduction of aircraft fuel burn, emissions, and noise. Currently, NASA and other organizations are actively working to identify and mature technologies necessary to bring EAP designs to reality. A requirement for the development of any civil aircraft and its systems is to ensure that potential hazards in the design are identified and appropriately mitigated to ensure that the system is safe. During aircraft development, a system safety assessment that consists of a functional hazard assessment is conducted to identify all potential failure conditions of each function, and classify those failures according to the severity of their effects on the aircraft or its occupants. The more severe a function's failure condition classification, the greater the development assurance level required for the function to ensure that the probability of the hazard is acceptably low. Today, aircraft engines and their control systems receive type certificate approval as a stand-alone system to signify their airworthiness. However, the complex coupling and distributed nature of EAP designs are expected to place added challenges on the certification of these systems. This presentation will provide an initial high-level review of the potential failure modes and hazards posed by a generic EAP system along with potential mitigation strategies for those failures. The EAP system is assumed to be a hybrid design consisting of gas turbine engines, mechanical drives, electric machines, power electronics and distribution systems, energy storage devices, and motor driven propulsors. The functionality provided by each of these EAP subsystems will be discussed along with the potential failure modes they may encounter. This will include a discussion of coupled failure effects, where a fault in one EAP subsystem effects the operation of other subsystems in the architecture. Next, potential failure mitigation strategies are discussed including both software-based and hardware-based mitigation strategies. The presentation will conclude with an example evaluation of the potential failure modes and mitigation strategies for a concept EAP system proposed by NASA.

Simon, Donald L.

Enabling Reliable, Fault-Tolerant Autonomous Lunar Habitats with High-Performance Spaceflight Computing

The lunar surface presents unfavorable constraints and harsh living conditions. To address these challenges, autonomous habitats will require complex integrated systems that combine advanced software, high-performance hardware, and cutting-edge sensors to ensure sustainability, safety, and operational efficiency. Consequently, maintaining a sustainable presence on the Moon requires reliable infrastructure and efficient development, precise monitoring, and utilization of resources within a lunar installation. These elements are essential not only to ensure that lunar settlement can be long-term, self-sustaining, and resource-efficient, but also to serve as a foundation for future missions and eventual human habitation on Mars. Humans are not native to the Moon; therefore, our survival and ability to thrive will depend on autonomous systems that can foster safety and resilience through high-availability architectures, graceful degradation, and highly fault-tolerant spaceflight hardware capable of continuing operation during failures. This requires advanced human-rated distributed systems architectures with specialized electronics, scalable capabilities, and an integrated design approach. Unlike current practices focused on short-term missions and regularly maintained components, permanent lunar compute systems must be designed for extended operations beyond mission durations. This paper explores the necessity of transitioning toward fault- tolerant, highly autonomous hardware systems designed for multi-year missions. It also identifies critical subsystems that require high levels of autonomy, supported by radiation-hardened processors and extreme thermal loads, which are essential to mitigate long-term degradation and ensure sustainable lunar habitation. Finally, the paper aligns with NASA’s identified Civil Space Shortfalls, particularly in high-performance onboard computing, advanced data acquisition, extreme-environment avionics, radiation monitoring and countermeasures, and autonomous health management. It proposes NASA’s new High-Performance Spaceflight Computing (HPSC) processor as a turnkey solution, delivering 100 times the performance-per-watt of legacy rad-hard CPUs and enabling onboard AI, edge computing, and fault-tolerant features essential for sustained lunar autonomy and beyond.

Sarkis S Mikaelian

Electrical Ground Support Equipment for the Sampling Caching System of the Mars 2020 Rover

In this work we describe in detail the architecture, design, testing and operation of the Electrical Ground Support Equipment (EGSE) “Blue Box” used to test and validate the Sampling Caching System (SCS) of the Mars 2020 Perseverance rover. The Blue Box architecture is centered around COTS motor controllers and COTS input-output modules communicating over an EtherCAT bus. A custom, low-level safety subsystem ensures no harm can be done to the flight articles. The modular architecture of the EGSE reduces cost and complexity while expediting assembly time. The Blue Box drives the 19 actuators of the SCS which span the main robotic arm, the corer system, the internal sample handling arm, the sample tube sealing system and the gas dust removal tool; mimicking the Rover Motor Control Assembly (RMCA). Due to the limited availability of RMCA’s, the EGSE enabled and performed the bulk of testing activities for SCS. The majority of the SCS actuators are composed of a 3-phase DC brushless motors, hall sensors for commutation, dual resolvers for output angular measurement, brakes, heaters and platinum thermistors. Additionally, the EGSE read 12 strain gauges forming part of a force torque sensor, and switches used for external positioning references. Over the 3-year span of the V&V campaign for the SCS, over 32 EGSE systems were built, tested and deployed to test venues at JPL and externally. The EGSE tested several families of the SCS subsystem, ranging from engineering units, life test units and two flight units. Test venues that this EGSE supported included lab benches, ultra-clean cleanrooms, ATLO facilities, and thermal vacuum chambers. Together with the test software systems, SSDEV and SSDEV-ECAT, the Blue Box EGSE enabled the team to efficiently test flight hardware and flight software together. We go over the safety features and fault management techniques employed to protect flight hardware. The effects of the long, 50-feet, EGSE harnesses on motor performance, EMI, electrical noise, and motor control performance are explained. Mitigations to these unwanted effects, including shielding strategy and inductance compensation, are summarized. We go over an excerpt of notable anomalies that this EGSE suffered through its operation, along with investigations and resolutions. Lessons learned, areas of improvement as part of future work, and recommendations for future implementations for similar EGSE’s, are shared.

Levine, Dan

R2U2 in Space: System and Software Health Management for Small Satellites

In order for small but complex systems like rovers, SmallSats, or Unmanned Aircraft (UAS) to operate autonomously, they must have a real-time solution for assessing their own system health. System and Software Health Management (SHM) enables better detection of faulty sensors and software problems, and enables better fault management including mitigation of unpredicted fault scenarios in the absence of a human on-board. In recent work, we have developed a Responsive, Realizable, Unobtrusive Unit (R2U2) for on-board SHM of autonomous UAS and demonstrated its ability to detect faults during flight time. These faults, from sensor failures, to software problems, to malicious security attacks, can present as transient temporal faults that even humans are challenged to find. An R2U2 congfiuration is a modular combination of multiple types of temporal logic runtime observers with fault-specic Bayesian Nets and sensor filters. R2U2 reasons about both on-board hardware and software components; R2U2 itself can be instantiated as an independent FPGA (Field-Programmable Gate Array)-based conguration or as a software component running independently from other software on-board. Small satellites, such as CubeSats, also require on-board SHM and failure mitigation, as limited telemetry bandwidth does not allow the transmission of the entire system state for ground-based health management. However, the autonomous operation of satellites brings a set of challenges different from UAS, including the effects of radiation on non-rad-hard, low-cost components, and the harsher environment of space. We surmise that a new extension of R2U2 could be adapted to help better detect, for example, radiation errors in cheaper COTS (Commercial Off the Shelf) (not rad-hard) components often used in small space systems. Since small satellites often operate in coordination, we will also examine new ways of distributed monitoring of their communication and cooperation and real-time detection of off-nominal situations utilizing multiple satellites. This talk will discuss preliminary work and ideas for building on terrestrial success of system and software health management for the harsher, and differently challenging, environment of space.

Runtime Verification & Validation

Quantum utility-scale error mitigation for quantum quench dynamics in Heisenberg spin chains

Here, we implement a quantum error mitigation method termed self-mitigation, which is comparable to zero-noise extrapolation, at large scales to achieve quantum utility on near-term, noisy quantum computers. We investigate the effectiveness of several quantum error mitigation strategies, including self-mitigation, by simulating quantum quench dynamics for Heisenberg spin chains with system sizes up to 104 qubits using IBM quantum processors. In particular, we discuss the limitations of zero-noise extrapolation and the advantages offered by self-mitigation at large scales. The self-mitigation method demonstrates stable accuracy with large systems of 104 qubits comprising more than 3,000 CNOT gates. Also, we combine the discussed quantum error mitigation methods with practical entanglement entropy measuring methods, and it shows a good agreement with the theoretical estimation. Our study illustrates the usefulness of near-term noisy quantum hardware in examining the quantum quench dynamics of many-body systems at large scales and lays the groundwork for surpassing classical simulations with quantum methods prior to the development of fault-tolerant quantum computers.

97 MATHEMATICS AND COMPUTING