Search NASASearch

SEARCH · Search NASA

Results for “multiple faults”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 217 records · Page 12

Measuring quasiparticle dynamics for particle impact reconstruction in a superconducting qubit chip

Quasiparticle poisoning following particle impacts poses a significant challenge to the development of fault-tolerant superconducting quantum computers, as a sudden excess of quasiparticles can simultaneously degrade the coherence of multiple qubits across large device arrays. In this work, we present a statistical analysis that models the time evolution of radiation-induced qubit energy relaxation through quasiparticle density dynamics. This study provides insight into quasiparticle loss processes by distinguishing between recombination and trapping decay channels and assessing their respective impact on qubit performance. We precisely measure quasiparticle recombination in multiple transmon qubits and uncover an unexpected dependence of qubit relaxation dynamics on deposited energy. By linking correlated relaxation events across qubits to ballistic phonon propagation, we introduce a statistical localization approach to extract the energy deposited in the substrate, which is in good agreement with Monte Carlo simulation. This work establishes the quantitative framework for using an arbitrary subset of superconducting transmon qubits in a QPU as energy-resolving witness particle detectors.

Celi, E. [Northwestern U.]

Characterization of fault recovery through fault injection on FTMP

The development of fault-injection procedures and statistical analysis techniques to characterize the fault recovery of fault-tolerant systems is described. Pin-level fault-injection was conducted on a fault-tolerant microprocessor computer in order to generate data to assess the utility of current fault-injection sampling methods. The validity of common reliability-modeling assumptions concerning the statistical distribution of recovery times is investigated. A multiple comparison analysis for detecting behavior variations, and a distribution fitting for determining the best fit for the data were conducted. It is observed that the detection behavior is not homogeneous across all data sets, and that none of the factors under experimental control can account for the observed groupings of behavior. It is determined that no single distribution fits all the data sets, and that stratified random sampling and statistically robust parameter-estimation techniques are required to characterize fault detection time.

Finelli, George B.

Spaceborne Processor Array

A Spaceborne Processor Array in Multifunctional Structure (SPAMS) can lower the total mass of the electronic and structural overhead of spacecraft, resulting in reduced launch costs, while increasing the science return through dynamic onboard computing. SPAMS integrates the multifunctional structure (MFS) and the Gilgamesh Memory, Intelligence, and Network Device (MIND) multi-core in-memory computer architecture into a single-system super-architecture. This transforms every inch of a spacecraft into a sharable, interconnected, smart computing element to increase computing performance while simultaneously reducing mass. The MIND in-memory architecture provides a foundation for high-performance, low-power, and fault-tolerant computing. The MIND chip has an internal structure that includes memory, processing, and communication functionality. The Gilgamesh is a scalable system comprising multiple MIND chips interconnected to operate as a single, tightly coupled, parallel computer. The array of MIND components shares a global, virtual name space for program variables and tasks that are allocated at run time to the distributed physical memory and processing resources. Individual processor- memory nodes can be activated or powered down at run time to provide active power management and to configure around faults. A SPAMS system is comprised of a distributed Gilgamesh array built into MFS, interfaces into instrument and communication subsystems, a mass storage interface, and a radiation-hardened flight computer.

Chow, Edward T.

Update on Development of SiC Multi-Chip Power Modules

Progress has been made in a continuing effort to develop multi-chip power modules (SiC MCPMs). This effort at an earlier stage was reported in 'SiC Multi-Chip Power Modules as Power-System Building Blocks' (LEW-18008-1), NASA Tech Briefs, Vol. 31, No. 2 (February 2007), page 28. The following recapitulation of information from the cited prior article is prerequisite to a meaningful summary of the progress made since then: 1) SiC MCPMs are, more specifically, electronic power-supply modules containing multiple silicon carbide power integrated-circuit chips and silicon-on-insulator (SOI) control integrated-circuit chips. SiC MCPMs are being developed as building blocks of advanced expandable, reconfigurable, fault-tolerant power-supply systems. Exploiting the ability of SiC semiconductor devices to operate at temperatures, breakdown voltages, and current densities significantly greater than those of conventional Si devices, the designs of SiC MCPMs and of systems comprising multiple SiC MCPMs are expected to afford a greater degree of miniaturization through stacking of modules with reduced requirements for heat sinking; 2) The stacked SiC MCPMs in a given system can be electrically connected in series, parallel, or a series/parallel combination to increase the overall power-handling capability of the system. In addition to power connections, the modules have communication connections. The SOI controllers in the modules communicate with each other as nodes of a decentralized control network, in which no single controller exerts overall command of the system. Control functions effected via the network include synchronization of switching of power devices and rapid reconfiguration of power connections to enable the power system to continue to supply power to a load in the event of failure of one of the modules; and, 3) In addition to serving as building blocks of reliable power-supply systems, SiC MCPMs could be augmented with external control circuitry to make them perform additional power-handling functions as needed for specific applications. Because identical SiC MCPM building blocks could be utilized in such a variety of ways, the cost and difficulty of designing new, highly reliable power systems would be reduced considerably. This concludes the information from the cited prior article. The main activity since the previously reported stage of development was the design, fabrication, and testing a 120- VDC-to-28-VDC modular power-converter system composed of eight SiC MCPMs in a 4 (parallel)-by-2 (series) matrix configuration, with normally-off controllable power switches. The SiC MCPM power modules include closed-loop control subsystems and are capable of operating at high power density or high temperature. The system was tested under various configurations, load conditions, load-transient conditions, and failure-recovery conditions. Planned future work includes refinement of the demonstrated modular system concept and development of a new converter hardware topology that would enable sharing of currents without the need for communication among modules. Toward these ends, it is also planned to develop a new converter control algorithm that would provide for improved sharing of current and power under all conditions, and to implement advanced packaging concepts that would enable operation at higher power density.

Lostetter, Alexander

Robot Position Sensor Fault Tolerance

Robot systems in critical applications, such as those in space and nuclear environments, must be able to operate during component failure to complete important tasks. One failure mode that has received little attention is the failure of joint position sensors. Current fault tolerant designs require the addition of directly redundant position sensors which can affect joint design. A new method is proposed that utilizes analytical redundancy to allow for continued operation during joint position sensor failure. Joint torque sensors are used with a virtual passive torque controller to make the robot joint stable without position feedback and improve position tracking performance in the presence of unknown link dynamics and end-effector loading. Two Cartesian accelerometer based methods are proposed to determine the position of the joint. The joint specific position determination method utilizes two triaxial accelerometers attached to the link driven by the joint with the failed position sensor. The joint specific method is not computationally complex and the position error is bounded. The system wide position determination method utilizes accelerometers distributed on different robot links and the end-effector to determine the position of sets of multiple joints. The system wide method requires fewer accelerometers than the joint specific method to make all joint position sensors fault tolerant but is more computationally complex and has lower convergence properties. Experiments were conducted on a laboratory manipulator. Both position determination methods were shown to track the actual position satisfactorily. A controller using the position determination methods and the virtual passive torque controller was able to servo the joints to a desired position during position sensor failure.

Aldridge, Hal A.

Evidence for a Single Holocene Paleoseismic Event on the Pajarito Fault, Northern New Mexico

Low-slip rate fault systems tend to be less studied than their high-slip rate counterparts, and paleoseismic techniques used to study them may pose challenges in interpretation that differ from high-slip rate systems. A good example of this is the Pajarito fault system (PFS), a normal fault complex within the Rio Grande rift. Despite numerous previous paleoseismic trenching studies conducted on the PFS between 1990 and 2003, considerable uncertainty remains regarding its Holocene paleoseismic history, particularly for the primary Pajarito fault (PF). To further clarify the PF paleoseismic history, we present data from paleoseismic investigations of 6 trenches at 3 distinct locations along the PF. Though the totality of the age and structural data obtained in this study is complex and not entirely consistent with any one interpretation, a single Holocene paleoearthquake occurring younger than ∼1,600 to 2,300 kcal yr BP is the simplest interpretation. It is possible that the PF records two Holocene events, with a penultimate event 6.9–2.4 kcal yr BP event and the aforementioned most recent event (MRE) between 2.3 and 1.6 kcal yr BP. However, only a single wall of one trench, out of a total of 12 walls in our 6 trenches, provides evidence supporting that interpretation. This study finds evidence of a single late Holocene paleoseismic event on the PF and sparse evidence for 2 Holocene paleoseismic events on the PF and highlights the benefits of logging multiple trench walls to better understand the complexity that results from this low-slip rate, low-deposition-rate fault system.

58 GEOSCIENCES

Detection of High-impedance Arcing Faults in Radial Distribution DC Systems

High voltage, low current arcing faults in DC power systems have been researched at the NASA Glenn Research Center in order to develop a method for detecting these 'hidden faults', in-situ, before damage to cables and components from localized heating can occur. A simple arc generator was built and high-speed and low-speed monitoring of the voltage and current waveforms, respectively, has shown that these high impedance faults produce a significant increase in high frequency content in the DC bus voltage and low frequency content in the DC system current. Based on these observations, an algorithm was developed using a high-speed data acquisition system that was able to accurately detect high impedance arcing events induced in a single-line system based on the frequency content of the DC bus voltage or the system current. Next, a multi-line, radial distribution system was researched to see if the arc location could be determined through the voltage information when multiple 'detectors' are present in the system. It was shown that a small, passive LC filter was sufficient to reliably isolate the fault to a single line in a multi-line distribution system. Of course, no modification is necessary if only the current information is used to locate the arc. However, data shows that it might be necessary to monitor both the system current and bus voltage to improve the chances of detecting and locating high impedance arcing faults

Gonzalez, Marcelo C.

Designing application software in wide area network settings

Progress in methodologies for developing robust local area network software has not been matched by similar results for wide area settings. The design of application software spanning multiple local area environments is examined. For important classes of applications, simple design techniques are presented that yield fault tolerant wide area programs. An implementation of these techniques as a set of tools for use within the ISIS system is described.

Makpangou, Mesaac

Predicted performance of an Integrated Modular Engine system

Space vehicle propulsion systems are traditionally comprised of a cluster of discrete engines, each with its own set of turbopumps, valves, and a thrust chamber. The Integrated Modular Engine (IME) concept proposes a vehicle propulsion system comprised of multiple turbopumps, valves, and thrust chambers which are all interconnected. The IME concept has potential advantages in fault-tolerance, weight, and operational efficiency compared with the traditional clustered engine configuration. The purpose of this study is to examine the steady-state performance of an IME system with various components removed to simulate fault conditions. An IME configuration for a hydrogen/oxygen expander cycle propulsion system with four sets of turbopumps and eight thrust chambers has been modeled using the Rocket Engine Transient Simulator program. The nominal steady-state performance is simulated, as well as turbopump, thrust chamber, and duct failures. The impact of component failures on system performance is discussed in the context of the system's fault tolerant capabilities.

Binder, Michael

Predicted performance of an integrated modular engine system

Space vehicle propulsion systems are traditionally comprised of a cluster of discrete engines, each with its own set of turbopumps, valves, and a thrust chamber. The Integrated Modular Engine (IME) concept proposes a vehicle propulsion system comprised of multiple turbopumps, valves, and thrust chambers which are all interconnected. The IME concept has potential advantages in fault-tolerance, weight, and operational efficiency compared with the traditional clustered engine configuration. The purpose of this study is to examine the steady-state performance of an IME system with various components removed to simulate fault conditions. An IME configuration for a hydrogen/oxygen expander cycle propulsion system with four sets of turbopumps and eight thrust chambers has been modeled using the Rocket Engine Transient Simulator (ROCETS) program. The nominal steady-state performance is simulated, as well as turbopump thrust chamber and duct failures. The impact of component failures on system performance is discussed in the context of the system's fault tolerant capabilities.

Binder, Michael

Integrated Hardware and Software for No-Loss Computing

When an algorithm is distributed across multiple threads executing on many distinct processors, a loss of one of those threads or processors can potentially result in the total loss of all the incremental results up to that point. When implementation is massively hardware distributed, then the probability of a hardware failure during the course of a long execution is potentially high. Traditionally, this problem has been addressed by establishing checkpoints where the current state of some or part of the execution is saved. Then in the event of a failure, this state information can be used to recompute that point in the execution and resume the computation from that point. A serious problem arises when one distributes a problem across multiple threads and physical processors is that one increases the likelihood of the algorithm failing due to no fault of the scientist but as a result of hardware faults coupled with operating system problems. With good reason, scientists expect their computing tools to serve them and not the other way around. What is novel here is a unique combination of hardware and software that reformulates an application into monolithic structure that can be monitored in real-time and dynamically reconfigured in the event of a failure. This unique reformulation of hardware and software will provide advanced aeronautical technologies to meet the challenges of next-generation systems in aviation, for civilian and scientific purposes, in our atmosphere and in atmospheres of other worlds. In particular, with respect to NASA s manned flight to Mars, this technology addresses the critical requirements for improving safety and increasing reliability of manned spacecraft.

James, Mark

LADEE Preparations for Contingency Operations for the Lunar Orbit Insertion Maneuver

The Lunar Atmosphere and Dust Environment Explorer (LADEE) spacecraft was launched on September 6, 2013, and completed its mission on April 17, 2014 with a directed impact to the Lunar Surface. Its primary goals were to examine the lunar atmosphere, measure lunar dust, and to demonstrate high rate laser communications. The LADEE mission was a resounding success, achieving all mission objectives, much of which can be attributed to careful planning and preparation. This paper discusses the specific preparations for fault conditions that could occur during a highly-critical phase of the mission. To get to the Moon, the spacecraft traversed multiple phasing loops around the Earth, and then executed a breaking maneuver to achieve lunar orbit. This Lunar Orbit Insertion (LOI) maneuver was perhaps the most time-critical phase of the entire mission. The LOI maneuver had to occur within a twenty minute window in order to achieve lunar orbit with an acceptable amount of propellant remaining. Missing this window would have likely resulted in a loss of the entire mission. An additional challenge of the maneuver was that spacecraft was out of view for approximately one hour prior to the main thruster burn, with the burn needing to occur within five minutes after coming into view. These conditions resulted in unique challenges for ground operations and the fault management system. Early in the planning stages of the mission, the criticality and challenges of this maneuver were evident to the system designers. The major concern was that any triggering of the on-board fault management system, whether it is in response to a true fault or a false positive, would result in an unacceptable delay to the burn. Therefore the flight software was designed with a flexible fault management system, such that any or all of the fault management responses could be disabled for the lead up and execution of the maneuver. Later, a triage was conducted to develop a list of fault responses, mapped to various parts of the timeline of the maneuver. Some of these contingency responses were solely ground-based if the time to detect, diagnose, and respond were adequate. Other responses were automated on-board if the response time from the ground would have been inadequate. For instance, in order to recover from a system reboot, on-board automation would have automatically reconfigured the spacecraft for the burn and reoriented the spacecraft to the burn attitude.These contingency responses were practiced, over and over, during numerous rehearsals. Although the LOI maneuver was executed without having to use any of these contingencies, the LADEE team was adequately prepared for this highly critical phase of the mission.

Cannon, Howard

Usage of Fault Detection Isolation & Recovery (FDIR) in Constellation (CxP) Launch Operations

This paper will explore the usage of Fault Detection Isolation & Recovery (FDIR) in the Constellation Exploration Program (CxP), in particular Launch Operations at Kennedy Space Center (KSC). NASA's Exploration Technology Development Program (ETDP) is currently funding a project that is developing a prototype FDIR to demonstrate the feasibility of incorporating FDIR into the CxP Ground Operations Launch Control System (LCS). An architecture that supports multiple FDIR tools has been formulated that will support integration into the CxP Ground Operation's Launch Control System (LCS). In addition, tools have been selected that provide fault detection, fault isolation, and anomaly detection along with integration between Flight and Ground elements.

Ferrell, Rob

Space Station Freedom integrated fault model

A demonstration of an integrated fault propagation model for Space Station Freedom is described. The demonstration uses a HyperCard graphical interface to show how failures can propagate from one component to another, both within a system and between systems. It also shows how hardware failures can impact certain defined functions like reboost, atmosphere maintenance or collision avoidance. The demonstration enables the user to view block diagrams for the various space station systems using an overview screen, and interactively choose a component and see what single or dual failure combinations can cause it to fail. It also allows the user to directly view the fault model, which is a collection of drawing and text listings accessible from a guide screen. Fault modeling provides a useful technique for analyzing individual systems and also interactions between systems in the presence of multiple failures so that a complete picture of failure tolerance and component criticality can be achieved.

Becker, Fred J.

Single Event Test Methodologies and System Error Rate Analysis for Triple Modular Redundant Field Programmable Gate Arrays

We present a test methodology for estimating system error rates of Field Programmable Gate Arrays (FPGAs) mitigated with Triple Modular Redundancy (TMR). The test methodology is founded in a mathematical model, which is also presented. Accelerator data from 90 nm Xilins Military/Aerospace grade FPGA are shown to fit the model. Fault injection (FI) results are discussed and related to the test data. Design implementation and the corresponding impact of multiple bit upset (MBU) are also discussed.

error rate calculations

Primitive Quantum Gates for an $SU(3)$ Discrete Subgroup: $Σ(72\times3)$

We construct a primitive gate set for the digital quantum simulation of a discrete subgroup of $SU(3)$: the 216-element $Σ(72\times3)$. The necessary primitives are the inversion gate, the group multiplication gate, the trace gate, and the group Fourier transform, for which we provide qubit decompositions. The resulting fault-tolerant T gate costs for a fiducial calculation of shear viscosity would require about $10^{12}$ T gates which compares favorably to other modern estimates.

Perez, Sebastian Osorio [Fermilab; Maryland U.]

Gearbox bearing crack growth prognostics and uncertainty quantification with physics-informed machine learning

This paper introduces the extreme theory of functional connections (X-TFC), a physics-informed machine learning algorithm, and tailors it to estimate the remaining useful life (RUL) of wind turbine gearbox bearings experiencing fatigue crack growth. Unlike purely data-driven methods, X-TFC embeds a physics model, based on Head's theory in this work, into its training objective. The core of X-TFC is a random-projection single-layer neural network trained via an extreme learning machine, which requires only limited damage progression data and solves for output weights with a least-squares optimization algorithm. A composite loss function balances the network's fit to observed degradation data against the residuals of the governing crack growth differential equation, ensuring the learned damage trajectory remains physically plausible. When applied to a vibration-based health-index (HI) dataset measured during the growth of a crack on the inner ring of a high-speed bearing in a wind turbine gearbox (Bechhoefer and Dubé, 2020), X-TFC achieves near-zero prediction bias. Even when trained on only the first 10 %–20 % of the damage progression data, with sufficient physics weighting its predictions remain monotonic and smooth, delivering high prognosability and trendability. To quantify the epistemic uncertainty, we employ a Monte Carlo ensemble of independently initialized X-TFC models trained on noise-perturbed data, which yields confidence intervals around each RUL estimate and captures both model-parameter and epistemic uncertainty. In addition to a vibration-based HI, we demonstrate that the proposed framework can be directly applied to a supervisory control and data acquisition (SCADA) data-based HI (Eftekhari Milani et al., 2026) measured during similar wind turbine gearbox bearing crack faults, preserving its accuracy and interpretability. This extension shows the versatility of our approach, which is applicable to bearings of multiple gearbox manufacturers, models, and ratings using only SCADA data. By integrating domain knowledge with machine learning, X-TFC offers a rapid, reliable tool for crack prognostics. Its adaptability to other bearing failure modes, such as pitch bearing ring cracks, positions X-TFC as a powerful enabler of data-driven, physics-informed asset management in the wind energy sector and beyond.

17 WIND ENERGY

Automated generation of reliability models

The abstract semi-Markov specification interface to the SURE (Semi-Markov Range Evaluator) tool (ASSIST) program allows the user to describe the Markov model in a high-level language. Instead of listing the individual states of the model, the user specifies the rules governing the behavior of the system, and these are used to automatically generate the model. A small number of statements in the abstract language can describe a large, complex model. Becuase no assumptions are made about the system being modeled, ASSIST can be used to generate models describing the behavior of any type of system. The abstract model definition and the automatic model generation strategy are described. Analysis of an example fault-tolerant architecture, a triad of processor with cold spare processors, shows how the behavior of a system can be captured by a few general rules. The syntax of the ASSIST input language is then described and demonstrated by creating a model to describe the fault behavior of the example architecture. The flexibility of the abstract language is demonstrated by expanding the example to model multiple triads of processors sharing a pool of cold spare processors.

Johnson, Sally C.