Search NASASearch

SEARCH · Search NASA

Results for “fault mitigation methods”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

Fault Mitigation Schemes for Future Spaceflight Multicore Processors

Future planetary exploration missions demand significant advances in on-board computing capabilities over current avionics architectures based on a single-core processing element. The state-of-the-art multi-core processor provides much promise in meeting such challenges while introducing new fault tolerance problems when applied to space missions. Software-based schemes are being presented in this paper that can achieve system-level fault mitigation beyond that provided by radiation-hard-by-design (RHBD). For mission and time critical applications such as the Terrain Relative Navigation (TRN) for planetary or small body navigation, and landing, a range of fault tolerance methods can be adapted by the application. The software methods being investigated include Error Correction Code (ECC) for data packet routing between cores, virtual network routing, Triple Modular Redundancy (TMR), and Algorithm-Based Fault Tolerance (ABFT). A robust fault tolerance framework that provides fail-operational behavior under hard real-time constraints and graceful degradation will be demonstrated using TRN executing on a commercial Tilera(R) processor with simulated fault injections.

software based

Evolutionary Based Techniques for Fault Tolerant Field Programmable Gate Arrays

The use of SRAM-based Field Programmable Gate Arrays (FPGAs) is becoming more and more prevalent in space applications. Commercial-grade FPGAs are potentially susceptible to permanently debilitating Single-Event Latchups (SELs). Repair methods based on Evolutionary Algorithms may be applied to FPGA circuits to enable successful fault recovery. This paper presents the experimental results of applying such methods to repair four commonly used circuits (quadrature decoder, 3-by-3-bit multiplier, 3-by-3-bit adder, 440-7 decoder) into which a number of simulated faults have been introduced. The results suggest that evolutionary repair techniques can improve the process of fault recovery when used instead of or as a supplement to Triple Modular Redundancy (TMR), which is currently the predominant method for mitigating FPGA faults.

Larchev, Gregory V.

Extended Testability Analysis Tool

The Extended Testability Analysis (ETA) Tool is a software application that supports fault management (FM) by performing testability analyses on the fault propagation model of a given system. Fault management includes the prevention of faults through robust design margins and quality assurance methods, or the mitigation of system failures. Fault management requires an understanding of the system design and operation, potential failure mechanisms within the system, and the propagation of those potential failures through the system. The purpose of the ETA Tool software is to process the testability analysis results from a commercial software program called TEAMS Designer in order to provide a detailed set of diagnostic assessment reports. The ETA Tool is a command-line process with several user-selectable report output options. The ETA Tool also extends the COTS testability analysis and enables variation studies with sensor sensitivity impacts on system diagnostics and component isolation using a single testability output. The ETA Tool can also provide extended analyses from a single set of testability output files. The following analysis reports are available to the user: (1) the Detectability Report provides a breakdown of how each tested failure mode was detected, (2) the Test Utilization Report identifies all the failure modes that each test detects, (3) the Failure Mode Isolation Report demonstrates the system s ability to discriminate between failure modes, (4) the Component Isolation Report demonstrates the system s ability to discriminate between failure modes relative to the components containing the failure modes, (5) the Sensor Sensor Sensitivity Analysis Report shows the diagnostic impact due to loss of sensor information, and (6) the Effect Mapping Report identifies failure modes that result in specified system-level effects.

Melcher, Kevin

NASA Tech Briefs, March 2014

Topics include: Data Fusion for Global Estimation of Forest Characteristics From Sparse Lidar Data; Debris and Ice Mapping Analysis Tool - Database; Data Acquisition and Processing Software - DAPS; Metal-Assisted Fabrication of Biodegradable Porous Silicon Nanostructures; Post-Growth, In Situ Adhesion of Carbon Nanotubes to a Substrate for Robust CNT Cathodes; Integrated PEMFC Flow Field Design for Gravity-Independent Passive Water Removal; Thermal Mechanical Preparation of Glass Spheres; Mechanistic-Based Multiaxial-Stochastic-Strength Model for Transversely-Isotropic Brittle Materials; Methods for Mitigating Space Radiation Effects, Fault Detection and Correction, and Processing Sensor Data; Compact Ka-Band Antenna Feed with Double Circularly Polarized Capability; Dual-Leadframe Transient Liquid Phase Bonded Power Semiconductor Module Assembly and Bonding Process; Quad First Stage Processor: A Four-Channel Digitizer and Digital Beam-Forming Processor; Protective Sleeve for a Pyrotechnic Reefing Line Cutter; Metabolic Heat Regenerated Temperature Swing Adsorption; CubeSat Deployable Log Periodic Dipole Array; Re-entry Vehicle Shape for Enhanced Performance; NanoRacks-Scale MEMS Gas Chromatograph System; Variable Camber Aerodynamic Control Surfaces and Active Wing Shaping Control; Spacecraft Line-of-Sight Stabilization Using LWIR Earth Signature; Technique for Finding Retro-Reflectors in Flash LIDAR Imagery; Novel Hemispherical Dynamic Camera for EVAs; 360 deg Visual Detection and Object Tracking on an Autonomous Surface Vehicle; Simulation of Charge Carrier Mobility in Conducting Polymers; Observational Data Formatter Using CMOR for CMIP5; Propellant Loading Physics Model for Fault Detection Isolation and Recovery; Probabilistic Guidance for Swarms of Autonomous Agents; Reducing Drift in Stereo Visual Odometry; Future Air-Traffic Management Concepts Evaluation Tool; Examination and A Priori Analysis of a Direct Numerical Simulation Database for High-Pressure Turbulent Flows; and Resource-Constrained Application of Support Vector Machines to Imagery.

Source record

Fault Tree Analysis Application for Safety and Reliability

Many commercial software tools exist for fault tree analysis (FTA), an accepted method for mitigating risk in systems. The method embedded in the tools identifies a root as use in system components, but when software is identified as a root cause, it does not build trees into the software component. No commercial software tools have been built specifically for development and analysis of software fault trees. Research indicates that the methods of FTA could be applied to software, but the method is not practical without automated tool support. With appropriate automated tool support, software fault tree analysis (SFTA) may be a practical technique for identifying the underlying cause of software faults that may lead to critical system failures. We strive to demonstrate that existing commercial tools for FTA can be adapted for use with SFTA, and that applied to a safety-critical system, SFTA can be used to identify serious potential problems long before integrator and system testing.

Wallace, Dolores R.

Logical Shadow Tomography: Efficient Estimation of Error-mitigated Observables

In near-term quantum applications, reducing errors and improving device reliability is an essential task. Towards these ends, various techniques have been introduced in recent literature, collectively referred to as quantum error mitigation techniques, for reducing errors in pre-fault-tolerant devices. Here, we introduce logical shadow tomography as a versatile error mitigation method. Our technique uses a stabilizer code to encode information in a logical state. Instead of doing active error correction, quantum states will be measured at the end of computation via shadow tomography and non-logical errors are projected out in the classical post-processing. Relative to quantum subspace expansion which requires O(2(M-1)L) experiments to estimate an logical Pauli observable encoded by an [[M, L, d]] code, our technique only requires 2L experiments, an important practical reduction in resources.

Hong-Ye Hu

Design and Testing of a Hard-Fault Protection Circuit for a 1 kV SiC MOSFET Inverter

Due to increasingly high DC link voltages and further advancements in the current density of silicon carbide (SiC) MOSFETs, it has become evident that conventional IGBT protection methods are not sufficient to prevent exceeding the current rating of these devices during low-inductance fault events. This paper explores the use of an air core Rogowski coil topology to mitigate these hard fault events. The design of this circuit resulted in safe shutdown of a low impedance phase-to-phase fault in under one microsecond, tested up to DC link voltages of 1 kV. This paper details the theory, design, simulation, and successful test results of this method.

hard fault protection

Human Error Analysis for Human-Rated Space Systems

Humans bring unique capabilities to space systems and contribute to mission success in a manner that cannot be matched by machines. Nevertheless, from time to time, human error can present a threat to system performance, and system designers must anticipate and manage this risk. NASA’s Human-Rating Requirements for Space Systems call for program managers to conduct a human error analysis (HEA) during system development but does not specify how to do this. In 2018, NASA’s Engineering and Safety Center asked the authors to develop a guidance document on HEA. The resulting position paper outlines a suggested method for HEA and makes it clear that error analysis is about identifying and mitigating problems at a system level, and not about finding fault with individuals. Error management strategies must be directed at error-producing conditions, thereby reducing the likelihood of human error, while retaining the positive contribution that humans make to system operations.

human error human-rated space

NASA Space Flight Vehicle Fault Isolation Challenges

The Space Launch System (SLS) is the new NASA heavy lift launch vehicle and is scheduled for its first mission in 2017. The goal of the first mission, which will be uncrewed, is to demonstrate the integrated system performance of the SLS rocket and spacecraft before a crewed flight in 2021. SLS has many of the same logistics challenges as any other large scale program. Common logistics concerns for SLS include integration of discrete programs geographically separated, multiple prime contractors with distinct and different goals, schedule pressures and funding constraints. However, SLS also faces unique challenges. The new program is a confluence of new hardware and heritage, with heritage hardware constituting seventy-five percent of the program. This unique approach to design makes logistics concerns such as testability of the integrated flight vehicle especially problematic. The cost of fully automated diagnostics can be completely justified for a large fleet, but not so for a single flight vehicle. Fault detection is mandatory to assure the vehicle is capable of a safe launch, but fault isolation is another issue. SLS has considered various methods for fault isolation which can provide a reasonable balance between adequacy, timeliness and cost. This paper will address the analyses and decisions the NASA Logistics engineers are making to mitigate risk while providing a reasonable testability solution for fault isolation.

Bramon, Christopher

Hard Fault Protection for a Silicon Carbide-Based Aerospace Motor Drive

Due to increasingly high DC link voltages and further advancements in the current density of silicon carbide (SiC) MOSFETs, it has become evident that conventional IGBT protection methods are not sufficient to protect these devices from overcurrent during low-inductance fault events. The use of an air core Rogowski coil topology was explored to see if it could mitigate these hard fault events. The design of this circuit resulted in safe shutdown of a low impedance phase-tophase fault, tested up to DC link voltages of 1 kV.

High Voltage

Risk Mitigation for Managing On-Orbit Anomalies

This slide presentation reviews strategies for managing risk mitigation that occur with anomalies in on-orbit spacecraft. It reviews the risks associated with mission operations, a diagram of the method used to manage undesirable events that occur which is a closed loop fault analysis and until corrective action is successful. It also reviews the fish bone diagram which is used if greater detail is required and aids in eliminating possible failure factors.

La, Jim

Damage Characterization Using the Extended Finite Element Method for Structural Health Management

The development of validated multidisciplinary Integrated Vehicle Health Management (IVHM) tools, technologies, and techniques to enable detection, diagnosis, prognosis, and mitigation in the presence of adverse conditions during flight will provide effective solutions to deal with safety related challenges facing next generation aircraft. The adverse conditions include loss of control caused by environmental factors, actuator and sensor faults or failures, and damage conditions. A major concern in these structures is the growth of undetected damage/cracks due to fatigue and low velocity foreign impact that can reach a critical size during flight, resulting in loss of control of the aircraft. Hence, development of efficient methodologies to determine the presence, location, and severity of damage/cracks in critical structural components is highly important in developing efficient structural health management systems.

Krishnamurthy, Thiagarajan

An Indirect Adaptive Control Scheme in the Presence of Actuator and Sensor Failures

The problem of controlling a system in the presence of unknown actuator and sensor faults is addressed. The system is assumed to have groups of actuators, and groups of sensors, with each group consisting of multiple redundant similar actuators or sensors. The types of actuator faults considered consist of unknown actuators stuck in unknown positions, as well as reduced actuator effectiveness. The sensor faults considered include unknown biases and outages. The approach employed for fault detection and estimation consists of a bank of Kalman filters based on multiple models, and subsequent control reconfiguration to mitigate the effect of biases caused by failed components as well as to obtain stability and satisfactory performance using the remaining actuators and sensors. Conditions for fault identifiability are presented, and the adaptive scheme is applied to an aircraft flight control example in the presence of actuator failures. Simulation results demonstrate that the method can rapidly and accurately detect faults and estimate the fault values, thus enabling safe operation and acceptable performance in spite of failures.

Sun, Joy Z.

Damage Characterization Method for Structural Health Management Using Reduced Number of Sensor Inputs

The development of validated multidisciplinary Integrated Vehicle Health Management (IVHM) tools, technologies, and techniques to enable detection, diagnosis, prognosis, and mitigation in the presence of adverse conditions during flight will provide effective solutions to deal with safety related challenges facing next generation aircraft. The adverse conditions include loss of control caused by environmental factors, actuator and sensor faults or failures, and damage conditions. A major concern in these structures is the growth of undetected damage (cracks) due to fatigue and low velocity foreign impacts that can reach a critical size during flight, resulting in loss of control of the aircraft. Hence, development of efficient methodologies to determine the presence, location, and severity of damage in critical structural components is highly important in developing efficient structural health management systems.

Krishnamurthy, T.

Error Mitigation of Point-to-Point Communication for Fault-Tolerant Computing

Fault tolerant systems require the ability to detect and recover from physical damage caused by the hardware s environment, faulty connectors, and system degradation over time. This ability applies to military, space, and industrial computing applications. The integrity of Point-to-Point (P2P) communication, between two microcontrollers for example, is an essential part of fault tolerant computing systems. In this paper, different methods of fault detection and recovery are presented and analyzed.

Akamine, Robert L.

Fault-Tolerant, Radiation-Hard DSP

Commercial digital signal processors (DSPs) for use in high-speed satellite computers are challenged by the damaging effects of space radiation, mainly single event upsets (SEUs) and single event functional interrupts (SEFIs). Innovations have been developed for mitigating the effects of SEUs and SEFIs, enabling the use of very-highspeed commercial DSPs with improved SEU tolerances. Time-triple modular redundancy (TTMR) is a method of applying traditional triple modular redundancy on a single processor, exploiting the VLIW (very long instruction word) class of parallel processors. TTMR improves SEU rates substantially. SEFIs are solved by a SEFI-hardened core circuit, external to the microprocessor. It monitors the health of the processor, and if a SEFI occurs, forces the processor to return to performance through a series of escalating events. TTMR and hardened-core solutions were developed for both DSPs and reconfigurable field-programmable gate arrays (FPGAs). This includes advancement of TTMR algorithms for DSPs and reconfigurable FPGAs, plus a rad-hard, hardened-core integrated circuit that services both the DSP and FPGA. Additionally, a combined DSP and FPGA board architecture was fully developed into a rad-hard engineering product. This technology enables use of commercial off-the-shelf (COTS) DSPs in computers for satellite and other space applications, allowing rapid deployment at a much lower cost. Traditional rad-hard space computers are very expensive and typically have long lead times. These computers are either based on traditional rad-hard processors, which have extremely low computational performance, or triple modular redundant (TMR) FPGA arrays, which suffer from power and complexity issues. Even more frustrating is that the TMR arrays of FPGAs require a fixed, external rad-hard voting element, thereby causing them to lose much of their reconfiguration capability and in some cases significant speed reduction. The benefits of COTS high-performance signal processing include significant increase in onboard science data processing, enabling orders of magnitude reduction in required communication bandwidth for science data return, orders of magnitude improvement in onboard mission planning and critical decision making, and the ability to rapidly respond to changing mission environments, thus enabling opportunistic science and orders of magnitude reduction in the cost of mission operations through reduction of required staff. Additional benefits of COTS-based, high-performance signal processing include the ability to leverage considerable commercial and academic investments in advanced computing tools, techniques, and infra structure, and the familiarity of the science and IT community with these computing environments.

Czajkowski, David

Hot-spot investigations of utility scale panel configurations

The causes of array faults and efforts to mitigate their effects are examined. Research is concentrated on the panel for the 900 kw second phase of the Sacramento Municipal Utility District (SMUD) project. The panel is designed for hot spot tolerance without comprising efficiency under normal operating conditions. Series/paralleling internal to each module improves tolerance in the power quadrant to cell short or open circuits. Analtyical methods are developed for predicting worst case shade patterns and calculating the resultant cell temperature. Experiments conducted on a prototype panel support the analytical calculations.

Arnett, J. C.

Probabilistic Risk Assessment for Decision Making During Spacecraft Operations

Decisions made during the operational phase of a space mission often have significant and immediate consequences. Without the explicit consideration of the risks involved and their representation in a solid model, it is very likely that these risks are not considered systematically in trade studies. Wrong decisions during the operational phase of a space mission can lead to immediate system failure whereas correct decisions can help recover the system even from faulty conditions. A problem of special interest is the determination of the system fault protection strategies upon the occurrence of faults within the system. Decisions regarding the fault protection strategy also heavily rely on a correct understanding of the state of the system and an integrated risk model that represents the various possible scenarios and their respective likelihoods. Probabilistic Risk Assessment (PRA) modeling is applicable to the full lifecycle of a space mission project, from concept development to preliminary design, detailed design, development and operations. The benefits and utilities of the model, however, depend on the phase of the mission for which it is used. This is because of the difference in the key strategic decisions that support each mission phase. The focus of this paper is on describing the particular methods used for PRA modeling during the operational phase of a spacecraft by gleaning insight from recently conducted case studies on two operational Mars orbiters. During operations, the key decisions relate to the commands sent to the spacecraft for any kind of diagnostics, anomaly resolution, trajectory changes, or planning. Often, faults and failures occur in the parts of the spacecraft but are contained or mitigated before they can cause serious damage. The failure behavior of the system during operations provides valuable data for updating and adjusting the related PRA models that are built primarily based on historical failure data. The PRA models, in turn, provide insight into the effect of various faults or failures on the risk and failure drivers of the system and the likelihood of possible end case scenarios, thereby facilitating the decision making process during operations. This paper describes the process of adjusting PRA models based on observed spacecraft data, on one hand, and utilizing the models for insight into the future system behavior on the other hand. While PRA models are typically used as a decision aid during the design phase of a space mission, we advocate adjusting them based on the observed behavior of the spacecraft and utilizing them for decision support during the operations phase.

dynamic fault trees