Search NASASearch

SEARCH · Search NASA

Results for “multiple faults”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 181 records · Page 10

High-Performance Computing Based EMT Simulation: Power Grid with IBRs

Electromagnetic transient (EMT) simulation of power grids with high-fidelity models of inverter-based resources (IBRs) is time-consuming and difficult to scale. The necessity for high-fidelity models of IBRs that incorporate the dynamics of individual inverters within IBRs has been showcased in recent studies. These studies focused on events with partial power reduction in each IBR during a transmission line fault in the power grid. These types of events have been documented in multiple North American Electric Reliability Council (NERC) reports in the past decade. It is imperative then to find solutions to speed-up EMT simulations and scale the size of the region with IBRs studied in EMT simulations. In this paper, a combination of numerical simulation algorithms with high-performance computing techniques are employed in discretization and linear solvers employed in the proposed RE-INTEGRATE EMT simulation platform for power grid with IBRs. For ease of scalability, modular and object-oriented programming is used as these techniques are implemented. Additionally, automation software is developed to convert legacy software codes to the proposed RE-INTEGRATE EMT simulation platform. Thereafter, this platform is evaluated on multi-core central processing units (CPUs). Finally, scale-up tests are performed to showcase the scalability that is possible.

Marthi, Phani Ratna Vanamali [ORNL] (ORCID:0000000

Quantifying fault recovery in multiprocessor systems

Various aspects of reliable computing are formalized and quantified with emphasis on efficient fault recovery. The mathematical model which proves to be most appropriate is provided by the theory of graphs. New measures for fault recovery are developed and the value of elements of the fault recovery vector are observed to depend not only on the computation graph H and the architecture graph G, but also on the specific location of a fault. In the examples, a hypercube is chosen as a representative of parallel computer architecture, and a pipeline as a typical configuration for program execution. Dependability qualities of such a system is defined with or without a fault. These qualities are determined by the resiliency triple defined by three parameters: multiplicity, robustness, and configurability. Parameters for measuring the recovery effectiveness are also introduced in terms of distance, time, and the number of new, used, and moved nodes and edges.

Malek, Miroslaw

Multichannel Error Correction Code Decoder

NASA Lewis Research Center's Digital Systems Technology Branch has an ongoing program in modulation, coding, onboard processing, and switching. Recently, NASA completed a project to incorporate a time-shared decoder into the very-small-aperture terminal (VSAT) onboard-processing mesh architecture. The primary goal was to demonstrate a time-shared decoder for a regenerative satellite that uses asynchronous, frequency-division multiple access (FDMA) uplink channels, thereby identifying hardware and power requirements and fault-tolerant issues that would have to be addressed in a operational system. A secondary goal was to integrate and test, in a system environment, two NASA-sponsored, proof-of-concept hardware deliverables: the Harris Corp. high-speed Bose Chaudhuri-Hocquenghem (BCH) codec and the TRW multichannel demultiplexer/demodulator (MCDD). A beneficial byproduct of this project was the development of flexible, multichannel-uplink signal-generation equipment.

Source record

IBM powerPC 405 SEU mitigation using processor voting techniques in Xilinx Virtex-I1 pro FPGA

Not until recently, Xilinx has developed a new field programmable gate array (FPGA) device family, Virtex-I1 Pro. In this single device, not only dies it have density logic cells (3K to125K), gigabit connectivity, on chip memory, digital clock management, but also it can have up to four IBM PowerPC 405 Processor hard cores, running up to 400MHz and 633 Mbps. To utilize this cutting edge device in space applications, a few Single Event Upset (SEU) mitigation techniques need to be implemented to a design for the device. At Jet Propulsion Laboratory (JPL), we have successfully demonstrated the feasibility of running multiple processors running in a lock step fashion to accomplish SEU mitigation and fault tolerance.

single event upset (SEU)

NASA Tech Briefs, November 2008

Topics covered include: Digital Phase Meter for a Laser Heterodyne Interferometer; Vision System Measures Motions of Robot and External Objects; Advanced Precipitation Radar Antenna to Measure Rainfall From Space; Wide-Band Radar for Measuring Thickness of Sea Ice; Vertical Isolation for Photodiodes in CMOS Imagers; Wide-Band Microwave Receivers Using Photonic Processing; L-Band Transmit/Receive Module for Phase-Stable Array Antennas; Microwave Power Combiner/Switch Utilizing a Faraday Rotator; Compact Low-Loss Planar Magic-T; Using Pipelined XNOR Logic to Reduce SEU Risks in State Machines; Quasi-Optical Transmission Line for 94-GHz Radar; Next Generation Flight Controller Trainer System; Converting from DDOR SASF to APF; Converting from CVF to AAF; Documenting AUTOGEN and APGEN Model Files; Sequence History Update Tool; Extraction and Analysis of Display Data; MRO DKF Post-Processing Tool; Rig Diagnostic Tools; MRO Sequence Checking Tool; Science Activity Planner for the MER Mission; UAVSAR Flight-Planning System; Templates for Deposition of Microscopic Pointed Structures; Adjustable Membrane Mirrors Incorporating G-Elastomers; Hall-Effect Thruster Utilizing Bismuth as Propellant; High-Temperature Crystal-Growth Cartridge Tubes Made by VPS; Quench Crucibles Reinforced with Metal; Deep-Sea Hydrothermal-Vent Sampler; Mars Rocket Propulsion System; Two-Stage Passive Vibration Isolator; Improved Thermal Design of a Compression Mold; Enhanced Pseudo-Waypoint Guidance for Spacecraft Maneuvers; Altimetry Using GPS-Reflection/Occultation Interferometry; Thermally Driven Josephson Effect; Perturbation Effects on a Supercritical C7H16/N2 Mixing Layer; Gold Nanoparticle Labels Amplify Ellipsometric Signals; Phase Matching of Diverse Modes in a WGM Resonator; WGM Resonators for Terahertz-to-Optical Frequency Conversion; Determining Concentration of Nanoparticles from Ellipsometry; Microwave-to-Optical Conversion in WGM Resonators; Four-Pass Coupler for Laser-Diode-Pumped Solid-State Laser; Low-Resolution Raman-Spectroscopy Combustion Thermometry; Temperature Sensors Based on WGM Optical Resonators; Varying the Divergence of Multiple Parallel Laser Beams; Efficient Algorithm for Rectangular Spiral Search; Algorithm-Based Fault Tolerance Integrated with Replication; Targeting and Localization for Mars Rover Operations; Terrain-Adaptive Navigation Architecture; Self-Adjusting Hash Tables for Embedded Flight Applications; Schema for Spacecraft-Command Dictionary; Combined GMSK Communications and PN Ranging; System-Level Integration of Mass Memory; Network-Attached Solid-State Recorder Architecture; Method of Cross-Linking Aerogels Using a One-Pot Reaction Scheme; An Efficient Reachability Analysis Algorithm.

Source record

ASTERIA Operations Demonstrates the Value of Combining the Mission Assurance and Fault Protection Roles on CubeSats

On November 20, 2017, ASTERIA (Arcsecond Space Telescope Enabling Research in Astrophysics), a 6U CubeSat performing a technology demonstration of astrophysical measurements, deployed from the ISS. The technology demonstration goals to achieve precision photometry via arcsecond-level line-of-sight pointing error and highly stable focal plane temperature control were met by February 2018. Extended mission operations are ongoing, with the primary focus on observing nearby stars for transiting exoplanets. Throughout development and operations, the roles of mission assurance and fault protection have proven critical to achieving the primary technical goals and to maintaining a healthy spacecraft through multiple extended missions. Given the budget and schedule constraints typical of a CubeSat, innovative tailoring of processes has been critical to success throughout both development and operations of ASTERIA. Mission assurance plays an important role in identifying and evaluating risk and developing cost-effective mitigations. Flexibility in the fault protection design offers a variety of options for implementing risk mitigations as risks have been uncovered both in pre-delivery testing and in mission operations. This paper will discuss the approach taken on ASTERIA to implement mission assurance and fault protection and the resulting benefits to operational efficiency and success. It will briefly address the advantages of this approach during development, in which the combination of the roles provided mission assurance significant insight to system risks, which feeds back into testing methodologies and directly into fault protection design. Operations will be discussed in detail. During this phase, the roles merge to identify in-flight fault protection updates to efficiently respond to anomalies and improve the likelihood of successful technology demonstrations. The paper will also detail the tools that are used to analyse data, identify anomalies, and develop the updates to uplink to the spacecraft. Finally, the general operational approach will be discussed to highlight the usefulness of the ASTERIA processes and their applicability to future CubeSat missions.

Knapp, Mary

Performance assessment of near-fault buildings subjected to physics-based simulated earthquake ground motions with fling step

The effects of the co-seismic static offset (known as fling step) and associated velocity pulses on civil structures have been difficult to study because the static offset is typically removed during the processing of earthquake ground motion records. Simulated ground motions contain fling features and require no processing; therefore, they create new opportunities for representing fling features in seismic hazard analysis and assessing their influence on the seismic demands on near-fault structures. We use physics-based fault rupture simulations to study the characteristics of ground motions with fling step and the sensitivity of the near-fault structural demands to strong fling features. We uncover that simulated ground motions with a large fling step tend to have higher spectral intensity than those without a fling step at the same rupture distance, especially at periods longer than 2 s. As a result, the structural demands on flexible buildings tend to be the most sensitive to the fling features. Statistical analysis suggests that the ground motion spectral shape (represented by spectral accelerations at multiple periods) is—in most cases—a sufficient predictor of the structural demands on near-fault low-rise and mid-rise buildings at locations that are susceptible to strong fling effects. Finally, ground motion record selection experiments reveal that representing the spectral shape features at periods that are most relevant to a given structure may be an effective strategy to reduce the bias in the estimated demands on near-fault long-period structures when the available database of records is considered deficient in fling features.

Fling step

Qualitative Event-Based Fault Isolation under Uncertain Observations

For many systems, automatic fault diagnosis is critical to ensuring safe and efficient operation. Fault isolation is performed by analyzing measured signals from the system, and reasoning over the system behavior to determine which faults have occurred, based on models of predicted faulty behavior. For dynamic systems, reasoning may be performed using qualitative analysis of the differences between measured signals and their predicted values, in which observations take the form of qualitative symbols. Such an approach is quick to isolate faults, but depends critically on correct generation of the qualitative symbols from the signals. In this paper, we develop an approach to qualitative event-based fault isolation for dynamic systems that is robust to incorrect qualitative observations. Observations are treated as uncertain, where multiple interpretations of an observation, each with its own probability, are considered. By interpreting observed symbols in a probabilistic manner, the approach degrades gracefully as the number of incorrectly-generated symbols increases. The approach is demonstrated on an electrical power system testbed, and experiments using real data obtained from the hardware demonstrate the improved fault isolation performance in the presence of incorrect symbol generation.

Daigle, Matthew

Supporting Hazard Analysis for Wildfire Response Using fmdtools and MIKA

The System Wide Safety (SWS) Safety Demonstrator (SD) Series drives development of an increasingly capable In-Time Aviation Safety Management System (IASMS) focusing on humanitarian applications, starting with wildfire response (SD-1). The goals of this report are to (1) provide an early hazard analysis and mitigation evaluation of wildfire response to support these efforts and (2) provide a demonstration of capabilities of the Fault Model Design Tools (fmdtools) and Manager for Intelligent Knowledge Access (MIKA) tools. fmdtools provides a modeling, simulation, and resiliency analysis framework in which a wildfire response model, the System Modeling and Analysis of Resiliency in Scalable Traffic Management for Emergency Response Operations (SMARt-STEReO), is built. MIKA is an intelligent knowledge manager with several capabilities, including assisting in hazard analysis by extracting and analyzing hazards from historical incident reports. The following topics are covered in the report: Understanding Wildfire Hazard Dynamics. We provide a description and simulated examples of how hazards occur in the SMARt-STEReO model of wildfire response and their effect on its outcome. This provides a common mental model and focuses the analysis presented in the remainder of the report. Wildfire Hazard Identification. MIKA identifies wildfire hazards from three relevant datasets: the ICS-209-PLUS, SAFECOM, and SAFENET. Hazards are manually organized into a taxonomy and MIKA analyzes each hazard’s effects, likelihood, severity, and risk. Evaluating Mitigation Strategies. The SMARt-STEReO wildfire response model built in fmdtools evaluates a subset of identified hazards. Specifically, we simulate the effect of communications faults and equipment faults on operator safety, the effect of changing winds and flammability, and a scenario with multiple ignition points and heavy smoke. Tool Limitations and Usage Considerations. We provide a discussion of appropriate tool use cases as well as limitations and considerations for usage. The tool findings are used to synthesize recommendations for wildfire response operations, which can be captured as part of an IASMS. Key recommendations are as follows: Hazards are identified from a broad spectrum of sources including aircraft subsystems, operational sources, and ground crew operations. Highest risk operational environment hazards identified are Evacuations. The highest risk manned aerial operations hazard categorized is Jumper Operations Mishap. Ground crew hazards that are highest risk are Burns, Cargo Operations Overhead, Dehydration, Entrapment, Falling Objects, Heart Attacks, Heat Exhaustion, Inadequate Training or Certification, Vehicle Breakdown, and Vehicle Collision. Modelled containment failures arise from a mismatch between the difficulty of the firefighting scenario and the capacity (e.g., speed, effectiveness, awareness) of the response. In firefighting scenarios where containment is possible (e.g., because the fire does not spread too quickly), these mismatches can occur because of a change in environmental conditions (e.g., wind, flammability, etc) or because of planning, equipment, or communications faults. Improvements to communications increase the capacity of the firefighting response by reducing the time needed to respond to the fire. While surveillance does not increase this capacity by itself, it increases operator safety by increasing state awareness, enabling firefighters to evade approaching fires. Increasing both has a synergistic effect. In general, these performance and resilience increases generalize over fault scenarios as well as unforeseen changes to circumstances (i.e., wind, aridity, etc.). However, these improvements need to be designed so as not to make the system prone to persistent large-scale communications outages, which can reduce performance.

Hazard analysis

Reconfiguration algorithms for tree architectures using sub-tree oriented fault tolerance

An approach to reconfiguration in tree architectures has been developed in which redundant processors are allocated at the leaves. The scheme is called sub-tree oriented fault tolerances (SOFT) and is capable of tolerating both link failures as well as multiple processor failures. In this paper, the SOFT scheme is examined from the perspective of reconfigurability. Specific algorithms are presented for reconfiguration.

Lowrie, M. B.

An experimental evaluation of software redundancy as a strategy for improving reliability

The strategy of using multiple versions of independently developed software as a means to tolerate residual software design faults is suggested by the success of hardware redundancy for tolerating hardware failures. Although, as generally accepted, the independence of hardware failures resulting from physical wearout can lead to substantial increases in reliability for redundant hardware structures, a similar conclusion is not immediate for software. The degree to which design faults are manifested as independent failures determines the effectiveness of redundancy as a method for improving software reliability. Interest in multi-version software centers on whether it provides an adequate measure of increased reliability to warrant its use in critical applications. The effectiveness of multi-version software is studied by comparing estimates of the failure probabilities of these systems with the failure probabilities of single versions. The estimates are obtained under a model of dependent failures and compared with estimates obtained when failures are assumed to be independent. The experimental results are based on twenty versions of an aerospace application developed and certified by sixty programmers from four universities. Descriptions of the application, development and certification processes, and operational evaluation are given together with an analysis of the twenty versions.

Eckhardt, Dave E., Jr.

A report on SHARP (Spacecraft Health Automated Reasoning Prototype) and the Voyager Neptune encounter

The development and application of the Spacecraft Health Automated Reasoning Prototype (SHARP) for the operations of the telecommunications systems and link analysis functions in Voyager mission operations are presented. An overview is provided of the design and functional description of the SHARP system as it was applied to Voyager. Some of the current problems and motivations for automation in real-time mission operations are discussed, as are the specific solutions that SHARP provides. The application of SHARP to Voyager telecommunications had the goal of being a proof-of-capability demonstration of artificial intelligence as applied to the problem of real-time monitoring functions in planetary mission operations. AS part of achieving this central goal, the SHARP application effort was also required to address the issue of the design of an appropriate software system architecture for a ground-based, highly automated spacecraft monitoring system for mission operations, including methods for: (1) embedding a knowledge-based expert system for fault detection, isolation, and recovery within this architecture; (2) acquiring, managing, and fusing the multiple sources of information used by operations personnel; and (3) providing information-rich displays to human operators who need to exercise the capabilities of the automated system. In this regard, SHARP has provided an excellent example of how advanced artificial intelligence techniques can be smoothly integrated with a variety of conventionally programmed software modules, as well as guidance and solutions for many questions about automation in mission operations.

Martin, R. G.

An experimental evaluation of software redundancy as a strategy for improving reliability

The strategy of using multiple versions of independently developed software as a means to tolerate residual software design faults is suggested by the success of hardware redundancy for tolerating hardware failires. Although, as generally accepted, the independence of hardware failures resulting from physical wearout can lead to substantial increases in reliability for redundant hardware structures, a similar conclusion is not immediate for software. The degree to which design faults are manifested as independent failures determines the effectiveness of redundancy as a method for improving software reliability. Interest in multi-version software centers on whether it provides an adequate measure of increased reliability to warrant its use in critical applications. The effectiveness of multi-version software is studied by comparing estimates of the failure probabilities of these systems with the failure probabilities of single versions. The estimates are obtained under a model of dependent failures and compared with the estimates obtained when failures are assumed to be independent. The experimental results are based on twenty versions of an aerospace application developed and certified by sixty programmers from four universities. Descriptions of the application, development and certifications processes, and operational evaluation are given together with an analysis of the twenty versions.

Eckhardt, Dave E.

Symbolic discrete event system specification

Extending discrete event modeling formalisms to facilitate greater symbol manipulation capabilities is important to further their use in intelligent control and design of high autonomy systems. An extension to the DEVS formalism that facilitates symbolic expression of event times by extending the time base from the real numbers to the field of linear polynomials over the reals is defined. A simulation algorithm is developed to generate the branching trajectories resulting from the underlying nondeterminism. To efficiently manage symbolic constraints, a consistency checking algorithm for linear polynomial constraints based on feasibility checking algorithms borrowed from linear programming has been developed. The extended formalism offers a convenient means to conduct multiple, simultaneous explorations of model behaviors. Examples of application are given with concentration on fault model analysis.

Zeigler, Bernard P.

A Decentralized Adaptive Approach to Fault Tolerant Flight Control

This paper briefly reports some results of our study on the application of a decentralized adaptive control approach to a 6 DOF nonlinear aircraft model. The simulation results showed the potential of using this approach to achieve fault tolerant control. Based on this observation and some analysis, the paper proposes a multiple channel adaptive control scheme that makes use of the functionally redundant actuating and sensing capabilities in the model, and explains how to implement the scheme to tolerate actuator and sensor failures. The conditions, under which the scheme is applicable, are stated in the paper.

Wu, N. Eva

Actuator and Motor Control End-to-End V&V on the Mars 2020 Rover

The Mars 2020 Perseverance rover is the most advanced robotic exploration system ever sent to another planet. To support the complex scientific and mobility needs of the mission, the rover utilizes 33 actuators, three multi-degree-of-freedom force-torque sensors, fifteen single or dual-speed resolvers, two solenoid valves, and twelve contact switches. The control for these actuators and sensors is achieved by several levels of flight software, coordinated between two computers with varying bandwidth control loops. Furthermore, the actuators and sensors were integrated into multiple larger robotic mechanisms that were delivered by different organizations at various points in the Integration and Test (I&T) timeline. All of this created a very complex Verification and Validation (V&V) scenario involving multiple subsystems and teams, several hardware and software testbeds with varying levels of fidelity, and significant systems engineering to ensure the overall I&T schedule could be maintained while ensuring system hardware safety.This paper details the integrated V&V effort across multiple teams and venues to provide full coverage of all necessary functionality, performance, and fault protection. First, it provides an overview of how the V&V campaign was subdivided among teams and venues and provides descriptions of the various hardware configurations used to support the testing. The Mars 2020 implementation of the plan incorporates many of the lessons learned from Mars Science Laboratory’s test campaign, and these value-added modifications are discussed here. Also included in this section is the system-level environmental testing approach used for mechanisms. Second, the paper describes the phased approach used by the teams to support new hardware and software deliveries to testbed and Systems I&T. In this approach the test campaign was built upon higher-level mechanism needs for performance, functionality, and safety at specific times in the campaign. Finally, the paper discusses lessons learned from the V&V campaign that should be applied to future large-scale motion control testing efforts.

Borne, Davis

Fault Tolerance Middleware for a Multi-Core System

Fault Tolerance Middleware (FTM) provides a framework to run on a dedicated core of a multi-core system and handles detection of single-event upsets (SEUs), and the responses to those SEUs, occurring in an application running on multiple cores of the processor. This software was written expressly for a multi-core system and can support different kinds of fault strategies, such as introspection, algorithm-based fault tolerance (ABFT), and triple modular redundancy (TMR). It focuses on providing fault tolerance for the application code, and represents the first step in a plan to eventually include fault tolerance in message passing and the FTM itself. In the multi-core system, the FTM resides on a single, dedicated core, separate from the cores used by the application. This is done in order to isolate the FTM from application faults and to allow it to swap out any application core for a substitute. The structure of the FTM consists of an interface to a fault tolerant strategy module, a responder module, a fault manager module, an error factory, and an error mapper that determines the severity of the error. In the present reference implementation, the only fault tolerant strategy implemented is introspection. The introspection code waits for an application node to send an error notification to it. It then uses the error factory to create an error object, and at this time, a severity level is assigned to the error. The introspection code uses its built-in knowledge base to generate a recommended response to the error. Responses might include ignoring the error, logging it, rolling back the application to a previously saved checkpoint, swapping in a new node to replace a bad one, or restarting the application. The original error and recommended response are passed to the top-level fault manager module, which invokes the response. The responder module also notifies the introspection module of the generated response. This provides additional information to the introspection module that it can use in generating its next response. For example, if the responder triggers an application rollback and errors are still occurring, the introspection module may decide to recommend an application restart.

Some, Raphael R.

Locating Anomalies in Complex Data Sets Using Visualization and Simulation

The research goals are to create a simulation framework that can accept any combination of models written at the gate or behavioral level. The framework provides the ability to fault simulate and create scenarios of experiments using concurrent simulation. In order to meet these goals we have had to fulfill the following requirements. The ability to accept models written in VHDL, Verilog or the C languages. The ability to propagate faults through any model type. The ability to create experiment scenarios efficiently without generating every possible combination of variables. The ability to accept adversity of fault models beyond the single stuck-at model. Major development has been done to develop a parser that can accept models written in various languages. This work has generated considerable attention from other universities and industry for its flexibility and usefulness. The parser uses LEXX and YACC to parse Verilog and C. We have also utilized our industrial partnership with Alternative System's Inc. to import vhdl into our simulator. For multilevel simulation, we needed to modify the simulator architecture to accept models that contained multiple outputs. This enabled us to accept behavioral components. The next major accomplishment was the addition of "functional fault models". Functional fault models change the behavior of a gate or model. For example, a bridging fault can make an OR gate behave like an AND gate. This has applications beyond fault simulation. This modeling flexibility will make the simulator more useful for doing verification and model comparison. For instance, two or more versions of an ALU can be comparatively simulated in a single execution. The results will show where and how the models differed so that the performance and correctness of the models may be evaluated. A considerable amount of time has been dedicated to validating the simulator performance on larger models provided by industry and other universities.

Panetta, Karen