Search NASA⌕ Search

SEARCH · Search NASA

Results for “multiple faults”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 163 records · Page 9

ASTERIA Operations Demonstrates the Value of Combining the Mission Assurance and Fault Protection Roles on CubeSats

On November 20, 2017, ASTERIA (Arcsecond Space Telescope Enabling Research in Astrophysics), a 6U CubeSat performing a technology demonstration of astrophysical measurements, deployed from the ISS. The technology demonstration goals to achieve precision photometry via arcsecond-level line-of-sight pointing error and highly stable focal plane temperature control were met by February 2018. Extended mission operations are ongoing, with the primary focus on observing nearby stars for transiting exoplanets. Throughout development and operations, the roles of mission assurance and fault protection have proven critical to achieving the primary technical goals and to maintaining a healthy spacecraft through multiple extended missions. Given the budget and schedule constraints typical of a CubeSat, innovative tailoring of processes has been critical to success throughout both development and operations of ASTERIA. Mission assurance plays an important role in identifying and evaluating risk and developing cost-effective mitigations. Flexibility in the fault protection design offers a variety of options for implementing risk mitigations as risks have been uncovered both in pre-delivery testing and in mission operations. This paper will discuss the approach taken on ASTERIA to implement mission assurance and fault protection and the resulting benefits to operational efficiency and success. It will briefly address the advantages of this approach during development, in which the combination of the roles provided mission assurance significant insight to system risks, which feeds back into testing methodologies and directly into fault protection design. Operations will be discussed in detail. During this phase, the roles merge to identify in-flight fault protection updates to efficiently respond to anomalies and improve the likelihood of successful technology demonstrations. The paper will also detail the tools that are used to analyse data, identify anomalies, and develop the updates to uplink to the spacecraft. Finally, the general operational approach will be discussed to highlight the usefulness of the ASTERIA processes and their applicability to future CubeSat missions.

Knapp, Mary↗

Qualitative Event-Based Fault Isolation under Uncertain Observations

For many systems, automatic fault diagnosis is critical to ensuring safe and efficient operation. Fault isolation is performed by analyzing measured signals from the system, and reasoning over the system behavior to determine which faults have occurred, based on models of predicted faulty behavior. For dynamic systems, reasoning may be performed using qualitative analysis of the differences between measured signals and their predicted values, in which observations take the form of qualitative symbols. Such an approach is quick to isolate faults, but depends critically on correct generation of the qualitative symbols from the signals. In this paper, we develop an approach to qualitative event-based fault isolation for dynamic systems that is robust to incorrect qualitative observations. Observations are treated as uncertain, where multiple interpretations of an observation, each with its own probability, are considered. By interpreting observed symbols in a probabilistic manner, the approach degrades gracefully as the number of incorrectly-generated symbols increases. The approach is demonstrated on an electrical power system testbed, and experiments using real data obtained from the hardware demonstrate the improved fault isolation performance in the presence of incorrect symbol generation.

Daigle, Matthew↗

Supporting Hazard Analysis for Wildfire Response Using fmdtools and MIKA

The System Wide Safety (SWS) Safety Demonstrator (SD) Series drives development of an increasingly capable In-Time Aviation Safety Management System (IASMS) focusing on humanitarian applications, starting with wildfire response (SD-1). The goals of this report are to (1) provide an early hazard analysis and mitigation evaluation of wildfire response to support these efforts and (2) provide a demonstration of capabilities of the Fault Model Design Tools (fmdtools) and Manager for Intelligent Knowledge Access (MIKA) tools. fmdtools provides a modeling, simulation, and resiliency analysis framework in which a wildfire response model, the System Modeling and Analysis of Resiliency in Scalable Traffic Management for Emergency Response Operations (SMARt-STEReO), is built. MIKA is an intelligent knowledge manager with several capabilities, including assisting in hazard analysis by extracting and analyzing hazards from historical incident reports. The following topics are covered in the report: Understanding Wildfire Hazard Dynamics. We provide a description and simulated examples of how hazards occur in the SMARt-STEReO model of wildfire response and their effect on its outcome. This provides a common mental model and focuses the analysis presented in the remainder of the report. Wildfire Hazard Identification. MIKA identifies wildfire hazards from three relevant datasets: the ICS-209-PLUS, SAFECOM, and SAFENET. Hazards are manually organized into a taxonomy and MIKA analyzes each hazard’s effects, likelihood, severity, and risk. Evaluating Mitigation Strategies. The SMARt-STEReO wildfire response model built in fmdtools evaluates a subset of identified hazards. Specifically, we simulate the effect of communications faults and equipment faults on operator safety, the effect of changing winds and flammability, and a scenario with multiple ignition points and heavy smoke. Tool Limitations and Usage Considerations. We provide a discussion of appropriate tool use cases as well as limitations and considerations for usage. The tool findings are used to synthesize recommendations for wildfire response operations, which can be captured as part of an IASMS. Key recommendations are as follows: Hazards are identified from a broad spectrum of sources including aircraft subsystems, operational sources, and ground crew operations. Highest risk operational environment hazards identified are Evacuations. The highest risk manned aerial operations hazard categorized is Jumper Operations Mishap. Ground crew hazards that are highest risk are Burns, Cargo Operations Overhead, Dehydration, Entrapment, Falling Objects, Heart Attacks, Heat Exhaustion, Inadequate Training or Certification, Vehicle Breakdown, and Vehicle Collision. Modelled containment failures arise from a mismatch between the difficulty of the firefighting scenario and the capacity (e.g., speed, effectiveness, awareness) of the response. In firefighting scenarios where containment is possible (e.g., because the fire does not spread too quickly), these mismatches can occur because of a change in environmental conditions (e.g., wind, flammability, etc) or because of planning, equipment, or communications faults. Improvements to communications increase the capacity of the firefighting response by reducing the time needed to respond to the fire. While surveillance does not increase this capacity by itself, it increases operator safety by increasing state awareness, enabling firefighters to evade approaching fires. Increasing both has a synergistic effect. In general, these performance and resilience increases generalize over fault scenarios as well as unforeseen changes to circumstances (i.e., wind, aridity, etc.). However, these improvements need to be designed so as not to make the system prone to persistent large-scale communications outages, which can reduce performance.

Hazard analysis↗

Reconfiguration algorithms for tree architectures using sub-tree oriented fault tolerance

An approach to reconfiguration in tree architectures has been developed in which redundant processors are allocated at the leaves. The scheme is called sub-tree oriented fault tolerances (SOFT) and is capable of tolerating both link failures as well as multiple processor failures. In this paper, the SOFT scheme is examined from the perspective of reconfigurability. Specific algorithms are presented for reconfiguration.

Lowrie, M. B.↗

An experimental evaluation of software redundancy as a strategy for improving reliability

The strategy of using multiple versions of independently developed software as a means to tolerate residual software design faults is suggested by the success of hardware redundancy for tolerating hardware failures. Although, as generally accepted, the independence of hardware failures resulting from physical wearout can lead to substantial increases in reliability for redundant hardware structures, a similar conclusion is not immediate for software. The degree to which design faults are manifested as independent failures determines the effectiveness of redundancy as a method for improving software reliability. Interest in multi-version software centers on whether it provides an adequate measure of increased reliability to warrant its use in critical applications. The effectiveness of multi-version software is studied by comparing estimates of the failure probabilities of these systems with the failure probabilities of single versions. The estimates are obtained under a model of dependent failures and compared with estimates obtained when failures are assumed to be independent. The experimental results are based on twenty versions of an aerospace application developed and certified by sixty programmers from four universities. Descriptions of the application, development and certification processes, and operational evaluation are given together with an analysis of the twenty versions.

Eckhardt, Dave E., Jr.↗

A report on SHARP (Spacecraft Health Automated Reasoning Prototype) and the Voyager Neptune encounter

The development and application of the Spacecraft Health Automated Reasoning Prototype (SHARP) for the operations of the telecommunications systems and link analysis functions in Voyager mission operations are presented. An overview is provided of the design and functional description of the SHARP system as it was applied to Voyager. Some of the current problems and motivations for automation in real-time mission operations are discussed, as are the specific solutions that SHARP provides. The application of SHARP to Voyager telecommunications had the goal of being a proof-of-capability demonstration of artificial intelligence as applied to the problem of real-time monitoring functions in planetary mission operations. AS part of achieving this central goal, the SHARP application effort was also required to address the issue of the design of an appropriate software system architecture for a ground-based, highly automated spacecraft monitoring system for mission operations, including methods for: (1) embedding a knowledge-based expert system for fault detection, isolation, and recovery within this architecture; (2) acquiring, managing, and fusing the multiple sources of information used by operations personnel; and (3) providing information-rich displays to human operators who need to exercise the capabilities of the automated system. In this regard, SHARP has provided an excellent example of how advanced artificial intelligence techniques can be smoothly integrated with a variety of conventionally programmed software modules, as well as guidance and solutions for many questions about automation in mission operations.

Martin, R. G.↗

An experimental evaluation of software redundancy as a strategy for improving reliability

The strategy of using multiple versions of independently developed software as a means to tolerate residual software design faults is suggested by the success of hardware redundancy for tolerating hardware failires. Although, as generally accepted, the independence of hardware failures resulting from physical wearout can lead to substantial increases in reliability for redundant hardware structures, a similar conclusion is not immediate for software. The degree to which design faults are manifested as independent failures determines the effectiveness of redundancy as a method for improving software reliability. Interest in multi-version software centers on whether it provides an adequate measure of increased reliability to warrant its use in critical applications. The effectiveness of multi-version software is studied by comparing estimates of the failure probabilities of these systems with the failure probabilities of single versions. The estimates are obtained under a model of dependent failures and compared with the estimates obtained when failures are assumed to be independent. The experimental results are based on twenty versions of an aerospace application developed and certified by sixty programmers from four universities. Descriptions of the application, development and certifications processes, and operational evaluation are given together with an analysis of the twenty versions.

Eckhardt, Dave E.↗

Symbolic discrete event system specification

Extending discrete event modeling formalisms to facilitate greater symbol manipulation capabilities is important to further their use in intelligent control and design of high autonomy systems. An extension to the DEVS formalism that facilitates symbolic expression of event times by extending the time base from the real numbers to the field of linear polynomials over the reals is defined. A simulation algorithm is developed to generate the branching trajectories resulting from the underlying nondeterminism. To efficiently manage symbolic constraints, a consistency checking algorithm for linear polynomial constraints based on feasibility checking algorithms borrowed from linear programming has been developed. The extended formalism offers a convenient means to conduct multiple, simultaneous explorations of model behaviors. Examples of application are given with concentration on fault model analysis.

Zeigler, Bernard P.↗

A Decentralized Adaptive Approach to Fault Tolerant Flight Control

This paper briefly reports some results of our study on the application of a decentralized adaptive control approach to a 6 DOF nonlinear aircraft model. The simulation results showed the potential of using this approach to achieve fault tolerant control. Based on this observation and some analysis, the paper proposes a multiple channel adaptive control scheme that makes use of the functionally redundant actuating and sensing capabilities in the model, and explains how to implement the scheme to tolerate actuator and sensor failures. The conditions, under which the scheme is applicable, are stated in the paper.

Wu, N. Eva↗

Actuator and Motor Control End-to-End V&V on the Mars 2020 Rover

The Mars 2020 Perseverance rover is the most advanced robotic exploration system ever sent to another planet. To support the complex scientific and mobility needs of the mission, the rover utilizes 33 actuators, three multi-degree-of-freedom force-torque sensors, fifteen single or dual-speed resolvers, two solenoid valves, and twelve contact switches. The control for these actuators and sensors is achieved by several levels of flight software, coordinated between two computers with varying bandwidth control loops. Furthermore, the actuators and sensors were integrated into multiple larger robotic mechanisms that were delivered by different organizations at various points in the Integration and Test (I&T) timeline. All of this created a very complex Verification and Validation (V&V) scenario involving multiple subsystems and teams, several hardware and software testbeds with varying levels of fidelity, and significant systems engineering to ensure the overall I&T schedule could be maintained while ensuring system hardware safety.This paper details the integrated V&V effort across multiple teams and venues to provide full coverage of all necessary functionality, performance, and fault protection. First, it provides an overview of how the V&V campaign was subdivided among teams and venues and provides descriptions of the various hardware configurations used to support the testing. The Mars 2020 implementation of the plan incorporates many of the lessons learned from Mars Science Laboratory’s test campaign, and these value-added modifications are discussed here. Also included in this section is the system-level environmental testing approach used for mechanisms. Second, the paper describes the phased approach used by the teams to support new hardware and software deliveries to testbed and Systems I&T. In this approach the test campaign was built upon higher-level mechanism needs for performance, functionality, and safety at specific times in the campaign. Finally, the paper discusses lessons learned from the V&V campaign that should be applied to future large-scale motion control testing efforts.

Borne, Davis↗

Fault Tolerance Middleware for a Multi-Core System

Fault Tolerance Middleware (FTM) provides a framework to run on a dedicated core of a multi-core system and handles detection of single-event upsets (SEUs), and the responses to those SEUs, occurring in an application running on multiple cores of the processor. This software was written expressly for a multi-core system and can support different kinds of fault strategies, such as introspection, algorithm-based fault tolerance (ABFT), and triple modular redundancy (TMR). It focuses on providing fault tolerance for the application code, and represents the first step in a plan to eventually include fault tolerance in message passing and the FTM itself. In the multi-core system, the FTM resides on a single, dedicated core, separate from the cores used by the application. This is done in order to isolate the FTM from application faults and to allow it to swap out any application core for a substitute. The structure of the FTM consists of an interface to a fault tolerant strategy module, a responder module, a fault manager module, an error factory, and an error mapper that determines the severity of the error. In the present reference implementation, the only fault tolerant strategy implemented is introspection. The introspection code waits for an application node to send an error notification to it. It then uses the error factory to create an error object, and at this time, a severity level is assigned to the error. The introspection code uses its built-in knowledge base to generate a recommended response to the error. Responses might include ignoring the error, logging it, rolling back the application to a previously saved checkpoint, swapping in a new node to replace a bad one, or restarting the application. The original error and recommended response are passed to the top-level fault manager module, which invokes the response. The responder module also notifies the introspection module of the generated response. This provides additional information to the introspection module that it can use in generating its next response. For example, if the responder triggers an application rollback and errors are still occurring, the introspection module may decide to recommend an application restart.

Some, Raphael R.↗

Locating Anomalies in Complex Data Sets Using Visualization and Simulation

The research goals are to create a simulation framework that can accept any combination of models written at the gate or behavioral level. The framework provides the ability to fault simulate and create scenarios of experiments using concurrent simulation. In order to meet these goals we have had to fulfill the following requirements. The ability to accept models written in VHDL, Verilog or the C languages. The ability to propagate faults through any model type. The ability to create experiment scenarios efficiently without generating every possible combination of variables. The ability to accept adversity of fault models beyond the single stuck-at model. Major development has been done to develop a parser that can accept models written in various languages. This work has generated considerable attention from other universities and industry for its flexibility and usefulness. The parser uses LEXX and YACC to parse Verilog and C. We have also utilized our industrial partnership with Alternative System's Inc. to import vhdl into our simulator. For multilevel simulation, we needed to modify the simulator architecture to accept models that contained multiple outputs. This enabled us to accept behavioral components. The next major accomplishment was the addition of "functional fault models". Functional fault models change the behavior of a gate or model. For example, a bridging fault can make an OR gate behave like an AND gate. This has applications beyond fault simulation. This modeling flexibility will make the simulator more useful for doing verification and model comparison. For instance, two or more versions of an ALU can be comparatively simulated in a single execution. The results will show where and how the models differed so that the performance and correctness of the models may be evaluated. A considerable amount of time has been dedicated to validating the simulator performance on larger models provided by industry and other universities.

Panetta, Karen↗

Optimization of Second Fault Detection Thresholds to Maximize Mission POS

In order to support manned spaceflight safety requirements, the Space Launch System (SLS) has defined program-level requirements for key systems to ensure successful operation under single fault conditions. To accommodate this with regards to Navigation, the SLS utilizes an internally redundant Inertial Navigation System (INS) with built-in capability to detect, isolate, and recover from first failure conditions and still maintain adherence to performance requirements. The unit utilizes multiple hardware- and software-level techniques to enable detection, isolation, and recovery from these events in terms of its built-in Fault Detection, Isolation, and Recovery (FDIR) algorithms. Successful operation is defined in terms of sufficient navigation accuracy at insertion while operating under worst case single sensor outages (gyroscope and accelerometer faults at launch). In addition to first fault detection and recovery, the SLS program has also levied requirements relating to the capability of the INS to detect a second fault, tracking any unacceptable uncertainty in knowledge of the vehicle's state. This detection functionality is required in order to feed abort analysis and ensure crew safety. Increases in navigation state error and sensor faults can drive the vehicle outside of its operational as-designed environments and outside of its performance envelope causing loss of mission, or worse, loss of crew. The criteria for operation under second faults allows for a larger set of achievable missions in terms of potential fault conditions, due to the INS operating at the edge of its capability. As this performance is defined and controlled at the vehicle level, it allows for the use of system level margins to increase probability of mission success on the operational edges of the design space. Due to the implications of the vehicle response to abort conditions (such as a potentially failed INS), it is important to consider a wide range of failure scenarios in terms of both magnitude and time. As such, the Navigation team is taking advantage of the INS's capability to schedule and change fault detection thresholds in flight. These values are optimized along a nominal trajectory in order to maximize probability of mission success, and reducing the probability of false positives (defined as when the INS would report a second fault condition resulting in loss of mission, but the vehicle would still meet insertion requirements within system-level margins). This paper will describe an optimization approach using Genetic Algorithms to tune the threshold parameters to maximize vehicle resilience to second fault events as a function of potential fault magnitude and time of fault over an ascent mission profile. The analysis approach, and performance assessment of the results will be presented to demonstrate the applicability of this process to second fault detection to maximize mission probability of success.

Anzalone, Evan↗

Design, Integration, Certification and Testing of the Orion Crew Module Propulsion System

The Orion Multipurpose Crew Vehicle (MPCV) is NASA's next generation spacecraft for human exploration of deep space. Lockheed Martin is the prime contractor for the design, development, qualification and integration of the vehicle. A key component of the Orion Crew Module (CM) is the Propulsion Reaction Control System, a high‐flow hydrazine system used during re‐entry to orient the vehicle for landing. The system consists of a completely redundant helium (GHe) pressurization system and hydrazine fuel system with monopropellant thrusters. The propulsion system has been designed, integrated, and qualification tested in support of the Orion program's first orbital flight test, Exploration Flight Test One (EFT‐1), scheduled for 2014. A subset of the development challenges and lessons learned from this first flight test campaign will be discussed in this paper for consideration when designing future spacecraft propulsion systems. The CONOPS and human rating requirements of the CM propulsion system are unique when compared with a typical satellite propulsion reaction control system. The system requires a high maximum fuel flow rate. It must operate at both vacuum and sea level atmospheric pressure conditions. In order to meet Orion's human rating requirements, multiple parts of the system must be redundant, and capable of functioning after spacecraft system fault events.

McKay, Heather↗

Modular Stirling Radioisotope Generator

High efficiency radioisotope power generators will play an important role in future NASA space exploration missions. Stirling Radioisotope Generators (SRG) have been identified as a candidate generator technology capable of providing mission designers with an efficient, high specific power electrical generator. SRGs high conversion efficiency has the potential to extend the limited Pu-238 supply when compared with current Radioisotope Thermoelectric Generators (RTG). Due to budgetary constraints, the Advanced Stirling Radioisotope Generator (ASRG) was canceled in the fall of 2013. Over the past year a joint study by NASA and DOE called the Nuclear Power Assessment Study (NPAS) recommended that Stirling technologies continue to be explored. During the mission studies of the NPAS, spare SRGs were sometimes required to meet mission power system reliability requirements. This led to an additional mass penalty and increased isotope consumption levied on certain SRG-based missions. In an attempt to remove the spare power system, a new generator architecture is considered which could increase the reliability of a Stirling generator and provide a more fault-tolerant power system. This new generator called the Modular Stirling Radioisotope Generator (MSRG) employs multiple parallel Stirling convertor/controller strings, all of which share the heat from the General Purpose Heat Source (GPHS) modules. For this design, generators utilizing one to eight GPHS modules were analyzed, which provide about 50 to 450 watts DC to the spacecraft, respectively. Four Stirling convertors are arranged around each GPHS module resulting in from 4 to 32 Stirling/controller strings. The convertors are balanced either individually or in pairs, and are radiatively coupled to the GPHS modules. Heat is rejected through the housing/radiator which is similar in construction to the ASRG. Mass and power analysis for these systems indicate that specific power may be slightly lower than the ASRG and similar to the MMRTG. However, the reliability should be significantly increased compared to ASRG.

Stirling Cycle↗

Modular Stirling Radioisotope Generator

High-efficiency radioisotope power generators will play an important role in future NASA space exploration missions. Stirling Radioisotope Generators (SRGs) have been identified as a candidate generator technology capable of providing mission designers with an efficient, high-specific-power electrical generator. SRGs high conversion efficiency has the potential to extend the limited Pu-238 supply when compared with current Radioisotope Thermoelectric Generators (RTGs). Due to budgetary constraints, the Advanced Stirling Radioisotope Generator (ASRG) was canceled in the fall of 2013. Over the past year a joint study by NASA and the Department of Energy (DOE) called the Nuclear Power Assessment Study (NPAS) recommended that Stirling technologies continue to be explored. During the mission studies of the NPAS, spare SRGs were sometimes required to meet mission power system reliability requirements. This led to an additional mass penalty and increased isotope consumption levied on certain SRG-based missions. In an attempt to remove the spare power system, a new generator architecture is considered, which could increase the reliability of a Stirling generator and provide a more fault-tolerant power system. This new generator called the Modular Stirling Radioisotope Generator (MSRG) employs multiple parallel Stirling convertor/controller strings, all of which share the heat from the General Purpose Heat Source (GPHS) modules. For this design, generators utilizing one to eight GPHS modules were analyzed, which provided about 50 to 450 W of direct current (DC) to the spacecraft, respectively. Four Stirling convertors are arranged around each GPHS module resulting in from 4 to 32 Stirling/controller strings. The convertors are balanced either individually or in pairs, and are radiatively coupled to the GPHS modules. Heat is rejected through the housing/radiator, which is similar in construction to the ASRG. Mass and power analysis for these systems indicate that specific power may be slightly lower than the ASRG and similar to the Multi-Mission Radioisotope Thermoelectric Generator (MMRTG). However, the reliability should be significantly increased compared to ASRG.

Modules↗

Interface Supports Multiple Broadcast Transceivers for Flight Applications

A wireless avionics interface provides a mechanism for managing multiple broadcast transceivers. This interface isolates the control logic required to support multiple transceivers so that the flight application does not have to manage wireless transceivers. All of the logic to select transceivers, detect transmitter and receiver faults, and take autonomous recovery action is contained in the interface, which is not restricted to using wireless transceivers. Wired, wireless, and mixed transceiver technologies are supported. This design s use of broadcast data technology provides inherent cross strapping of data links. This greatly simplifies the design of redundant flight subsystems. The interface fully exploits the broadcast data link to determine the health of other transceivers used to detect and isolate faults for fault recovery. The interface uses simplified control logic, which can be implemented as an intellectual-property (IP) core in a field-programmable gate array (FPGA). The interface arbitrates the reception of inbound data traffic appearing on multiple receivers. It arbitrates the transmission of outbound traffic. This system also monitors broadcast data traffic to determine the health of transmitters in the network, and then uses this health information to make autonomous decisions for routing traffic through transceivers. Multiple selection strategies are supported, like having an active transceiver with the secondary transceiver powered off except to send periodic health status reports. Transceivers can operate in round-robin for load-sharing and graceful degradation.

Block, Gary L.↗

Reliability of Fault Tolerant Control Systems

This paper reports Part II of a two part effort that is intended to delineate the relationship between reliability and fault tolerant control in a quantitative manner. Reliability properties peculiar to fault-tolerant control systems are emphasized, such as the presence of analytic redundancy in high proportion, the dependence of failures on control performance, and high risks associated with decisions in redundancy management due to multiple sources of uncertainties and sometimes large processing requirements. As a consequence, coverage of failures through redundancy management can be severely limited. The paper proposes to formulate the fault tolerant control problem as an optimization problem that maximizes coverage of failures through redundancy management. Coverage modeling is attempted in a way that captures its dependence on the control performance and on the diagnostic resolution. Under the proposed redundancy management policy, it is shown that an enhanced overall system reliability can be achieved with a control law of a superior robustness, with an estimator of a higher resolution, and with a control performance requirement of a lesser stringency.

Wu, N. Eva↗