Search NASA⌕ Search

SEARCH · Search NASA

Results for “failure recovery”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

Failure recovery control for space robotic systems

The problem of controlling a failed joint of a space manipulator is addressed. It is shown that failure-recovery control is possible when dynamic coupling exists between the link whose joint has failed and some other link whose joint is working and when the system inertia matrix is invariant with respect to the failed joint angle. A failure-recovery control technique is developed and applied to two simple examples.

Papadopoulos, Evangelos↗

ISS Ammonia Pump Failure, Recovery, and Lesson Learned A Hydrodynamic Bearing Perspective

The design, development, and operation of long duration spaceflight hardware has become an evolutionary process in which meticulous attention to details and lessons learned from previous experiences play a critical role. Invaluable to this process is the ability to retrieve and examine spaceflight hardware that has experienced a premature failure. While these situations are rare and unfortunate, the failure investigation and recovery from the event serve a valuable purpose in advancing future space mechanism development. Such a scenario began on July 31, 2010 with the premature failure of an ammonia pump on the external active thermal control system of the International Space Station. The ground-based inspections of the returned pump and ensuing failure investigation revealed five potential bearing forces that were un-accounted for in the design phase and qualification testing of the pump. These forces could combine in a number of random orientations to overload the pump bearings leading to solid-surface contact, wear, and premature failure. The recovery plan identified one of these five forces as being related to the square of the operating speed of the pump and this fact was used to recover design life through a change in flight rules for the operation of the pump module. Through the course of the failure investigation, recovery, and follow-on assessment of pump wear life, design guidance has been developed to improve the life of future mechanically pumped thermal control systems for both human and robotic exploration missions.

Bruckner, Robert J.↗

A failure recovery planning prototype for Space Station Freedom

NASA is investigating the use of advanced automation to enhance crew productivity for Space Station Freedom in numerous areas, including failure management. A prototype is described that uses various advanced automation techniques to generate courses of action whose intents are to recover from a diagnosed failure, and to do so within the constraints levied by the failure and by Freedom's configuration and operating conditions.

Hammen, David G.↗

On scheduling tasks with a quick recovery from failure

Multiprocessors used in life-critical real-time systems must recover quickly from failure. Part of this recovery consists of switching to a new task schedule which ensures that hard deadlines for critical tasks continue to be met. A dynamic programming algorithm is presented that ensures that backup, or contingency, schedules can be efficiently embedded within the original, 'primary' schedule to ensure that hard deadlines continue to be met in the face of up to a given maximum number of processor failures. Several illustrative examples are included.

Krishna, C. M.↗

Memory management and compiler support for rapid recovery from failures in computer systems

This paper describes recent developments in the use of memory management and compiler technology to support rapid recovery from failures in computer systems. The techniques described include cache coherence protocols for user transparent checkpointing in multiprocessor systems, compiler-based checkpoint placement, compiler-based code modification for multiple instruction retry, and forward recovery in distributed systems utilizing optimistic execution.

Fuchs, W. K.↗

Extravehicular Mobility Unit (EMU) / International Space Station (ISS) Coolant Loop Failure and Recovery

Following the Colombia accident, the Extravehicular Mobility Units (EMU) onboard ISS were unused for several months. Upon startup, the units experienced a failure in the coolant system. This failure resulted in the loss of Extravehicular Activity (EVA) capability from the US segment of ISS. With limited on-orbit evidence, a team of chemists, engineers, metallurgists, and microbiologists were able to identify the cause of the failure and develop recovery hardware and procedures. As a result of this work, the ISS crew regained the capability to perform EVAs from the US segment of the ISS.

Lewis, John F.↗

EUREX D: An expert system for failure diagnosis and recovery in the TCS of the European retrievable carrier EURECA

An expert system for diagnosis and recovery of failures in the Freon cooling loop of the European retrievable experiment carrier EURECA is described. The system demonstrates the feasibility of a functional scope of expert diagnostic systems which appears to be essential for practical applications of such systems in space technology. The scope includes early warning and treatment of incomplete information, fault tolerance, identification of failure superpositions, intelligent reaction to unforeseen events, and detailed status display for optimal recovery action.

Kellner, A.↗

On the design of fault-tolerant robotic manipulator systems

Robotic systems are finding increasing use in space applications. Many of these devices are going to be operational on board the Space Station Freedom. Fault tolerance has been deemed necessary because of the criticality of the tasks and the inaccessibility of the systems to maintenance and repair. Design for fault tolerance in manipulator systems is an area within robotics that is without precedence in the literature. In this paper, we will attempt to lay down the foundations for such a technology. Design for fault tolerance demands new and special approaches to design, often at considerable variance from established design practices. These design aspects, together with reliability evaluation and modeling tools, are presented. Mechanical architectures that employ protective redundancies at many levels and have a modular architecture are then studied in detail. Once a mechanical architecture for fault tolerance has been derived, the chronological stages of operational fault tolerance are investigated. Failure detection, isolation, and estimation methods are surveyed, and such methods for robot sensors and actuators are derived. Failure recovery methods are also presented for each of the protective layers of redundancy. Failure recovery tactics often span all of the layers of a control hierarchy. Thus, a unified framework for decision-making and control, which orchestrates both the nominal redundancy management tasks and the failure management tasks, has been derived. The well-developed field of fault-tolerant computers is studied next, and some design principles relevant to the design of fault-tolerant robot controllers are abstracted. Conclusions are drawn, and a road map for the design of fault-tolerant manipulator systems is laid out with recommendations for a 10 DOF arm with dual actuators at each joint.

Tesar, Delbert↗

ISS Solar Array Alpha Rotary Joint (SARJ) Bearing Failure and Recovery: Technical and Project Management Lessons Learned

The photovoltaic solar panels on the International Space Station (ISS) track the Sun through continuous rotating motion enabled by large bearings on the main truss called solar array alpha rotary joints (SARJs). In late 2007, shortly after installation, the starboard SARJ had become hard to turn and had to be shut down after exceeding drive current safety limits. The port SARJ, of the same design, had been working well for over 2 years. An exhaustive failure investigation ensued that included multiple extravehicular activities to collect information and samples for engineering forensics, detailed structural and thermal analyses, and a careful review of the build records. The ultimate root cause was determined to be kinematic design vulnerability coupled with inadequate lubrication, and manufacturing flaws; this was corroborated through ground tests, metallurgical studies, and modeling. A highly successful recovery plan was developed and implemented that included replacing worn and damaged components in orbit and applying space-compatible grease to improve lubrication. Beyond the technical aspects, however, lie several key programmatic lessons learned. These lessons, such as running ground tests to intentional failure to experimentally verify failure modes, are reviewed and discussed so they can be applied to future projects to avoid such problems.

DellaCorte, Christopher↗

Piloted Evaluation of a Fault Recovery System for an Aircraft with Distributed Electric Propulsion

Electrified aircraft powertrains contain multiple tightly coupled subsystems, making them much more complex than traditional aircraft propulsion systems, both in terms of integration and control. Electrification enables aircraft to have multiple distributed thrust-producing fans that the flight control system can utilize for enhanced maneuverability, further increasing the control complexity. The SUbsonic Single Aft eNgine (SUSAN) Electrofan is a NASA concept aircraft that leverages this technology. SUSAN is a series/parallel partial hybrid electric single-aisle transport aircraft that takes advantage of its electrified powertrain to provide fuel burn and emissions benefits when compared to the state-of-the-art. Achieving these benefits requires an appropriately designed control architecture that coordinates the various powertrain and flight control subsystems. As such, the SUSAN aircraft is designed with a high level of automation, allowing it to properly manage coupled subsystems and react rapidly to failures and anomalies. To do this effectively, algorithms that perform component health management, fault detection, isolation, and accommodation, and continuous optimization, must be developed, tested, validated, and implemented. This paper describes a piloted evaluation of such an algorithm in scenarios with multiple fan failures, performed in a flight simulator, demonstrating failure recovery and continued safe operation up to the limits of the powertrain. These scenarios are subsequently related to certification requirements.

Electrified Aircraft Propulsion↗

Piloted Evaluation of a Fault Recovery System for an Aircraft with Distributed Electric Propulsion

Electrified aircraft powertrains contain multiple tightly coupled subsystems, making them much more complex than traditional aircraft propulsion systems, both in terms of integration and control. Electrification enables aircraft to have multiple distributed thrust-producing fans that the flight control system can utilize for enhanced maneuverability, further increasing the control complexity. The SUbsonic Single Aft eNgine (SUSAN) Electrofan is a NASA concept aircraft that leverages this technology. SUSAN is a series/parallel partial hybrid electric single-aisle transport aircraft that takes advantage of its electrified powertrain to provide fuel burn and emissions benefits when compared to the state-of-the-art. Achieving these benefits requires an appropriately designed control architecture that coordinates the various powertrain and flight control subsystems. As such, the SUSAN aircraft is designed with a high level of automation, allowing it to properly manage coupled subsystems and react rapidly to failures and anomalies. To do this effectively, algorithms that perform component health management, fault detection, isolation, and accommodation, and continuous optimization, must be developed, tested, validated, and implemented. This paper describes a piloted evaluation of such an algorithm in scenarios with multiple fan failures, performed in a flight simulator, demonstrating failure recovery and continued safe operation up to the limits of the powertrain. These scenarios are subsequently related to certification requirements.

Electrified Aircraft Propulsion↗

Piloted Evaluation of a Fault Recovery System for an Aircraft with Distributed Electric Propulsion

Electrified aircraft powertrains contain multiple tightly coupled subsystems, making them much more complex than traditional aircraft propulsion systems, both in terms of integration and control. Electrification enables aircraft to have multiple distributed thrust-producing fans that the flight control system can utilize for enhanced maneuverability, further increasing the control complexity. The SUbsonic Single Aft eNgine (SUSAN) Electrofan is a NASA concept aircraft that leverages this technology. SUSAN is a series/parallel partial hybrid electric single-aisle transport aircraft that takes advantage of its electrified powertrain to provide fuel burn and emissions benefits when compared to the state-of-the-art. Achieving these benefits requires an appropriately designed control architecture that coordinates the various powertrain and flight control subsystems. As such, the SUSAN aircraft is designed with a high level of automation, allowing it to properly manage coupled subsystems and react rapidly to failures and anomalies. To do this effectively, algorithms that perform component health management, fault detection, isolation, and accommodation, and continuous optimization, must be developed, tested, validated, and implemented. This paper describes a piloted evaluation of such an algorithm in scenarios with multiple fan failures, performed in a flight simulator, demonstrating failure recovery and continued safe operation up to the limits of the powertrain. These scenarios are subsequently related to certification requirements.

Electrified Aircraft Propulsion↗

NASA Space Nuclear Propulsion (SNP) MBSE Initiatives

NASA’s Space Nuclear Propulsion (SNP) program is developing several MagicDraw SysML models to support the development of high performance Nuclear Thermal Rocket Engines (NTRE). Currently, the Demonstration Rocket for Agile Cislunar Operations (DRACO) project is aiming to perform the first ever flight demonstration of an NTRE, and NASA is developing a DRACO Insight Project Model Based Systems Engineering (MBSE) model to capture, define, analyze, and report on the flight and ground test system architecture, functional behavior, requirements, risks, and lessons learned. Additional models are in work for engine component trade trees, fault detection sensor coverage analysis using a Goal Function Tree (GFT) plugin, stakeholder engagement, and technology maturation projects. The GFT plugin is the Galois, Inc. Failure Recovery Instruction Generation using Automata derived from Traditional Engineering models (FRIGATE) tool. A new capability for Jira to MagicDraw data sharing using the OpenPDM collaboration platform is under development with partner Victory Solutions, Inc. to enhance risk impact analysis.

Space Nuclear Propulsion (SNP)↗

A.I.-based real-time support for high performance aircraft operations

Artificial intelligence (AI) based software and hardware concepts are applied to the handling system malfunctions during flight tests. A representation of malfunction procedure logic using Boolean normal forms are presented. The representation facilitates the automation of malfunction procedures and provides easy testing for the embedded rules. It also forms a potential basis for a parallel implementation in logic hardware. The extraction of logic control rules, from dynamic simulation and their adaptive revision after partial failure are examined. It uses a simplified 2-dimensional aircraft model with a controller that adaptively extracts control rules for directional thrust that satisfies a navigational goal without exceeding pre-established position and velocity limits. Failure recovery (rule adjusting) is examined after partial actuator failure. While this experiment was performed with primitive aircraft and mission models, it illustrates an important paradigm and provided complexity extrapolations for the proposed extraction of expertise from simulation, as discussed. The use of relaxation and inexact reasoning in expert systems was also investigated.

Vidal, J. J.↗

Nonblocking and orphan free message logging protocols

Currently existing message logging protocols demonstrate a classic pessimistic vs. optimistic tradeoff. We show that the optimistic-pessimistic tradeoff is not inherent to the problem of message logging. We construct a message-logging protocol that has the positive features of both optimistic and pessimistic protocol: our protocol prevents orphans and allows simple failure recovery; however, it requires no blocking in failure-free runs. Furthermore, this protocol does not introduce any additional message overhead as compared to one implemented for a system in which messages may be lost but processes do not crash.

Alvisi, Lorenzo↗

Progressive retry for software error recovery in distributed systems

In this paper, we describe a method of execution retry for bypassing software errors based on checkpointing, rollback, message reordering and replaying. We demonstrate how rollback techniques, previously developed for transient hardware failure recovery, can also be used to recover from software faults by exploiting message reordering to bypass software errors. Our approach intentionally increases the degree of nondeterminism and the scope of rollback when a previous retry fails. Examples from our experience with telecommunications software systems illustrate the benefits of the scheme.

Wang, Yi-Min↗

Robotic and Human-Tended Collaborative Drilling Automation for Subsurface Exploration

Future in-situ lunar/martian resource utilization and characterization, as well as the scientific search for life on Mars, will require access to the subsurface and hence drilling. Drilling on Earth is hard - an art form more than an engineering discipline. Human operators listen and feel drill string vibrations coming from kilometers underground. Abundant mass and energy make it possible for terrestrial drilling to employ brute-force approaches to failure recovery and system performance issues. Space drilling will require intelligent and autonomous systems for robotic exploration and to support human exploration. Eventual in-situ resource utilization will require deep drilling with probable human-tended operation of large-bore drills, but initial lunar subsurface exploration and near-term ISRU will be accomplished with lightweight, rover-deployable or standalone drills capable of penetrating a few tens of meters in depth. These lightweight exploration drills have a direct counterpart in terrestrial prospecting and ore-body location, and will be designed to operate either human-tended or automated. NASA and industry now are acquiring experience in developing and building low-mass automated planetary prototype drills to design and build a pre-flight lunar prototype targeted for 2011-12 flight opportunities. A successful system will include development of drilling hardware, and automated control software to operate it safely and effectively. This includes control of the drilling hardware, state estimation of both the hardware and the lithography being drilled and state of the hole, and potentially planning and scheduling software suitable for uncertain situations such as drilling. Given that Humans on the Moon or Mars are unlikely to be able to spend protracted EVA periods at a drill site, both human-tended and robotic access to planetary subsurfaces will require some degree of standalone, autonomous drilling capability. Human-robotic coordination will be important, either between a robotic drill and humans on Earth, or a human-tended drill and its visiting crew. The Mars Analog Rio Tinto Experiment (MARTE) is a current project that studies and simulates the remote science operations between an automated drill in Spain and a distant, distributed human science team. The Drilling Automation for Mars Exploration (DAME) project, by contrast: is developing and testing standalone automation at a lunar/martian impact crater analog site in Arctic Canada. The drill hardware in both projects is a hardened, evolved version of the Advanced Deep Drill (ADD) developed by Honeybee Robotics for the Mars Subsurface Program. The current ADD is capable of 20m, and the DAME project is developing diagnostic and executive software for hands-off surface operations of the evolved version of this drill. The current drill automation architecture being developed by NASA and tested in 2004-06 at analog sites in the Arctic and Spain will add downhole diagnosis of different strata, bit wear detection, and dynamic replanning capabilities when unexpected failures or drilling conditions are discovered in conjunction with simulated mission operations and remote science planning. The most important determinant of future 1unar and martian drilling automation and staffing requirements will be the actual performance of automated prototype drilling hardware systems in field trials in simulated mission operations. It is difficult to accurately predict the level of automation and human interaction that will be needed for a lunar-deployed drill without first having extensive experience with the robotic control of prototype drill systems under realistic analog field conditions. Drill-specific failure modes and software design flaws will become most apparent at this stage. DAME will develop and test drill automation software and hardware under stressful operating conditions during several planned field campaigns. Initial results from summer 2004 tests show seven identifi distinct failure modes of the drill: cuttings-removal issues with low-power drilling into permafrost, and successful steps at executive control and initial automation.

Glass, Brian↗