Search NASA⌕ Search

SEARCH · Search NASA

Results for “Spacecraft Anomaly Recovery”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

Galileo spacecraft anomaly and safing recovery

A high-level anomaly recovery plan which identifies the steps necessary to recover from a spacecraft 'Safing' incident was developed for the Galileo spacecraft prior to launch. Since launch, a total of four in-flight anomalies have lead to entry into a system fault protection 'Safing' routine which has required the Galileo flight team to refine and execute the recovery plan. These failures have allowed the flight team to develop an efficient recovery process when permanent spacecraft capability degradation is minimal and the cause of the anomaly is quickly diagnosed. With this previous recovery experience and the very focused boundary conditions of a specific potential failure, a Gaspra asteroid recovery plan was designed to be implemented in as quickly as forty hours (desired goal). This paper documents the work performed above, however, the Galileo project remains challenged to develop a generic detailed recovery plan which can be implemented in a relatively short time to configure the spacecraft to a nominal state prior to future high priority mission objectives.

Basilio, Ralph R.↗

SORCE Daylight-Only Operations

The recent experience of the SORCE flight operations team offers an excellent example of innovative engineering using limited resources. The goal of this paper is to extend to the space operations community the lessons learned during this critical redesign in order to aid other missions facing equally daunting challenges. The end result is a mission extended well beyond its designed life continuing to return important data to the science community to extend the climate record.

Spacecraft Anomaly Recovery↗

The Recovery of TOMS-EP

On December 13th 1998, the Total Ozone Mapping Spectrometer - Earth Probe (TOMS-EP) spacecraft experienced a Single Event Upset which caused the system to reconfigure and enter a Safe Mode. This incident occurred two and a half years after the launch of the spacecraft which was designed for a two year life. A combination of factors, including changes in component behavior due to age and extended use, very unfortunate initial conditions and the safe mode processing logic prevented the spacecraft from entering its nominal long term storage mode. The spacecraft remained in a high fuel consumption mode designed for temporary use. By the time the onboard fuel was exhausted, the spacecraft was Sun pointing in a high rate flat spin. Although the uncontrolled spacecraft was initially in a power and thermal safe orientation, it would not stay in this state indefinitely due to a slow precession of its momentum vector. A recovery team was immediately assembled to determine if there was time to develop a method of despinning the vehicle and return it to normal science data collection. A three stage plan was developed that used the onboard magnetic torque rods as actuators. The first stage was designed to reduce the high spin rate to within the linear range of the gyros. The second stage transitioned the spacecraft from sun pointing to orbit reference pointing. The final stage returned the spacecraft to normal science operation. The entire recovery scenario was simulated with a wide range of initial conditions to establish the expected behavior. The recovery sequence was started on December 28th 1998 and completed by December 31st. TOMS-EP was successfully returned to science operations by the beginning of 1999. This paper describes the TOMS-EP Safe Mode design and the factors which led to the spacecraft anomaly and loss of fuel. The recovery and simulation efforts are described. Flight data are presented which show the performance of the spacecraft during its return to science. Finally, lessons learned are presented.

Robertson, Brent↗

Validating system-level error recovery for spacecraft

The system-level software onboard a spacecraft is responsible for recovery from communication, thermal, power, and computer-health anomalies that may occur. The recovery must occur without disrupting any critical scientific or engineering activity that is executing at the time of the error. Thus, the error-recovery software may have to execute concurrently with the ongoing acquisition of scientific data or with spacecraft maneuvers. This paper provides a technique by which the rules that constrain the concurrent execution of these processes can be modeled in a graph. An algorithm is described that uses this model to validate that the constraints hold for all concurrent executions of the error-recovery softwave with the softwave that controls the science and engineering events on the spacecraft.

Lutz, Robyn R.↗

Constraint checking during error recovery

The system-level software onboard a spacecraft is responsible for recovery from communication, power, thermal, and computer-health anomalies that may occur. The recovery must occur without disrupting any critical scientific or engineering activity that is executing at the time of the error. Thus, the error-recovery software may have to execute concurrently with the ongoing acquisition of scientific data or with spacecraft maneuvers. This work provides a technique by which the rules that constrain the concurrent execution of these processes can be modeled in a graph. An algorithm is described that uses this model to validate that the constraints hold for all concurrent executions of the error-recovery software with the software that controls the science and engineering activities of the spacecraft. The results are applicable to a variety of control systems with critical constraints on the timing and ordering of the events they control.

Lutz, Robyn R.↗

Lessons Learned During the Transition of SORCE Science Operations to Daylight Only Operations

In July 2013, NASA's Solar Radiation and Climate Experiment experienced a battery anomaly which placed it into safemode halting all science observations. Initial attempts to recover the spacecraft to an operational configuration failed due to the reduced capacity of the battery. As the keystone mission for measuring total solar irradiance, and the cornerstone mission for measuring the solar spectral irradiance there was a strong motivation for developing a new operations concept that would allow SORCE to resume daily measurements of the Sun. The operations team faced many challenges over the next several months. For a five-day period in late 2013 the operations team was able resume science observations to cross-calibrate SORCE data with a new instrument launched in November 2013. After the cross-calibration campaign was completed a new operations concept was deployed which allowed SORCE to perform daylight only operations. In this mode of operations all non-essential components are powered off at each eclipse entry and then turned back on at sunrise. In March 2014 SORCE resumed making daily measurements of the Sun. This paper will review the events and lessons learned from the six-month recovery effort.

Spacecraft Anomaly↗

Huygens Probe Relay Data Subsystem Anomaly and Recovery

European Space Agency Mission is designed to study the atmosphere and surface of Saturn's largest satellite, Titan carried by the Cassini spacecraft which provides: a) Power for support equipment; b) S-band antenna system; and c) Data storage and playback. Instruments/investigations include: 1) Aerosol Collector Pyrolyzer (ACP). Study of clouds and aerosols in the Titan atmosphere. 2) Descent Imager and Spectral Radiometer (DISR). Aerosol and cloud optical properties and spectroscopy measurements of Titan's atmosphere and surface. 3) Doppler Wind Experiment (DWE). Study of winds from their effect on the Probe during Titan descent. 4) Gas Chromatograph and Mass Spectrometer (GCMS). Chemical composition of gases and aerosols in Titan's atmosphere. 5) Huygens Atmospheric Structure Instrument (HASI). In-situ study of Titan atmospheric physical and electrical properties. 6) Surface Science Package (SSP). Physical properties of Titan's surface and related atmospheric properties.

Huygens↗

Enhancing the Cassini Mission Through FP Applications After Launch

Although rigorous pre-emptive measures are taken to preclude failures and anomalous conditions from occurring in JPL spacecraft missions prior to launch, unforeseeable problems can still surface after liftoff. In the case of the Cassini/Huygens Mission-to-Saturn spacecraft, several problems were observed post-launch: 1) immediately after takeoff, the collected engineering/science data stored on the Solid State Recorders (SSR) contained a significantly higher number of corrupted bits than was expected (considerably over spec) due to human error in the memory mapping of these devices, 2) numerous Solid State Power Switches (SSPS) sporadically tripped off throughout the mission due to cosmic ray bombardment from the unique space environment, and 3) false assumptions in the pressure regulator design in combination with missing heritage test data led to inaccurate design conclusions, causing the issuance of two waivers for the regulator to close properly (a potentially mission catastrophic single-point failure which occurred 24 days after launch) - amongst other problems. For Cassini, some of these anomalies led to arduous work-arounds or required continuous monitoring of telemetry variables by the ground-based Spacecraft Operations Flight Support (SOFS) team in order to detect and fix fault occurrences as they happened. Fortunately, sufficient funding and schedule margin allowed several Fault Protection (FP) solutions to be implemented into post-launch Flight Software (FSW) uploads to help resolve these issues autonomously, reducing SOFS ground support efforts while improving anomaly recovery time in order to preserve maximum science capture. This paper details the FP applications used to resolve the above issues as well as to optimize solutions for several other problems experienced by the Cassini spacecraft during its fight, in order to enhance the spacecraft's overall mission success throughout the 18 years of its 20 year expedition to and within the Saturnian system.

fault protection↗

Autonomous RPOD for Arbitrarily Configured Spacecraft with Anomaly Detection

Autonomous GN&C is a necessary component for a sustainable deep-space logistics architecture. The challenges for establishing robust autonomy are numerous, from state uncertainty, to anomaly detection and recovery. In this work, previous work investigating autonomous GN&C for arbitrary thruster configurations and mass properties is expanded to include state uncertainty and anomaly detection. Logistics vehicles with off-center-of-mass thruster configurations and in the presence of large but realistic state uncertainties are simulated in a Rendezvous, Proximity Operations and Docking scenario. Furthermore, stuck and non-functional thrusters are simulated, demonstrating the vehicle's ability to identify and overcome thruster anomalies. The simulations demonstrate that even with these realistic ambiguities, the vehicle is able to converge to the desired pose.

GN&C↗

CloudSat Anomaly Recovery and Operational Lessons Learned

In April 2011, NASA's pioneering cloud profiling radar satellite, CloudSat, experienced a battery anomaly that placed it into emergency mode and rendered it operations incapable. All initial attempts to recover the spacecraft failed as the resultant power limitations could not support even the lowest power mode. Originally part of a six-satellite constellation known as the "A-Train", CloudSat was unable to stay within its assigned control box, posing a threat to other A-Train satellites. CloudSat needed to exit the constellation, but with the tenuous power profile, conducting maneuvers was very risky. The team was able to execute a complex sequence of operations which recovered control, conducted an orbit lower maneuver, and returned the satellite to safe mode, within one 65 minute sunlit period. During the course of the anomaly recovery, the team developed several bold, innovative operational strategies. Details of the investigation into the root-cause and the multiple approaches to revive CloudSat are examined. Satellite communication and commanding during the anomaly are presented. A radical new system of "Daylight Only Operations" (DO-OP) was developed, which cycles the payload and subsystem components off in tune with earth eclipse entry and exit in order to maintain positive power and thermal profiles. The scientific methodology and operational results behind the graduated testing and ramp-up to DO-OP are analyzed. In November 2011, the CloudSat team successfully restored the vehicle to consistent operational collection of cloud radar data during sunlit portions of the orbit. Lessons learned throughout the six-month return-to-operations recovery effort are discussed and offered for application to other R&D satellites, in the context of on-orbit anomaly resolution efforts.

cloud profiling radar satellite↗

Terra Mission Operations: Launch to the Present (and Beyond)

The Terra satellite, flagship of NASA's long-term Earth Observing System (EOS) Program, continues to provide useful earth science observations well past its 5-year design lifetime. This paper describes the evolution of Terra operations, including challenges and successes and the steps taken to preserve science requirements and prolong spacecraft life. Working cooperatively with the Terra science and instrument teams, including NASA's international partners, the mission operations team has successfully kept the Terra operating continuously, resolving challenges and adjusting operations as needed. Terra retains all of its observing capabilities (except Short Wave Infrared) despite its age. The paper also describes concepts for future operations. This paper will review the Terra spacecraft mission successes and unique spacecraft component designs that provided significant benefits extending mission life and science. In addition, it discusses special activities as well as anomalies and corresponding recovery efforts. Lastly, it discusses future plans for continued operations.

Terra↗

SOHO Mission Interruption Joint NASA/ESA Investigation Board

Contact with the SOlar Heliospheric Observatory (SOHO) spacecraft was lost in the early morning hours of June 25, 1998, Eastern Daylight Time (EDT), during a planned period of calibrations, maneuvers, and spacecraft reconfigurations. Prior to this the SOHO operations team had concluded two years of extremely successful science operations. A joint European Space Agency (ESA)/National Aeronautics and Space Administration (NASA) engineering team has been planning and executing recovery efforts since loss of contact with some success to date. ESA and NASA management established the SOHO Mission Interruption Joint Investigation Board to determine the actual or probable cause(s) of the SOHO spacecraft mishap. The Board has concluded that there were no anomalies on-board the SOHO spacecraft but that a number of ground errors led to the major loss of attitude experienced by the spacecraft. The Board finds that the loss of the SOHO spacecraft was a direct result of operational errors, a failure to adequately monitor spacecraft status, and an erroneous decision which disabled part of the on-board autonomous failure detection. Further, following the occurrence of the emergency situation, the Board finds that insufficient time was taken by the operations team to fully assess the spacecraft status prior to initiating recovery operations. The Board discovered that a number of factors contributed to the circumstances that allowed the direct causes to occur. The Board strongly recommends that the two Agencies proceed immediately with a comprehensive review of SOHO operations addressing issues in the ground procedures, procedure implementation, management structure and process, and ground systems. This review process should be completed and process improvements initiated prior to the resumption of SOHO normal operations.

Source record↗

Effects of Communication Delay on Human Spaceflight Missions

Missions onboard the International Space Station rely on the real-time availability of a large ground team of system experts to command the vehicle, solve safety-critical problems, and guide the crew during complex operations. Also, in Low Earth Orbit (LEO), supplies can be sent and crews evacuated quite quickly if needed. Future missions Beyond Low Earth Orbit (BLEO) will not have this 24/7, real-time safety net as communication latency increases, resupply difficulty increases, and evacuation opportunities diminish. There are few, if any, terrestrial analogs for human spaceflight missions BLEO that reflect the conditions—including extreme environments, long mission durations, and small crew sizes – that make these missions so high risk. Studies on specific conditions, such as communication delays and asynchronous interactions, have been performed in NASA Earth-based analog missions and have found that communication delays can disrupt ground-crew interactions and adversely impact team performance. However, there are gaps and limitations in studies conducted to date, notably on human spacecraft system failure response and recovery, the impacts of shorter lunar-relevant communication delays on complex operations, and the effectiveness of countermeasures. The work presented here breaks down real anomalies that occurred on ISS and Apollo missions and creates example scenarios fort Lunar Surface and Mars missions to explore the impact of communication delays of varying length on onboard operations and mission outcomes. Our analyses indicate that short communication delays (e.g., seconds to a minute) adversely impact the ability for ground to provide real-time oversight and guidance and to catch quickly emerging problems in time. Longer communication delays (e.g., up to 40 minutes on Mars missions) call for a shift of responsibility for tactical operations from ground to crew; crew must make time-critical decisions independently and respond to time-critical vehicle anomalies to prevent consequences.

human-systems integration↗

Effects of Communication Delay on Human Spaceflight Missions

Missions onboard the International Space Station rely on the real-time availability of a large ground team of system experts to command the vehicle, solve safety-critical problems, and guide the crew during complex operations. Also, in Low Earth Orbit (LEO), supplies can be sent and crews evacuated quite quickly if needed. Future missions Beyond Low Earth Orbit (BLEO) will not have this 24/7, real-time safety net as communication latency increases, resupply difficulty increases, and evacuation opportunities diminish. There are few, if any, terrestrial analogs for human spaceflight missions BLEO that reflect the conditions—including extreme environments, long mission durations, and small crew sizes – that make these missions so high risk. Studies on specific conditions, such as communication delays and asynchronous interactions, have been performed in NASA Earth-based analog missions and have found that communication delays can disrupt ground-crew interactions and adversely impact team performance. However, there are gaps and limitations in studies conducted to date, notably on human spacecraft system failure response and recovery, the impacts of shorter lunar-relevant communication delays on complex operations, and the effectiveness of countermeasures. The work presented here breaks down real anomalies that occurred on ISS and Apollo missions and creates example scenarios fort Lunar Surface and Mars missions to explore the impact of communication delays of varying length on onboard operations and mission outcomes. Our analyses indicate that short communication delays (e.g., seconds to a minute) adversely impact the ability for ground to provide real-time oversight and guidance and to catch quickly emerging problems in time. Longer communication delays (e.g., up to 40 minutes on Mars missions) call for a shift of responsibility for tactical operations from ground to crew; crew must make time-critical decisions independently and respond to time-critical vehicle anomalies to prevent consequences.

human-systems integration↗

Swift BAT Instrument Thermal Control System Recovery after Spacecraft Safehold in August 2007

The Swift mission Burst Alert Telescope (BAT) Detector Array thermal control system includes two propylene loop heat pipes (LHPs), eight ammonia constant conductance heat pipes (CCHPs), a radiator that has AZ-Tek's AZW-LA-II low alpha white paint, and precision heater controllers that have adjustable set points in flight. The Power Converter Box (PCB) and Image Processor Electronics (IPE) boxes (a primary and a redundant) of the BAT have Z93P white paint radiators. Swift was successfully launched into orbit on November 20, 2004. The spacecraft (S/C) was placed into a safehold mode on August 10, 2007 after an anomaly on inertial reference unit (IRU) #3. It was returned to inertial pointing on August 16 and instrument power up followed. This paper presents a thermal assessment of the BAT instrument thermal control system (TCS) shut down and recovery as a result of the S/C safehold mode. The recovery required starting up the LHPs manually.

Choi, Michael K.↗

Strategies for estimating the marine geoid from altimeter data

Altimeter data from a spacecraft borne altimeter was processed to estimate the fine structure of the marine geoid. Simulation studies show that, among several competing parameterizations, the mean free air gravity anomaly model exhibited promising geoid recovery characteristics. Using covariance analysis techniques, quantitative measures of the orthogonality properties are investigated.

Argentiero, P.↗

Geoscience Laser Altimeter System (GLAS) Loop Heat Pipes: An Eventual First Year On-Orbit

Goddard Space Flight Center's Geoscience Laser Altimeter System (GLAS) is the sole scientific instrument on the Ice, Cloud and land Elevation Satellite (ICESat) that was launched on January 12, 2003 from Vandenberg AFB. A thermal control architecture based on propylene Loop Heat Pipe technology was developed to provide selectable/stable temperature control for the lasers and other electronics over the widely varying mission environment. Following a nominal LHP and instrument start-up, the mission was interrupted with the failure of the first laser after only 36 days of operation. During the 5-month failure investigation, the two GLAS LHPs and the electronics operated nominally, using heaters as a substitute for the laser heat load. Just prior to resuming the mission, following a seasonal spacecraft yaw maneuver, one of the LHPs deprimed and created a thermal runaway condition that resulted in an emergency shutdown of the GLAS instrument. This paper presents details of the LHP anomaly, the resulting investigation and recovery, along with on-orbit flight data during these critical events.

Grob, E.↗