A study of first day space malfunctions
Unmanned spacecraft first day failures, discussing launch environment, duration tests in simulated space and performance improvement
SEARCH · Search NASA
Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.
Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.
Unmanned spacecraft first day failures, discussing launch environment, duration tests in simulated space and performance improvement
The command module television camera monitor exhibited loss of horizontal synchronization during the initial usage. This condition cleared and performance of the monitor was normal until the press conference telecast during the transearth coast phase. At that time, the monitor had the same horizontal synchronization problem reported during the initial usage. The horizontal hold control adjustment would not correct the condition. The monitor was turned off for approximately 5 minutes, then turned back on, after which the monitor's picture was normal. The loss of horizontal synchronization was most likely caused by either an intermittent condition in the precision voltage regulator circuit assembly in the low voltage power supply, or the shift of the stable range of the horizontal potentiometer setting when warming up after turn-on.
Explore the source record for details and available documents.
Explore the source record for details and available documents.
Reasoning about physical systems in operation is a difficult task, and any attempt to automate the process must overcome the problems of modeling normal behavior, diagnosing faults, and predicting future behavior. This paper describes a prototypical case-based reasoner (CBR) that operates in the domain of in-flight fault diagnosis and prognosis of aviation subsystems, particularly jet engines. The reasoner operates on the observation that the ability of a CBR program to reason about physical systems can be significantly enhanced by the addition to the CBR program of a model of the physical system to describe the system's structural, functional, and causal behavior.
A strategy for detecting control law calculation errors in critical flight control computers during laboratory validation testing is presented. This paper addresses Part I of the detection strategy which involves the use of modeling of the aircraft control laws and the design of Kalman filters to predict the correct control commands. Part II of the strategy which involves the use of the predicted control commands to detect control command errors is presented in the companion paper.
A system for on-board anomaly resolution for a vehicle has a data repository. The data repository stores data related to different systems, subsystems, and components of the vehicle. The data stored is encoded in a tree-based structure. A query engine is coupled to the data repository. The query engine provides a user and automated interface and provides contextual query to the data repository. An inference engine is coupled to the query engine. The inference engine compares current anomaly data to contextual data stored in the data repository using inference rules. The inference engine generates a potential solution to the current anomaly by referencing the data stored in the data repository.
The Trick Simulation Environment is a generic simulation toolkit used for constructing and running simulations. This release includes a Monte Carlo analysis simulation framework and a data analysis package. It produces all auto documentation in XML. Also, the software is capable of inserting a malfunction at any point during the simulation. Trick 07 adds variable server output options and error messaging and is capable of using and manipulating wide characters for international support. Wide character strings are available as a fundamental type for variables processed by Trick. A Trick Monte Carlo simulation uses a statistically generated, or predetermined, set of inputs to iteratively drive the simulation. Also, there is a framework in place for optimization and solution finding where developers may iteratively modify the inputs per run based on some analysis of the outputs. The data analysis package is capable of reading data from external simulation packages such as MATLAB and Octave, as well as the common comma-separated values (CSV) format used by Excel, without the use of external converters. The file formats for MATLAB and Octave were obtained from their documentation sets, and Trick maintains generic file readers for each format. XML tags store the fields in the Trick header comments. For header files, XML tags for structures and enumerations, and the members within are stored in the auto documentation. For source code files, XML tags for each function and the calling arguments are stored in the auto documentation. When a simulation is built, a top level XML file, which includes all of the header and source code XML auto documentation files, is created in the simulation directory. Trick 07 provides an XML to TeX converter. The converter reads in header and source code XML documentation files and converts the data to TeX labels and tables suitable for inclusion in TeX documents. A malfunction insertion capability allows users to override the value of any simulation variable, or call a malfunction job, at any time during the simulation. Users may specify conditions, use the return value of a malfunction trigger job, or manually activate a malfunction. The malfunction action may consist of executing a block of input file statements in an action block, setting simulation variable values, call a malfunction job, or turn on/off simulation jobs.
The NASA Human System Risk Board (HSRB) has the overall responsibility for tracking the evolution of the top ~30 human system risks that it has identified to be associated with human spaceflight. As part of this process, the Board is charged with maintaining a consistent, integrated process to mitigate those risks, and developing evidence-based risk posture recommendations. One of the identified risks is due to inadequate human systems integration architecture (HSIA) and a driving factor of this risk is that given decreasing real-time ground support for execution of complex operations during future exploration missions, there is a possibility of adverse performance outcomes including that crew are unable to adequately respond to unanticipated critical malfunctions or detect safety critical procedural errors. The HSRB uses Directed Acyclic Graphs (DAGs) as a communication tool for describing how astronaut exposure to spaceflight hazards leads to meaningful mission-level health and performance outcomes and as the basis for understanding intermediate causal relationships between risk contributing factors and countermeasures that link hazards to outcomes. The HSIA risk DAG will be presented and described. Historically, critical malfunctions requiring Crew/MCC management occurred at a rate of 1.7 times per year for ISS averaged over the lifetime and 3-4 times per year in the burn in phase for the vehicle. These averages do not include EVA data, which greatly increases the incident rate. Prior experience from the Apollo program showed 10/11 crewed missions experienced significant anomalies where crew relied heavily on MCC expertise in real-time. These failure patterns are in line with those observed in other complex engineered systems (e.g., oil rigs, launch systems, commercial aviation, etc.) It is likely that general malfunction and error rates are > 10% for short duration missions (<30 days), based on past and current spaceflight operations data. Likelihood of adverse outcomes has the potential to increase as crew conduct work with new, complex systems and with less ground support. For Low Earth Orbit missions and Lunar missions less than 30 days, assuming minimal comm delays, disruptions and bandwidth limitations, malfunctions and errors can affect mission objectives and crew health but may be mitigated by ground support. For Lunar missions greater than 30 days and any potential Mars mission malfunctions and errors can have Loss of Crew and Loss of Mission consequences due to reduced ground support (communication delays, constraints and blackouts) for more complex operations, as well as reduced resupply and evacuation options.
Malfunction data from the thermal-vacuum tests of 39 flight-model spacecraft were analyzed. The results are interpreted in terms of the test variables, and in terms of the spacecraft performance. The malfunction data are correlated with the test time as a single variable, and also with the composite variable of time plus temperature. The improvement in spacecraft performance is examined by means of malfunction rates, malfunctions per spacecraft, and the probability of no failure related to test time. The minimum thermal-vacuum test profile required for Goddard Space Flight Center spacecraft is verified, and the probability of a defect remaining undetected is estimated.
Malfunction data from the thermal-vacuum tests of 39 flight-model spacecraft have been analyzed. The results are interpreted in terms of the test variables and the spacecraft performance. The malfunction data are correlated with the test time as a single variable, and also with the composite variable of time plus temperature. The improvement in spacecraft performance is examined by means of malfunction rates, malfunctions per spacecraft, and the probability of no failure related to test time. The minimum thermal-vacuum test profile required for Goddard Space Flight Center spacecraft is verified, and the probability of a defect remaining undetected is estimated.
Malfunction data from the thermal-vacuum tests of 39 flight-model spacecraft have been analyzed. These data are compared to the data listed in the 1968 Technical Note, NASA TN D-4908, 'Time Required for an Adequate Thermal-Vacuum Test of Flight Model Spacecraft', by A. R. Timmins. The present analyses include the relationship of malfunctions to time and temperature of the test, malfunction rates, effect of retest data, and the probability of a malfunction in each of four thermal-vacuum environments. Data are presented that relate a thermal-vacuum test profile to the risk involved using that profile. The minimum thermal-vacuum test profile is verified, and no change is recommended for the present thermal-vacuum test profile for flight model spacecraft.
The SBUV instrument, on Nimbus-7, measures the backscatter ultraviolet radiance at 12 wavelengths. The radiance data from these wavelengths was used to deduce the ozone profile and the total column ozone. In February 1987, there was an instrument malfunction. The purpose of this paper is to describe the malfunction, to determine the effect of the malfunction on the data quality, and if possible, to correct for the effects of the malfunction on the data from the SBUV instrument.
A report describes the history and the continuing evolution of an avionic system aboard the space shuttle, denoted the caution and warning system, that generates visual and auditory displays to alert astronauts to malfunctions. The report focuses mainly on planned human-factors-oriented upgrades of an alphanumeric fault-summary display generated by the system. Such upgrades are needed because the display often becomes cluttered with extraneous messages that contribute to the difficulty of diagnosing malfunctions. In the first of two planned upgrades, the fault-summary display will be rebuilt with a more logical task-oriented graphical layout and multiple text fields for malfunction messages. In the second upgrade, information displayed will be changed, such that text fields will indicate only the sources (that is, root causes) of malfunctions; messages that are not operationally useful will no longer appear on the displays. These and other aspects of the upgrades are based on extensive collaboration among astronauts, engineers, and human-factors scientists. The report describes the human-factors principles applied in the upgrades.
This study was the first in a series of planned tests to use physics-based subsystem simulations to investigate the interactions between a spacecraft's crew and a ground-based mission control center for vehicle subsystem operations across long communication delays. The simulation models the life support system of a deep space habitat. It contains models of an environmental control and life support system, an electrical power system, an active thermal control systems, and crew metabolic functions. The simulation has three interfaces: 1) a real-time crew interface that can be use to monitor and control the subsystems; 2) a mission control center interface with data transport delays up to 15 minute each way; and 3) a real-time simulation test conductor interface used to insert subsystem malfunctions and observe the interactions between the crew, ground, and simulated vehicle. The study was conducted at the 21st NASA Extreme Environment Mission Operations (NEEMO) mission. The NEEMO crew and ground support team performed a number of relevant deep space mission scenarios that included both nominal activities and activities with system malfunctions. While this initial test sequence was focused on test infrastructure and procedures development, the data collected in the study already indicate that long communication delays have notable impacts on the operation of deep space systems. For future human missions beyond cis-lunar, NASA will need to design systems and support tools to meet these challenges. These will be used to train the crew to handle critical malfunctions on their own, to predict malfunctions and assist with vehicle operations. Subsequent more detailed and involved studies will be conducted to continue advancing NASA's understanding of space systems operations across long communications delays.
This study was the first in a series of planned tests to use physics-based subsystem simulations to investigate the interactions between a spacecraft's crew and a ground-based mission control center for vehicle subsystem operations across long communication delays. The simulation models the life support system of a deep space habitat. It contains models of an environmental control and life support system, an electrical power system, an active thermal control system, and crew metabolic functions. The simulation has three interfaces: 1) a real-time crew interface that can be use to monitor and control the subsystems; 2) a mission control center interface with data transport delays up to 15 minute each way; and 3) a real-time simulation test conductor interface used to insert subsystem malfunctions and observe the interactions between the crew, ground, and simulated vehicle. The study was conducted at the 21st NASA Extreme Environment Mission Operations (NEEMO) mission. The NEEMO crew and ground support team performed a number of relevant deep space mission scenarios that included both nominal activities and activities with system malfunctions. While this initial test sequence was focused on test infrastructure and procedures development, the data collected in the study already indicate that long communication delays have notable impacts on the operation of deep space systems. For future human missions beyond cis-lunar, NASA will need to design systems and support tools to meet these challenges. These will be used to train the crew to handle critical malfunctions on their own, to predict malfunctions, and to assist with vehicle operations. Subsequent more detailed and involved studies will be conducted to continue advancing NASA's understanding of space systems operations across long communications delays.
Resilience of cyber-physical networks to unexpected failures is a critical need widely recognized across domains. For instance, power grids, telecommunication networks, transportation infrastructures, and water treatment systems have all been subject to disruptive malfunctions and catastrophic cyberattacks. Following such adverse events, we investigate scenarios where a node of a linear network suffers a loss of control authority over some of its actuators. These actuators are not following the controller's commands and are instead producing undesirable outputs. The repercussions of such a loss of control can propagate and destabilize the whole network despite the malfunction occurring at a single node. To assess system vulnerability, we establish resilience conditions for networks with a subsystem enduring a loss of control authority over some of its actuators. Furthermore, we quantify the destabilizing impact on the overall network when such a malfunction perturbs a nonresilient subsystem. We illustrate our resilience conditions on two academic examples, on an islanded microgrid and on the linearized IEEE 39-bus system.