Research on failure free systems Quarterly report no. 8, Mar. 23 - Jun. 23, 1966
Failure free electronic systems - computer simulation program documentation and redundant system test point allocation and reliability estimation
SEARCH · Search NASA
Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.
Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.
Failure free electronic systems - computer simulation program documentation and redundant system test point allocation and reliability estimation
This study explores two distinct Balance of Plant (BOP) configurations: the Rankine cycle for a sodium-cooled fast reactor (SFR) and the Brayton cycle for a gas-cooled reactor (GCR). As representative designs, the Power Reactor Innovative Small Module (PRISM) by GE Hitachi Nuclear Energy was selected for the SFR, while the Gas Turbine Modular Helium Reactor (GT MHR) by General Atomics was chosen for the gas-cooled reactor. Both configurations were adapted to deliver high-quality heat for industrial applications. A Failure Modes and Effects Analysis (FMEA) was conducted for each system to identify critical failure modes affecting key components. This study marks the first phase of a two-step design optimization approach, integrating FMEA into the design process. Based on the analysis, design modifications and mitigation strategies were proposed to enhance system resilience. The second phase, to be detailed in a subsequent report, will focus on the role of the control system in mitigating these failures. The FMEA serves as a foundation for defining the control topology, ensuring system resilience against component failures that could compromise essential functions, such as electricity and heat production.
This paper presents a scheme for on-line failure detection in systems with appreciable structural dynamics. The design is suboptimal because of extensive computational requirements of optimal schemes. To accomplish failure detection, the innovations sequence of a finite order Kalman filter is examined. Because of the heavy dependence of the system on the zero-mean character of the innovations sequence of the filter much attention has been given to the design and evaluation of the Kalman filters used. The filter designs are based on modal models of the structural dynamics. Two modal models were considered, one based on an analytic finite element model and the other based on empirically derived frequency and damping. Experiments using a grid structure are presented which illustrate operation and performance of the filter designs based on these models. The general character of the results presented is that appreciable errors exist in the filter design based on the finite element model. Substantial improvement results if the design model is modified to include empirically derived frequency and damping.
A complex system has many parts and interactions and so is difficult to understand. Systems with higher complexity generally have higher costs and failure rates. A System Complexity Metric (SCM) is defined to be the sum of the number of nodes, N, in the system block diagram plus the number of one-way interactions, I, between the nodes. SCM = N + I. SCMs are easily determined by direct inspection of high level block diagrams of life support systems. System cost was found to be directly proportional to SCM. The system MTBF (Mean Time Before Failure) is the inverse of the system failure rate. MTBF = 1/f. The system MTBF was found to be proportional to SCM^(-2.2) for estimated preflight MTBFs. As is typical for systems that are not extensively tested and redesigned to eliminate unexpected failure modes, the life support flight failure rates were about ten times higher than the preflight estimates and the MTBFs one-tenth the preflight estimates. The system MTBF was found to be proportional to SCM^(-2.6) for observed flight MTBFs.
Common cause failures occur when several malfunctions are produced by a single event or process. They are especially damaging when they eliminate an entire set of redundant systems and disable their intended function. Redundancy is used when the individual system failure probability is unacceptably high. Redundancy can improve the overall system failure probability if the failures are independent, but the reliability gain is limited if there are dependent failures having a common cause. No amount of redundancy can reduce the total failure probability below the common cause failure probability. Common cause failures defeat redundancy. Systems with high reliability requirements often use extensive redundancy. These highly redundant systems rarely fail unless all the redundant components providing a particular function fail. Complete failures of such highly redundant systems are then usually common cause failures. Common cause failures are prevalent in highly redundant, high reliability systems. Common cause failures dominate redundancy. Redundant systems may fail due to specification, design, manufacturing, operations, or maintenance problems that disable all the identical redundant systems. Common cause failures typically account for one tenth of all failures. If the failure probability is relatively low and common cause failures are significant, adding more than two or three redundant identical units usually gives little added reliability improvement. Common cause failures can be reduced by using diverse components with different technologies and manufacturers, by separating and shielding subsystems, and by avoiding shared control, power, or location. External events and shared vulnerabilities may still cause common cause failures.
Common cause failures occur when several malfunctions are produced by a single event or process. They are especially damaging when they eliminate an entire set of redundant systems and disable their intended function. Redundancy is used when the individual system failure probability is unacceptably high. Redundancy can improve the overall system failure probability if the failures are independent, but the reliability gain is limited if there are dependent failures having a common cause. No amount of redundancy can reduce the total failure probability below the common cause failure probability. Common cause failures defeat redundancy. Systems with high reliability requirements often use extensive redundancy. These highly redundant systems rarely fail unless all the redundant components providing a particular function fail. Complete failures of such highly redundant systems are then usually common cause failures. Common cause failures are prevalent in highly redundant, high reliability systems. Common cause failures dominate redundancy. Redundant systems may fail due to specification, design, manufacturing, operations, or maintenance problems that disable all the identical redundant systems. Common cause failures typically account for one tenth of all failures. If the failure probability is relatively low and common cause failures are significant, adding more than two or three redundant identical units usually gives little added reliability improvement. Common cause failures can be reduced by using diverse components with different technologies and manufacturers, by separating and shielding subsystems, and by avoiding shared control, power, or location. External events and shared vulnerabilities may still cause common cause failures.
A Markov model of a highly reliable triplex system was constructed to evaluate the probability of system failure as a function of the propagation of latent faults. It is found that if the propagation rate of latent faults is extremely high, they do not significantly affect the probability of system failure, while if the propagation rate is extremely low, the survivability of the system is improved. The propagation rate that is most harmful to the survivability of the system is determined as a function of the duration of the flight. A decrease in the probability of system failure due to latency is noted if the probability of any two faults giving the same output is extremely low.
Advanced concepts for detecting, isolating, and accommodating sensor failures were studied to determine their applicability to the gas turbine control problem. Five concepts were formulated based upon such techniques as Kalman filters and a screening process led to the selection of one advanced concept for further evaluation. The selected advanced concept uses a Kalman filter to generate residuals, a weighted sum square residuals technique to detect soft failures, likelihood ratio testing of a bank of Kalman filters for isolation, and reconfiguring of the normal mode Kalman filter by eliminating the failed input to accommodate the failure. The advanced concept was compared to a baseline parameter synthesis technique. The advanced concept was shown to be a viable concept for detecting, isolating, and accommodating sensor failures for the gas turbine applications.
An expert system for diagnosis and recovery of failures in the Freon cooling loop of the European retrievable experiment carrier EURECA is described. The system demonstrates the feasibility of a functional scope of expert diagnostic systems which appears to be essential for practical applications of such systems in space technology. The scope includes early warning and treatment of incomplete information, fault tolerance, identification of failure superpositions, intelligent reaction to unforeseen events, and detailed status display for optimal recovery action.
Aerobee rocket propulsion system failure
This paper is concerned with the reinitialization of fault tolerant systems in which detection and isolation (FDI) techniques are used, on-line, to identify and compensate for system failures. Specifically, it will focus on FDI techniques which utilize analytic redundancy, arising from a knowledge of the plant dynamics, by analyzing the residuals of a no-fail filter designed on the assumption of no failures. In these types of fault tolerant systems, system failures have to propagate through the no-fail filter dynamics in order to get detected. Therefore, the no-fail filter must be reinitialized after the isolation of a failure so that the accumulated effects of the failure are removed. In this paper, various approaches to this reinitialization problem will be discussed.
A conceptual design of a model based failure detection and diagnosis system is developed for the space shuttle main engine. This design relies on the accurate and reliable identification of the parameters of the highly nonlinear and very complex engine. The design approach is presented in some detail and results for a failed valve are presented. These preliminary results verify that the developed parameter identification technique together with a neural network classifier can be used for this purpose.
A conceptual design of a model based failure deeteection and diagnosis system is deeveloped for the Space Shuttle Main Engine. This design relies on the accurate and reliable identification of the parameters of the highly nonlinear and very complex engine. The design approach is presented in some detail and results for a failed valve are presented. These preliminary results verify that the developed parameter identification technique, together with a neural network classifier, can be used for this purpose.
This white paper explores how to increase the success and operation of critical, complex, national systems by effectively capturing knowledge management requirements within the federal acquisition process. Although we focus on aerospace flight systems, the principles outlined within may have a general applicability to other critical federal systems as well. Fundamental design deficiencies in federal, mission-critical systems have contributed to recent, highly visible system failures, such as the V-22 Osprey and the Delta rocket family. These failures indicate that the current mechanisms for knowledge management and risk management are inadequate to meet the challenges imposed by the rising complexity of critical systems. Failures of aerospace system operations and vehicles may have been prevented or lessened through utilization of better knowledge management and information management techniques.
Following the Long Duration Exposure Facility (LDEF) retrieval, the Systems Special Investigation Group (SIG) participated in an extensive series of tests of various electronic systems, including the NASA provided data and initiate systems, and some experiment systems. Overall, these were found to have performed remarkably well, even though most were designed and tested under limited budgets and used at least some nonspace qualified components. However, several anomalies were observed, including a few which resulted in some loss of data. The postflight test program objectives, observations, and lessons learned from these examinations are discussed. All analyses are not yet complete, but observations to date will be summarized, including the Boeing experiment component studies and failure analysis results related to the Interstellar Gas Experiment. Based upon these observations, suggestions for avoiding similar problems on future programs are presented.
We consider how a jet transport airplane interface supports the flight crew in managing airplane system failures (or non-normals) for continued safe flight and landing. The existing state of the art starts with a list of airplane system component failures and asks the flight crew to determine, with the help of non-normal procedures, the operational consequences of those failures. As airplane systems become more complex and interconnected, the flight crew's ability to determine operational consequences will become inadequate. We describe an approach that attempts to translate airplane system failures directly into airplane "capabilities," which is a set of basic airplane functions, such as the ability to stop after landing. This paper describes the overall framework for supporting flight crews in operational decision making and the initial efforts to develop a language and display concepts.
All failure detection methods are based, either explicitly or implicitly, on the use of redundancy, i.e. on (possibly dynamic) relations among the measured variables. The robustness of the failure detection process consequently depends to a great degree on the reliability of the redundancy relations, which in turn is affected by the inevitable presence of model uncertainties. In this paper the problem of determining redundancy relations that are optimally robust is addressed in a sense that includes several major issues of importance in practical failure detection and that provides a significant amount of intuition concerning the geometry of robust failure detection. A procedure is given involving the construction of a single matrix and its singular value decomposition for the determination of a complete sequence of redundancy relations, ordered in terms of their level of robustness. This procedure also provides the basis for comparing levels of robustness in redundancy provided by different sets of sensors.
Several topics are presented in viewgraph form which together encompass the preliminary assessment of nuclear thermal rocket engine clustering. The study objectives, schedule, flow, and groundrules are covered. This is followed by the NASA groundrules mission and our interpretation of the associated operational scenario. The NASA reference vehicle is illustrated, then the four propulsion system options are examined. Each propulsion system's preliminary design, fluid systems, operating characteristics, thrust structure, dimensions, and mass properties are detailed as well as the associated key propulsion system/vehicle interfaces. A brief series of systems analysis is also covered including: thrust vector control requirements, engine out possibilities, propulsion system failure modes, surviving system requirements, and technology requirements. An assessment of vehicle/propulsion system impacts due to the lessons learned are presented.