Search NASA⌕ Search

SEARCH · Search NASA

Results for “common cause failures”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

Common Cause Failure Modeling

Common Cause Failures (CCFs) are a known and documented phenomenon that defeats system redundancy. CCFS are a set of dependent type of failures that can be caused by: system environments; manufacturing; transportation; storage; maintenance; and assembly, as examples. Since there are many factors that contribute to CCFs, the effects can be reduced, but they are difficult to eliminate entirely. Furthermore, failure databases sometimes fail to differentiate between independent and CCF (dependent) failure and data is limited, especially for launch vehicles. The Probabilistic Risk Assessment (PRA) of NASA's Safety and Mission Assurance Directorate at Marshall Space Flight Center (MFSC) is using generic data from the Nuclear Regulatory Commission's database of common cause failures at nuclear power plants to estimate CCF due to the lack of a more appropriate data source. There remains uncertainty in the actual magnitude of the common cause risk estimates for different systems at this stage of the design. Given the limited data about launch vehicle CCF and that launch vehicles are a highly redundant system by design, it is important to make design decisions to account for a range of values for independent and CCFs. When investigating the design of the one-out-of-two component redundant system for launch vehicles, a response surface was constructed to represent the impact of the independent failure rate versus a common cause beta factor effect on a system's failure probability. This presentation will define a CCF and review estimation calculations. It gives a summary of reduction methodologies and a review of examples of historical CCFs. Finally, it presents the response surface and discusses the results of the different CCFs on the reliability of a one-out-of-two system.

Hark, Frank↗

Common Cause Failure Modeling

Common Cause Failures (CCFs) are a known and documented phenomenon that defeats system redundancy. CCFS are a set of dependent type of failures that can be caused by: system environments; manufacturing; transportation; storage; maintenance; and assembly, as examples. Since there are many factors that contribute to CCFs, the effects can be reduced, but they are difficult to eliminate entirely. Furthermore, failure databases sometimes fail to differentiate between independent and CCF (dependent) failure and data is limited, especially for launch vehicles. The Probabilistic Risk Assessment (PRA) of NASA's Safety and Mission Assurance Directorate at Marshal Space Flight Center (MFSC) is using generic data from the Nuclear Regulatory Commission's database of common cause failures at nuclear power plants to estimate CCF due to the lack of a more appropriate data source. There remains uncertainty in the actual magnitude of the common cause risk estimates for different systems at this stage of the design. Given the limited data about launch vehicle CCF and that launch vehicles are a highly redundant system by design, it is important to make design decisions to account for a range of values for independent and CCFs. When investigating the design of the one-out-of-two component redundant system for launch vehicles, a response surface was constructed to represent the impact of the independent failure rate versus a common cause beta factor effect on a system's failure probability. This presentation will define a CCF and review estimation calculations. It gives a summary of reduction methodologies and a review of examples of historical CCFs. Finally, it presents the response surface and discusses the results of the different CCFs on the reliability of a one-out-of-two system.

Hark, Frank↗

Common Cause Failures and Ultra Reliability

A common cause failure occurs when several failures have the same origin. Common cause failures are either common event failures, where the cause is a single external event, or common mode failures, where two systems fail in the same way for the same reason. Common mode failures can occur at different times because of a design defect or a repeated external event. Common event failures reduce the reliability of on-line redundant systems but not of systems using off-line spare parts. Common mode failures reduce the dependability of systems using off-line spare parts and on-line redundancy.

reliability↗

Common Cause Failures Dominate and Defeat Redundancy

Common cause failures occur when several malfunctions are produced by a single event or process. They are especially damaging when they eliminate an entire set of redundant systems and disable their intended function. Redundancy is used when the individual system failure probability is unacceptably high. Redundancy can improve the overall system failure probability if the failures are independent, but the reliability gain is limited if there are dependent failures having a common cause. No amount of redundancy can reduce the total failure probability below the common cause failure probability. Common cause failures defeat redundancy. Systems with high reliability requirements often use extensive redundancy. These highly redundant systems rarely fail unless all the redundant components providing a particular function fail. Complete failures of such highly redundant systems are then usually common cause failures. Common cause failures are prevalent in highly redundant, high reliability systems. Common cause failures dominate redundancy. Redundant systems may fail due to specification, design, manufacturing, operations, or maintenance problems that disable all the identical redundant systems. Common cause failures typically account for one tenth of all failures. If the failure probability is relatively low and common cause failures are significant, adding more than two or three redundant identical units usually gives little added reliability improvement. Common cause failures can be reduced by using diverse components with different technologies and manufacturers, by separating and shielding subsystems, and by avoiding shared control, power, or location. External events and shared vulnerabilities may still cause common cause failures.

common cause failures↗

Common Cause Failures Dominate and Defeat Redundancy

Common cause failures occur when several malfunctions are produced by a single event or process. They are especially damaging when they eliminate an entire set of redundant systems and disable their intended function. Redundancy is used when the individual system failure probability is unacceptably high. Redundancy can improve the overall system failure probability if the failures are independent, but the reliability gain is limited if there are dependent failures having a common cause. No amount of redundancy can reduce the total failure probability below the common cause failure probability. Common cause failures defeat redundancy. Systems with high reliability requirements often use extensive redundancy. These highly redundant systems rarely fail unless all the redundant components providing a particular function fail. Complete failures of such highly redundant systems are then usually common cause failures. Common cause failures are prevalent in highly redundant, high reliability systems. Common cause failures dominate redundancy. Redundant systems may fail due to specification, design, manufacturing, operations, or maintenance problems that disable all the identical redundant systems. Common cause failures typically account for one tenth of all failures. If the failure probability is relatively low and common cause failures are significant, adding more than two or three redundant identical units usually gives little added reliability improvement. Common cause failures can be reduced by using diverse components with different technologies and manufacturers, by separating and shielding subsystems, and by avoiding shared control, power, or location. External events and shared vulnerabilities may still cause common cause failures.

common cause failures↗

Common Cause Failure Modeling in Space Launch Vehicles

Common Cause Failures (CCFs) are a known and documented phenomenon that defeats system redundancy. CCFs are a set of dependent type of failures that can be caused for example by system environments, manufacturing, transportation, storage, maintenance, and assembly. Since there are many factors that contribute to CCFs, they can be reduced, but are difficult to eliminate entirely. Furthermore, failure databases sometimes fail to differentiate between independent and dependent CCF. Because common cause failure data is limited in the aerospace industry, the Probabilistic Risk Assessment (PRA) Team at Bastion Technology Inc. is estimating CCF risk using generic data collected by the Nuclear Regulatory Commission (NRC). Consequently, common cause risk estimates based on this database, when applied to other industry applications, are highly uncertain. Therefore, it is important to account for a range of values for independent and CCF risk and to communicate the uncertainty to decision makers. There is an existing methodology for reducing CCF risk during design, which includes a checklist of 40+ factors grouped into eight categories. Using this checklist, an approach to produce a beta factor estimate is being investigated that quantitatively relates these factors. In this example, the checklist will be tailored to space launch vehicles, a quantitative approach will be described, and an example of the method will be presented.

Hark, Frank↗

Common Cause Failure Modeling: Aerospace Versus Nuclear

Aggregate nuclear plant failure data is used to produce generic common-cause factors that are specifically for use in the common-cause failure models of NUREG/CR-5485. Furthermore, the models presented in NUREG/CR-5485 are specifically designed to incorporate two significantly distinct assumptions about the methods of surveillance testing from whence this aggregate failure data came. What are the implications of using these NUREG generic factors to model the common-cause failures of aerospace systems? Herein, the implications of using the NUREG generic factors in the modeling of aerospace systems are investigated in detail and strong recommendations for modeling the common-cause failures of aerospace systems are given.

Stott, James E.↗

Modeling Common Cause Failures of Thrusters on ISS Visiting Vehicles

This paper discusses the methodology used to model common cause failures of thrusters on the International Space Station (ISS) Visiting Vehicles. The ISS Visiting Vehicles each have as many as 32 thrusters, whose redundancy makes them susceptible to common cause failures. The Global Alpha Model (as described in NUREG/CR‐5485) can be used to represent the system common cause contribution, but NUREG/CR‐5496 supplies global alpha parameters for groups only up to size six. Because of the large number of redundant thrusters on each vehicle, regression is used to determine parameter values for groups of size larger than six. An additional challenge is that Visiting Vehicle thruster failures must occur in specific combinations in order to fail the propulsion system; not all failure groups of a certain size are critical.

Haught, Megan↗

Modeling Common Cause Failures of Thrusters on ISS Visiting Vehicles

This paper discusses the methodology used to model common cause failures of thrusters on the International Space Station (ISS) Visiting Vehicles. The ISS Visiting Vehicles each have as many as 32 thrusters, whose redundancy and similar design make them susceptible to common cause failures. The Global Alpha Model (as described in NUREG/CR-5485) can be used to represent the system common cause contribution, but NUREG/CR-5496 supplies global alpha parameters for groups only up to size six. Because of the large number of redundant thrusters on each vehicle, regression is used to determine parameter values for groups of size larger than six. An additional challenge is that Visiting Vehicle thruster failures must occur in specific combinations in order to fail the propulsion system; not all failure groups of a certain size are critical.

Haught, Megan↗

Common Cause Failure Modeling

Space Launch System (SLS) Agenda: Objective; Key Definitions; Calculating Common Cause; Examples; Defense against Common Cause; Impact of varied Common Cause Failure (CCF) and abortability; Response Surface for various CCF Beta; Takeaways.

Hark, Frank↗

Common Cause Failure Modes

High technology industries with high failure costs commonly use redundancy as a means to reduce risk. Redundant systems, whether similar or dissimilar, are susceptible to Common Cause Failures (CCF). CCF is not always considered in the design effort and, therefore, can be a major threat to success. There are several aspects to CCF which must be understood to perform an analysis which will find hidden issues that may negate redundancy. This paper will provide definition, types, a list of possible causes and some examples of CCF. Requirements and designs from NASA projects will be used in the paper as examples.

Wetherholt, Jon↗

Historical Aerospace Software Errors Categorized to Influence Fault Tolerance

- Motivation - Very little literature exists characterizing software errors in real-time avionic systems - How, where, and why is software most likely to fail? - Purpose - Raise awareness of how software fails through historical study - Recommend improvements to software fault tolerant design based on historical study - Outline - Discuss Software Failures - Common Cause, Failure Classes, Mitigation strategies - Review NASA requirements for Software Fault Tolerance - Review Historical Software Failures - Analyze failures and provide statistics - Erroneous vs. fail-Silent - Reboot recoverability likelihood - Code Location - Missing or unknown code?

Flilght↗

We Can't Count on Repairing All Failures Going to Mars

Reliability analysis often assumes that a complex system can be kept operating indefinitely with scheduled maintenance and emergency repair using a stock of spare parts, as long as the spare parts are not depleted. This assumption seems justified for well-tested, widely used, long operational systems with a multigenerational history of failure, redesign, and reliability growth. It seems doubtful that newer, relatively untried, high technology space systems can always be repaired. We cannot assume space systems will have a low rate of random failures that can all be repaired with a few identical spares. New untried systems usually have a high initial failure rate, called infant mortality, due to errors in requirements, design, parts, materials, and operations planning. These problems can cause groups of related failures called Common Cause Failures (CCFs). The practical definition of a CCF is any failure mode that cannot be cured using identical redundant systems or spare parts. Systems with CCFs may fail repeatedly for the same reason. Can a life support system be kept operating on the way to Mars using only redundant systems and spare parts? The failure history of International Space Station (ISS) life support systems suggests that CCFs are likely to occur and will probably require design changes rather than being reparable with spare parts.

Mars↗

An Efficient Approach for the Reliability Analysis of Phased-Mission Systems with Dependent Failures

We consider the reliability analysis of phased-mission systems with common-cause failures in this paper. Phased-mission systems (PMS) are systems supporting missions characterized by multiple, consecutive, and nonoverlapping phases of operation. System components may be subject to different stresses as well as different reliability requirements throughout the course of the mission. As a result, component behavior and relationships may need to be modeled differently from phase to phase when performing a system-level reliability analysis. This consideration poses unique challenges to existing analysis methods. The challenges increase when common-cause failures (CCF) are incorporated in the model. CCF are multiple dependent component failures within a system that are a direct result of a shared root cause, such as sabotage, flood, earthquake, power outage, or human errors. It has been shown by many reliability studies that CCF tend to increase a system's joint failure probabilities and thus contribute significantly to the overall unreliability of systems subject to CCF.We propose a separable phase-modular approach to the reliability analysis of phased-mission systems with dependent common-cause failures as one way to meet the above challenges in an efficient and elegant manner. Our methodology is twofold: first, we separate the effects of CCF from the PMS analysis using the total probability theorem and the common-cause event space developed based on the elementary common-causes; next, we apply an efficient phase-modular approach to analyze the reliability of the PMS. The phase-modular approach employs both combinatorial binary decision diagram and Markov-chain solution methods as appropriate. We provide an example of a reliability analysis of a PMS with both static and dynamic phases as well as CCF as an illustration of our proposed approach. The example is based on information extracted from a Mars orbiter project. The reliability model for this orbiter considers the various phases of Launch, Cruise, Mars Orbit Insertion, and Orbit. Some of the CCF for the orbiter in this mission include environmental effects, such as micrometeoroids, human operator errors, and software errors.

reliability analysis↗

Diverse Redundant Systems for Reliable Space Life Support

Reliable life support systems are required for deep space missions. The probability of a fatal life support failure should be less than one in a thousand in a multi-year mission. It is far too expensive to develop a single system with such high reliability. Using three redundant units would require only that each have a failure probability of one in ten over the mission. Since the system development cost is inverse to the failure probability, this would cut cost by a factor of one hundred. Using replaceable subsystems instead of full systems would further cut cost. Using full sets of replaceable components improves reliability more than using complete systems as spares, since a set of components could repair many different failures instead of just one. Replaceable components would require more tools, space, and planning than full systems or replaceable subsystems. However, identical system redundancy cannot be relied on in practice. Common cause failures can disable all the identical redundant systems. Typical levels of common cause failures will defeat redundancy greater than two. Diverse redundant systems are required for reliable space life support. Three, four, or five diverse redundant systems could be needed for sufficient reliability. One system with lower level repair could be substituted for two diverse systems to save cost.

life support↗

Fault Management Algorithm Risk Assessment for the NASA Space Launch System

This paper presents the false positive (FP) and false negative (FN) risk assessment process currently being conducted for the Space Launch System (SLS) Artemis II Fault Management (FM) detection functions. The analysis scope, general assumptions and guide rules, and key modeling concepts were discussed to establish the basis of the risk assessments conducted. Initial analyses indicated a dominance in the total risk by software and firmware failures. This paper presents efforts applied to refine the software risks and the overall impact of implementing those modifications. Current analyses conducted on the detection functions implemented for the SLS Artemis II mission indicate primary risk drivers for the individual FM detection functions are flight software failures, firmware design failures, and hardware Common Cause Failures (CCFs). There still remains issues of how to account for time and redundancy in the software risk estimations.

probability risk analysis↗