Search NASA⌕ Search

SEARCH · Search NASA

Results for “common cause failures”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 37 records · Page 2

Fault Management Algorithm Risk Assessment for the NASA Space Launch System

This paper presents the false positive (FP) and false negative (FN) risk assessment process currently being conducted for the Space Launch System (SLS) Artemis II Fault Management (FM) detection functions. The analysis scope, general assumptions and guide rules, and key modeling concepts were discussed to establish the basis of the risk assessments conducted. Initial analyses indicated a dominance in the total risk by software and firmware failures. This paper presents efforts applied to refine the software risks and the overall impact of implementing those modifications. Current analyses conducted on the detection functions implemented for the SLS Artemis II mission indicate primary risk drivers for the individual FM detection functions are flight software failures, firmware design failures, and hardware Common Cause Failures (CCFs). There still remains issues of how to account for time and redundancy in the software risk estimations.

probability risk analysis↗

Fault Management Algorithm Risk Assessment for the NASA Space Launch System

This presentation describes the false positive (FP) and false negative (FN) risk assessment process currently being conducted for the Space Launch System (SLS) Artemis II Fault Management (FM) detection functions. The analysis scope, general assumptions and guide rules, and key modeling concepts were discussed to establish the basis of the risk assessments conducted. Initial analyses indicated a dominance in the total risk by software and firmware failures. This paper presents efforts applied to refine the software risks and the overall impact of implementing those modifications. Current analyses conducted on the detection functions implemented for the SLS Artemis II mission indicate primary risk drivers for the individual FM detection functions are flight software failures, firmware design failures, and hardware Common Cause Failures (CCFs). There still remains issues of how to account for time and redundancy in the software risk estimations.

probability risk analysis↗

Operating Experience Data Analysis for Digital Instrumentation and Control System Reliability and Risk Assessment in Nuclear Power Plants

The implementation of advanced digital instrumentation and control (DI&C) systems in U.S. nuclear power plants (NPPs) can bring significant advancements in reliability, monitoring, and control capabilities. However, these systems also introduce new challenges, particularly in assessing risks such as common-cause failures (CCFs) and establishing robust reliability estimates for DI&C components. Addressing these challenges is critical for ensuring the safe and efficient operation of NPPs. Recently, Idaho National Laboratory was tasked by the U.S. Nuclear Regulatory Commission (NRC) to conduct a DI&C reliability study using operating experience data from the nuclear industry. The two operating experience data sources for the study are the Institute of Nuclear Power Operations’ Industry Reporting and Information System (IRIS) and the NRC’s Licensee Event Report database which is hosted at Idaho National Laboratory at https://lersearch.inl.gov/LERSearchCriteria.aspx. This report provides a comprehensive examination of DI&C systems, including their architecture, operational advantages, and associated challenges. It reviews existing industry DI&C studies and failure mode taxonomies, along with reliability data from various industries. Through a detailed analysis of these databases, the study provides insights into DI&C system performance. Considerations should be given to incorporate DI&C failure data into the NRC's Integrated Data Collection and Coding System and updating the Reliability and Availability Data System to support ongoing DI&C reliability studies. Recommendations are also provided for modeling DI&C reliability and CCF in probabilistic risk assessment, thereby supporting risk-informed decision-making and enhancing the reliability and safety of NPPs.

22 GENERAL STUDIES OF NUCLEAR REACTORS↗

VIPER project

The VIPER project has so far produced a formal specification of a 32 bit RISC microprocessor, an implementation of that chip in radiation-hard SOS technology, a partial proof of correctness of the implementation which is still being extended, and a large body of supporting software. The time has now come to consider what has been achieved and what directions should be pursued in the future. The most obvious lesson from the VIPER project was the time and effort needed to use formal methods properly. Most of the problems arose in the interfaces between different formalisms, e.g., between the (informal) English description and the HOL spec, between the block-level spec in HOL and the equivalent in ELLA needed by the low-level CAD tools. These interfaces need to be made rigorous or (better) eliminated. VIPER 1A (the latest chip) is designed to operate in pairs, to give protection against breakdowns in service as well as design faults. We have come to regard redundancy and formal design methods as complementary, the one to guard against normal component failures and the other to provide insurance against the risk of the common-cause failures which bedevil reliability predictions. Any future VIPER chips will certainly need improved performance to keep up with increasingly demanding applications. We have a prototype design (not yet specified formally) which includes 32 and 64 bit multiply, instruction pre-fetch, more efficient interface timing, and a new instruction to allow a quick response to peripheral requests. Work is under way to specify this device in MIRANDA, and then to refine the spec into a block-level design by top-down transformations. When the refinement is complete, a relatively simple proof checker should be able to demonstrate its correctness. This paper is presented in viewgraph form.

Kershaw, John↗

Bayesian And Human Reliability Analysis (hra)-aided Method For The Reliability Analysis Of Software (bahamas)

The purpose of the BAHAMAS code is to provide a simplified process for performing quantitative evaluations of software reliability. The Bayesian and Human Reliability Analysis (HRA)-Aided method for the Reliability Analysis of software (BAHAMAS) was developed specifically to perform quantification under limited data conditions, i.e., when limited testing or operational data are available, such as during early development stages. BAHAMAS essentially examines the quality of a software development life cycle to determine the probability of specific types of software failure. BAHAMAS will have modules to support user input for detailed and simplified analyses. The user interface will also support software common cause failure analysis.

Wang, Congjian (0000000207789927)↗

Comparative Analysis of Static and Dynamic Probabilistic Risk Assessment

This study examines three different methodologies for producing loss-of-mission (LOM) and loss-of-crew (LOC) risks estimates for probabilistic risk assessments (PRA) of crewed spacecraft. The three bottom-up, component-based PRA approaches examined are a traditional static fault tree, a dynamic Monte Carlo simulation, and a fault tree hybrid that incorporates some dynamic elements. These approaches were used to model the reaction control system thruster pod of a generic crewed spacecraft and mission, and a comparative analysis of the methods is presented. The methodologies are assessed in terms of the process of modeling a system, the actionable information produced for the design team, and the overall fidelity of the quantitative risk evaluation generated. The system modeling process is compared in terms of the effort required to generate the initial model, update the model in response to design changes, and support mass-versus-risk trade studies. The results are compared by examining the top-level LOM/LOC estimates and the relative risk driver rankings at the failure mode level. The fidelity of each modeling methodology is discussed in terms of its capability to handle real-world system dynamics such as cold-sparing, changes in mission operations due to loss of redundancy, and common cause failure modes. The paper also discusses the applicability of each methodology to different phases of system development and shows that a single methodology may not be suitable for all of the many purposes of a spacecraft PRA. The fault tree hybrid approach is shown to be best suited to the needs of early assessments during conceptual design phases. As the design begins to mature, the level of detail represented in the risk model must go beyond redundancy and nominal mission operations to include dynamic, time- and state-dependent system responses as well as diverse system capabilities. This is best accomplished using the dynamic simulation approach, since these phenomena are not easily captured by static methods. Ultimately, once the design has been finalized and the goal of the PRA is to provide design validation and requirement verification, more traditional, static fault tree approaches may become as appropriate as the simulation method.

Mattenberger, Christopher J.↗

Conical Seat Shut-Off Valve

A moveable valve for controlling flow of a pressurized working fluid was designed. This valve consists of a hollow, moveable floating piston pressed against a stationary solid seat, and can use the working fluid to seal the valve. This open/closed, novel valve is able to use metal-to-metal seats, without requiring seat sliding action; therefore there are no associated damaging effects. During use, existing standard high-pressure ball valve seats tend to become damaged during rotation of the ball. Additionally, forces acting on the ball and stem create large amounts of friction. The combination of these effects can lead to system failure. In an attempt to reduce damaging effects and seat failures, soft seats in the ball valve have been eliminated; however, the sliding action of the ball across the highly loaded seat still tends to scratch the seat, causing failure. Also, in order to operate, ball valves require the use of large actuators. Positioning the metal-to-metal seats requires more loading, which tends to increase the size of the required actuator, and can also lead to other failures in other areas such as the stem and bearing mechanisms, thus increasing cost and maintenance. This novel non-sliding seat surface valve allows metal-to-metal seats without the damaging effects that can lead to failure, and enables large seating forces without damaging the valve. Additionally, this valve design, even when used with large, high-pressure applications, does not require large conventional valve actuators and the valve stem itself is eliminated. Actuation is achieved with the use of a small, simple solenoid valve. This design also eliminates the need for many seals used with existing ball valve and globe valve designs, which commonly cause failure, too. This, coupled with the elimination of the valve stem and conventional valve actuator, improves valve reliability and seat life. Other mechanical liftoff seats have been designed; however, they have only resulted in increased cost, and incurred other reliability issues. With this novel design, the seat is lifted by simply removing the working fluid pressure that presses it against the seat and no external force is required. By eliminating variables associated with existing ball and globe configurations that can have damaging effects upon a valve, this novel design reduces downtime in rocket engine test schedules and maintenance costs.

Farner, Bruce↗

Evaluation of Hardware and Software Bill of Materials (HBOMs/SBOMs) Extraction Methods

Hardware and software bills of materials (HBOMs and SBOMs) provide important visibility into the components, dependencies, and supply chain relationships within programmable digital devices. This visibility is critical for advanced nuclear reactor applications, where use of common or shared hardware components, software libraries, suppliers, or manufacturing processes may create common cause failure (CCF) vulnerabilities despite apparent diversity. This paper evaluates current approaches for obtaining and analyzing HBOMs and SBOMs in support of CCF, diversity and defense-in-depth (D3) assessments, and begins to explore potential methods for artificial intelligence/machine learning-based analysis. The availability of BOM information from advanced reactor manufacturers and vendors, representative hardware and software categories found in advanced reactor systems continues to limit research [13]. This paper compares commonly used BOM formats, including CycloneDX, SPDX, and SWID. It also surveys publicly available tools for generating BOMs from source code, compiled binaries, and hardware-related information, noting limitations in language coverage, system age, and format interoperability. Finally, this paper evaluates methods for correlating BOM data with vulnerability and exploitability information, including VEX, CVE, and CWE resources. The findings indicate that publicly available nuclear-vendor BOMs are limited, making third-party extraction and research into novel analysis techniques necessary.

Cybersecurity↗

Programmable Digital Devices used in Advanced Reactors

This paper introduces the concepts of common cause failure, diversity, and defense-in-depth used by the nuclear industry to analyze resilience in reactors. A survey of publicly traded and private companies building advanced reactors and their licensing status is presented. Safety and non-safety systems found in the NuScale Power design are summarized and the likely hardware and software categories used by those systems are enumerated. The importance of industry partners is highlighted. This paper also identifies an alternate path forward without industry partners to advance the knowledge needed to use artificial intelligence to analyze HBOMs and SBOMs to better understand reactor resiliency.

cybersecurity↗

Aircraft cabin water spray disbenefits study

The concept of utilizing a cabin water spray system (CWSS) as a means of increasing passenger evacuation and survival time following an accident has received considerable publicity and has been the subject of testing by the regulatory agencies in both the United States and Europe. A test program, initiated by the CAA in 1987, involved the regulatory bodies in both Europe and North America in a collaborative research effort to determine the benefits and 'disbenefits' (disadvantages) of a CWSS. In order to obtain a balanced opinion of an onboard CWSS, NASA, and FAA requested the Boeing Commercial Airplane Group to investigate the potential 'disbenefits' of the proposed system from the perspective of the manufacturer and an operator. This report is the result of a year-long, cost-sharing contract study between the Boeing Commercial Airplane Group, NASA, and FAA. Delta Air Lines participated as a subcontract study team member and investigated the 'return to service' costs for an aircraft that would experience an uncommanded operation of a CWSS without the presence of fire. Disbenefits identified include potential delays in evacuation, introduction of 'common cause failure' in redundant safety of flight systems, physiological problems for passengers, high cost of refurbishment for inadvertent discharge, and potential to negatively affect other safety systems.

Reynolds, Thomas L.↗

Factors which Limit the Value of Additional Redundancy in Human Rated Launch Vehicle Systems

The National Aeronautics and Space Administration (NASA) has embarked on an ambitious program to return humans to the moon and beyond. As NASA moves forward in the development and design of new launch vehicles for future space exploration, it must fully consider the implications that rule-based requirements of redundancy or fault tolerance have on system reliability/risk. These considerations include common cause failure, increased system complexity, combined serial and parallel configurations, and the impact of design features implemented to control premature activation. These factors and others must be considered in trade studies to support design decisions that balance safety, reliability, performance and system complexity to achieve a relatively simple, operable system that provides the safest and most reliable system within the specified performance requirements. This paper describes conditions under which additional functional redundancy can impede improved system reliability. Examples from current NASA programs including the Ares I Upper Stage will be shown.

Anderson, Joel M.↗

Techniques for Assuring NASA Mission Success Using Redundancy and Multi-Functionality Designs

Topics include NASA centers around the country; 2009 highlights of significant successes in space transportation, exploration, and science; significant accomplishments; places to explore include Lagrange points, near-Earth objects, Mars and the Moon, and International Space Station research; Marshall's missions include propulsion and transportation systems, life support systems, and earth and space science spacecraft, systems, and operations; project lifecycle management model; motivation of avionics fault-tolerance, redundancy needs and concerns, redundancy versus reliability; parallel-series configurations; effect of adding redundancy on mission success; example of rules-based approach where reliability and safety interaction impacts design; impact of common cause failure; approach ot bottom-up reliability analysis; three factors that lead to redundant system failure; Apollo 13 multi-functional reliability and example; and mitigating the risk of single string spacecraft architecture;.

Shivers, Herb↗

Practical Application of PRA as an Integrated Design Tool for Space Systems

This paper presents the application of the first comprehensive Probabilistic Risk Assessment (PRA) during the design phase of a joint NASA/NOAA weather satellite program, Geostationary Operational Environmental Satellite Series R (GOES-R). GOES-R is the next generation weather satellite primarily to help understand the weather and help save human lives. PRA has been used at NASA for Human Space Flight for many years. PRA was initially adopted and implemented in the operational phase of manned space flight programs and more recently for the next generation human space systems. Since its first use at NASA, PRA has become recognized throughout the Agency as a method of assessing complex mission risks as part of an overall approach to assuring safety and mission success throughout project lifecycles. PRA is now included as a requirement during the design phase of both NASA next generation manned space vehicles as well as for high priority robotic missions. The influence of PRA on GOES-R design and operation concepts are discussed in detail. The GOES-R PRA is unique at NASA for its early implementation. It also represents a pioneering effort to integrate risks from both Spacecraft (SC) and Ground Segment (GS) to fully assess the probability of achieving mission objectives. PRA analysts were actively involved in system engineering and design engineering to ensure that a comprehensive set of technical risks were correctly identified and properly understood from a design and operations perspective. The analysis included an assessment of SC hardware and software, SC fault management system, GS hardware and software, common cause failures, human error, natural hazards, solar weather and infrastructure (such as network and telecommunications failures, fire). PRA findings directly resulted in design changes to reduce SC risk from micro-meteoroids. PRA results also led to design changes in several SC subsystems, e.g. propulsion, guidance, navigation and control (GNC), communications, mechanisms, and command and data handling (C&DH). The fault tree approach assisted in the development of the fault management system design. Human error analysis, which examined human response to failure, indicated areas where automation could reduce the overall probability of gaps in operation by half. In addition, the PRA brought to light many potential root causes of system disruptions, including earthquakes, inclement weather, solar storms, blackouts and other extreme conditions not considered in the typical reliability and availability analyses. Ultimately the PRA served to identify potential failures that, when mitigated, resulted in a more robust design, as well as to influence the program's concept of operations. The early and active integration of PRA with system and design engineering provided a well-managed approach for risk assessment that increased reliability and availability, optimized lifecyc1e costs, and unified the SC and GS developments.

Kalia, Prince↗

Scaling Impacts in Life Support Architecture and Technology Selection

For long-duration space missions outside of Earth orbit, reliability considerations will drive higher levels of redundancy and/or on-board spares for life support equipment. Component scaling will be a critical element in minimizing overall launch mass while maintaining an acceptable level of system reliability. Building on an earlier reliability study (AIAA 2012-3491), this paper considers the impact of alternative scaling approaches, including the design of technology assemblies and their individual components to maximum, nominal, survival, or other fractional requirements. The optimal level of life support system closure is evaluated for deep-space missions of varying duration using equivalent system mass (ESM) as the comparative basis. Reliability impacts are included in ESM by estimating the number of component spares required to meet a target system reliability. Common cause failures are included in the analysis. ISS and ISS-derived life support technologies are considered along with selected alternatives. This study focusses on minimizing launch mass, which may be enabling for deep-space missions.

Lange, Kevin↗

Methods and Costs to Achieve Ultra Reliable Life Support

A published Mars mission is used to explore the methods and costs to achieve ultra reliable life support. The Mars mission and its recycling life support design are described. The life support systems were made triply redundant, implying that each individual system will have fairly good reliability. Ultra reliable life support is needed for Mars and other long, distant missions. Current systems apparently have insufficient reliability. The life cycle cost of the Mars life support system is estimated. Reliability can be increased by improving the intrinsic system reliability, adding spare parts, or by providing technically diverse redundant systems. The costs of these approaches are estimated. Adding spares is least costly but may be defeated by common cause failures. Using two technically diverse systems is effective but doubles the life cycle cost. Achieving ultra reliability is worth its high cost because the penalty for failure is very high.

deep space life support↗

Safety Expertise and the Perils of Novelty

Emerging aviation markets such as urban air mobility are giving rise to new technologies and means of operation. However, novelty may hide ‘unknown unknowns,’ raising new hazards. This paper examines how expertise and safety techniques enable transformative technologies such as reduced crew operations, hybrid wing-borne and rotor-born flight, federated air traffic services, and urban operations. We explore how analysts use expertise to address common-cause failures, collect and interpret safety data, and perform exacting tradeoffs between dissimilarity, redundancy, independence, and diversity (human, process lifecycle, or otherwise) to ensure safety. When novelty is present, analysts might not possess the expertise needed to fully understand the implications of design decisions and tradeoffs being made, especially in early lifecycle phases, on emergent properties such as safety. Safety expertise must be carefully cultivated. The conflicting views of safety experts must be unpacked to identify the divergence in fundamental assumptions, models, means, and methods that may be causing them. Once systems venture beyond the basis of what safety expertise can reliably guarantee, projects take on risk that must be managed. The paper contains key takeaways and actionable recommendations for novel OEMs and regulators touching on topics such as robust monitoring; clear and transparent reporting; incremental approaches to fielding novel systems in hazard-rich, risk-tolerant environments; the cultivation of safety culture and expertise in an organization; and the use of scientific study to reduce epistemic uncertainty in novel operations with new technologies. Since excessive novelty in aviation can undermine the current foundation of safety, humility and incrementalism are necessary to enable emerging aviation markets safely.

safety expertise↗

Safety Expertise and the Perils of Novelty

Emerging aviation markets such as urban air mobility are giving rise to new technologies and means of operation. However, novelty may hide ‘unknown unknowns,’ raising new hazards. This paper examines how expertise and safety techniques enable transformative technologies such as reduced crew operations, hybrid wing-borne and rotor-born flight, federated air traffic services, and urban operations. We explore how analysts use expertise to address common-cause failures, collect and interpret safety data, and perform exacting tradeoffs between dissimilarity, redundancy, independence, and diversity (human, process lifecycle, or otherwise) to ensure safety. When novelty is present, analysts might not possess the expertise needed to fully understand the implications of design decisions and tradeoffs being made, especially in early lifecycle phases, on emergent properties such as safety. Safety expertise must be carefully cultivated. The conflicting views of safety experts must be unpacked to identify the divergence in fundamental assumptions, models, means, and methods that may be causing them. Once systems venture beyond the basis of what safety expertise can reliably guarantee, projects take on risk that must be managed. The paper contains key takeaways and actionable recommendations for novel OEMs and regulators touching on topics such as robust monitoring; clear and transparent reporting; incremental approaches to fielding novel systems in hazard-rich, risk-tolerant environments; the cultivation of safety culture and expertise in an organization; and the use of scientific study to reduce epistemic uncertainty in novel operations with new technologies. Since excessive novelty in aviation can undermine the current foundation of safety, humility and incrementalism are necessary to enable emerging aviation markets safely.

safety expertise↗

Nuclear Safety [Vol. 36, No. 1, January-June 1995]

Nuclear Safety is a journal that covers significant issues in the field of nuclear safety. Its primary scope is safety in the design, construction, operation, and decommissioning of nuclear power reactors worldwide and the research and analysis activities that promote this goal, but it also encompasses the safety aspects of the entire nuclear fuel cycle, including fuel fabrication, spent-fuel processing and handling, and nuclear waste disposal, the handling of fissionable materials and radioisotopes, and the environmental effects of all these activities. Table of Contents for this issue follows. THE CHORNOBYL ACCIDENT: 1 The Chornobyl Accident Revisited, Part II: The State of the Nuclear Fuel Located Within the Chornobyl Sarcophagus, A A. Borovoi and A. R. Sich; GENERAL SAFETY CONSIDERATIONS: 33 Nuclear Power Safety in Central and Eastern Europe, R. Wilson; 46 Safety of Nuclear Power Reactors in the Former Eastern European Countries, S. Chakraborty; 53 Technical Note: On the Definition of Common-Cause Failures, H. Paula; ACCIDENT ANALYSIS: 58 Modeling and Analysis of Core-Debris Recriticality During Hypothetical Severe Accidents in the Advanced Neutron Source Reactor, S.-H. Kim, V. Georgevich, D. B. Simpson, C. O. Slater, and R. P. Taleyarkhan; 68 Ignitability of Hydrogen/Oxygen/Diluent Mixtures in the Presence of Hot Surfaces, R. K. Kumar and G. W. Koroll; 94 Coupled RELAP5 and CONTAIN Accident Analysis Using PVM, K. A. Smith, A. J. Baratta, and G. E. Robinson; CONTROL AND INSTRUMENTATION: 109 Application of Fuzzy Logic in Nuclear Reactor Control Part I: An Assessment of State-of-the-Art, A. S. Heger, N. K. Alang-Rashid, and M. Jamshidi; DESIGN FEATURES: 122 Twenty-Third DOE/NRC Nuclear Air-Cleaning and Treatment Conference, R. R. Bellamy, J. J. Hayes, and M. W. First; ENVIRONMENTAL EFFECTS: 135 Atmospheric Dispersion and the Radiological Consequences of Normal Airborne Effluents from a Nuclear Power Plant, D. Fang, C. Z. Sun, and L. Yang; 142 Calculation of Distribution Coefficients for Radionuclides in Soils and Sediments, I. Puigdomenech and U. Bergstrom: OPERATING EXPERIENCES: 155 Reactor Shutdown Experience, Compiled by J. W. Cletcher; U.S. NUCLEAR REGULATORY COMMISSION INFORMATION AND ANALYSES: 158 Operating Experience Feedback Report—Reliability of Safety-Related Steam Turbine-Driven Standby Pumps Used in U.S. Commercial Nuclear Power Plants, J. R. Boardman; 166 Turbine Building Hazards, H. L Ornstein; RECENT DEVELOPMENTS: 169 Reports, Standards, and Safety Guides, D. S. Queener; 175 Proposed Rule Changes as of Dec. 31,1994; ANNOUNCEMENTS: 32 Harvard School of Public Health In-Place Filter Testing Workshop; 134 International Conference on Advances in the Operational Safety of Nuclear Power Plants; 193 30th Tennessee Industries Week; 193 DOE Technical Standards Program 1995 Workshop; 194 Multiphase Flow Experiments and Instrumentation; 180 The Authors; 185 Indexes to Nuclear Safety, Volumes 34 and 35.

11 NUCLEAR FUEL CYCLE AND FUEL MATERIALS↗