Search NASA⌕ Search

SEARCH · Search NASA

Results for “common cause failures”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 73 records · Page 4

Productivity in an evolutionary space station

Space station productivity is treated from a systems point of view, considering the functions and attributes of space station development, formation, and operation that affect productivity. An optimum planning method is needed to assure that the station will have mission flexibility, technology advancement, maintainability, and evolutionary capability. Advanced technology will be designed into the housekeeping and utility functions of the station. Greater risk taking may be allowed into designs if the potential benefits of the advanced system support the risk, and if the system can be buffered from causing a failure cascade throughout the station. A common data base is needed to store and track all designs, developments, and changes in the station subsystems. Systems that can be automated and free the human inhabitants for more productive work are favored, as are modular components that are highly fault-free. Human control must also be possible, especially during check-out and verification, and also for teaching the automated systems new or modified tasks.

Anderson, J. L.↗

Proposed system safety design and test requirements for the microlaser ordnance system

Safety for pyrotechnic ignition systems is becoming a major concern for the military. In the past twenty years, stray electromagnetic fields have steadily increased during peacetime training missions and have dramatically increased during battlefield missions. Almost all of the ordnance systems in use today depend on an electrical bridgewire for ignition. Unfortunately, the bridgewire is the cause of the majority of failure modes. The common failure modes include the following: broken bridgewires; transient RF power, which induces bridgewire heating; and cold temperatures, which contracts the explosive mix away from the bridgewire. Finding solutions for these failure modes is driving the costs of pyrotechnic systems up. For example, analyses are performed to verify that the system in the environment will not see more energy than 20 dB below the 'No-fire' level. Range surveys are performed to determine the operational, storage, and transportation RF environments. Cryogenic tests are performed to verify the bridgewire to mix interface. System requirements call for 'last minute installation,' 'continuity checks after installation,' and rotating safety devices to 'interrupt the explosive train.' As an alternative, MDESC has developed a new approach based upon our enabling laser diode technology. We believe that Microlaser initiated ordnance offers a unique solution to the bridgewire safety concerns. For this presentation, we will address, from a system safety viewpoint, the safety design and the test requirements for a Microlaser ordnance system. We will also review how this system could be compliant to MIL-STD-1576 and DOD-83578A and the additional necessary requirements.

Stoltz, Barb A.↗

Facilitating Data Collection of Maintenance Events to Populate the Hydrogen Component Reliability Database (HyCReD)

The Hydrogen Component Reliability Database (HyCReD) is a collaborative project between the National Renewable Energy Laboratory, the University of Maryland, and hydrogen stakeholders to improve safety and reliability for hydrogen facilities by implementing component reliability data taxonomies that support hydrogen infrastructure failure rate analysis. The project aims to quantify failure rates of hydrogen components through high-quality data collection and analysis on root causes and maintenance needed. HyCReD provides a common database for cataloging hydrogen component failures which exists for reliability research in many other mature industries [2]. The database fills a gap for the hydrogen community by providing a scientifically rigorous approach to quantitative risk assessment (QRA), prognostic health management (PHM), and reliability-centered maintenance (RCM) analysis. High level results will be aggregated and anonymized to protect company sensitive information; detailed results will be used to help address issues of hydrogen components. These advanced analytics will support accelerated deployment of hydrogen infrastructure by enabling better: design and safety of projects (safety codes and standards development), infrastructure reliability and cost (component failure rates, maintenance protocols), and component R&D needs (robust supply chain). A key to a successful HyCReD implementation is facilitating the ease of reporting and data quality in the database that can be used for analysis. Maintenance data was a previously identified gap in initial efforts to populate and validate the database taxonomies [3]. Collection of maintenance data will be instrumental in identifying failure modes and rates, identifying incipient component failures or reduced performance, cataloging best practices for maintenance routines and methods for prognostic health management, and quantifying the risk and effect of different failure modes. Several key priorities are identified for streamlined data collection to achieve quality and detailed failure data: Applicability, Ease of Use, Accessibility, and Information Security. The HyCReD team has now begun deployment of the database to several companies and groups that have signed non-disclosure agreements to facilitate the data collection of failures in industry hydrogen refueling station infrastructure. This paper will provide an update into the process of HyCReD deployment including the development of a coding guide for facility personnel to reference and ensure data quality and consistency from one station to another as well as implementation of contextually dependent data fields of system taxonomy and formatted entries to provide ease of use. The goal is to communicate the lessons learned from the roll-out to technicians and engineers in the field, and the addition of need for high level of security to protect all stakeholders.

29 ENERGY PLANNING, POLICY, AND ECONOMY↗

​​Hydrogen Transportation and Distributed Energy Systems Seismic Risk Assessment for Cascadia Subduction Zone Airport Facilities​

Portland International Airport in Oregon is exploring operating a fleet of 28 fuel cell electric buses to support airport operations and possibly provide backup power during outages. In this report, we evaluate the risk of hydrogen deployment at an airport considering the potential for seismic activity. We present a seismic risk assessment for pressurized piping for a generic hydrogen refueling station at the airport. Only pressurized piping was considered in the risk analysis because of its vulnerability and availability of fragility information unlike the other components. Seismic capacity of other components should be included as more information becomes available. To characterize the seismic hazard, we considered the initiating event frequency of a Cascadia Subduction Zone earthquake event based on data from the literature. Seismic stressors that translate to the site from an earthquake event are expressed as peak ground acceleration. Site disturbance is characterized as a function of soil conditions, represented by different shear wave velocity values, which affect how seismic waves propagate and impact structures. Pipe fragility curves as a function of seismic capacity were used to correlate ground acceleration and failure probability. These fragility characteristics were combined with site-specific soil conditions to calculate the probability of pipe rupture during a seismic event. It should be noted that multiple simultaneous failures due to an earthquake event as a common cause are not considered. Some epistemic uncertainties such as self-exciting shocks, aftershocks, and the time variant nature of the ground motion are not considered either.

08 HYDROGEN↗

Spaceflight Over the Last Ten Years: Failures and Fix-ups 2013 - 2022 RAMS XV Conference

This project is an overview and analysis of the past ten years (2013 – present) in global spaceflight, highlighting the orbital launches of countries, particularly the failures that occurred, the reasons they occurred, and a breakdown of the related statistics and background information. The analysis is conducted from a perspective of reliability and maintainability engineering and Probabilistic Risk Assessment (PRA) to formulate a quantifiable understanding of the data and how it is pertinent to Safety and Mission Assurance (SMA) in spaceflight. There is a breakdown by country or group, timelines, and number of launches. The failures over the years are categorized, all given a broad analysis of subsystem failures and details of events. A few failures are given a more in-depth analysis of root causes and failure modes determined by their unique or common nature. The data is retrieved from online, publicly available sources.

Quinn Slaugenhoupt↗

A Multifunctional Coating for Autonomous Corrosion Control

Corrosion is a destructive process that often causes failure in metallic components and structures. Protective coatings are the most commonly used method of corrosion control. However, progressively stricter environmental regulations have resulted in the ban of many commercially available corrosion protective coatings due to the harmful effects of their solvents or corrosion inhibitors. This work concerns the development of a multifunctional, smart coating for the autonomous control of corrosion. This coating is being developed to have the inherent ability to detect the chemical changes associated with the onset of corrosion and respond autonomously to control it. The multi-functionality of the coating is based on microencapsulation technology specifically designed for corrosion control applications. This design has, in addition to all the advantages of other existing microcapsules designs, the corrosion controlled release function that allows the delivery of corrosion indicators and inhibitors on demand only when and where they are needed. Corrosion indicators as well as corrosion inhibitors have been incorporated into the microcapsules, blended into several paint systems, and tested for corrosion detection and protection efficacy.

Calle, Luz M.↗

A Multifunctional Smart Coating for Autonomous Corrosion Control

Corrosion is a destructive process that often causes failure in metallic components and structures. Protective coatings are the most commonly used method of corrosion control. However, progressively stricter environmental regulations have resulted in the ban of many commercially available corrosion protective coatings due to the harmful effects of their solvents or corrosion inhibitors. This work concerns the development of a multifunctional, smart coating for the autonomous control of corrosion. This coating is being developed to have the inherent ability to detect the chemical changes associated with the onset of corrosion and respond autonomously to control it. The multi-functionality of the coating is based on micro-encapsulation technology specifically designed for corrosion control applications. This design has, in addition to all the advantages of other existing microcapsules designs, the corrosion controlled release function that allows the delivery of corrosion indicators and inhibitors on demand only when and where needed. Corrosion indicators as well as corrosion inhibitors have been incorporated into microcapsules, blended into several paint systems, and tested for corrosion detection and protection efficacy. This

Calle, Luz Marina↗

A Technical Evaluation of Self-Protection Dose Rates in U.S. Domestic Nuclear Security Policy

U.S. nuclear security policy uses radiation self-protection as a basis for reducing material attractiveness, but existing dose-rate criteria may not reflect the time scales of theft or sabotage. This report evaluates whether the commonly cited 1 Gy/h at 1 m criterion can plausibly cause adversary task failure during short-duration malicious acts.

Fritchie, Jacob Wesley [Sandia National Laboratori↗

Delamination durability of composite materials for rotorcraft

Delamination is the most commonly observed failure mode in composite rotorcraft dynamic components. Although delamination may not cause immediate failure of the composite part, it often precipitates component repair or replacement, which inhibits fleet readiness, and results in increased life cycle costs. A fracture mechanics approach for analyzing, characterizing, and designing against delamination will be outlined. Examples of delamination problems will be illustrated where the strain energy release rate associated with delamination growth was found to be a useful generic parameter, independent of thickness, layup, and delamination source, for characterizing delamination failure. Several analysis techniques for calculating strain energy release rates for delamination from a variety of sources will be outlined. Current efforts to develop ASTM standard test methods for measuring interlaminar fracture toughness and developing delamination failure criteria will be reviewed. A technique for quantifying delamination durability due to cyclic loading will be presented. The use of this technique for predicting fatigue life of composite laminates and developing a fatigue design philosophy for composite structural components will be reviewed.

Obrien, T. Kevin↗

Implementation and Qualifications Lessons Learned for Space Flight Photonic Components

This slide presentation reviews the process for implementation and qualification of space flight photonic components. It discusses the causes for most common anomalies for the space flight components, design compatibility, a specific failure analysis of optical fiber that occurred in a cable in 1999-2000, and another ExPCA connector anomaly involving pins that broke off. It reviews issues around material selection, quality processes and documentation, and current projects that the Photonics group is involved in. The importance of good documentation is stressed.

Ott, Melanie N.↗

Reliability Effects of Surge Current Testing of Solid Tantalum Capacitors

Solid tantalum capacitors are widely used in space applications to filter low-frequency ripple currents in power supply circuits and stabilize DC voltages in the system. Tantalum capacitors manufactured per military specifications (MIL-PRF-55365) are established reliability components and have less than 0.001% of failures per 1000 hours (the failure rate is less than 10 FIT) for grades D or S, thus positioning these parts among electronic components with the highest reliability characteristics. Still, failures of tantalum capacitors do happen and when it occurs it might have catastrophic consequences for the system. This is due to a short-circuit failure mode, which might be damaging to a power supply, and also to the capability of tantalum capacitors with manganese cathodes to self-ignite when a failure occurs in low-impedance applications. During such a failure, a substantial amount of energy is released by exothermic reaction of the tantalum pellet with oxygen generated by the overheated manganese oxide cathode, resulting not only in destruction of the part, but also in damage of the board and surrounding components. A specific feature of tantalum capacitors, compared to ceramic parts, is a relatively large value of capacitance, which in contemporary low-size chip capacitors reaches dozens and hundreds of microfarads. This might result in so-called surge current or turn-on failures in the parts when the board is first powered up. Such a failure, which is considered as the most prevalent type of failures in tantalum capacitors [I], is due to fast changes of the voltage in the circuit, dV/dt, producing high surge current spikes, I(sub sp) = Cx(dV/dt), when current in the circuit is unrestricted. These spikes can reach hundreds of amperes and cause catastrophic failures in the system. The mechanism of surge current failures has not been understood completely yet, and different hypotheses were discussed in relevant literature. These include a sustained scintillation breakdown model [1-3]; electrical oscillations in circuits with a relatively high inductance [4-6]; local overheating of the cathode [5,7, 8]; mechanical damage to tantalum pentoxide dielectric caused by the impact of MnO2 crystals [2,9, 10]; or stress-induced-generation of electron traps caused by electromagnetic forces developed during current spikes [11]. A commonly accepted explanation of the surge current failures is that at unlimited current supply during surge current conditions, the self-healing mechanism in tantalum capacitors does not work, and what would be a minor scintillation spike if the current were limited, becomes a catastrophic failure of the part [l, 12]. However, our data show that the scintillation breakdown voltages are significantly greater that the surge current breakdown voltages, so it is still not clear why the part, which has no scintillations, would fail at the same voltage during surge current testing (SCT).

Teverovsky, Alexander↗

On the Use of Resilience Models as Digital Twins for Operational Support and In time Decision Making

Human error is a major contributor to accidents and performance losses in complex engineered systems. If one examines these human error caused failures further, a specific cause, the lack of situation awareness, has dominated as a major cause of human errors that instigate latent or catastrophic failures in complex systems. Studies of aviation accidents involving major air carriers revealed that situation awareness was the root cause of around 90% of accidents involving pilot error. Another study explored offshore drilling accidents involving human error and found that 40% of accidents were directly attributed to the loss of situation awareness. Studies of human errors in other domains such as nuclear power, air traffic control, process industry, and advanced driving show that loss of SA was a root cause in a majority of the events. Situation awareness-related failures are not only common but also costly and fatal (e.g., Bhopal Gas Leak, Air France 447 Flight Crash). Thus, the concept of situation awareness has emerged as an important construct in human factors, resulting in numerous models and measurement methods to aid in promoting appropriate levels of situation awareness.

Lukman Irshad↗

Software For Fault-Tree Diagnosis Of A System

Fault Tree Diagnosis System (FTDS) computer program is automated-diagnostic-system program identifying likely causes of specified failure on basis of information represented in system-reliability mathematical models known as fault trees. Is modified implementation of failure-cause-identification phase of Narayanan's and Viswanadham's methodology for acquisition of knowledge and reasoning in analyzing failures of systems. Knowledge base of if/then rules replaced with object-oriented fault-tree representation. Enhancement yields more-efficient identification of causes of failures and enables dynamic updating of knowledge base. Written in C language, C++, and Common LISP.

Iverson, Dave↗

Failure Modes and Effects Analysis (FMEA) Assistant Tool Feasibility Study

An effort to determine the feasibility of a software tool to assist in Failure Modes and Effects Analysis (FMEA) has been completed. This new and unique approach to FMEA uses model based systems engineering concepts to recommend failure modes, causes, and effects to the user after they have made several selections from pick lists about a component s functions and inputs/outputs. Recommendations are made based on a library using common failure modes identified over the course of several major human spaceflight programs. However, the tool could be adapted for use in a wide range of applications from NASA to the energy industry.

Flores, Melissa↗

Oversimplification of Systems Engineering Goals, Processes, and Criteria in NASA Space Life Support

This paper investigates the oversimplification of the inherently complex systems engineering process in space life support. The standard systems engineering process steps are described. The International Space Station (ISS) life support system is explained with its goals and performance criteria. Although it is not usually emphasized, the essential function of developing a hierarchy of systems and subsystems is to simplify the design process. The System Complexity Metric (SCM) shows how this di-vide-and-conquer approach also reduces the system complexity. The complete systems engineering process has many detailed steps. It is often simplified because of the effort required and the human limitations on working memory and decision span. Systems analysis demands slow, logical, and fo-cused thinking but is often bypassed in favor of quick, intuitive, subconscious “gut feel.” A study of 100 system designs found examples of 12 specific mental mistakes, such as ignoring stakeholder needs, and these mistakes are essentially oversimplifications of the systems engineering process. An analysis of space life support goals, options, criteria, and processes found 11 examples of oversimplifications in systems engineering, such as neglecting safety and cost. All these 11 oversimplifications could be traced to one or more of the 12 previously identified mental mistakes or other well-known ones, such as ig-noring sunk costs. Oversimplification of the systems engineering process is rarely noticed but is a common and harmful problem. A study of failures in 50 different space systems found that problems in systems engineering caused failures and often led to errors in design, development, and test that further contributed to failure. It seems that more diligent systems engineering could prevent many project problems and failures, but projects seem to be more guided by “gut feel” based on tradition, authority, and consensus than on the logical, rational systems engineering approach.

Simplified systems engineering↗

Common Cause Case Study: An Estimated Probability of Four Solid Rocket Booster Hold-Down Post Stud Hang-ups

Until Solid Rocket Motor ignition, the Space Shuttle is mated to the Mobil Launch Platform in part via eight (8) Solid Rocket Booster (SRB) hold-down bolts. The bolts are fractured using redundant pyrotechnics, and are designed to drop through a hold-down post on the Mobile Launch Platform before the Space Shuttle begins movement. The Space Shuttle program has experienced numerous failures where a bolt has hung up. That is, it did not clear the hold-down post before liftoff and was caught by the SRBs. This places an additional structural load on the vehicle that was not included in the original certification requirements. The Space Shuttle is currently being certified to withstand the loads induced by up to three (3) of eight (8) SRB hold-down experiencing a "hang-up". The results of loads analyses performed for (4) stud hang-ups indicate that the internal vehicle loads exceed current structural certification limits at several locations. To determine the risk to the vehicle from four (4) stud hang-ups, the likelihood of the scenario occurring must first be evaluated. Prior to the analysis discussed in this paper, the likelihood of occurrence had been estimated assuming that the stud hang-ups were completely independent events. That is, it was assumed that no common causes or factors existed between the individual stud hang-up events. A review of the data associated with the hang-up events, showed that a common factor (timing skew) was present. This paper summarizes a revised likelihood evaluation performed for the four (4) stud hang-ups case considering that there are common factors associated with the stud hang-ups. The results show that explicitly (i.e. not using standard common cause methodologies such as beta factor or Multiple Greek Letter modeling) taking into account the common factor of timing skew results in an increase in the estimated likelihood of four (4) stud hang-ups of an order of magnitude over the independent failure case.

Cross, Robert↗

Common Cause Case Study: An Estimated Probability of Four Solid Rocket Booster Hold-down Post Stud Hang-ups

Until Solid Rocket Motor ignition, the Space Shuttle is mated to the Mobil Launch Platform in part via eight (8) Solid Rocket Booster (SRB) hold-down bolts. The bolts are fractured using redundant pyrotechnics, and are designed to drop through a hold-down post on the Mobile Launch Platform before the Space Shuttle begins movement. The Space Shuttle program has experienced numerous failures where a bolt has "hung-up." That is, it did not clear the hold-down post before liftoff and was caught by the SRBs. This places an additional structural load on the vehicle that was not included in the original certification requirements. The Space Shuttle is currently being certified to withstand the loads induced by up to three (3) of eight (8) SRB hold-down post studs experiencing a "hang-up." The results af loads analyses performed for four (4) stud-hang ups indicate that the internal vehicle loads exceed current structural certification limits at several locations. To determine the risk to the vehicle from four (4) stud hang-ups, the likelihood of the scenario occurring must first be evaluated. Prior to the analysis discussed in this paper, the likelihood of occurrence had been estimated assuming that the stud hang-ups were completely independent events. That is, it was assumed that no common causes or factors existed between the individual stud hang-up events. A review of the data associated with the hang-up events, showed that a common factor (timing skew) was present. This paper summarizes a revised likelihood evaluation performed for the four (4) stud hang-ups case considering that there are common factors associated with the stud hang-ups. The results show that explicitly (i.e. not using standard common cause methodologies such as beta factor or Multiple Greek Letter modeling) taking into account the common factor of timing skew results in an increase in the estimated likelihood of four (4) stud hang-ups of an order of magnitude over the independent failure case.

Cross, Robert↗

What Reliability Engineers Should Know about Space Radiation Effects

Space radiation in space systems present unique failure modes and considerations for reliability engineers. Radiation effects is not a one size fits all field. Threat conditions that must be addressed for a given mission depend on the mission orbital profile, the technologies of parts used in critical functions and on application considerations, such as supply voltages, temperature, duty cycle, and redundancy. In general, the threats that must be addressed are of two types-the cumulative degradation mechanisms of total ionizing dose (TID) and displacement damage (DD). and the prompt responses of components to ionizing particles (protons and heavy ions) falling under the heading of single-event effects. Generally degradation mechanisms behave like wear-out mechanisms on any active components in a system: Total Ionizing Dose (TID) and Displacement Damage: (1) TID affects all active devices over time. Devices can fail either because of parametric shifts that prevent the device from fulfilling its application or due to device failures where the device stops functioning altogether. Since this failure mode varies from part to part and lot to lot, lot qualification testing with sufficient statistics is vital. Displacement damage failures are caused by the displacement of semiconductor atoms from their lattice positions. As with TID, failures can be either parametric or catastrophic, although parametric degradation is more common for displacement damage. Lot testing is critical not just to assure proper device fi.mctionality throughout the mission. It can also suggest remediation strategies when a device fails. This paper will look at these effects on a variety of devices in a variety of applications. This paper will look at these effects on a variety of devices in a variety of applications. (2) On the NEAR mission a functional failure was traced to a PIN diode failure caused by TID induced high leakage currents. NEAR was able to recover from the failure by reversing the current of a nearby Thermal Electric Cooler (turning the TEC into a heater). The elevated temperature caused the PIN diode to anneal and the device to recover. It was by lot qualification testing that NEAR knew the diode would recover when annealed. This paper will look at these effects on a variety of devices in a variety of applications. Single Event Effects (SEE): (1) In contrast to TID and displacement damage, Single Event Effects (SEE) resemble random failures. SEE modes can range from changes in device logic (single-event upset, or SEU). temporary disturbances (single-event transient) to catastrophic effects such as the destructive SEE modes, single-event latchup (SEL). single-event gate rupture (SEGR) and single-event burnout (SEB) (2) The consequences of nondestructive SEE modes such as SEU and SET depend critically on their application--and may range from trivial nuisance errors to catastrophic loss of mission. It is critical not just to ensure that potentially susceptible devices are well characterized for their susceptibility, but also to work with design engineers to understand the implications of each error mode. -For destructive SEE, the predominant risk mitigation strategy is to avoid susceptible parts, or if that is not possible. to avoid conditions under which the part may be susceptible. Destructive SEE mechanisms are often not well understood, and testing is slow and expensive, making rate prediction very challenging. (3) Because the consequences of radiation failure and degradation modes depend so critically on the application as well as the component technology, it is essential that radiation, component. design and system engineers work togetherpreferably starting early in the program to ensure critical applications are addressed in time to optimize the probability of mission success.

DiBari, Rebecca↗