Search NASA⌕ Search

SEARCH · Search NASA

Results for “common cause failures”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 55 records · Page 3

Stress failure of pulmonary capillaries: role in lung and heart disease

Pulmonary capillaries have extremely thin walls to allow rapid exchange of respiratory gases across them. Recently it has been shown that the wall stresses become very large when the capillary pressure is raised, and in anaesthetised rabbits, ultrastructural damage to the walls is seen at pressures of 40 mm Hg and above. The changes include breaks in the capillary endothelial layer, alveolar epithelial layer, and sometimes all layers of the wall. The strength of the thin part of the capillary wall can be attributed to the type IV collagen in the extracellular matrix. Stress failure of pulmonary capillaries results in a high-permeability form of oedema, or even frank haemorrhage, and is apparently the mechanism of neurogenic pulmonary oedema and high-altitude pulmonary oedema. It also explains the exercise-induced pulmonary haemorrhage that occurs in all racehorses. Several features of mitral stenosis are consistent with stress failure. Overinflation of the lung also leads to stress failure, a common cause of increased capillary permeability in the intensive care environment. Stress failure also occurs if the type IV collagen of the capillary wall is weakened by autoantibodies as in Goodpasture's syndrome. Neutrophil elastase degrades type IV collagen and this may be the starting point of the breakdown of alveolar walls that is characteristic of emphysema. Stress failure of pulmonary capillaries is a hitherto overlooked and potentially important factor in lung and heart disease.

Review↗

Performance Losses and Current-Driven Recovery from Cation Contaminants in PEM Water Electrolysis

Water contaminants are a common cause of failure for polymer electrolyte membrane (PEM) electrolyzers in the field as well as a confounding factor in research on cell performance and durability. In this study, we investigated the performance impacts of feed water containing representative tap water cations at concentrations ranging from 0.5–500 μ M, with conductivities spanning from ASTM Type II to tap-water levels. We present multiple diagnostic signatures to help identify the presence of contaminants in PEM electrolysis cells. Through analysis of polarization curves and impedance spectroscopy to understand the origins of performance losses, we found that a switch from the acidic to alkaline hydrogen evolution mechanism is a key factor in contaminated cell behavior. Finally, we demonstrated that this mechanism switching can be harnessed to remove cation contaminants and recover cell performance without the use of an acid wash. We demonstrated near-complete recovery of cells contaminated with sodium and calcium, and partial recovery of a cell contaminated with iron, which was further investigated by post-mortem microscopy. The improved understanding of contaminant impacts from this work can inform development of strategies to mitigate or recover performance losses as well as improve the consistency and rigor of electrolysis research.

30 DIRECT ENERGY CONVERSION↗

An Acid-Free, Temperature-Based Cation Contamination Removal Strategy for PEM Water Electrolysis

It is widely understood that the durability and reliability of polymer electrolyte membrane (PEM) water electrolyzers are heavily dependent on feedwater purity, with cation contaminants that originate from incomplete water purification and balance of plant materials significantly harming electrolyzer performance. However, contamination remains a challenge and a common cause of failure at the stack level, indicating the need for strategies to recover the performance of contaminated cells. In this study, we investigate the effects of temperature on the uptake, electrochemical impacts, and removal of contaminant calcium and iron cations. Lower operating temperatures increase the sensitivity of the cell performance to contaminant cations, while also decreasing cation uptake and promoting contaminant removal. Computational charge transfer modelling shows that lower temperature increases the concentration of contaminant at the cathode and facilitates their removal from the cell. By testing single cells under scenarios designed to mimic stack temperature dynamics, we investigate low-temperature operation as an approach to stack-relevant contaminant recovery. Together, these results demonstrate that the low-temperature recovery approach is a promising approach for acid-free contamination recovery for PEM water electrolysis to promote stack reliability and durability.

08 HYDROGEN↗

An Integrated Framework for Risk Assessment of Safety-related Digital Instrumentation and Control Systems in Nuclear Power Plants: Methodology Advancement and Application

This report documents activities performed by Idaho National Laboratory (INL) during fiscal year (FY) 2024 for the U.S. Department of Energy (DOE) Light Water Reactor Sustainability (LWRS) Program, Risk Informed Systems Analysis (RISA) Pathway, Digital Instrumentation and Control (DI&C) Risk Assessment project. The goal of the RISA Pathway is to optimize safety margins and minimize uncertainties to achieve economic efficiencies while maintaining high levels of safety. This is accomplished by providing scientific basis to better represent safety margins and factors that contribute to cost and safety, and by developing new technologies that reduce operating costs. The research efforts for FY 2024 encompass methodology refinement and exploration. The efforts include: (1) The implementation of a natural language processing tool to expedite key aspects of the reliability analysis methods developed by INL; (2) advances to support intersystem CCF analysis by providing guidance for and identification of coupling mechanisms that may contribute to CCF; (3) the investigation of how generative artificial intelligence tools can aid in hazard analysis and diversity and defense in depth (i.e., D3) assessments; (4) Industry collaboration, allowing the demonstration of and INL's risk assessment tools to support risk assessment of DI&C systems at early and late stages of development; (4) a roadmap for the development of a software for each of INL's risk assessment tools; (5) The development of a theory and methodology manual for a risk quantification methodology; (6) the development of a reliability analysis for machine learning (ML)-integrated control systems.

22 GENERAL STUDIES OF NUCLEAR REACTORS↗

Two is One, One is None: A Discussion on Redundancy

Redundancy is the duplication of critical components or functions of a system with the intention of increasing reliability of the system, usually in the form of a backup or fail-safe. But the topic of redundancy can be a contentious topic; what is too little, too much, and just right may not be readily apparent. This presentation will be a discussion on redundancy situations from different industries as well as the evolving nature of the “best” amount of redundancy.

Redundancy↗

Artificial Intelligence Thermostat to Detect Faults

Residential air conditioners and heat pumps often experience faults due to inadequate maintenance, which can severely reduce efficiency or even cause system failure. Common issues include dirty or clogged air filters and refrigerant leaks. These problems degrade performance and increase energy use and operating costs. This study presents a smart thermostat with embedded artificial intelligence to detect such faults and alert homeowners when maintenance is needed. The thermostat uses low-cost measurements—including return-air temperature, relative humidity, supply-air temperature, outdoor-air temperature, and condenser subcooling—to identify abnormal operations. Because different faults produce distinct response patterns, tailored algorithms are developed to recognize characteristic fault signatures. The investigation is built on a detailed co-simulation platform that couples EnergyPlus with the DOE/ORNL Heat Pump Design Model (HPDM). EnergyPlus represents the building’s dynamic environment, while HPDM is a high-fidelity, hardware-based model that can simulate fault-free performance as well as a wide range of faults, including gradual degradation such as minor refrigerant leakage. This platform provides a virtual training and testing environment that helps distinguish fault-induced behavior from normal operation and supports development of robust diagnostic algorithms. Using this framework, a Dynamic Bayesian Network was developed to identify two common faults—gradual refrigerant charge loss and indoor airflow blockage—and the AI-embedded thermostat was verified through annual building simulations.

Shen, Bo [ORNL] (ORCID:0000000336600393)↗

Evaluation of control parameters for Spray-In-Air (SIA) aqueous cleaning for shuttle RSRM hardware

HD-2 grease is deliberately applied to Shuttle Redesigned Solid Rocket Motor (RSRM) D6AC steel hardware parts as a temporary protective coating for storage and shipping. This HD-2 grease is the most common form of surface contamination on RSRM hardware and must be removed prior to subsequent surface treatment. Failure to achieve an acceptable level of cleanliness (HD-2 calcium grease removal) is a common cause of defect incidence. Common failures from ineffective cleaning include poor adhesion of surface coatings, reduced bond performance of structural adhesives, and failure to pass cleanliness inspection standards. The RSRM hardware is currently cleaned and refurbished using methyl chloroform (1,1,1-trichloroethane). This chlorinated solvent is mandated for elimination due to its ozone depleting characteristics. This report describes an experimental study of an aqueous cleaning system (which uses Brulin 815 GD) as a replacement for methyl chloroform. Evaluation of process control parameters for this cleaner are discussed as well as cleaning mechanisms for a spray-in-air process.

Davis, S. J.↗

32-Bit-Wide Memory Tolerates Failures

Electronic memory system of 32-bit words corrects bit errors caused by some common type of failures - even failure of entire 4-bit-wide random-access-memory (RAM) chip. Detects failure of two such chips, so user warned that ouput of memory may contain errors. Includes eight 4-bit-wide DRAM's configured so each bit of each DRAM assigned to different one of four parallel 8-bit words. Each DRAM contributes only 1 bit to each 8-bit word.

Buskirk, Glenn A.↗

Four faces of baroreflex failure: hypertensive crisis, volatile hypertension, orthostatic tachycardia, and malignant vagotonia

BACKGROUND: The baroreflex normally serves to buffer blood pressure against excessive rise or fall. Baroreflex failure occurs when afferent baroreceptive nerves or their central connections become impaired. In baroreflex failure, there is loss of buffering ability, and wide excursions of pressure and heart rate occur. Such excursions may derive from endogenous factors such as stress or drowsiness, which result in quite high and quite low pressures, respectively. They may also derive from exogenous factors such as drugs or environmental influences. METHODS AND RESULTS: Impairment of the baroreflex may produce an unusually broad spectrum of clinical presentations; with acute baroreflex failure, a hypertensive crisis is the most common presentation. Over succeeding days to weeks, or in the absence of an acute event, volatile hypertension with periods of hypotension occurs and may continue for many years, usually with some attenuation of pressor surges and greater prominence of depressor valleys during long-term follow-up. With incomplete loss of baroreflex afferents, a mild syndrome of orthostatic tachycardia or orthostatic intolerance may appear. Finally, if the baroreflex failure occurs without concomitant destruction of the parasympathetic efferent vagal fibers, a resting state may lead to malignant vagotonia with severe bradycardia and hypotension and episodes of sinus arrest. CONCLUSIONS: Although baroreflex failure is not the most common cause of the above conditions, correct differentiation from other cardiovascular disorders is important, because therapy of baroreflex failure requires specific strategies, which may lead to successful control.

Review↗

Failure environment analysis tool applications

Understanding risks and avoiding failure are daily concerns for the women and men of NASA. Although NASA's mission propels us to push the limits of technology, and though the risks are considerable, the NASA community has instilled within, the determination to preserve the integrity of the systems upon which our mission and, our employees lives and well-being depend. One of the ways this is being done is by expanding and improving the tools used to perform risk assessment. The Failure Environment Analysis Tool (FEAT) was developed to help engineers and analysts more thoroughly and reliably conduct risk assessment and failure analysis. FEAT accomplishes this by providing answers to questions regarding what might have caused a particular failure; or, conversely, what effect the occurrence of a failure might have on an entire system. Additionally, FEAT can determine what common causes could have resulted in other combinations of failures. FEAT will even help determine the vulnerability of a system to failures, in light of reduced capability. FEAT also is useful in training personnel who must develop an understanding of particular systems. FEAT facilitates training on system behavior, by providing an automated environment in which to conduct 'what-if' evaluation. These types of analyses make FEAT a valuable tool for engineers and operations personnel in the design, analysis, and operation of NASA space systems.

Pack, Ginger L.↗

Failure environment analysis tool applications

Understanding risks and avoiding failure are daily concerns for the women and men of NASA. Although NASA's mission propels us to push the limits of technology, and though the risks are considerable, the NASA community has instilled within it, the determination to preserve the integrity of the systems upon which our mission and, our employees lives and well-being depend. One of the ways this is being done is by expanding and improving the tools used to perform risk assessment. The Failure Environment Analysis Tool (FEAT) was developed to help engineers and analysts more thoroughly and reliably conduct risk assessment and failure analysis. FEAT accomplishes this by providing answers to questions regarding what might have caused a particular failure; or, conversely, what effect the occurrence of a failure might have on an entire system. Additionally, FEAT can determine what common causes could have resulted in other combinations of failures. FEAT will even help determine the vulnerability of a system to failures, in light of reduced capability. FEAT also is useful in training personnel who must develop an understanding of particular systems. FEAT facilitates training on system behavior, by providing an automated environment in which to conduct 'what-if' evaluation. These types of analyses make FEAT a valuable tool for engineers and operations personnel in the design, analysis, and operation of NASA space systems.

Pack, Ginger L.↗

Regulation of sarcomere formation and function in the healthy heart requires a titin intronic enhancer

Heterozygous truncating variants in the sarcomere protein titin (TTN) are the most common genetic cause of heart failure. To understand mechanisms that regulate abundant cardiomyocyte (CM) TTN expression, we characterized highly conserved intron 1 sequences that exhibited dynamic changes in chromatin accessibility during differentiation of human CMs from induced pluripotent stem cells (hiPSC-CMs). Homozygous deletion of these sequences in mice caused embryonic lethality, whereas heterozygous mice showed an allele-specific reduction in Ttn expression. A 296 bp fragment of this element, denoted E1, was sufficient to drive expression of a reporter gene in hiPSC-CMs. Deletion of E1 downregulated TTN expression, impaired sarcomerogenesis, and decreased contractility in hiPSC-CMs. Site-directed mutagenesis of predicted binding sites of NK2 homeobox 5 (NKX2-5) and myocyte enhancer factor 2 (MEF2) within E1 abolished its transcriptional activity. In embryonic mice expressing E1 reporter gene constructs, we validated in vivo cardiac-specific activity of E1 and the requirement for NKX2-5- and MEF2-binding sequences. Moreover, isogenic hiPSC-CMs containing a rare E1 variant in the predicted MEF2-binding motif that was identified in a patient with unexplained dilated cardiomyopathy (DCM) showed reduced TTN expression. Together, these discoveries define an essential, functional enhancer that regulates TTN expression. Manipulation of this element may advance therapeutic strategies to treat DCM caused by TTN haploinsufficiency.

Kim, Yuri↗

High Reliability Requires More than Providing Spares

It is sometimes optimistically hoped that a space life support system can be kept working throughout a long duration mission by repairing failed components, as long as sufficient spares are flown. It is usually assumed that the components have constant known failure rates. Then the needed numbers of spares can be computed to have any particular probability that all failed components can be replaced by available spares. This approach can provide high reliability if its favorable assumptions, including constant known failure rates, are satisfied. Other favorable assumptions are that the failures are statistically independent, repair will be successful without causing further failures, and all failures are due to internal component failures. These assumptions are not usually justified. The failure rates may be estimates that are inadequately verified because of insufficient testing. Failure rates may change due to materials substitutions, manufacturing changes, redesigns to fix failures, and new failures caused by redesigns. Failures that are not statistically independent may result from one common cause, such as a design or manufacturing error or a cascade of cause and effect, possibly caused by an external event such as a power outage. Repair may be unsuccessful or cause damage. Many failures occur at component interfaces or at the overall systems level, not within isolated components. Other failures causes are completely external to the system, due to assembly, maintenance, and operational errors or to unexpected environmental challenges. Replacement with sufficient spares can compensate for expected internal component failures but may not be able to cope with unpredictable design and manufacturing flaws, human errors, and environmental impacts. Reliability estimates based on providing sufficient spares to compensate for expected failures may be far too high. They are essentially upper bounds on reliability that might be approached if many frequent but often unconsidered failure causes can be eliminated.

spares↗

Apollo experience report: The problem of stress-corrosion cracking

Stress-corrosion cracking has been the most common cause of structural-material failures in the Apollo Program. The frequency of stress-corrosion cracking has been high and the magnitude of the problem, in terms of hardware lost and time and money expended, has been significant. In this report, the significant Apollo Program experiences with stress-corrosion cracking are discussed. The causes of stress-corrosion cracking and the corrective actions are discussed, in terminology familiar to design engineers and management personnel, to show how stress-corrosion cracking can be prevented.

Johnson, R. E.↗

Cyber-Threat Assessment for the Air Traffic Management System: A Network Controls Approach

Air transportation networks are being disrupted with increasing frequency by failures in their cyber- (computing, communication, control) systems. Whether these cyber- failures arise due to deliberate attacks or incidental errors, they can have far-reaching impact on the performance of the air traffic control and management systems. For instance, a computer failure in the Washington DC Air Route Traffic Control Center (ZDC) on August 15, 2015, caused nearly complete closure of the Centers airspace for several hours. This closure had a propagative impact across the United States National Airspace System, causing changed congestion patterns and requiring placement of a suite of traffic management initiatives to address the capacity reduction and congestion. A snapshot of traffic on that day clearly shows the closure of the ZDC airspace and the resulting congestion at its boundary, which required augmented traffic management at multiple locations. Cyber- events also have important ramifications for private stakeholders, particularly the airlines. During the last few months, computer-system issues have caused several airlines fleets to be grounded for significant periods of time: these include United Airlines (twice), LOT Polish Airlines, and American Airlines. Delays and regional stoppages due to cyber- events are even more common, and may have myriad causes (e.g., failure of the Department of Homeland Security systems needed for security check of passengers, see [3]). The growing frequency of cyber- disruptions in the air transportation system reflects a much broader trend in the modern society: cyber- failures and threats are becoming increasingly pervasive, varied, and impactful. In consequence, an intense effort is underway to develop secure and resilient cyber- systems that can protect against, detect, and remove threats, see e.g. and its many citations. The outcomes of this wide effort on cyber- security are applicable to the air transportation infrastructure, and indeed security solutions are being implemented in the current system. While these security solutions are important, they only provide a piecemeal solution. Particular computers or communication channels are protected from particular attacks, without a holistic view of the air transportation infrastructure. On the other hand, the above-listed incidents highlight that a holistic approach is needed, for several reasons. First, the air transportation infrastructure is a large scale cyber-physical system with multiple stakeholders and diverse legacy assets. It is impractical to protect every cyber- asset from known and unknown disruptions, and instead a strategic view of security is needed. Second, disruptions to the cyber- system can incur complex propagative impacts across the air transportation network, including its physical and human assets. Also, these implications of cyber- events are exacerbated or modulated by other disruptions and operational specifics, e.g. severe weather, operator fatigue or error, etc. These characteristics motivate a holistic and strategic perspective on protecting the air transportation infrastructure from cyber- events. The analysis of cyber- threats to the air traffic system is also inextricably tied to the integration of new autonomy into the airspace. The replacement of human operators with cyber functions leaves the network open to new cyber threats, which must be modeled and managed. Paradoxically, the mitigation of cyber events in the airspace will also likely require additional autonomy, given the fast time scale and myriad pathways of cyber-attacks which must be managed. The assessment of new vulnerabilities upon integration of new autonomy is also a key motivation for a holistic perspective on cyber threats.

Complex Networks↗

Life prediction of aging aircraft wiring systems

The program goal is to develop a computerized life prediction model capable of identifying present aging progress and predicting end of life for aircraft wiring. A summary is given in viewgraph format of progress made on phase 1 objectives, which were to identify critical aircraft wiring problems; relate most common failures identified to the wire mechanism causing the failure; assess wiring requirments, materials, and stress environment for fighter aircraft; and demonstrate the feasibility of a time-temperature-environment model.

Slenski, George↗

Contamination and Radiation Effects on Nonlinear Crystals for Space Laser Systems

Space Lasers are vital tools for NASA s space missions and military applications. Although, lasers are highly reliable on the ground, several past space laser missions proved to be short-lived and unreliable. In this communication, we are shedding more light on the contamination and radiation issues, which are the most common causes for optical damages and laser failures in space. At first, we will present results based on the study of liquids and subsequently correlate these results to the particulates of the laser system environment. We present a model explaining how the laser beam traps contaminants against the optical surfaces and cause optical damages and the role of gravity in the process. We also report the results of the second harmonic generation efficiency for nonlinear optical crystals irradiated with high-energy beams of protons. In addition, we are proposing to employ the technique of adsorption to minimize the presence of adsorbing molecules present in the laser compartment.

Abdeldayem, Hossain A.↗

A Prognostic Launch Vehicle Probability of Failure Assessment Methodology for Conceptual Systems Predicated on Human Causal Factors

Create an improved method to calculate reliability of a conceptual launch vehicle system prior to fabrication by using historic data of actual root causes of failures. While failures have unique "proximate causes", there are typically a finite amount of common "root causes". Heretofore launch vehicle reliability evaluation typically hardware-centric statistical analyses, while most root causes of failures are been shown to be human-centric. A method based on human-centric root causes can be used to quantify reliability assessments and focus proposed actions to mitigate problems. Existing methods have been optimistic in their projections of launch vehicle reliability compared to actuals. Hypothesis: reliability of a conceptual launch vehicle can be more accurately evaluated based on a rational, probabilistic approach using past failure assessment teams' findings predicated on human-centric causes."Human Reliability Analysis Methods Selection Guidance for NASA"Chandler F.T., et al., NASA HQ/OSMA study group, July 2006. Outside HRA experts from academia, other federal labs, and the private sector. 50 system reliability methods considered, fourteen selected for further study, four finally selected as best suited for human spaceflight. Probabilistic Risk Analysis (PRA) + Human Reliability Analysis (HRA) enabled incorporating effects and probabilities of human errors. While four down-selected methods deemed appropriate for failure assessment, it did not appear that these methods could be concisely applied to perform major system-wide assessment of probability of failure of a conceptual design without becoming unwieldy."Engineering a Safer World", Detailed, comprehensive study external to NASA Leveson N. G., MIT, 2011.Systems-Theoretic Accident Model and Processes (STAMP). All-encompassing accident model based on systems theory analyzed accidents after they occurred and created approaches to prevent occurrence in developing systems not focused on failure prevention per se, but rather reducing hazards by influencing human behavior through use of constraints, hierarchical control structures, and process models to improve system safetySystem Theoretic Process Analysis (STPA) addresses predictive part of problem (a "hazard analysis"). Includes all causal factors identified in STAMP: "...design errors, software flaws, component interaction accidents, cognitively complex human decision-making errors, and social organizational and management factors contributing to accidents" can guide design process rather than require it to exist before-hand did not appear capable of concise application for system-wide assessment of probability of failure of a conceptual design without becoming unwieldy.

Williams, Craig H.↗