Search NASA⌕ Search

SEARCH · Search NASA

Results for “common cause failures”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 91 records · Page 5

A diagnosis system using object-oriented fault tree models

Spaceborne computing systems must provide reliable, continuous operation for extended periods. Due to weight, power, and volume constraints, these systems must manage resources very effectively. A fault diagnosis algorithm is described which enables fast and flexible diagnoses in the dynamic distributed computing environments planned for future space missions. The algorithm uses a knowledge base that is easily changed and updated to reflect current system status. Augmented fault trees represented in an object-oriented form provide deep system knowledge that is easy to access and revise as a system changes. Given such a fault tree, a set of failure events that have occurred, and a set of failure events that have not occurred, this diagnosis system uses forward and backward chaining to propagate causal and temporal information about other failure events in the system being diagnosed. Once the system has established temporal and causal constraints, it reasons backward from heuristically selected failure events to find a set of basic failure events which are a likely cause of the occurrence of the top failure event in the fault tree. The diagnosis system has been implemented in common LISP using Flavors.

Iverson, David L.↗

Phenomena associated with bench and thermal-vacuum testing of super conductors - Heat pipes.

Test failures of heat pipes occur when the functional performance is unable to match the expected design limits or when the power applied to the heat pipe (in the form of heat) is distributed unevenly through the system, yielding a large thermal gradient. When a thermal gradient larger than expected is measured, it normally occurs in the evaporator or condenser sections of the pipe. Common causes include evaporator overheating, condenser dropout, noncondensable gas formation, surge and partial recovery of evaporator temperatures, masking of thermal profiles, and simple malfunctions due to leaks and mechanical failures or flaws. Examples of each of these phenomena are described along with corresponding failure analyses and corrective measures.

Marshburn, J. P.↗

Cryo-EM confirms a common fibril fold in the heart of four patients with ATTRwt amyloidosis

ATTR amyloidosis results from the conversion of transthyretin into amyloid fibrils that deposit in tissues causing organ failure and death. This conversion is facilitated by mutations in ATTRv amyloidosis, or aging in ATTRwt amyloidosis. ATTRv amyloidosis exhibits extreme phenotypic variability, whereas ATTRwt amyloidosis presentation is consistent and predictable. Previously, we found unique structural variabilities in cardiac amyloid fibrils from polyneuropathic ATTRv-I84S patients. In contrast, cardiac fibrils from five genotypically different patients with cardiomyopathy or mixed phenotypes are structurally homogeneous. To understand fibril structure’s impact on phenotype, it is necessary to study the fibrils from multiple patients sharing genotype and phenotype. Here we show the cryo-electron microscopy structures of fibrils extracted from four cardiomyopathic ATTRwt amyloidosis patients. Our study confirms that they share identical conformations with minimal structural variability, consistent with their homogenous clinical presentation. Our study contributes to the understanding of ATTR amyloidosis biopathology and calls for further studies.

59 BASIC BIOLOGICAL SCIENCES↗

Detecting and Characterizing Patterns of Failure in Complex Engineered Systems: an Ontology Development and Clustering Approach

While the causes of failures in complex engineered systems are often clear in hindsight, it can be challenging to predict failures proactively during the design of novel engineered products or systems. Identifying patterns can be useful for capturing common characteristics that may lead to failure. In this paper, we present a methodology for identifying patterns of failure from NASA’s publicly available Lessons Learned Information System (LLIS). We apply an ontology development and clustering approach to identify representative patterns leading to failures in historical lessons learned. A joint inductive-deductive approach reveals the key themes in lessons that lead to failure, which are formalized and recorded as an ontology of complex systems failure causes. Documents from the LLIS are manually tagged with relevant characteristics from the ontology. From the tagged set, clustering is used to capture co-occurring sets of characteristics that lead to failure. The primary contribution of this work is a method for extracting a set of generic failure patterns in complex engineered systems and characteristics for these patterns that can be identified at design time, knowledge of which can be used to plan mitigation strategies.

Systems Engineering↗

Detecting and Characterizing Patterns of Failure in Complex Systems: An Ontology Development and Clustering Approach

While the causes of failures in complex engineered systems are often clear in hindsight, it can be challenging to predict failures proactively during the design of novel engineered products or systems. Identifying patterns can be useful for capturing common characteristics that may lead to failure. In this paper, we present a methodology for identifying patterns of failure from NASA’s publicly available Lessons Learned Information System (LLIS). We apply an ontology development and clustering approach to identify representative patterns leading to failures in historical lessons learned. A joint inductive-deductive approach reveals the key themes in lessons that lead to failure, which are formalized and recorded as an ontology of complex systems failure causes. Documents from the LLIS are manually tagged with relevant characteristics from the ontology. From the tagged set, clustering is used to capture co-occurring sets of characteristics that lead to failure. The primary contribution of this work is a method for extracting a set of generic failure patterns in complex engineered systems and characteristics for these patterns that can be identified at design time, knowledge of which can be used to plan mitigation strategies.

Systems Engineering↗

Parts, Materials, and Processes Experience Summary

The ALERT program, a system for communicating common problems with parts, materials, and processes, is condensed and catalogued. Expanded information on selected topics is provided by relating the problem area (failure) to the cause, the investigations and findings, the suggestions for avoidance (inspections, screening tests, proper part applications), and failure analysis procedures. The basic objective of ALERT is the avoidance of the recurrence of parts, materials, and processed problems, thus improving the reliability of equipment produced for and used by the government.

Source record↗

Z-2 Threaded Insert Design and Testing

NASA's Z-2 prototype space suit contains several components fabricated from an advanced hybrid composite laminate consisting of IM10 carbon fiber and fiber glass. One requirement was to have removable, replaceable helicoil inserts to which other suit components would be fastened. An approach utilizing bonded in inserts with helicoils inside of them was implemented. During initial assembly, cracking sounds were heard followed by the lifting of one of the blind inserts out of its hole when the screws were torqued. A failure investigation was initiated to understand the mechanism of the failure. Ultimately, it was determined that the pre-tension caused by torqueing the fasteners is a much larger force than induced from the pressure loads of the suit which was not considered in the insert design. Bolt tension is determined by dividing the torque on the screw by a k value multiplied by the thread diameter of the bolt. The k value is a factor that accounts for friction in the system. A common value used for k for a non-lubricated screw is 0.2. The k value can go down by as much as 0.1 if the screw is lubricated which means for the same torque, a much larger tension could be placed on the bolt and insert. This paper summarizes the failure investigation that was performed to identify the root cause of the suit failure and details how the insert design was modified to resist a higher pull out tension.

Ross, Amy↗

Failure Modes Experienced on Spacecraft Nicd Batteries

A review was made of failures and irregularities experienced on nickel cadmium batteries for 31 spacecraft. Only rarely did batteries fail completely. In many cases, poorly performing batteries were compensated for by a reduction in loads or by continuing to operate in spite of out-of-voltage conditions. Low discharge voltage was the most common problem observed in flight spacecraft (42%). Spacecraft batteries are often designed to protect against cell shorts, but cell shorts accounted for only 16% of the failures. Other causes of problems were high charge voltage (16%), battery problems caused by other elements of the spacecraft (10%), and open circuit failures (6%). Problems of miscellaneous or unknown causes occurred in 10% of the cases.

Gross, S.↗

Investigation of Abnormal Level Control Oscillations in a BWR Feedwater System

In the long-term operation of nuclear power plants, the aging of systems, structures, and components can lead to maintenance issues that must be dealt with to maintain cost-effective plant operations. One common issue affecting the currently operated boiling water reactors is the onset of unexpected level oscillations in feedwater heaters. This phenomenon can cause excessive cycling of drain valves and lead to premature failures. In this work, we develop a dynamic model of a set of feedwater heaters to determine the root cause of oscillations observed in an operating plant. Simulation results of various transient scenarios were used to investigate the effects of the controller parameters, boundary conditions, and possible valve and instrument issues. The analysis led to the conclusion that the most likely causes of the observed self-sustained oscillations in the system are the nonlinear behaviors of the drain valve and the level transmitter induced by degraded equipment condition. In conclusion, a partial plug of the pressure line used for level sensing in the system can account for a significant deadtime in the level transmitter, a nonlinear effect shown to induce self-sustained oscillatory behaviors.

Boiling water reactors↗

Multiversion software reliability through fault-avoidance and fault-tolerance

In this project we have proposed to investigate a number of experimental and theoretical issues associated with the practical use of multi-version software in providing dependable software through fault-avoidance and fault-elimination, as well as run-time tolerance of software faults. In the period reported here we have working on the following: We have continued collection of data on the relationships between software faults and reliability, and the coverage provided by the testing process as measured by different metrics (including data flow metrics). We continued work on software reliability estimation methods based on non-random sampling, and the relationship between software reliability and code coverage provided through testing. We have continued studying back-to-back testing as an efficient mechanism for removal of uncorrelated faults, and common-cause faults of variable span. We have also been studying back-to-back testing as a tool for improvement of the software change process, including regression testing. We continued investigating existing, and worked on formulation of new fault-tolerance models. In particular, we have partly finished evaluation of Consensus Voting in the presence of correlated failures, and are in the process of finishing evaluation of Consensus Recovery Block (CRB) under failure correlation. We find both approaches far superior to commonly employed fixed agreement number voting (usually majority voting). We have also finished a cost analysis of the CRB approach.

Vouk, Mladen A.↗

Improved Processing Techniques for Inclusion-Free Steel for Bearing and Mechanical Component Applications

High hardness, high carbide powder metallurgy tools steels such as M62 enable the operation of ball bearings at extremely high load and stress levels. Operation under such conditions increases the potential for rolling contact fatigue failure attributed to ceramic particle inclusions. To address this challenge, industry has sought steel made from ever increasing levels of cleanliness but the results have been uneven owing to the random nature of the occurrence of material flaws. One common approach is to rely upon careful ingot inspections prior to bearing manufacture. By selecting the cleanest portion of an ingot, it is expected that bearings relatively free from material flaws will result. This approach is not always successful because detrimental flaws that exist deep within an ingot can pass inspection undetected potentially causing subsequent failure. Recent efforts to commercialize an intermetallic material, 60NiTi, for rolling element bearings demonstrates a pathway to produce bearing steel that is free from unwanted ceramic particle inclusions. In this paper, the process used to make bearing grade ceramic-free NiTi alloys is described and applied to steelmaking. At its core, the NiTi process differs from steel making in one key aspect. NiTi alloys are made from elementally pure starting materials that are melted, blended and processed in equipment absolutely free from exposure to oxygen and ceramics ensuring a ceramic particle-free product. In contrast, the predominant method to make bearing steel is to employ a successive series of purification steps to reduce contamination levels below required thresholds. This paper describes the processes developed and applied to high carbide tool steel, M62. The resulting material and microstructures are evaluated and compared to M62 prepared by conventional powder metallurgy techniques. It is hoped that the application of materials manufacturing techniques used for fracture sensitive ceramics and intermetallic materials like NiTi can provide a pathway to Ultra-Clean, Ceramic-Inclusion free steels for rolling element bearings and other failure critical applications.

Steel↗

Electrochemical Impedance Spectroscopy of Alloys in a Simulated Space Shuttle Launch Environment

Type 304L stainless steel (304L SS) tubing is currently used in various supply lines that service the Orbiter at NASA's John F. Kennedy Space Center Launch Pads in Florida (USA). The atmosphere at the Space Shuffle launch site is very corrosive due to a combination of factors, such as the proximity of the Atlantic Ocean and the concentrated hydrochloric acid produced by the fuel combustion reaction in the solid rocket boosters. The acidic chloride environment is aggressive to most metals and causes severe pitting in many of the common stainless steel alloys such as 304L SS. Stainless steel tubing is susceptible to pitting corrosion that can cause cracking and rupture of both high-pressure gas and fluid systems. Outages in the systems where failures occur can impact the normal operation of the shuttle and launch schedules. The use of a more corrosion resistant tubing alloy for launch pad applications would greatly reduce the probability of failure, improve safety, lessen maintenance costs, and reduce downtime. A study which included ten alloys was undertaken to find a more corrosion resistant material to replace the existing 304L SS tubing. The study included atmospheric exposure at NASA's John F. Kennedy Space Center outdoor corrosion test site near the launch pads and electrochemical measurements in the laboratory which included DC techniques and electrochemical impedance spectroscopy (EIS). This paper presents the results from EIS measurements on three of the alloys: AL6XN (UN N08367), 254SMO (UNS S32l54), and 304L SS (UNS S30403). Type 304L SS was included in the study as a control. The alloys were tested in three electrolyte solutions which consisted of neutral 3.55% NaC1, 3.55% NaCl in O.1N HC1, and 3.55% NaCl in 1.ON HC1. The solutions were chosen to simulate environments that were expected to be less, similar, and more aggressive, respectively, than those present at the Space Shuttle launch pads. The results from the EIS measurements were analyzed to evaluate the corrosion susceptibility of the alloys and to predict the long-term corrosion performance of the subject materials. The results from the EIS measurements for the three alloys indicated that the higher-alloyed 254SMO and AL6XN exhibited a significantly improved resistance to corrosion than the 304L SS as the concentration of hydrochloric acid in the 3.55% NaC1 solution was increased. The polarization resistance values obtained from the EIS measurements were consistent with those from linear polarization measurements, and were indicative of the actual long-term corrosion performance of the alloys during a two-year atmospheric exposure study.

Calle, L. M.↗

Temperature Effects in Elastohydrodynamically Lubricated Contacts

This paper gives an overview of our current understanding of thermal phenomena in elastohydrodynamic contacts and suggests some avenues for fruitful research in the next decade. Typical measured temperatures are presented for representative conditions and ranges of operating parameters. Temperatures can range from bulk ambient temperature to several hundred degrees centigrade in fully separated elastohydrodynamic films. Although attention in the past decade has been on the full film for the purposes of understanding film thickness and traction phenomena, the more interesting conditions are in the mixed elastohydrodynamic films. These mixed conditions are both common in tribological systems and they are the conditions that border on unsuccessful run-in and failure of the elastohydrodynamic contact. In mixed film conditions local hotspots can have temperatures of the order of 1888 C which cause increased reactivity of the surfaces with surrounding materials as well as changes of the surface physical properties so important to the operation of concentrated contacts. An additional area discussed is that of the bulk system thermal transients which occur in tribological systems. These transients are frequently long in duration and have a direct bearing on the elastohydrodynamic film thickness and traction.

Ward O Winer↗

Understanding How Kurtosis Is Transferred from Input Acceleration to Stress Response and Its Influence on Fatigue Llife

High cycle fatigue of metals typically occurs through long term exposure to time varying loads which, although modest in amplitude, give rise to microscopic cracks that can ultimately propagate to failure. The fatigue life of a component is primarily dependent on the stress amplitude response at critical failure locations. For most vibration tests, it is common to assume a Gaussian distribution of both the input acceleration and stress response. In real life, however, it is common to experience non-Gaussian acceleration input, and this can cause the response to be non-Gaussian. Examples of non-Gaussian loads include road irregularities such as potholes in the automotive world or turbulent boundary layer pressure fluctuations for the aerospace sector or more generally wind, wave or high amplitude acoustic loads. The paper first reviews some of the methods used to generate non-Gaussian excitation signals with a given power spectral density and kurtosis. The kurtosis of the response is examined once the signal is passed through a linear time invariant system. Finally an algorithm is presented that determines the output kurtosis based upon the input kurtosis, the input power spectral density and the frequency response function of the system. The algorithm is validated using numerical simulations. Direct applications of these results include improved fatigue life estimations and a method to accelerate shaker tests by generating high kurtosis, non-Gaussian drive signals.

Kihm, Frederic↗

Validation of a Custom Ball-on-Ring Apparatus and Consideration of Common Issues for Use in Further Testing (SULI Deliverables)

Ceramic materials are well-known for their high hardness and strength but are limited in their application due to low toughness and sudden failure. As a potential solution, inspiration can be taken from dental enamel nanostructure, where undulating rods cause cracks to branch or deflect, increasing the energy needed to cause total fracture of a ceramic part. Following the dental enamel structure, a novel ceramic which uses 3D printed Yttria stabilized Zirconia rods in an alumina matrix was developed. To test this bio-inspired ceramic material, a proper testing apparatus needed to be created and tested to verify its accuracy. For this project, a bespoke ball on ring testing apparatus was created and tested using both conventionally sintered alumina disks and purchased alumina disks to validate its accuracy. By comparing the Weibull distribution of rupture strengths measured by the tests to literature values, it was shown that the testing frame had a wide distribution of strength which did not align with literature values on the lower end. Through fractography, it was found that some samples fractured from the contact stress induced by the ball indenter, which could not be used to calculate rupture strength. This fracture was often linked to low stress to failure, which was initiated by a flaw on the surface near the indenter which acted as a stress concentrator. Removal of these samples from the data set increased the accuracy of the reported rupture strength values for the ceramic. Considering the equations for the magnitude of contact and flexural strength, along with observations of initiating flaws, several measures can be taken for testing the bio-inspired composite. These measures include proper polishing of both sides of the sample, reducing sample thickness, and potentially using a softer indenter material.

36 - MATERIALS SCIENCE↗

Faults Discovery By Using Mined Data

Fault discovery in the complex systems consist of model based reasoning, fault tree analysis, rule based inference methods, and other approaches. Model based reasoning builds models for the systems either by mathematic formulations or by experiment model. Fault Tree Analysis shows the possible causes of a system malfunction by enumerating the suspect components and their respective failure modes that may have induced the problem. The rule based inference build the model based on the expert knowledge. Those models and methods have one thing in common; they have presumed some prior-conditions. Complex systems often use fault trees to analyze the faults. Fault diagnosis, when error occurs, is performed by engineers and analysts performing extensive examination of all data gathered during the mission. International Space Station (ISS) control center operates on the data feedback from the system and decisions are made based on threshold values by using fault trees. Since those decision-making tasks are safety critical and must be done promptly, the engineers who manually analyze the data are facing time challenge. To automate this process, this paper present an approach that uses decision trees to discover fault from data in real-time and capture the contents of fault trees as the initial state of the trees.

Lee, Charles↗

Integrated Systems Engineering, Safety, Reliability and Risk Management – Minimizing Black Swan Events

This paper examines key barriers that can possibly inhibit safe and reliable mission execution and, in the worst case, result in loss of human life due to many unknown contributory factors that can lead to Black Swan events. Some of the representative contributory factors include decision errors, overconfidence and a host of common causes including cultural and human factors. Decisions are always easy to criticize in hindsight when more information is available after a major accident. Depending on the type and complexity of the project and/or mission, the catastrophic risks of drifting into failure can be alleviated by implementing uniquely and strategically tailored Integrated-System-of-Systems, dynamic, risk-informed decision management processes. This paper presents some of the lessons learned from James Webb Space Telescope (JWST), NASA’s Human Space Flight program, and industry that provide motivation to organizations working on mega-complex missions to prudently accomplish targeted mission success. These lessons are important for future human Lunar, Mars and Beyond missions planned to be pursued by NASA through a public-private partnership using nimble but effective safety-conscious, proven sound engineering practices including implementation of integrated risk mitigation practices.

SLS↗

HRA Aerospace Challenges

Compared to equipment designed to perform the same function over and over, humans are just not as reliable. Computers and machines perform the same action in the same way repeatedly getting the same result, unless equipment fails or a human interferes. Humans who are supposed to perform the same actions repeatedly often perform them incorrectly due to a variety of issues including: stress, fatigue, illness, lack of training, distraction, acting at the wrong time, not acting when they should, not following procedures, misinterpreting information or inattention to detail. Why not use robots and automatic controls exclusively if human error is so common? In an emergency or off normal situation that the computer, robotic element, or automatic control system is not designed to respond to, the result is failure unless a human can intervene. The human in the loop may be more likely to cause an error, but is also more likely to catch the error and correct it. When it comes to unexpected situations, or performing multiple tasks outside the defined mission parameters, humans are the only viable alternative. Human Reliability Assessments (HRA) identifies ways to improve human performance and reliability and can lead to improvements in systems designed to interact with humans. Understanding the context of the situation that can lead to human errors, which include taking the wrong action, no action or making bad decisions provides additional information to mitigate risks. With improved human reliability comes reduced risk for the overall operation or project.

DeMott, Diana↗