Search NASASearch

SEARCH · Search NASA

Results for “root cause analysis”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 55 records · Page 3

Achieving Improved Reliability with Failure Analysis

Reliability is the ability of a product to properly function, within specified performance limits, for a specified period of time, under the life cycle application conditions. Failure analysis is a vital tool in the effort to ensure reliability of electronic products and systems throughout their product lifecycle. Today, organizations involved in activities within the electronics supply chain are facing new challenges, not just from complex assembly styles, harsher lifecycle environments, and sophisticated supply chains, but also from customers who are demanding a quicker turn-around. Unfortunately, root cause failure analysis is often performed incompletely, leading to a poor understanding of failure mechanisms and causes and, customer dissatisfaction due to recurring failures. The PDC starts with an introduction to reliability concepts, physics of failure and an overview of failure mechanisms that affect PCBs, PCBAs and components. The PDC then dives into root cause hypothesizing techniques (Pareto, FMEA, fishbone, FTA), non-destructive and destructive analysis and, materials characterization will be discussed. Numerous failure analysis case studies will be used to illustrate the techniques and analysis principles to arrive at the root cause(s) of field failures on printed circuit boards, active components, and assemblies. What Will You Learn: Topics include: Overview of Reliability Concepts Failure mechanisms of electronic products Root cause analysis Failure analysis techniques -Non-destructive techniques (optical, CSAM etc.) -Destructive analysis (DPA, Decap, FIB etc.) -Materials characterization (XRF, EDS, TMA/DSC etc.) Who Will Benefit: Reliability engineers, failure analysis engineers, engineering managers, design engineers, component engineers, quality assurance functions and, personnel involved with reliability activities within their company.

non-destructive techniques

Memory Circuit Fault Simulator

Spacecraft are known to experience significant memory part-related failures and problems, both pre- and postlaunch. These memory parts include both static and dynamic memories (SRAM and DRAM). These failures manifest themselves in a variety of ways, such as pattern-sensitive failures, timingsensitive failures, etc. Because of the mission critical nature memory devices play in spacecraft architecture and operation, understanding their failure modes is vital to successful mission operation. To support this need, a generic simulation tool that can model different data patterns in conjunction with variable write and read conditions was developed. This tool is a mathematical and graphical way to embed pattern, electrical, and physical information to perform what-if analysis as part of a root cause failure analysis effort.

Sheldon, Douglas J.

A System for Fault Management for NASA's Deep Space Habitat

NASA's exploration program envisions the utilization of a Deep Space Habitat (DSH) for human exploration of the space environment in the vicinity of Mars and/or asteroids. Communication latencies with ground control of as long as 20+ minutes make it imperative that DSH operations be highly autonomous, as any telemetry-based detection of a systems problem on Earth could well occur too late to assist the crew with the problem. A DSH-based development program has been initiated to develop and test the automation technologies necessary to support highly autonomous DSH operations. One such technology is a fault management tool to support performance monitoring of vehicle systems operations and to assist with real-time decision making in connection with operational anomalies and failures. Toward that end, we are developing Advanced Caution and Warning System (ACAWS), a tool that combines dynamic and interactive graphical representations of spacecraft systems, systems modeling, automated diagnostic analysis and root cause identification, system and mission impact assessment, and mitigation procedure identification to help spacecraft operators (both flight controllers and crew) understand and respond to anomalies more effectively. In this paper, we describe four major architecture elements of ACAWS: Anomaly Detection, Fault Isolation, System Effects Analysis, and Graphic User Interface (GUI), and how these elements work in concert with each other and with other tools to provide fault management support to both the controllers and crew. We then describe recent evaluations and tests of ACAWS on the DSH testbed. The results of these tests support the feasibility and strength of our approach to failure management automation and enhanced operational autonomy.

Fault management

A System for Fault Management and Fault Consequences Analysis for NASA's Deep Space Habitat

NASA's exploration program envisions the utilization of a Deep Space Habitat (DSH) for human exploration of the space environment in the vicinity of Mars and/or asteroids. Communication latencies with ground control of as long as 20+ minutes make it imperative that DSH operations be highly autonomous, as any telemetry-based detection of a systems problem on Earth could well occur too late to assist the crew with the problem. A DSH-based development program has been initiated to develop and test the automation technologies necessary to support highly autonomous DSH operations. One such technology is a fault management tool to support performance monitoring of vehicle systems operations and to assist with real-time decision making in connection with operational anomalies and failures. Toward that end, we are developing Advanced Caution and Warning System (ACAWS), a tool that combines dynamic and interactive graphical representations of spacecraft systems, systems modeling, automated diagnostic analysis and root cause identification, system and mission impact assessment, and mitigation procedure identification to help spacecraft operators (both flight controllers and crew) understand and respond to anomalies more effectively. In this paper, we describe four major architecture elements of ACAWS: Anomaly Detection, Fault Isolation, System Effects Analysis, and Graphic User Interface (GUI), and how these elements work in concert with each other and with other tools to provide fault management support to both the controllers and crew. We then describe recent evaluations and tests of ACAWS on the DSH testbed. The results of these tests support the feasibility and strength of our approach to failure management automation and enhanced operational autonomy

System Effects Analysis

Facilitating Data Collection of Maintenance Events to Populate the Hydrogen Component Reliability Database (HyCReD)

The Hydrogen Component Reliability Database (HyCReD) is a collaborative project between the National Renewable Energy Laboratory, the University of Maryland, and hydrogen stakeholders to improve safety and reliability for hydrogen facilities by implementing component reliability data taxonomies that support hydrogen infrastructure failure rate analysis. The project aims to quantify failure rates of hydrogen components through high-quality data collection and analysis on root causes and maintenance needed. HyCReD provides a common database for cataloging hydrogen component failures which exists for reliability research in many other mature industries [2]. The database fills a gap for the hydrogen community by providing a scientifically rigorous approach to quantitative risk assessment (QRA), prognostic health management (PHM), and reliability-centered maintenance (RCM) analysis. High level results will be aggregated and anonymized to protect company sensitive information; detailed results will be used to help address issues of hydrogen components. These advanced analytics will support accelerated deployment of hydrogen infrastructure by enabling better: design and safety of projects (safety codes and standards development), infrastructure reliability and cost (component failure rates, maintenance protocols), and component R&D needs (robust supply chain). A key to a successful HyCReD implementation is facilitating the ease of reporting and data quality in the database that can be used for analysis. Maintenance data was a previously identified gap in initial efforts to populate and validate the database taxonomies [3]. Collection of maintenance data will be instrumental in identifying failure modes and rates, identifying incipient component failures or reduced performance, cataloging best practices for maintenance routines and methods for prognostic health management, and quantifying the risk and effect of different failure modes. Several key priorities are identified for streamlined data collection to achieve quality and detailed failure data: Applicability, Ease of Use, Accessibility, and Information Security. The HyCReD team has now begun deployment of the database to several companies and groups that have signed non-disclosure agreements to facilitate the data collection of failures in industry hydrogen refueling station infrastructure. This paper will provide an update into the process of HyCReD deployment including the development of a coding guide for facility personnel to reference and ensure data quality and consistency from one station to another as well as implementation of contextually dependent data fields of system taxonomy and formatted entries to provide ease of use. The goal is to communicate the lessons learned from the roll-out to technicians and engineers in the field, and the addition of need for high level of security to protect all stakeholders.

29 ENERGY PLANNING, POLICY, AND ECONOMY

Unsupervised Process Anomaly Detection and Identification Using the Leave-One-Variable-Out Approach

Automated anomaly detection and identification can signal equipment issues and pinpoint causes in large-scale industrial systems. For systems with limited failure history, unsupervised machine learning methods can be utilized as they do not require past failures. This study introduces the leave-one-variable-out (LOVO) model, which masks one variable at a time to predict the others, learning underlying process correlations. Detection performance was assessed with synthetic and experimental data, while identification performance used only synthetic data due to its ability to generate labeled anomaly types. For detection using synthetic data, the LOVO model generally outperformed comparative models; while using experimental data, the comparative methods outperformed the LOVO model. However, the comparative methods required selecting a latent size, and these conclusions pertain to using the optimal size. In practice, it would not be feasible to always select the optimal value, and incorrect selections impacted performance. In contrast, the LOVO model does not require a latent space. For identification using synthetic data, the LOVO model was slightly outperformed in interpretability and repeatability but still demonstrated impressive results. These outcomes suggest that the LOVO model is an effective model and may be more easily implemented without the challenging tuning process of selecting a latent size.

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND

Building Reliable Printed Circuit Boards - the Lessons Learned

"Printed circuit boards (PCBs) are the baseline for electronics manufacturing upon which electronic components are mounted and formed into electronic systems. PCBs are used in a variety of electronic circuits from simple one-transistor amplifiers to large super computers. A PCB serves three main functions: 1) it provides the necessary mechanical support for the components in the circuit 2) it provides the necessary electrical interconnections, and 3) it bears some form of legend which identifies the components it carries. The failure modes on the PCBs can be categorized in a hierarchical structure, in which the mechanisms and causes are site or location dependant. This one day workshop will discuss a variety of failure mechanisms that effect the functionality of PCBs. These mechanisms can be related to how PCB materials are selected, PCBs are designed, manufactured, tested and used in the field conditions. The workshop will begin with an overview of PCB manufacturing, materials and processes. With the help of examples and case studies, a wide range of failure mechanisms will be discussed, case studies are focused on digital circuits, however some failures in analog, double sided boards are also presented. The workshop then provides the guideline for selection of methodologies for identifying potential failure mechanisms based on the failure history and how a systematic root cause failure analysis of the PCB can result in prevention of future issues."

Sood, Bhanu

Creating Formal Characterizations of Routine Contingency Management in Commercial Aviation

The identification, modelling, and analysis of root causes of accidents and incidents dominate conventional safety management approaches. However, the effect of humans’ safety-producing behavior on the overall resilience of the system is often neglected. Additionally, emerging aviation markets are giving rise to concepts of operation, such as urban air mobility and optionally piloted air cargo operations, that are leading to a shift in locus of control between humans and automation. Without an understanding of the human contribution to safety, it is difficult to assess the effects of these novel role allocations on overall system safety. In this work, safety-producing behaviors are identified and abstracted into resilient performance strategies. Production rules that encapsulate these strategies are then generated and classified in the Soar cognitive architecture. The strategies are then applied to a remotely-operated air cargo example to demonstrate how safe learning is facilitated. The learned rules and strategies are then formally verified.

Safety Critical Systems

Spaceflight Over the Last Ten Years: Failures and Fix-ups 2013 - 2022 RAMS XV Conference

This project is an overview and analysis of the past ten years (2013 – present) in global spaceflight, highlighting the orbital launches of countries, particularly the failures that occurred, the reasons they occurred, and a breakdown of the related statistics and background information. The analysis is conducted from a perspective of reliability and maintainability engineering and Probabilistic Risk Assessment (PRA) to formulate a quantifiable understanding of the data and how it is pertinent to Safety and Mission Assurance (SMA) in spaceflight. There is a breakdown by country or group, timelines, and number of launches. The failures over the years are categorized, all given a broad analysis of subsystem failures and details of events. A few failures are given a more in-depth analysis of root causes and failure modes determined by their unique or common nature. The data is retrieved from online, publicly available sources.

Quinn Slaugenhoupt

Understanding Peelle’s Pertinent Puzzle bias in generalized least squares regression through eigenspectrum analysis

Certain correlation structures in the data covariance matrix (DCM) used for generalized least squares (GLS) regression can result in biased estimates, commonly known in the field of nuclear data evaluation as Peele’s Pertinent Puzzle (PPP). This article introduces a generative, forward modeling framework within which the PPP bias is characterized through an eigenspectrum analysis of the DCM. This analysis highlights the root cause of the bias, generalizes the problem beyond the nuclear data field, and provides insight to the problem regimes where it can occur. What follows is an understanding that the bias can show up for any experimental neutron time-of-flight data for which systematic uncertainties have been quantified. Lastly, a discussion of the adaptation of cross validation approaches that require pre-whitening to incorporate the known ‘fix’ to the PPP bias in the GLS estimator.

73 NUCLEAR PHYSICS AND RADIATION PHYSICS

Helios High Altitude Long Endurance Mission Mishap

This slide presentation reviews the failure of the Helios solar aircraft failure. Included are pictures of the aircraft, inflight, and after the mishap, analysis of the root causes of the mishap, contributing factors, recommendations and lessons learned in respect to crew training, and assessing the level of risk.

Henwood, Barton E.

Space Shuttle Stiffener Ring Foam Failure Analysis, a Non-Conventional Approach

The Space Shuttle Program made use of the excellent properties of rigid polyurethane foam for cryogenic tank insulation and as structural protection on the solid rocket boosters. When foam applications de-bond, classical methods of failure analysis did not provide root cause of the failure of the foam. Realizing that foam is the ideal media to document and preserve its own mode of failure, thin sectioning was seen as a logical approach for foam failure analysis to observe the three dimensional morphology of the foam cells. The cell foam morphology provided a much greater understanding of the failure modes than previously achieved.

cryogenic tank insulation

Space Shuttle Stiffener Ring Foam Failure Analysis, a Non-Conventional Approach.

The Space Shuttle Program made use of the excellent properties of rigid polyurethane foam for cryogenic tank insulation and as structural protection on the solid rocket boosters. When foam applications de-bond, classical methods of failure analysis did not provide root cause of the failure of the foam. Realizing that foam is the ideal media to document and preserve its own mode of failure, thin sectioning was seen as a logical approach for foam failure analysis to observe the three dimensional morphology of the foam cells. The cell foam morphology provided a much greater understanding of the failure modes than previously achieved.

Foam

Rapid Characterization and Statistical Analysis of High-Volume Field-Harvested Photovoltaic Connectors

Photovoltaic (PV) installations heavily depend on connectors for efficient module and string interconnections without requiring skilled labor. Yet this seemingly innocuous component of PV systems is a leading cause of module failures, multiple high-profile fires, and lawsuits in the PV industry. This work aims to answer critical questions regarding why connectors fail and the contributing factors to their failure. The study involves collecting and analyzing more than 17,000 field-harvested connectors from various solar installations across the United States. The vast dataset, which includes connector metadata, visual inspections, and resistance measurements, provides unprecedented insight into the state of health of PV connectors across the US, including the geographic locations, connector types, and installation practices most prone to failures. The work presented here describes a novel rapid characterization method for processing large numbers of connectors and is supported by parallel forensic analysis to discern the root causes of failures as well as a levelized cost of lifetime model to determine the economic ramifications of connector failure. Ultimately, the findings may inform PV developers about the best practices to extend connector longevity and lead to more resilient and reliable PV systems.

connectors

5 Inch Thick 2219-T87 Plate Low Ductility Investigation - Material Analysis on Broken 2219 Tensile Samples

2219-T87 material used on a LOX tank baffle support on Artemis-I showed unusually low fracture elongation. Recent material data indicated 5” thick 2219-T87 plate stock had very low ductility (<1%) in the short transverse direction at the half-thickness location. Metallurgical analysis performed to understand the root cause that led to the drastic reduction in tensile ductility. The PowerPoint report details the key findings from material analysis by optical and scanning electron microscopy.

aluminum alloy 2219