Search NASASearch

SEARCH · Search NASA

Results for “root cause analysis”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 37 records · Page 2

Coincident learning for beam-based rf station fault identification using phase information at the SLAC linac coherent light source

Anomalies in radio-frequency (rf) stations can result in unplanned downtime and performance degradation in linear accelerators such as SLAC’s Linac Coherent Light Source (LCLS). Detecting these anomalies is challenging due to the complexity of accelerator systems, high data volume, and scarcity of labeled fault data. Prior work identified faults using beam-based detection, combining rf amplitude and beam position monitor data. Due to the simplicity of the rf amplitude data, classical methods are sufficient to identify faults, but the recall is constrained by the low-frequency and asynchronous characteristics of the data. In this work, we leverage high-frequency, time-synchronous rf phase data to enhance anomaly detection in the LCLS accelerator. Due to the complexity of phase data, classical methods fail, and we instead train deep neural networks within the Coincident Anomaly Detection (CoAD) framework. We find that applying CoAD to phase data detects nearly 3 times as many anomalies as when applied to amplitude data, while achieving broader coverage across rf stations. Furthermore, the rich structure of phase data enables us to cluster anomalies into distinct physical categories. Through the integration of auxiliary system status bits, we link clusters to specific fault signatures, providing additional granularity for uncovering the root cause of faults. We also investigate interpretability via Shapley values, confirming that the learned models focus on the most informative regions of the data and providing insight for cases where the model makes mistakes. This work demonstrates that phase-based anomaly detection for rf stations improves both diagnostic coverage and root cause analysis in accelerator systems and that deep neural networks are essential for effective analysis.

Accelerator Physics (physics.acc-ph)

Mitigating field emission in SSR2 cryomodules for PIP-II

The SSR2 cavities for PIP-II have consistently been affected by field emission since the first cold tests with the unity coupler. A comprehensive root cause analysis was conducted to investigate the origin of this issue and to identify the fabrication, processing, and handling factors that have the greatest impact on field emission onset. New techniques were developed and effectively implemented to achieve field emission-free SSR2 cavities. In addition, effort was dedicated both to relating the radiation level measured at the test stand with the expected levels in the LINAC tunnel and to understand the evolution of field emission through the assembly steps. The challenge of overcoming field emission also led to a reassessment of design choices, enhancing our understanding of their effects on cavity performance.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS

Preparation and Qualification of Preproduction SSR2 Jacketed Cavities for PIP-II

The qualification of 325 MHz Single Spoke Resonators type 2 (SSR2) jacketed cavities to meet technical requirements represents a significant milestone in the development of the SSR2 cryomodules for the PIP-II Project at Fermilab. This poster reports the procedures and lessons learned in processing and preparing these cavities for horizontal cold testing prior to integration into a cavity string assembly, with a focus on addressing the field emission issues observed during the cold testing. A comprehensive root cause analysis identified critical fabrication, processing, and handling factors impacting field emission onset. New techniques were successfully developed and implemented to achieve field emission-free SSR2 cavities, and efforts were made to correlate radiation levels measured at the test stand with expected levels in the LINAC tunnel. Additionally, the evolution of field emission through assembly steps was thoroughly investigated, leading to a reassessment of design choices and enhancing our understanding of their effects on cavity performance.

Grassellino, L. [Fermilab]

Enhanced Preparation for Intelligent Cybermanufacturing Systems (EPICS)

Opportunities exist for realizing transformative advances in productivity and reductions in energy footprint through ubiquitous sensing in manufacturing environments. Enhanced Preparation for Intelligent Cybermanufacturing Systems (EPICS) is a 21-month (4 academic semesters, plus one summer) experience for graduate students that focuses on scaling the knowledge, understanding and leadership skills in the cyber manufacturing area. Masters students (8/year, 32 total) complete 2-year projects on industrially-driven project topics, rotating to internships in summer semester to work on scoping and implementation at project partners. Students complete academic training in embedded systems, process modeling, data science, and cloud-based systems design. Their projects are targeted toward sensor retrofit, process monitoring, root cause analysis, and sensor fusion.

Advanced Manufacturing

Software Sleuth

NASA's need to trace mistakes to their source to try and eliminate them in the future has resulted in software known as Root Cause Analysis (RoCA). Fair, Isaac & Co., Inc. has applied RoCA software, originally developed under an SBIR contract with Kennedy, to its predictive software technology. RoCA can generate graphic reports to make analysis of problems easier and more efficient.

Source record

Applying Model-Based Reasoning to the FDIR of the Command and Data Handling Subsystem of the International Space Station

All of the International Space Station (ISS) systems which require computer control depend upon the hardware and software of the Command and Data Handling System (C&DH) system, currently a network of over 30 386-class computers called Multiplexor/Dimultiplexors (MDMs)[18]. The Caution and Warning System (C&W)[7], a set of software tasks that runs on the MDMs, is responsible for detecting, classifying, and reporting errors in all ISS subsystems including the C&DH. Fault Detection, Isolation and Recovery (FDIR) of these errors is typically handled with a combination of automatic and human effort. We are developing an Advanced Diagnostic System (ADS) to augment the C&W system with decision support tools to aid in root cause analysis as well as resolve differing human and machine C&DH state estimates. These tools which draw from sources in model-based reasoning[ 16,291, will improve the speed and accuracy of flight controllers by reducing the uncertainty in C&DH state estimation, allowing for a more complete assessment of risk. We have run tests with ISS telemetry and focus on those C&W events which relate to the C&DH system itself. This paper describes our initial results and subsequent plans.

Robinson, Peter

Intelligent Elements for ISHM

There are a number of architecture models for implementing Integrated Systems Health Management (ISHM) capabilities. For example, approaches based on the OSA-CBM and OSA-EAI models, or specific architectures developed in response to local needs. NASA s John C. Stennis Space Center (SSC) has developed one such version of an extensible architecture in support of rocket engine testing that integrates a palette of functions in order to achieve an ISHM capability. Among the functional capabilities that are supported by the framework are: prognostic models, anomaly detection, a data base of supporting health information, root cause analysis, intelligent elements, and integrated awareness. This paper focuses on the role that intelligent elements can play in ISHM architectures. We define an intelligent element as a smart element with sufficient computing capacity to support anomaly detection or other algorithms in support of ISHM functions. A smart element has the capabilities of supporting networked implementations of IEEE 1451.x smart sensor and actuator protocols. The ISHM group at SSC has been actively developing intelligent elements in conjunction with several partners at other Centers, universities, and companies as part of our ISHM approach for better supporting rocket engine testing. We have developed several implementations. Among the key features for these intelligent sensors is support for IEEE 1451.1 and incorporation of a suite of algorithms for determination of sensor health. Regardless of the potential advantages that can be achieved using intelligent sensors, existing large-scale systems are still based on conventional sensors and data acquisition systems. In order to bring the benefits of intelligent sensors to these environments, we have also developed virtual implementations of intelligent sensors.

Schmalzel, John L.

Streamlining Payload Integration

Payload integration onto space transport vehicles and the International Space Station (ISS) is a complex process. Yet, cargo transport is the sole reason for any space mission, be it for ferrying humans, science, or hardware. As the largest such effort in history, the ISS offers a wide variety of payload experience. However, for any payload to reach the Space Station under the current process, Payload Developers face a list of daunting tasks that go well beyond just designing the payload to the constraints of the transport vehicle and its stowage topology. Payload customers are required to prove their payload s functionality, structural integrity, and safe integration - including under less than nominal situations. They must also plan for or provide training, procedures, hardware labeling, ground support, and communications. In addition, they must deal with negotiating shared consumables, integrating software, obtaining video, and coordinating the return of data and hardware. All the while, they must meet export laws, launch schedules, budget limits, and the consensus of more than 12 panel and board reviews. Despite the cost and infrastructure overhead, payload proposals have increased. Just in the span from FY08 to FY09, the NASA Payload Space Station Support Office budget rose from $78M to $96M in attempt to manage the growing manifest, but the potential number of payloads still exceeds available Payload Integration Management manpower. The growth has also increased management difficulties due to the fact that payloads are more frequently added to a flight schedule late in the flow. The current standard ISS template for payload integration from concept to payload turn-over is 36 months, or 18 months if the payload already has a preliminary design. Customers are increasingly requiring a turn-around of 3 to 6-months to meet market needs. The following paper suggests options for streamlining the current payload integration process in order to meet customer schedule needs and reduce costs for both the integration support teams and the developers, without reducing quality or compromising safety. Issues for the key integration areas of planning, training, verification, and safety are presented in a Root-Cause Analysis study, with plausible solutions provided that involve technology and tools already available to the ISS community. Although based upon the ISS process, the payload integration techniques outlined herein also offer an integration template for any space transport endeavor.

Lufkin, Susan N.

Smart Sensor Demonstration Payload

Sensors are a critical element to any monitoring, control, and evaluation processes such as those needed to support ground based testing for rocket engine test. Sensor applications involve tens to thousands of sensors; their reliable performance is critical to achieving overall system goals. Many figures of merit are used to describe and evaluate sensor characteristics; for example, sensitivity and linearity. In addition, sensor selection must satisfy many trade-offs among system engineering (SE) requirements to best integrate sensors into complex systems [1]. These SE trades include the familiar constraints of power, signal conditioning, cabling, reliability, and mass, and now include considerations such as spectrum allocation and interference for wireless sensors. Our group at NASA s John C. Stennis Space Center (SSC) works in the broad area of integrated systems health management (ISHM). Core ISHM technologies include smart and intelligent sensors, anomaly detection, root cause analysis, prognosis, and interfaces to operators and other system elements [2]. Sensor technologies are the base fabric that feed data and health information to higher layers. Cost-effective operation of the complement of test stands benefits from technologies and methodologies that contribute to reductions in labor costs, improvements in efficiency, reductions in turn-around times, improved reliability, and other measures. ISHM is an active area of development at SSC because it offers the potential to achieve many of those operational goals [3-5].

Schmalzel, John

Model-Based Fault Diagnosis: Performing Root Cause and Impact Analyses in Real Time

Generic, object-oriented fault models, built according to causal-directed graph theory, have been integrated into an overall software architecture dedicated to monitoring and predicting the health of mission- critical systems. Processing over the generic fault models is triggered by event detection logic that is defined according to the specific functional requirements of the system and its components. Once triggered, the fault models provide an automated way for performing both upstream root cause analysis (RCA), and for predicting downstream effects or impact analysis. The methodology has been applied to integrated system health management (ISHM) implementations at NASA SSC's Rocket Engine Test Stands (RETS).

Figueroa, Jorge F.

Photogrammetry Measurements During a Tanking Test on the Space Shuttle External Tank, ET-137

On November 5, 2010, a significant foam liberation threat was observed as the Space Shuttle STS-133 launch effort was scrubbed because of a hydrogen leak at the ground umbilical carrier plate. Further investigation revealed the presence of multiple cracks at the tops of stringers in the intertank region of the Space Shuttle External Tank. As part of an instrumented tanking test conducted on December 17, 2010, a three dimensional digital image correlation photogrammetry system was used to measure radial deflections and overall deformations of a section of the intertank region. This paper will describe the experimental challenges that were overcome in order to implement the photogrammetry measurements for the tanking test in support of STS-133. The technique consisted of configuring and installing two pairs of custom stereo camera bars containing calibrated cameras on the 215-ft level of the fixed service structure of Launch Pad 39-A. The cameras were remotely operated from the Launch Control Center 3.5 miles away during the 8 hour duration test, which began before sunrise and lasted through sunset. The complete deformation time history was successfully computed from the acquired images and would prove to play a crucial role in the computer modeling validation efforts supporting the successful completion of the root cause analysis of the cracked stringer problem by the Space Shuttle Program. The resulting data generated included full field fringe plots, data extraction time history analysis, section line spatial analyses and differential stringer peak ]valley motion. Some of the sample results are included with discussion. The resulting data showed that new stringer crack formation did not occur for the panel examined, and that large amounts of displacement in the external tank occurred because of the loads derived from its filling. The measurements acquired were also used to validate computer modeling efforts completed by NASA Marshall Space Flight Center (MSFC).

Littell, Justin D.

Arecibo Observatory Auxiliary M4N Socket Termination Failure Investigation

The NASA Engineering and Safety Center (NESC) was requested to support the Arecibo Observatory failure investigation in determining the root cause of the Auxiliary M4N cable failure. The NESC and Kennedy Space Center led an integrated NASA investigation with support from Marshall Space Flight Center and The Aerospace Corporation that included forensic investigation of failed hardware, finite element modeling, materials characterization, and root cause analysis. This document contains the outcome of the NESC Assessment.

Zinc Spelter Sockets

A Visual Analytics Approach to Debugging Cooperative, Autonomous Multi-Robot Systems’ Worldviews

Autonomous multi-robot systems, where a team of robots shares information to perform tasks that are beyond an individual robot’s abilities, hold great promise for a number of applications, such as planetary exploration missions. Each robot in a multi- robot system autonomously schedules which robots should perform a given task and when, using its worldview–the robot’s internal representation of its belief about the environment and other robots’ states. A key problem for operators is that robots’ worldviews can fall out of sync (often due to weak communication links), leading to desynchronization of the robots’ scheduling decisions and inconsistent emergent behavior (e.g., tasks not performed, or performed by multiple robots). Operators face the time-consuming and difficult task of making sense of the robots’ scheduling decisions, detecting de-synchronizations, and pinpointing their cause by comparing every robot’s worldview. To address these challenges, we introduce MOSAIC Viewer, a visual analytics system that helps operators (i) make sense of the robots’ schedules and (ii) detect and conduct a root cause analysis of the robots’ desynchronized worldviews. Over a year-long partnership with roboticists at the NASA Jet Propulsion Laboratory, a formative study was performed to identify the necessary system design requirements, which supports the design of the system. A qualitative study with 12 roboticists reveals that MOSAIC Viewer is faster- and easier-to-use than the users’ current approaches, and it allows them to stitch low-level details to formulate a high-level understanding of the robots’ schedules and detect and pin-point the cause of desynchronized worldviews.

Ma, Kwan-Liu

BERT-Based Topic Modeling and Information Retrieval to Support Fishbone Diagramming for Safe Integration of Unmanned Aircraft Systems in Wildfire Response

Recent concepts for emerging wildfire response operations have included unmanned aircraft systems (UAS) due to their increasing accessibility and capabilities. To integrate UAS into wildfire response safely, researchers have studied the use of large repositories of historic incident reports to improve the scope of root cause analysis. Recent work has emphasized applying state-of-the-art natural language processing techniques to extract useful information from these repositories. However, it has not yet been studied how these results can be interpreted and integrated into the systems engineering process. In this work, we propose a process in which Bidirectional Encoder Representations from Transformers (BERT)-based topic modeling and information retrieval are applied to a relevant set of documents in order to support the development of a fishbone diagram in a semiautomated process. High-level themes in the document set are identified using topic modeling, which are then refined and interpreted by a human analyst. Then, the themes are used to guide a finer search using information retrieval, which returns specific incident reports of relevance. This provides traceability to specific incidents as well as broader categorizations that comprise the fishbone branches. We apply the proposed process to relevant documents from NASA’s Aviation Safety Reporting System (ASRS). The proposed process is widely applicable when relevant documents are available, and the results from this study will be useful to identifying potential causes of wildfire response UAS incidents.

Hazard analysis

BERT-based Topic Modeling and Information Retrieval to Support Fishbone Diagramming for Safe Integration of Unmanned Aircraft Systems in Wildfire Response

Recent concepts for emerging wildfire response operations have included unmanned aircraft systems (UAS) due to their increasing accessibility and capabilities. To integrate UAS into wildfire response safely, researchers have studied the use of large repositories of historic incident reports to improve the scope of root cause analysis. Recent work has emphasized applying state-of-the-art natural language processing techniques to extract useful information from these repositories. However, it has not yet been studied how these results can be interpreted and integrated into the systems engineering process. In this work, we propose a process in which Bidirectional Encoder Representations from Transformers (BERT)-based topic modeling and information retrieval are applied to a relevant set of documents in order to support the development of a fishbone diagram in a semiautomated process. High-level themes in the document set are identified using topic modeling, which are then refined and interpreted by a human analyst. Then, the themes are used to guide a finer search using information retrieval, which returns specific incident reports of relevance. This provides traceability to specific incidents as well as broader categorizations that comprise the fishbone branches. We apply the proposed process to relevant documents from NASA’s Aviation Safety Reporting System (ASRS). The proposed process is widely applicable when relevant documents are available, and the results from this study will be useful to identifying potential causes of wildfire response UAS incidents.

hazard analysis

BERT-based Topic Modeling and Information Retrieval to Support Fishbone Diagramming for Safe Integration of Unmanned Aircraft Systems in Wildfire Response

Recent concepts for emerging wildfire response operations have included unmanned aircraft systems (UAS) due to their increasing accessibility and capabilities. To integrate UAS into wildfire response safely, researchers have studied the use of large repositories of historic incident reports to improve the scope of root cause analysis. Recent work has emphasized applying state-of-the-art natural language processing techniques to extract useful information from these repositories. However, it has not yet been studied how these results can be interpreted and integrated into the systems engineering process. In this work, we propose a process in which Bidirectional Encoder Representations from Transformers (BERT)-based topic modeling and information retrieval are applied to a relevant set of documents in order to support the development of a fishbone diagram in a semiautomated process. High-level themes in the document set are identified using topic modeling, which are then refined and interpreted by a human analyst. Then, the themes are used to guide a finer search using information retrieval, which returns specific incident reports of relevance. This provides traceability to specific incidents as well as broader categorizations that comprise the fishbone branches. We apply the proposed process to relevant documents from NASA’s Aviation Safety Reporting System (ASRS). The proposed process is widely applicable when relevant documents are available, and the results from this study will be useful to identifying potential causes of wildfire response UAS incidents.

hazard analysis

Space Shuttle Operations and Infrastructure: A Systems Analysis of Design Root Causes and Effects

This NASA Technical Publication explores and documents the nature of Space Shuttle operations and its supporting infrastructure and addresses fundamental questions often asked of the Space Shuttle program why does it take so long to turnaround the Space Shuttle for flight and why does it cost so much? Further, the report provides an overview of the cause-and effect relationships between generic flight and ground system design characteristics and resulting operations by using actual cumulative maintenance task times as a relative measure of direct work content. In addition, this NASA TP provides an overview of how the Space Shuttle program's operational infrastructure extends and accumulates from these design characteristics. Finally, and most important, the report derives a set of generic needs from which designers can revolutionize space travel from the inside out by developing and maturing more operable and supportable systems.

McCleskey, Carey M.

Achieving Improved Reliability with Failure Analysis

Reliability is the ability of a product to properly function, within specified performance limits, for a specified period of time, under the life cycle application conditions. Failure analysis is a vital tool in the effort to ensure reliability of electronic products and systems throughout their product lifecycle. Today, organizations involved in activities within the electronics supply chain are facing new challenges, not just from complex assembly styles, harsher lifecycle environments, and sophisticated supply chains, but also from customers who are demanding a quicker turn-around. Unfortunately, root cause failure analysis is often performed incompletely, leading to a poor understanding of failure mechanisms and causes and, customer dissatisfaction due to recurring failures. The PDC (Professional Development Course) starts with an introduction to reliability concepts, physics of failure and an overview of failure mechanisms that affect PCBs (Printed Circuit Boards), PCBAs (Printed Circuit Board Assembly) and components. The PDC then dives into root cause hypothesizing techniques (Pareto, FMEA (Failure Modes and Effects Analysis), fishbone (Cause-And-Effect Diagram), FTA (Fault Tree Analysis)), non-destructive and destructive analysis and, materials characterization will be discussed. Numerous failure analysis case studies will be used to illustrate the techniques and analysis principles to arrive at the root cause(s) of field failures on printed circuit boards, active components, and assemblies. What Attendees will Learn: Topics include: Overview of Reliability Concepts Failure mechanisms of electronic products Root cause analysis Failure analysis techniques -Non-destructive techniques (optical, CSAM (Confocal Scanning Electron Microscopy) etc.) -Destructive analysis (DPA (Destructive Physical Analysis), Decap (Decapsulation), FIB (Focused Ion Beam) etc.) -Materials characterization (XRF (X-Ray Fluorescence) , EDS (Error Detection Sequential), TMA/DSC (Thermal Mechanical Analysis/Differential Scanning Calorimetry) etc.)

PCB quality