Search NASASearch

SEARCH · Search NASA

Results for “root cause analysis”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

Root cause analysis of a molten salt pump in FLUSTFA

The primary salt pump installed in the high-temperature FLUoride Salt Test Facility (FLUSTFA) was successfully operated for some time, but later ceased operation. To understand what occurred, a Root Cause Analysis (RCA) was performed. Steps taken to try to get the pump operational include adjusting the shaft position, increasing the heating power of the tape heaters on the pump volute, and manually rotating the pump shaft. While removing the insulation, corrosion was noted on the outside of the pump volute, and decolorization of the insulation and tape heaters was observed. Significant corrosion products were also observed in the pump itself and the piping connected to the pump. The nitrogen cover gas was maintained from before salt was introduced into the loop until the pump was dismounted and continues to be maintained even after the pump was removed. After considering probable scenarios, causes were assigned and corrective actions were developed to prevent those causes. Then, the RCA was presented to an advisory committee for review, the “Review Committee,” consisting of experts in large molten salt systems: Brandon Haugh, David Holcomb, Kevin Robb, and Vicente Rojas. As a result, the advisory committee provided comprehensive feedback, which have been incorporated into a revised RCA. Findings have then been summarized and reported in this publication.

22 - GENERAL STUDIES OF NUCLEAR REACTORS

Feedback and Oscillations: Constructing Feedback Systems for Root Cause Analysis of Oscillations in Power Grids

Dynamic phenomena linked to inverter-based resources (IBRs) have gained global attention. Several IBR-induced dynamics have caused bulk power system-connected wind or solar power plants to trip, and some have even led to widespread outages. In addition, many oscillations have been observed involving IBR power plants. In 2023, the IEEE Power & Energy Society (PES) IBR Subsynchronous Oscillations (SSO) task force published a journal article, “Real-World Subsynchronous Oscillation Events in Power Grids With High Penetrations of Inverter-Based Resources,” in which 19 IBR oscillation events were examined for their causation. Earlier in 2020, another PES task force article, “Definition and Classification of Power System Stability-Revisited & Extended,” authored by prominent academics, introduced converter-driven stability as a new category of stability. The international power grid industry community also took action by publishing the CIGRE Green Book, Power System Dynamic Modelling and Analysis in Evolving Networks (led by Babak Badrzadeh and Zia Emin) in 2024. In August 2024, the Energy Systems Integration Group (ESIG) released a practical guide led by Nick Miller, “Diagnosis and Mitigation of Observed Oscillations in IBR-Dominant Power System: A Practical Guide.” The goal of the guide is to assist practicing engineers in making initial judgments and conducting detailed analyses about oscillations. Finally, when addressing the classification of stability and oscillations, the guide emphasizes a causality-based taxonomy for grouping, such as voltage control-induced oscillations, synchronization-induced oscillations, and frequency or active power control-induced oscillations.

Fan, Lingling [Univ. of South Florida, Tampa, FL (

Fault localization in a microfabricated surface ion trap using diamond nitrogen-vacancy center magnetometry

Here, as quantum computing hardware becomes more complex with ongoing design innovations and growing capabilities, the quantum computing community needs increasingly powerful techniques for fabrication failure root-cause analysis. This is especially true for trapped-ion quantum computing. As trapped-ion quantum computing aims to scale to thousands of ions, the electrode numbers are growing to several hundred, with likely integrated photonic components also adding to the electrical and fabrication complexity, making faults even harder to locate. In this work, we used a high-resolution quantum magnetic imaging technique, based on nitrogen-vacancy centers in diamond, to investigate short-circuit faults in an ion trap chip. We imaged currents from these short-circuit faults to ground and compared them to intentionally created faults, finding that the root cause of the faults was failures in the on-chip trench capacitors. This work, where we exploited the performance advantages of a quantum magnetic sensing technique to troubleshoot a piece of quantum computing hardware, is a unique example of the evolving synergy between emerging quantum technologies to achieve capabilities that were previously inaccessible.

71 CLASSICAL AND QUANTUM MECHANICS, GENERAL PHYSIC

Coincident learning for beam-based rf station fault identification using phase information at the SLAC linac coherent light source

Anomalies in radio-frequency (rf) stations can result in unplanned downtime and performance degradation in linear accelerators such as SLAC’s Linac Coherent Light Source (LCLS). Detecting these anomalies is challenging due to the complexity of accelerator systems, high data volume, and scarcity of labeled fault data. Prior work identified faults using beam-based detection, combining rf amplitude and beam position monitor data. Due to the simplicity of the rf amplitude data, classical methods are sufficient to identify faults, but the recall is constrained by the low-frequency and asynchronous characteristics of the data. In this work, we leverage high-frequency, time-synchronous rf phase data to enhance anomaly detection in the LCLS accelerator. Due to the complexity of phase data, classical methods fail, and we instead train deep neural networks within the Coincident Anomaly Detection (CoAD) framework. We find that applying CoAD to phase data detects nearly 3 times as many anomalies as when applied to amplitude data, while achieving broader coverage across rf stations. Furthermore, the rich structure of phase data enables us to cluster anomalies into distinct physical categories. Through the integration of auxiliary system status bits, we link clusters to specific fault signatures, providing additional granularity for uncovering the root cause of faults. We also investigate interpretability via Shapley values, confirming that the learned models focus on the most informative regions of the data and providing insight for cases where the model makes mistakes. This work demonstrates that phase-based anomaly detection for rf stations improves both diagnostic coverage and root cause analysis in accelerator systems and that deep neural networks are essential for effective analysis.

Accelerator Physics (physics.acc-ph)

Mitigating field emission in SSR2 cryomodules for PIP-II

The SSR2 cavities for PIP-II have consistently been affected by field emission since the first cold tests with the unity coupler. A comprehensive root cause analysis was conducted to investigate the origin of this issue and to identify the fabrication, processing, and handling factors that have the greatest impact on field emission onset. New techniques were developed and effectively implemented to achieve field emission-free SSR2 cavities. In addition, effort was dedicated both to relating the radiation level measured at the test stand with the expected levels in the LINAC tunnel and to understand the evolution of field emission through the assembly steps. The challenge of overcoming field emission also led to a reassessment of design choices, enhancing our understanding of their effects on cavity performance.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS

Preparation and Qualification of Preproduction SSR2 Jacketed Cavities for PIP-II

The qualification of 325 MHz Single Spoke Resonators type 2 (SSR2) jacketed cavities to meet technical requirements represents a significant milestone in the development of the SSR2 cryomodules for the PIP-II Project at Fermilab. This poster reports the procedures and lessons learned in processing and preparing these cavities for horizontal cold testing prior to integration into a cavity string assembly, with a focus on addressing the field emission issues observed during the cold testing. A comprehensive root cause analysis identified critical fabrication, processing, and handling factors impacting field emission onset. New techniques were successfully developed and implemented to achieve field emission-free SSR2 cavities, and efforts were made to correlate radiation levels measured at the test stand with expected levels in the LINAC tunnel. Additionally, the evolution of field emission through assembly steps was thoroughly investigated, leading to a reassessment of design choices and enhancing our understanding of their effects on cavity performance.

Grassellino, L. [Fermilab]

Enhanced Preparation for Intelligent Cybermanufacturing Systems (EPICS)

Opportunities exist for realizing transformative advances in productivity and reductions in energy footprint through ubiquitous sensing in manufacturing environments. Enhanced Preparation for Intelligent Cybermanufacturing Systems (EPICS) is a 21-month (4 academic semesters, plus one summer) experience for graduate students that focuses on scaling the knowledge, understanding and leadership skills in the cyber manufacturing area. Masters students (8/year, 32 total) complete 2-year projects on industrially-driven project topics, rotating to internships in summer semester to work on scoping and implementation at project partners. Students complete academic training in embedded systems, process modeling, data science, and cloud-based systems design. Their projects are targeted toward sensor retrofit, process monitoring, root cause analysis, and sensor fusion.

Advanced Manufacturing

Facilitating Data Collection of Maintenance Events to Populate the Hydrogen Component Reliability Database (HyCReD)

The Hydrogen Component Reliability Database (HyCReD) is a collaborative project between the National Renewable Energy Laboratory, the University of Maryland, and hydrogen stakeholders to improve safety and reliability for hydrogen facilities by implementing component reliability data taxonomies that support hydrogen infrastructure failure rate analysis. The project aims to quantify failure rates of hydrogen components through high-quality data collection and analysis on root causes and maintenance needed. HyCReD provides a common database for cataloging hydrogen component failures which exists for reliability research in many other mature industries [2]. The database fills a gap for the hydrogen community by providing a scientifically rigorous approach to quantitative risk assessment (QRA), prognostic health management (PHM), and reliability-centered maintenance (RCM) analysis. High level results will be aggregated and anonymized to protect company sensitive information; detailed results will be used to help address issues of hydrogen components. These advanced analytics will support accelerated deployment of hydrogen infrastructure by enabling better: design and safety of projects (safety codes and standards development), infrastructure reliability and cost (component failure rates, maintenance protocols), and component R&D needs (robust supply chain). A key to a successful HyCReD implementation is facilitating the ease of reporting and data quality in the database that can be used for analysis. Maintenance data was a previously identified gap in initial efforts to populate and validate the database taxonomies [3]. Collection of maintenance data will be instrumental in identifying failure modes and rates, identifying incipient component failures or reduced performance, cataloging best practices for maintenance routines and methods for prognostic health management, and quantifying the risk and effect of different failure modes. Several key priorities are identified for streamlined data collection to achieve quality and detailed failure data: Applicability, Ease of Use, Accessibility, and Information Security. The HyCReD team has now begun deployment of the database to several companies and groups that have signed non-disclosure agreements to facilitate the data collection of failures in industry hydrogen refueling station infrastructure. This paper will provide an update into the process of HyCReD deployment including the development of a coding guide for facility personnel to reference and ensure data quality and consistency from one station to another as well as implementation of contextually dependent data fields of system taxonomy and formatted entries to provide ease of use. The goal is to communicate the lessons learned from the roll-out to technicians and engineers in the field, and the addition of need for high level of security to protect all stakeholders.

29 ENERGY PLANNING, POLICY, AND ECONOMY

Unsupervised Process Anomaly Detection and Identification Using the Leave-One-Variable-Out Approach

Automated anomaly detection and identification can signal equipment issues and pinpoint causes in large-scale industrial systems. For systems with limited failure history, unsupervised machine learning methods can be utilized as they do not require past failures. This study introduces the leave-one-variable-out (LOVO) model, which masks one variable at a time to predict the others, learning underlying process correlations. Detection performance was assessed with synthetic and experimental data, while identification performance used only synthetic data due to its ability to generate labeled anomaly types. For detection using synthetic data, the LOVO model generally outperformed comparative models; while using experimental data, the comparative methods outperformed the LOVO model. However, the comparative methods required selecting a latent size, and these conclusions pertain to using the optimal size. In practice, it would not be feasible to always select the optimal value, and incorrect selections impacted performance. In contrast, the LOVO model does not require a latent space. For identification using synthetic data, the LOVO model was slightly outperformed in interpretability and repeatability but still demonstrated impressive results. These outcomes suggest that the LOVO model is an effective model and may be more easily implemented without the challenging tuning process of selecting a latent size.

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND

Understanding Peelle’s Pertinent Puzzle bias in generalized least squares regression through eigenspectrum analysis

Certain correlation structures in the data covariance matrix (DCM) used for generalized least squares (GLS) regression can result in biased estimates, commonly known in the field of nuclear data evaluation as Peele’s Pertinent Puzzle (PPP). This article introduces a generative, forward modeling framework within which the PPP bias is characterized through an eigenspectrum analysis of the DCM. This analysis highlights the root cause of the bias, generalizes the problem beyond the nuclear data field, and provides insight to the problem regimes where it can occur. What follows is an understanding that the bias can show up for any experimental neutron time-of-flight data for which systematic uncertainties have been quantified. Lastly, a discussion of the adaptation of cross validation approaches that require pre-whitening to incorporate the known ‘fix’ to the PPP bias in the GLS estimator.

73 NUCLEAR PHYSICS AND RADIATION PHYSICS

Rapid Characterization and Statistical Analysis of High-Volume Field-Harvested Photovoltaic Connectors

Photovoltaic (PV) installations heavily depend on connectors for efficient module and string interconnections without requiring skilled labor. Yet this seemingly innocuous component of PV systems is a leading cause of module failures, multiple high-profile fires, and lawsuits in the PV industry. This work aims to answer critical questions regarding why connectors fail and the contributing factors to their failure. The study involves collecting and analyzing more than 17,000 field-harvested connectors from various solar installations across the United States. The vast dataset, which includes connector metadata, visual inspections, and resistance measurements, provides unprecedented insight into the state of health of PV connectors across the US, including the geographic locations, connector types, and installation practices most prone to failures. The work presented here describes a novel rapid characterization method for processing large numbers of connectors and is supported by parallel forensic analysis to discern the root causes of failures as well as a levelized cost of lifetime model to determine the economic ramifications of connector failure. Ultimately, the findings may inform PV developers about the best practices to extend connector longevity and lead to more resilient and reliable PV systems.

connectors

JPL/NASA/IEEE Test Effectiveness Workshop

(none given)From OBJECTIVES: Specific objectives of the working group are to support the innovation, development, evaluation and implementation of test methods, metrics and tools based on failure engineering/physics and/or root cause evaluations. Data sources systems and tools shall be developed and implemented that: 1) provide improved preventions, controls, analyses and tests (PACT) & field failure data collection, 2) facilities data analysis, archiving, retrieval, failure physics and/or root cause evaluations and 3) enable new and existing technology suitability evaluations to be performed.

effectiveness concurrent engineering metrics test

Dynamic tooth loads and stressing for high contact ratio spur gears

An analysis and computer program were developed for calculating the dynamic gear tooth loading and root stressing for high contact ratio gearing (HCRG) as well as LCRG. The analysis includes the effects of the variable tooth stiffness during the mesh, tooth profile modification, and gear errors. The calculation of the tooth root stressing caused by the dynamic gear tooth loads is based on a modified Heywood gear tooth stress analysis, which appears more universally applicable to both LCRG and HCRG. The computer program is presently being expanded to calculate the tooth contact stressing and PV values. Sample application of the gear program to equivalent LCRG (1.566 contact ratio) and HCRG (2.40 contact ratio) revealed the following: (1) the operating conditions and dynamic characteristics of the gear system an affect the gear tooth loading and root stressing, and therefore, life significantly; (2) the length of the profile modification affect the tooth loading and root stressing significantly, the amount depending on the applied load, speed, and contact ratio; and (3) the effect of variable tooth stiffness is small, shifting and increasing the response peaks slightly from those for constant tooth stiffness.

Cornell, R. W.

Automated Programmable Logic Controller Memory Forensics Using RGB Image Analysis and Deep Learning

The introduction of Industry 4.0 and Internet-based technologies has enhanced industrial control system operations but have inadvertently increased their vulnerabilities to cyber attacks. When an industrial control system is compromised, security analysts need to identify the root cause quickly to start the recovery process and develop mitigation strategies. Memory forensics is critical in the incident analysis process to ascertain what occurred. Approaches for analyzing the persistent memory in industrial control devices are limited and almost nonexistent for volatile memory. This chapter proposes an automated methodology for programmable logic controller memory dump analysis using computer vision and deep learning techniques. The methodology converts the sequences of bytes in a programmable logic controller memory dump to red-green-blue pixels and employs a deep learning model that learns the underlying patterns and features of pre-labeled forensic artifacts in images and segments them into distinct regions. The trained model is employed to automatically segment new memory images and identify forensic artifacts. Evaluation of the methodology on a Schneider Electric Modicon M221 programmable logic controller under code injection and code modification attacks demonstrates its ability to detect attack artifacts in memory dumps.

Asmar Awad, Rima [ORNL] (ORCID:0000000233407742)

Active multi-mode data analysis to improve fault diagnosis in AHUs

Faults in heating, ventilation and air conditioning systems can lead to increased energy consumption, occupant comfort issues, and reduced equipment lifetime. Commercial fault detection and diagnosis (FDD) tools has been increasingly deployed in U.S. commercial buildings. While they are helping to achieve energy efficiency and operational reliability, there remain gaps in their fault diagnostic capabilities. The diagnostic results often contain multiple distinct candidate root causes (CRCs) or offer no insight into CRCs. This study developed a novel active rule-based multi-mode data analysis method to enhance diagnostic resolution by applying proven rule sets and additional new rules to data from multiple known operational modes. The proposed method was demonstrated using enhanced air handling unit performance assessment rule sets and validated with the simulated data of two air handling units. New metrics, namely, reduced number of CRCs and improvement ratio, were developed to quantify the improvement of fault diagnostic resolution. The validation results showed that the proposed method effectively reduced the number of CRCs in contrast to analyzing data solely for a single mode of operation. It achieved a median improvement ratio of 80% in 19 test cases.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI

Soil Moisture Buffers the Impact of Precipitation Variability on Ecosystem Productivity

Water availability governs ecosystem productivity, yet estimates of vegetation sensitivity to water can differ greatly depending on whether the sensitivity is examined spatially or temporally. In particular, the spatial sensitivity is often reported to be much stronger than temporal sensitivities, leading to highly uncertain projections of ecosystem responses to future climate change when using space-for-time substitution. The large difference between spatial and temporal sensitivities remains unexplained. Prior research, however, primarily relied on precipitation as the water availability proxy, whereas vegetation responds to soil moisture. Here, we combined satellite estimates of vegetation productivity with soil moisture data across water-limited ecosystems of the continental United States (CONUS) to identify a convergent sensitivity of productivity to water availability. Using precipitation, we show that temporal sensitivity is 66% lower than spatial sensitivity overall. Our analysis identified the cause of the difference to be primarily driven by the seasonal variability of water availability, rooting depth, and soil properties. When using soil moisture instead of precipitation, we observed widespread convergence in the spatial and temporal sensitivities—that is, the two sensitivities became much more similar in magnitude across all water-limited ecosystems within CONUS. These results show that overlooking soil hydrology can inflate perceived discrepancies between spatial and temporal vegetation sensitivities, leading to biased projections of ecosystem dynamics under future hydro-climatic change.

Wang, Huiqi [University of California, Berkeley, C

Under Pressure: The Artemis I Heatshield Char Loss Investigation and Artemis II’s Successful Entry

Following the successful Artemis I skip reentry on December 11, 2022, unexpected char liberation from the Orion heat- shield prompted formation of an Anomaly Response Team to determine root cause and establish flight rationale and corrective actions for subsequent missions. Through a comprehensive investigation including detailed hardware analysis, modeling & simulation, and ground testing, the team determined that the low permeability Avcoat experienced extreme internal gas pressure buildup during the skip entry that could not adequately outgas, leading to crack formation and char liberation. Key discoveries included significant performance differences between permeable and impermeable heatshield regions and successful replication of the anomaly through ground testing. Based on extensive ground testing and analysis, Artemis II flew a modified trajectory without the skip entry to minimize crack-inducing conditions, and successfully splashed down on April 10, 2026 with significantly reduced char loss. For Artemis III and beyond, a corrective action was implemented to produce a more permeable version of Avcoat that will substantially reduce internal pressure accumulation during entry.

Artemis II

Impacts of PV Module Connector Failures on Cost and Performance of Utility Scale Photovoltaic Systems

The reliability, cost and performance of electrical connectors are a concern in all types of electrical systems, and demands on connectors used on photovoltaic (PV) systems include that connectors maintain electrical conductivity and physical strength, endure ultraviolet sunlight and high ambient temperature, and resist moisture and chemical intrusion over a very long (>25 year) performance period. Connector failures increase operation and maintenance (O&M) costs and reduce plant production, but connector failure can also cause safety and liability problems, which are of greater concern. This work results from a three-year collaboration between Sandia National Laboratories (SNL), the Electric Power Research Institute (EPRI), and the National Renewable Energy Laboratory (NREL) and funded by the U.S. Department of Energy (DOE) Solar Energy Technology Office (SETO) under Agreements #39035 and #38531 "Connector Reliability Across the US Solar Sector." a multi-pronged investigation of PV connector health across the US (see https://energy.sandia.gov/pvconnectors/). This report presents derivation of a Techno-Economic Analysis (TEA) that models failure modes and frequencies (how often failure occurs), estimates O&M costs and lost production associated with connector failures, and then calculates the effect that PV module connectors can have on Levelized Cost of Energy (LCOE). The model is informed with initial data from quantitative assessment of failure rates, root causes and mechanisms, in-situ diagnostics and data collection, lab-based forensics, and interviews with PV connector manufacturers and plant operators. SNL conducted site inspections at multiple utility-scale sites in different climates and subjected field samples of new, used, and degraded connectors to visual and electrical characterization. EPRI conducted metallurgical analysis of the pin and sleeve conductors to study failure-induced morphological and compositional changes. There is in general a shortage of statistically valid data, but data from PVROM database maintained by SNL was sufficient to ascertain failure rates and lost production as well as provide qualitative insight in its curated maintenance records. This report details the structure of the mathematical model but the sources of data to inform the model will continue to evolve. Analysis of a 100 MW PV plant is provided as an example of the use of the model, with results indicating that connectors are responsible for Annualized O&M Costs of $\$$71,933/year; Annualized Unit O&M Costs of $\$$0.72/kW/year; that a Reserve Account of $\$$187,220 should be available to fund repairs related to connectors; that connectors add $\$$1,494,004 to the Net Present Value of the O&M Costs (project life); and that O&M related to connectors adds about $\$$0.00088/kWh to the Levelized Cost of Energy. The impact of this model is to provide a tool to make the US solar sector more robust by quantifying and monetizing the reliability risks to utility-scale PV systems posed by poorly installed, mismatched and/or poorly designed and manufactured connectors. The TEA provides a model incorporating failure statistics, O&M cost data, and lost production into a single figure of merit, informing decisions and enabling practitioners to optimize cost and performance trade-offs. Stakeholders include connector manufacturers, system designers and equipment specifiers, standards bodies, installers and O&M providers, investors and insurance underwriters. This report supports continued growth of PV predicated on assurances that properly installed and maintained PV system connectors are safe and reliable. The project team is proposing future work including accelerated testing of connectors and expanding the approach taken here to other PV system components, such as TEA for rapid shut-down devices.

14 SOLAR ENERGY