Search NASASearch

SEARCH · Search NASA

Results for “Failure Modes”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 37 records · Page 2

Risk Assessment of EIC Central Detector (ePIC) Solenoid Magnet (MARCO)

As part of the BNL-JLab-CEA Electron Ion Collider (EIC) collaboration, the design of a 2 T, 2.8 m bore diameter, 3.8 m long conduction cooled superconducting detector magnet design is completed. Such magnet will be employed at the interaction region of ePIC for physics experiment. The magnet is a passive shielded solenoidal magnet system consisting of 3 coils wound with specially designed conductor using NbTi Rutherford type cable in copper stabilized channel. This paper describes the risk analysis as the part of Failure Modes and Effects Analysis (FMEA) that was carried out as a team to identify their various failure modes and risks associated with the magnets system. In conclusion, this FMEA is intended to become the content of the designed document as an integral part of the engineering assessment and the statement of work for the potential vendors towards design and built.

Ghoshal, Probir K. [Thomas Jefferson National Acce

Highly accelerated life testing (HALT): A review from a statistical perspective

Despite its use in one form or another for at least four decades, HALT and related techniques [e.g., highly accelerated-stress screening (HASS) and stress audits (HASA)] are not well understood within the statistical community and remain controversial. This largely reflects a conflict in motivation between engineers, testing under harsh conditions to discover and eliminate failure modes, and statisticians, taking a more cautious approach to develop quantitative estimates of parameters such as mean time between failures (MTBF). Here, this review article will clarify HALT concepts and methods and explain where it fits within the universe of methods that involve the application of accelerating factors to compress the time required to evaluate or enhance product reliability. A major distinction is between methods such as HALT, a high-stress test-analyze-fix-test iterative process directed at improving reliability by discovering and fixing weak points in a design, and quantitative accelerated life testing (QALT), whose goal is the estimation of product life for a fixed design. We discuss methods such as physics of failure that offer some hope of bridging the gap between the qualitative nature of HALT, and purely quantitative statistical methods. We present a variety of engineering applications of HALT including metal fatigue, piping and pressure vessels, structural damage, radiation damage, and rotating machinery. We also discuss potential synergies between HALT and QALT, such as rapid identification, through HALT, of failure modes requiring quantitative analysis. For further study, extensive references to the applicable literature are provided as well as an appendix that describes related methods.

97 MATHEMATICS AND COMPUTING

Success Path Method: Introduction to the Success Path Method Software Tool©

As part of its commitment to advancing safety and reliability assessment methodologies, Argonne National Laboratory pioneered the use of an evaluation method called the Success Path Method (SPM) to improve risk management for offshore oil and gas operations. The development of the SPM at Argonne has been driven by the need to improve existing risk assessment methodologies by focusing on the steps necessary for success rather than failure modes alone. This is particularly important for industrial environments like offshore facilities that perform multiple functions under a continuously evolving set of operational conditions – such as water depth and temperature, currents, and weather conditions. In these dynamic environments, the traditional Probabilistic Risk Assessment (PRA) approach is far too complex as it focuses on what can go wrong – which comprises an infinite failure space that must be fully explored and understood. By shifting the focus to a finite space of success paths, the SPM enables operators and decision makers to prioritize a manageable number of steps that must go right to ensure success. Building on its five decades of experience in safety assessments for the nuclear industry, Argonne made major adaptations to existing risk assessment methods utilizing features similar to fault trees that are traditionally used in PRA to map all pathways in which the system can malfunction. In contrast, SPM identifies the components and processes that must function correctly to achieve specific outcomes – such as preventing the uncontrolled release of hydrocarbons during drilling operations. The SPM framework integrates equipment, procedures, software, processes, and human actions to ensure that physical barriers meet critical safety functions in dynamic operational conditions. This approach helps identify failure modes and improve operational risk management by narrowing the focus to key success elements, which in turn reduces uncertainty and helps users understand, manage, and respond to failures.

97 MATHEMATICS AND COMPUTING

Finite element procedure for thermomechanical and structural integrity analysis of beam intercepting devices subjected to free electron laser

High-energy particles, including photons (x-ray, γ-ray, bremsstrahlung), electrons, and protons, possess the capability to penetrate materials and deposit energy within them. The degree of absorption depends on both the energy and type of particles, as well as the properties of the materials with which they interact. This energy deposition can manifest either at the material's surface or throughout its volume, potentially resulting in various failure modes. The primary aim of this paper is to establish a structured analysis methodology for evaluating the structural integrity of beam-intercepting devices when subjected to high-energy particles. Here, the paper also reviews some of the underlying physics, pertinent to the scope of the thermomechanical analysis, potential failure modes, and introduces verification and validation methodologies. Engineers and researchers can utilize the guidelines presented in this paper to effectively plan the development of beam intercepting devices, thereby ensuring their reliability and performance in the presence of high-energy particle exposure.

Finite element analysis

Micromechanical response of SiC-OPyC layers in TRISO fuel particles

Tristructural isotropic (TRISO)–coated particle fuel is a proposed fuel for multiple advanced reactor concepts. The performance of the particle depends on whether the silicon carbide (SiC) layer remains intact to prevent the release of metallic and gaseous fission products. Mechanical fracture of the SiC layer is a potential failure mode under various fuel configurations and operating environments, including the potential transmission of matrix-originating cracks through TRISO particles. Furthermore, this study uses instrumented indentation techniques on cross-sectioned surrogate particles to examine the mechanical stability of the critical interface between SiC and the outer pyrolytic carbon (OPyC) layer. The observed behavior at the interface is rationalized by examining the radially dependent fracture behavior of the SiC layer and performing a numerical analysis to quantify the residual stresses that develop during the processing and cross-sectioning of the as-fabricated particle. Characterizing the SiC-OPyC interface of surrogate TRISO particles using nanoindentation provides unique insight into the interface's room-temperature residual stress and mechanical stability. The modeling efforts were used to investigate the experimental procedure further, and the results are presented herein to validate this fuel form's potential mechanical failure modes.

11 NUCLEAR FUEL CYCLE AND FUEL MATERIALS

Microreactor Automated Control System Test Bed Digital Architecture for Real-Time, Hardware-in-the-Loop Simulation

This work describes progress made towards the development of a real-time hardware-in-the-loop (HIL) test bed for non-nuclear testing of microreactor control schemes and failure modes. Non-nuclear testing is a crucial step in developing robust control algorithms for managing microreactor dynamics. The creation of an HIL simulation harnesses the realistic dynamics of physical analogue systems while additionally considering the challenges of variable communication delay. This collaborative effort between Oak Ridge National Laboratory and Idaho National Laboratory has resulted in a LabVIEW-based gRPC communication protocol which couples a TRANSFORM Modelica simulation of nuclear components to the ViBRANT physical hardware for realistic feedback and visual representation of control action in real time. A modular python client structure is developed to manage FMU-based Modelica simulation and real-time gRPC communication. HIL testing suggests that the modeled reactor with natural convection molten salt loop coolant configuration responds well to PID control of drum positioning for modulation of reactor core power, however, future efforts will be made to explore the added thermal inertial delay of system level control and downstream demand changes. Development of this platform with a generalized methodology provides a foundation for exploring a variety of reactor configurations and failure modes in rapid order to provide insight into the most effective avenues of study for further research and development.

McConnell, Jono [ORNL] (ORCID:0000000238984741)

DROP DURABILITY ASSESSMENT OF ELECTRONIC ASSEMBLIES UNDER OFF-AXIS LOADING WITH SKEWED FIXTURES

This thesis studies drop durability of electronic assemblies when the acceleration vector is oriented at 45° to the out-of-plane direction of the circuit card. The off-axis drop tests are accomplished with a skewed fixture and are conducted as a proxy for multiaxial drop testing. Advanced shock testing and vibration test methods have been developed over the last few decades to better represent real-world field environments during ground-based laboratory testing. However, many of these test methods require expensive and specialized equipment not available in most laboratories. An alternative approach for approximating simultaneous loading along multiple axes on conventional equipment utilizes skewed fixtures which have seen use in off-axis random vibration and drop impact testing. These methods generally rely on the conversion of a uniaxial input load from the test equipment (using a uniaxial drop tower or shaker) into a multiaxial load when resolved in the reference frame of the test article (mounted on a skewed fixture). Skewed fixture design is presented and recommendations for conducting skewed angle drop testing are introduced based on local measurements along the skewed face of the fixture to accurately monitor the impact event. Characterization tests were performed with a skewed fixture, at simultaneous acceleration loads from 500 to 3,000 g in two (in-plane and out-of-plane) directions, while meeting standard time domain tolerances. Upon experimental characterization, drop shock durability tests were conducted on a printed circuit assembly (PCA). Mean drops-to-failure were measured and quantified with Weibull statistics. Dominant solder joint failure modes were identified via failure analysis. Prior work on inclined angle impact testing is limited, and the majority of solder joint interconnect level fatigue studies are conducted considering perpendicular loading normal the circuit card. Low-cycle fatigue curves are generated based on plastic strain and plastic work density within the solder joint. A multiscale nonlinear finite element model is used to relate board-level flexure to solder joint interconnect level plastic strain. A high strain rate solder constitutive model allows for accurate modeling of solder plasticity resulting from high-impact drop shock. Fatigue parameters are computed from the Coffin-Manson relation and Palmgren-Miner damage accumulation. This work serves to apply established low-cycle fatigue methods for conventional drop shock loading (impact normal to circuit card) to non-perpendicular loading with a skewed fixture.

Hower, Jonathan [Kansas City National Security Cam

Reliability Assessment of Cooling Fans for PV Inverters: Testing, Modeling, and Case Studies

The reliability of photovoltaic (PV) inverters is critical for long-term solar system performance, with cooling fan failures frequently leading to costly downtime. While much research exists on general cooling fan reliability, little attention has been given to fans operating within PV inverters and their unique environmental challenges. Here, this article proposes a comprehensive methodology to address this gap. First, a failure mode and effects analysis is performed on fans to identify the key failure mechanisms in PV applications, their corresponding stressors, and the models necessary for lifetime prediction. Second, an accelerated life test is designed and conducted to collect valuable experimental data for PV inverter fans in a reasonable amount of time. Third, a mathematical conversion of dynamic mission profiles into effective constant stress levels is derived. Fourth, case studies are given, showcasing lifetime estimates that account for geographic variations in mission profile data. The results demonstrate that this integrated approach leads to an accurate reliability assessment for PV inverter cooling fans.

accelerated life testing (ALT)

Shock-induced twinning/detwinning and spall failure in Cu–Ta nanolaminates at atomic scales

Here, this study provides new insights into the role of interfaces on the deformation and failure mechanisms in shock-loaded Cu–Ta–Cu trilayer system. The thickness of the Ta layer, piston velocities, and shock pulse durations were varied to explore the impact of impedance mismatch and loading conditions on spallation behavior and twin formation. It was found that the interfaces play a crucial role in the dynamic response of these multilayered systems since secondary reflection waves generated at the interfaces significantly affected the peak stress and pressure profiles, influencing void nucleation and failure modes. In the trilayer systems, failure predominantly occurred at interfaces and within the Ta layer, with void nucleation sites and twinning behavior being markedly different compared to single-crystal Cu and Ta. Increasing the Ta layer thickness modified the wave interactions, leading to different failure locations. Higher piston velocities were associated with increased spall strength by enhancing wave interactions and void formation, particularly at the interfaces and within the Ta layer, under specific configurations. Additionally, shorter shock pulse durations facilitated earlier initiation of the release fan, reducing twin formation and altering the failure dynamics by accelerating twin annihilation and pressure release.

36 MATERIALS SCIENCE

Impacts of PV Module Connector Failures on Cost and Performance of Utility Scale Photovoltaic Systems

The reliability, cost and performance of electrical connectors are a concern in all types of electrical systems, and demands on connectors used on photovoltaic (PV) systems include that connectors maintain electrical conductivity and physical strength, endure ultraviolet sunlight and high ambient temperature, and resist moisture and chemical intrusion over a very long (>25 year) performance period. Connector failures increase operation and maintenance (O&M) costs and reduce plant production, but connector failure can also cause safety and liability problems, which are of greater concern. This work results from a three-year collaboration between Sandia National Laboratories (SNL), the Electric Power Research Institute (EPRI), and the National Renewable Energy Laboratory (NREL) and funded by the U.S. Department of Energy (DOE) Solar Energy Technology Office (SETO) under Agreements #39035 and #38531 "Connector Reliability Across the US Solar Sector." a multi-pronged investigation of PV connector health across the US (see https://energy.sandia.gov/pvconnectors/). This report presents derivation of a Techno-Economic Analysis (TEA) that models failure modes and frequencies (how often failure occurs), estimates O&M costs and lost production associated with connector failures, and then calculates the effect that PV module connectors can have on Levelized Cost of Energy (LCOE). The model is informed with initial data from quantitative assessment of failure rates, root causes and mechanisms, in-situ diagnostics and data collection, lab-based forensics, and interviews with PV connector manufacturers and plant operators. SNL conducted site inspections at multiple utility-scale sites in different climates and subjected field samples of new, used, and degraded connectors to visual and electrical characterization. EPRI conducted metallurgical analysis of the pin and sleeve conductors to study failure-induced morphological and compositional changes. There is in general a shortage of statistically valid data, but data from PVROM database maintained by SNL was sufficient to ascertain failure rates and lost production as well as provide qualitative insight in its curated maintenance records. This report details the structure of the mathematical model but the sources of data to inform the model will continue to evolve. Analysis of a 100 MW PV plant is provided as an example of the use of the model, with results indicating that connectors are responsible for Annualized O&M Costs of $\$$71,933/year; Annualized Unit O&M Costs of $\$$0.72/kW/year; that a Reserve Account of $\$$187,220 should be available to fund repairs related to connectors; that connectors add $\$$1,494,004 to the Net Present Value of the O&M Costs (project life); and that O&M related to connectors adds about $\$$0.00088/kWh to the Levelized Cost of Energy. The impact of this model is to provide a tool to make the US solar sector more robust by quantifying and monetizing the reliability risks to utility-scale PV systems posed by poorly installed, mismatched and/or poorly designed and manufactured connectors. The TEA provides a model incorporating failure statistics, O&M cost data, and lost production into a single figure of merit, informing decisions and enabling practitioners to optimize cost and performance trade-offs. Stakeholders include connector manufacturers, system designers and equipment specifiers, standards bodies, installers and O&M providers, investors and insurance underwriters. This report supports continued growth of PV predicated on assurances that properly installed and maintained PV system connectors are safe and reliable. The project team is proposing future work including accelerated testing of connectors and expanding the approach taken here to other PV system components, such as TEA for rapid shut-down devices.

14 SOLAR ENERGY

Emergence of Diverse Failure Patterns in Weathering‐Induced Landslides: Insights From Particle Finite Element Simulations

Weathering is a fundamental driver of landslide evolution over geological timescales. Despite its ubiquity and importance, quantifying how weathering drives the progressive destabilization of rock slopes remains challenging. In this work, we develop a unified computational framework based on the particle finite element method to investigate the evolution of weathering‐induced landslides, from long‐term weathering to short‐term slope failure and runout dynamics. The framework integrates key processes, including weathering front propagation, time‐dependent strength degradation, rupture surface development, and post‐failure runout dynamics. Through numerical simulation experiments, we elucidate how interactions among weathering characteristics (type, intensity, and rate law), bedrock strength, fracture distribution, and slope geometry govern the failure modes and kinematics of weathering‐induced landslides. Simulations show that matrix‐dominated weathering leads to shallow translational failures, whereas fracture‐dominated weathering produces deep‐seated rotational and compound landslides. Pre‐existing fractures and slope morphology also strongly influence the movement of destabilized landmasses, affecting the failure pattern (e.g., kinematic mode and rupture surface geometry) and post‐failure behavior (e.g., runout velocity). We further demonstrate that the failure time and volume of weathered slopes are governed by the competition between gravitational driving forces and cohesive resisting forces during progressive destabilization. These findings provide new insights into the fundamental mechanisms that drive the emergence of diverse failure patterns of weathering‐induced landslides with important implications for landslide hazard assessment.

Wang, Liang [Eidgenoessische Technische Hochschule

Managing Marine Energy Risks for Project Success

This presentation reviews recommended practices for marine energy risk management based on NLR's recent risk management framework publication (https://www.nrel.gov/docs/fy24osti/90212.pdf). This presentation will include a demonstration of risk management processes and techniques that everyone in the marine energy industry can use to successfully meet their project objectives. This presentation will include a description of methods to identify and manage risks that are specific to marine energy, while demonstrating this through tools such as risk registers, failure modes effects and criticality analysis (FMECA), and more tools that are currently being developed. The goal of this presentation is for the participants to have knowledge and access to tools to help them manage the risks specific to their marine energy projects.

13 HYDRO ENERGY

Hazard and risk analysis framework for nuclear power plant–based integrated energy systems

Employing integrated energy systems (IESs) with nuclear power plants (NPPs) can improve NPP utilization by leveraging dedicated thermal and electric power delivery, but it may also increase operational safety risks. This paper presents a framework to identify and quantify hazards and risks for such IESs. The framework combines accidentology to review past industrial accidents with failure modes and effects analysis (FMEA) to identify potential future incidents. Hydrogen explosion and toxic chemical release hazards are of particular concern. Explosion consequences are quantified using the Bauwens-Dorofeev (Bauwens) and trinitrotoluene equivalent mass (TNT-EM) methods, while chemical release consequences are computed using the Gaussian atmospheric dispersion method. Operational disturbances from direct electrical and thermal integration that may affect NPP safety are modeled using probabilistic risk analysis (PRA). Hazards and risks are then evaluated for regulatory compliance. The framework is applied to IESs comprising pressurized or boiling water reactors supplying three levels of thermal and electrical power to industrial customers. Case studies include high-temperature steam electrolysis hydrogen plants of varying capacities and a synthetic fuel production plant. Sensitivity analysis examines piping component failures in the PRA model as a precursor to cost estimation for thermal extraction line design. Additionally, Fussel-Vessely (FV) and risk increase importance (RII) measures identify risk-informed design improvements for the thermal extraction system. FMEA highlights hazards such as loss of offsite power, prompt loss of electrical load, loss of thermal output, and immediate steam diversion, in addition to hydrogen explosions and toxic chemical releases. Both Bauwens and TNT-EM methods suggest maintaining several hundred meters of separation between the NPP and hydrogen facility to mitigate explosion risks. PRA results show a maximum initiating event frequency increase of 1.15% and an overall risk increase of 0.28%. Importance measure analysis identifies upstream pipe leak isolation components as critical. Evaluating the results against safety regulations, it is concluded that hazards and risks can be managed to comply with regulations through risk-informed thermal and electrical connection designs, component selection, maintenance programs, and safe separation distances between NPPs and integrated industrial facilities.

08 - HYDROGEN

Nuclear Reactor Heat Extraction for Synthetic Fuel Plants

This report looks at the viability and best approach when producing synthetic fuels using heat and power supplied by advanced nuclear energy systems. The report looks at a low temperature integration pathway with four types of advanced reactors: a pressurized water reactor (PWR), an advanced light water reactor (A-LWR), a sodium fast reactor (SFR), and a high-temperature gas-cooled reactor (HTGR). A failure modes and effects analysis (FMEA) of the coupling system between the nuclear plant and the synthetic fuel systems is performed with the goal of identifying the reliability of such a thermal delivery system. Furthermore, a high temperature pathway is investigated for synfuel coupling to determine if this is more efficient and more cost effective as a coupling approach.

10 - SYNTHETIC FUELS

Testing Recommendations for Thermal Failure Analysis

Minimizing ranges of uncertainty in failure QOIs is vital for making informed engineering decisions. To improve our decision-making ability, it is necessary to have a QOI representation that captures more complex failure modes. Capturing the physics of complex time-dependent temperature failure in explicit detail can be computationally and experimentally expensive. To avoid this blooming of complication and costs, an empirical QOI component-level failure/damage model may be sufficient to help identify component-level failure. In this context, failure and damage are abstractions of more complicated processes that lead to the cessation of normal operating behavior, not a representation of mechanical damage, such as void formation. Such a model requires minimal increases in computational cost and experimental burden. The proposed QOI serves as an indicator that something has likely failed, similar to the previous critical failure temperatures or pressures. This QOI will not replace any existing failure QOIs. It will be calculated alongside them, allowing for preservation of legacy style QOI measurements with new context from the new QOI.

36 MATERIALS SCIENCE

An Approach to Automate tools for the Risk Assessment of Digital Instrumentation and Control Systems

Reliable digital instrumentation and control systems (DI&C) are integral for sustaining the continued operation of nuclear power plants. These systems ensure that nuclear reactors operate safely, efficiently, and within regulatory requirements. Yet, the cost of designing and licensing new nuclear DI&C can be prohibitively expensive. Under the U.S. Department of Energy Light Water Reactor Sustainability Program, Idaho National Laboratory has developed a framework for supporting the risk-informed design of DI&C systems by offering methods to support the identification, quantification, and evaluation of risks for various DI&C design architectures. The framework indicates potential software failure modes and provides pathways for quantifying the potential for these software failures, including common cause failures. Using the framework’s systematic approach, challenges for assessing risks within new and existing nuclear DI&C systems can be reduced. Nevertheless, the current framework can be further improved using the convenience of automation. This paper introduces the development of Software for the Hazard Identification and Evaluation of Digital Systems (SHIELDS). SHIELDS is an engineering software package that enables the identification, elimination, and mitigation of potential risks and reduces the burden of deploying reliable DI&C systems. This work introduces plans and techniques to digitize and improve the manual risk assessment modules of the framework. These improvements will save time and increase the repeatability and usability of the framework, making it more accessible to a wider range of users. Ultimately, this introduces SHIELDS and how its modules support efficient development of safe and reliable DI&C systems.

46 - INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AN

Toward Drilling the Perfect Geothermal Well: An International Research Coordination Network for Geothermal Drilling Optimization Supported by Deep Machine Learning and Cloud Based Data Aggregation

The EDGE project, supported by the U.S. Department of Energy Geothermal Technologies Office under award DE-EE0008793, established a data-driven framework for improving the efficiency, cost-effectiveness, and reliability of geothermal well drilling. The project focused on developing scalable data infrastructure, advanced machine learning and probabilistic models, and integrated analytics tools to support continuous drilling optimization. A central objective was to reduce geothermal drilling costs by up to seventy percent while minimizing the risk of well failure through predictive diagnostics and adaptive planning. Over the project period, a comprehensive data repository was designed and deployed, incorporating records from over one hundred geothermal wells across varied geological settings. This repository supported both structured and unstructured data and adhered to FAIR data principles, enabling provenance tracking, quality control, and standardized metadata. The project introduced automated ingestion pipelines and a cloud-hosted platform that facilitated access to raw, processed, and derived datasets. This infrastructure served as the foundation for model development and analysis. Machine learning workflows were developed to predict key drilling metrics including rate of penetration, non-productive time, and total drilling costs. Self-organizing maps and dimensionality reduction methods were used to uncover operational patterns and outliers, while supervised learning algorithms such as random forests and deep neural networks were applied to forecast performance outcomes. The models were validated on heterogeneous datasets from both U.S. and Icelandic fields, demonstrating variable but significant predictive accuracy. The results indicated that finer temporal resolution, inclusion of lithological data, and consistency in operational annotations could substantially improve model performance. The project also implemented process mining techniques to reconstruct state-transition models from drilling event logs. These models enabled the identification of deviations from optimal workflows and provided insights into recurring failure modes. Analysis of non-productive time highlighted the impact of equipment failures, geological challenges, and human factors, offering opportunities for targeted mitigation strategies. The EDGE Dashboard was developed as a web-based expert system integrating data visualization, model outputs, and user-driven queries. It provided an accessible interface for operators to explore historical data, evaluate predicted outcomes, and compare drilling scenarios. Initial feedback from project partners suggested that the dashboard could serve as a foundation for more advanced advisory and optimization tools. Overall, the EDGE project demonstrated the feasibility and value of applying modern data science techniques to geothermal drilling. It delivered a set of interoperable tools and models that can support more efficient, lower-risk well development. The findings point toward a viable path for transitioning from advisory analytics to semi-autonomous drilling systems, contingent on continued collaboration, expanded datasets, and field validation. The project results have immediate relevance for drilling operations, data management practices, and future geothermal R&D efforts aimed at achieving reliable, cost-competitive geothermal energy at scale.

15 GEOTHERMAL ENERGY

Adaptive Cybersecurity for Distributed Energy Resources (AdCyDER): Online Reinforcement Learning with Stackelberg-Optimized Defenses — Pipeline Architecture, Evaluation Methodology, and Findings from a Synthetic-Data Evaluation

This report documents the design and evaluation of an integrated online-learning pipeline developed within the AdCyDER project for Distributed Energy Resource (DER) cybersecurity. The pipeline couples a Reinforcement Learning (RL) attack classifier — which produces an attack-type probability distribution — with a Stackelberg game-theoretic (GT) defense selector that consumes those distributions alongside SME-encoded priors over (defense, attack) effectiveness pairings and perdefense costs to choose grid-health-preserving defenses. The objective is not attack classification per se but production of distributions that drive effective defense selection through the Stackelberg layer, learned from delayed grid-health feedback rather than labeled attack data. AdCyDER as a whole is broader than the work presented here; this report covers the specific RL/GT loop integration and its evaluation. We present the integrated pipeline (SCADA telemetry with Fronius inverter physics, Suricata IDS, time-windowed aggregation, per-facility LSTM classifier, Stackelberg optimizer, OpenC2 actuators), an experimental campaign of 28 eight-hour iterations across three baseline modes, and a pipeline-ordered diagnostic protocol. The protocol identifies two distinct failure modes within the loop: paired supervised ceilings on the same features establish that the deployed online RL classifier (macro F1 ≈ 0.07) sits at least 4.7× below a same-architecture supervised LSTM (≈ 0.34) and 10–11× below a linear feature-signal ceiling (≈ 0.70–0.79 depending on per-facility isolation), localizing the dominant failure to the training procedure; and the reward signal driving online updates carries weak directional coupling with classifier correctness in the methodology-expected direction (multi-lens convergent: top-decile P(true) records produce more frequent state changes and slightly larger improvements, top-vs-bot Cohen’s 𝑑 ≈ −0.19), but at effect magnitudes too small to drive gradient-based learning at the campaign sample size. The original learning hypothesis is not supported by the data. The primary contributions are the diagnostic methodology — proposed as a transferable falsification protocol for online RL/GT defense pipelines learning from delayed environmental reward — and the open, reproducible experimental infrastructure. We outline reward reformulation as the highest-priority aspirational next step given the underpowered-but-aligned Q6 reading, with hardware-in-the-loop evaluation as the broadest scope-expansion option.

Blakely, Benjamin [Argonne National Laboratory (AN