Search NASA⌕ Search

SEARCH · Search NASA

Results for “REINFORCEMENT”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 145 records · Page 8

Reinforcement Learning-Based Secondary Control Strategy for Voltage and Frequency Regulation in Islanded Inverter-Based Microgrids

This paper presents a reinforcement learning (RL) approach for secondary voltage and frequency control in islanded inverter-based microgrids. The proposed control strategy aims to restore voltage and frequency deviations caused by the primary droop control while ensuring proper power sharing between distributed generators. The RL agent is designed to provide correction signals to the primary control, considering communication delays and system constraints. The effectiveness of the proposed control strategy is validated through simulation results in MATLAB/Simulink environment, demonstrating superior performance in maintaining voltage and frequency within the nominal values.

Rodriguez Martinez, Omar Felipe [University of Pue↗

Topology-Aware Reinforcement Learning for Voltage Control: Centralized and Decentralized Strategies

Volt-VAR control (VVC) methods based on deep reinforcement learning (DRL) can effectively control distribution grid voltage and minimize power loss by implementing corrective and preventive control measures on the reactive power output of inverter-based distributed energy resources (DERs). However, model-free DRL-based VVC approaches usually cannot capture the important topological feature of the power system since they use a fully-connected network (FCN) to deliver the action. Therefore, this paper proposes a graph convolutional network (GCN)-based DRL approach that can employ the topological information of the network to take better control action for regulating the voltage. Our implementation allows for both centralized and decentralized configurations, utilizing a single agent and multiple agents respectively. Although the centralized GCN-based DRL approach has its advantages of minimizing voltage fluctuation and power loss, it is not suitable for large scale power systems due to its challenges in terms of scalability, computation speed and potential single points of failure. Therefore, these problems can be resolved using the decentralized GCN-based DRL approach. Moreover, to ensure the safe operation of the model, our proposed approach incorporates an exponential barrier function while formulating the reward function for each agent. To validate performance of the proposed approaches, the proposed model is tested on modified IEEE test systems and the performances are measured in terms on voltage fluctuation reduction, minimization of power loss and computational speed. Finally, the results show that the proposed topology-aware approach outperforms the FCN-based DRL approach in terms of reducing voltage fluctuation and minimizing power loss of the network. Moreover, it is shown that the decentralized GCN-based DRL has faster computational speed than other approaches.

42 ENGINEERING↗

Safe Deep Reinforcement Learning for Robust Frequency and Voltage-Constrained Networked Microgrid Restoration

Here, this paper proposes a safe soft actor-critic reinforcement learning (RL) algorithm–based controller for networked microgrid restoration. It formulates the post black-start start as a finite-horizon constrained Markov decision process. The RL agent co-optimizes real and reactive power set-points for both grid-forming and grid-following inverters under explicit voltage and frequency constraints, while enforcing proper power sharing via the Mean Active Power Sharing Index (MPSI) and Mean Reactive Power Sharing Index (MQSI). Numerical results obtained on the IEEE 123-bus distribution system show that the proposed method achieves a mean voltage build-up time of 0.01 s without breaching the 5% sharing-violation budget under various load scenarios, considering MPSI and MQSI indices. These findings demonstrate that the proposed method yields fast and safe black-start schedules without resorting to heuristic penalties.

Selim, Alaa [Dartmouth College, Hanover, NH (Unite↗

Deep Reinforcement Learning-Based Control of Energy Storage for Interarea Oscillation Damping

With the increasing electricity consumption and lack of transmission investment, today's power systems are operated much closer to their limits, raising concerns of inter-area oscillations that deteriorate the system stability. Here, this article presents a novel energy storage placement and control approach for enhanced damping of interarea oscillations. Combining the residual analysis and dominant mode analysis, we are able to identify the advantageous locations for placing energy storage that achieve improved damping performance. To overcome the challenges, such as fixed control parameters and insufficient damping, we propose to use a deep reinforcement learning-based approach for energy storage control. A state-of-the-art guided surrogate-gradient-based evolutionary strategy is used to train a learning agent in a robust, efficient, and reproducible manner. Parallel computing is also adopted to speed up the training process. The proposed strategy has been tested on both medium and large-scale systems. The proposed methods have demonstrated their effectiveness in mitigating various interarea oscillations within a timeframe of 20 s, thereby averting system collapse and enhancing power grid stability effectively.

25 ENERGY STORAGE↗

CyRRL (Cyber Resilient Reinforcement Learning for grid voltage control) [SWR-24-115]

This codebase contains a multi-agent, actor-critic reinforcement learning implementation for cyber-resilient grid voltage control. It uses a 123-bus OpenDSS system as the environment, with three-phase power flow translating nodal power injections into solved nodal voltages. The reward function penalizes deviations from nominal voltage as well as reactive power dispatch, while encouraging agents to take actions that result in fast convergence to nominal conditions. The codebase models false data injection attacks and includes functionality for training, testing, hyper-parameter tuning, and visualization.

Murphy, Sinnott [National Renewable Energy Laborat↗

Vertical z-axis discontinuous carbon fibers for improved lightning strike performance of continuous fiber-reinforced polymer composites

Effective lightning strike protection for critical aerospace and wind applications requires high electrical conductivity to dissipate current efficiently. However, polymer matrix composites face a challenge due to their inherently insulating nature. While conventional carbon fiber-reinforced composites (CFRP) exhibit electrical conductivity in the planar direction, achieving through-thickness conductivity remains an ongoing challenge. In this work, we have undertaken the fabrication of CFRP interleaved with vertically oriented carbon fibers (Z-fiber) to impart higher electrical conductivity along the thickness direction. Two Z-fiber composite variations are prepared: Z-1 with a single layer of Z-fiber and Z-5 with five interleaved layers and compared with no Z-fiber layer (Z-0) composite. The composite panels were subjected to lab-scale lightning strike tests with a current magnitude of 100 kA. To emulate real-world service conditions, an aerospace-grade paint coating was applied to the composite laminates. Comparative analysis shows Z-1 reduces damage diameter to ∼22 mm compared to Z-0 (∼26 mm), while Z-5 exhibits the least damage (∼16.7 mm), confirmed by optical microscopy. Z-5 demonstrates nine times higher through-thickness electrical conductivity than Z-0, reducing electrical anisotropy substantially. Thermal-electric finite element damage modeling predicts surface damage within 6% of experimental values for both Z-0 and Z-5 composites. Flexural tests post-lightning reveal Z-5 retains 66% flexural strength and 86% modulus, significantly better than Z-0, which retains less than 40% for both properties. This study highlights the efficacy of Z-fiber composites in lightning strike protection, offering improved through-thickness conductivity and mechanical property retention.

36 MATERIALS SCIENCE↗

ReLIC: Full-Scale Realization of Reinforcement Learning for Infrastructure Control

Prior efforts have shown that deep reinforcement learning (DRL) may provide a new method for controlling networked power systems. Though successful, prior approaches have not yet demonstrated their behavior on systems of realistic scale. This effort examined multiple theoretical and technical approaches to allow a DRL model to operate over a system of 2,000 buses or more. We find that allowing the DRL models to run training episodes in parallel provides near limitless efficiency gains, allowing us to train successful agents to behave on our Kuramoto transmission model of up to 4,000 buses. We further show that we can expand our PowerWorld DRL implementation to systems of up to 25 buses but struggle to go beyond this limit due to PowerWorld’s inability to run multiple instances at once. Finally, we examine a multi-agent approach and find that it performs as well if not better than our existing centralized approach.

97 MATHEMATICS AND COMPUTING↗

Reducing Mass of Steel Auto Bodies using Thin Advanced High Strength Steel with Carbon-Fiber Reinforced Epoxy

Diversitak, a company based in Detroit, MI, has developed a proprietary, low specific gravity, carbon fiber reinforced epoxy (CFRE) under U.S. patent number 9,963,58832. Preliminary testing on this new material conducted in collaboration with ArcelorMittal Steel Company proved out the CFRE concept. A thin layer of this CFRE was applied to a stamped sheet of steel with residual stamping oils from a mill, in a time corresponding to automotive processing (e.g., ~15 seconds), and processed following automotive e-coat procedures (phosphating + 175–200°C heating), to complete the curing. No problems with adherence or performance were noted. While the CFRE does add weight to a thin gauge steel panel, it weighs much less than what is displaced by using thicker conventional mild steel gauges. The application of the coating showed a significant increased dent resistance, oil canning resistance, and part stiffness.

36 MATERIALS SCIENCE↗

A Novel Manufacturing Process of Lightweight Automotive Seats (Integration of Additive Manufacturing and Reinforced Polymer Composite)

Lightweight automotive seats offer multiple benefits to original equipment manufacturers in terms of cost savings from various aspects, including less material usage, more integrated processes, and compliance with Corporate Average Fuel Economy Standards. Original equipment manufacturers have been focusing on innovative ways to produce light weight automotive seats. The commercially available automotive seats are currently made of multiple metal components combined through welding and fasteners. The use of additive manufacturing and composite structures is particularly useful for light weighting the automotive components. Additive manufacturing (AM) offers multiple advantages over traditional manufacturing processes such as freedom of design thereby enabling complex structural geometries, mass customization and waste minimization, and control over the fiber alignment through deposition in a predetermined pattern. Combining metal inserts with polymer composites through a novel manufacturing process allows design of lightweight and high-performance materials for automotive components. However, fabricating these metal polymer composite structures through traditional manufacturing processes limits their mechanical properties due to limited design freedom, lack of control over fiber orientation in composite parts, and poor interfacial bonding between the constituent materials. It is essential to develop a novel manufacturing process to enable high throughput production of lightweight automotive seats using metal and polymer composites. As such it is important to design the automotive seat suitable for manufacturing via this process and perform mechanical characterization on various subcomponents of the seat to ensure that the design and performance requirements provided by the auto manufacturer are met. The aim of this project is to develop a novel manufacturing technique to produce lightweight automotive seat by combining AM with conventional manufacturing processes. The car seat back panel will be designed via topology optimization and numerical simulations to minimize the overall weight while ensuring it meets all the performance requirements. The optimization of the seat back structure will be based on computational stress analysis to maximize the stiffness and minimize the weight. Materials currently used by Ford Motor Company will be adopted for a few subcomponents while the in-house composite materials will be used for the rest of the seat back. The composite and metallic materials will be tested to determine their mechanical properties as these are necessary for simulations. A novel manufacturing process will be developed to integrate AM metal inserts with discontinuous reinforced composite through large scale additive manufacturing and compression overmolding processes. The developed manufacturing technique will be used to fabricated various subcomponents suitable for the seat back design and mechanically tested to determine their properties. The manufacturing of the lightweight seat back design through this process involves integrated AM metal inserts with the composite structure for recliner connection. The manufacturing of the entire seat back which is lightweight through the novel manufacturing process will be discussed. The performance of the designed seat back will be investigated through numerical simulations and shown to meet all the requirements provided by the auto manufacturer. The final goal of developing a novel manufacturing process for lightweight automotive seats is met through design optimization of seat back, manufacturing of subcomponents, mechanical characterization, and validation through numerical simulations. The routes to achieve the final goal of the project and the depth in which they were investigated changed throughout the project due to personnel changes and the COVID-19 pandemic. The project resulted in the development of a novel manufacturing process to integrate metal inserts with tailored polymer composite preforms through overmolding. Leveraging this proven manufacturing process, a lightweight seat back was designed through topology optimization and numerical simulations. The designed seat back uses AM metal inserts and compression overmolding of tailored polymer composite preforms obtained via large scale additive manufacturing. The metal polymer composite structures fabricated through this process exhibited enhancement in stiffness and improved ductility upon testing. Overall, the project provided an alternative design and manufacturing technique for automotive seat back that enables weight saving while meeting the safety and performance requirements.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗

Nondestructive Evaluation of Carbon Fiber Reinforced Polymers

The American Society of Mechanical Engineers Boiler and Pressure Vessel Code requires repair and replacement of safety-related piping materials to meet the original Construction Code; however, no construction criteria currently exist for carbon fiber reinforced polymer (CFRP) materials in the nuclear industry. Although nondestructive examination (NDE) techniques for cast and wrought ferritic and austenitic steels are well established, licensees are increasingly deploying novel materials such as CFRP for which assessment of the use of NDE is required, as the inspectability of such materials and the influence of manufacturing processes on inspectability remain insufficiently understood. To address these gaps, Pacific Northwest National Laboratory (PNNL) evaluated the fabrication of several CFRP repair mockups and the effectiveness of NDE methods to support regulatory review of CFRP repairs in nuclear power plant applications. Representative flat-plate mockups containing realistic fabrication defects were manufactured and inspected using manual and automated tap testing, dynamic response spectroscopy (DRS), and ultrasonic testing (UT). The study identified significant fabrication variability, particularly in controlling defect size and epoxy thickness, which strongly influenced defect detectability by NDE methods. Tap testing was effective for shallow defects in thin epoxy configurations but unreliable for thicker repairs and deeper flaws. DRS demonstrated higher sensitivity to dry spots but produced unconfirmed indications for thicker plates, requiring further validation. Conventional UT, particularly at 1.0 MHz, showed the strongest overall capability for detecting a range of defects, although performance degraded with increasing repair thickness and complexity. The results highlight the need to better define critical defect sizes, develop reliable mockup fabrication methods, and validate NDE methods needed to support CFRP repairs.

36 MATERIALS SCIENCE↗

Adaptive Cybersecurity for Distributed Energy Resources (AdCyDER): Online Reinforcement Learning with Stackelberg-Optimized Defenses — Pipeline Architecture, Evaluation Methodology, and Findings from a Synthetic-Data Evaluation

This report documents the design and evaluation of an integrated online-learning pipeline developed within the AdCyDER project for Distributed Energy Resource (DER) cybersecurity. The pipeline couples a Reinforcement Learning (RL) attack classifier — which produces an attack-type probability distribution — with a Stackelberg game-theoretic (GT) defense selector that consumes those distributions alongside SME-encoded priors over (defense, attack) effectiveness pairings and perdefense costs to choose grid-health-preserving defenses. The objective is not attack classification per se but production of distributions that drive effective defense selection through the Stackelberg layer, learned from delayed grid-health feedback rather than labeled attack data. AdCyDER as a whole is broader than the work presented here; this report covers the specific RL/GT loop integration and its evaluation. We present the integrated pipeline (SCADA telemetry with Fronius inverter physics, Suricata IDS, time-windowed aggregation, per-facility LSTM classifier, Stackelberg optimizer, OpenC2 actuators), an experimental campaign of 28 eight-hour iterations across three baseline modes, and a pipeline-ordered diagnostic protocol. The protocol identifies two distinct failure modes within the loop: paired supervised ceilings on the same features establish that the deployed online RL classifier (macro F1 ≈ 0.07) sits at least 4.7× below a same-architecture supervised LSTM (≈ 0.34) and 10–11× below a linear feature-signal ceiling (≈ 0.70–0.79 depending on per-facility isolation), localizing the dominant failure to the training procedure; and the reward signal driving online updates carries weak directional coupling with classifier correctness in the methodology-expected direction (multi-lens convergent: top-decile P(true) records produce more frequent state changes and slightly larger improvements, top-vs-bot Cohen’s 𝑑 ≈ −0.19), but at effect magnitudes too small to drive gradient-based learning at the campaign sample size. The original learning hypothesis is not supported by the data. The primary contributions are the diagnostic methodology — proposed as a transferable falsification protocol for online RL/GT defense pipelines learning from delayed environmental reward — and the open, reproducible experimental infrastructure. We outline reward reformulation as the highest-priority aspirational next step given the underpowered-but-aligned Q6 reading, with hardware-in-the-loop evaluation as the broadest scope-expansion option.

Blakely, Benjamin [Argonne National Laboratory (AN↗

Reinforcement Learning Control for Enhancing Marine Hydrokinetic Turbine Energy Generation

This paper proposes a reinforcement learning-based method to maximize power generation for a direct-drive marine hydrokinetic turbine. A high levelized cost of energy (LCOE) is preventative in the widespread adoption of many marine energy conversion technologies. A straightforward way to reduce LCOE is to increase conversion efficiency and ensure maximum energy generation. The proposed method utilizes a damping control methodology, varying applied generator torque via a linear relationship between the applied damping coefficient and rotor speed. A state-action-reward-state-action (SARSA) algorithm has been used to learn the optimal control action for a given flow velocity. The proposed SARSA methodology uses Gaussian radial basis functions to create a three-dimensional surface to estimate the relationship between damping coefficient, incoming flow velocity, and coefficient of power (C p ). Here, the SARSA algorithm was compared against a baseline optimal tip speed ratio controller over a year-long flow velocity case profile while considering the effects of biofouling on the turbine system, where the proposed RL method generated 0.92% more energy than the baseline.

Damp↗

A User-Friendly GUI Tool for Automated Microstructural Analysis of Fiber-Reinforced Composites and Porous Structures

Understanding and quantifying microstructural features such as fiber orientation and porosity is critical for predicting the mechanical behavior and performance of fiber-reinforced polymer composites. Traditional manual analysis is time-consuming, subjective, and unsuitable for high-throughput datasets. We present a graphical user interface (GUI) application that automates the analysis of microscopy images to extract key microstructural metrics, including fiber orientation tensors, fiber orientation distribution, porosity and pore size distribution. The app integrates multiple image segmentation techniques including global and local thresholding, clustering, and region-based approaches, offering flexibility for different types of image qualities and features. Users can load microstructural images, select regions of interest and segmentation techniques tailored to their image dataset. It also addresses a critical challenge in fiber orientation analysis: the ambiguities caused by touching, overlapping, or partially cut fibers. It supports autorun examples for standardized workflows, enabling reproducible analysis and facilitating training and benchmarking. This tool significantly reduces manual intervention, enhances consistency, and accelerates data generation for structure–property modeling, process optimization, and digital materials research. The tool is intended for use by materials scientists, engineers, and researchers engaged in composite characterization, quality control, and machine learning-based microstructural studies.

Chawla, Komal [ORNL] (ORCID:0000000190327565)↗

Enhancement of the Physical and Mechanical Properties of Cellulose Nanofibril-Reinforced Lignocellulosic Foams for Packaging and Building Applications

Biobased foams have the potential to serve as eco-friendly alternatives to petroleum-based foams, provided they achieve comparable thermomechanical and physical properties. We propose a facile approach to fabricate eco-friendly cellulose nanofibril (CNF)-reinforced thermomechanical pulp (TMP) fiber-based foams via an oven-drying process with thermal conductivity as low as 0.036 W/(m·K) at a 34.4 kg/m3 density. Acrodur®, iron chloride (FeCl3), and cationic polyacrylamide (CPAM) were used to improve the foam properties. Acrodur® did not have any significant effect on the foamability and density of the foams. Mechanical, thermal, cushioning, and water absorption properties of the foams were dependent on the density and interactions of the additives with the fibers. Due to their high density, foams with CPAM and FeCl3 at a 1% additive dosage had significantly higher compressive properties at the expense of slightly higher thermal conductivity. There was slight increase in compressive properties with the addition of Acrodur®. All additives improved the water stability of the foams, rendering them stable even after 24 h of water absorption.

Chemistry↗

GEPA: Reflective Prompt Evolution Can Outperform Reinforcement Learning

Large language models (LLMs) are increasingly adapted to downstream tasks via reinforcement learning (RL) methods like Group Relative Policy Optimization (GRPO), which often require thousands of rollouts to learn new tasks. We argue that the interpretable nature of language often provides a much richer learning medium for LLMs, compared to policy gradients derived from sparse, scalar rewards. To test this, we introduce GEPA (Genetic-Pareto), a prompt optimizer that thoroughly incorporates natural language reflection to learn high-level rules from trial and error. Given any AI system containing one or more LLM prompts, GEPA samples trajectories (e.g., reasoning, tool calls, and tool outputs) and reflects on them in natural language to diagnose problems, propose and test prompt updates, and combine complementary lessons from the Pareto frontier of its own attempts. As a result of GEPA's design, it can often turn even just a few rollouts into a large quality gain. Across six tasks, GEPA outperforms GRPO by 6% on average and by up to 20%, while using up to 35x fewer rollouts. GEPA also outperforms the leading prompt optimizer, MIPROv2, by over 10% (e.g., +12% accuracy on AIME-2025), and demonstrates promising results as an inference-time search strategy for code optimization. We release our code at https://github.com/gepa-ai/gepa.

97 MATHEMATICS AND COMPUTING↗

MSD CoP Webinar: Modeling the Operations of Reservoir Systems with LLMs and Inverse Reinforcement Learning

Context: This panel featured three presentations centered on the common theme of applying LLMs and inverse reinforcement learning (IRL) to capture the complex human-environment interactions that are central to the operation of reservoir systems. Dr. Wyatt Arnold will kick off the webinar with a talk on how analyzing LLM chain-of-thought reasoning reveals sophisticated quantitative justification and risk awareness, showing promise as a bridge between quantitative models and value-driven water management decisions. Next, Dr. Matteo Giuliani will build on this with a discussion demonstrating that AI- and IRL-driven approaches can infer the trade-offs between flood control and water supply using historical observations. Finally, Dr. Rohan Singh Wilkho will close the webinar with a talk establishing IRL as a generalizable diagnostic tool for decoding decision-making in managed hydrologic and human-infrastructure systems. Across the three presentations, the application of LLMs and IRL opens new possibilities for the development of adaptive, transparent, and human-aware models supporting water management in an increasingly uncertain future. Presenters: Wyatt Arnold (Politecnico di Milano); Matteo Giuliani (Politecnico di Milano); Rohan Singh Wilkho (Cornell University) Moderator: Patrick M. Reed (MSD CoP Facilitation Team); Stefano Galelli (MSD CoP AI Working Group Co-Chair); David Gold (MSD CoP AI Working Group Co-Chair) This webinar was held on: June 23rd, 2026 from 12-1 PM EST.

Arnold, Wyatt [Politecnico di Milano]↗

Reinforcement Learning for Energy-Movement Optimization in Arduino-Based Robotics

Our goal is to develop a learning method for an Arduino-based robotthat maximizes travel distance and minimizes energy expenditure. • Will implement State ActionReward State Action (SARSA) reinforcement learning algorithm • Learning steps informed by state of environment • Rewards good decisions and punishes bad ones.

Gilmore, Blake↗

Reinforcement Learning for Distance Maximization in Energy-Constrained Robots

Our goal is to teach a robot to walk in an energy efficient way. Inspired by nature [1], we create a reinforcement learning method that balances distance traveled and energy consumption, which we call “utility.” Further, each step is symmetric to ensure equal strides while walking straight. We achieve our goal by rewarding decisions that lead to the best utility over several training iterations.

Pereira, Luiz Manella↗