Search NASA⌕ Search

SEARCH · Search NASA

Results for “REINFORCEMENT”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 127 records · Page 7

Incentivizing Cooperative Merging Control: Insights from Multi-Agent Deep Reinforcement Learning

Cooperative driving automation enables connected and automated vehicles (CAVs) to devise cooperative merging control, introducing great potentials to alleviate traffic congestion, reduce energy consumption, and enhance safety for highway on-ramp operations. Although numerous CAV cooperative merging algorithms have been developed to improve energy and traffic performance, the agreement-seeking among CAV users and their local benefits have been understudied. This can lead to rejections of cooperative merging plans and jeopardizing CAV performance, as a cooperation may entail certain CAVs to sacrifice their local benefits to achieve a system optimum. To address this issue, the study first leverages multi-agent deep reinforcement learning (MADRL) factoring both local reward and regional reward to demonstrate the discrepancies between CAV users’ local benefits and system optimum. Next, the existence of a correlated equilibrium is proved to characterize the convergence of MADRL training. This further facilitates the incorporation of incentives (computed based on reward discrepancies) to compensate for CAV users’ local benefits and facilitate system-optimal agreements in cooperative merging operations.

Zhou, Anye [ORNL] (ORCID:0000000301455579)↗

Harnessing biomass-derived reinforcements for sustainable flame-resistant composite

Bio-based composites containing natural fiber reinforcements have been integral to developing innovative, low-cost, and low-energy technologies for transportation applications in marine, automotive, and other industries. The low density of natural fibers gives these materials a high specific stiffness and strength, which can lead to significant weight savings in composites. This study focused on a specialty type of wood fiber treated with boric acid (≤ 7%) sourced from Northern White Pine in the USA, marketed as TimberFill. TimberFill is a high aspect ratio, fibrillated wood fiber that is produced commercially for wood fiber insulation; it is chemically treated to improve both its flame retardance and water resistance. In the present study, TimberFill was processed into non-woven mats using a process akin to papermaking and then infiltrated with epoxy resin. Biochar from wood waste was also incorporated into the epoxy resin as a low-cost filler and functional additive. The effects of biochar concentration on composites’ flame retardance (17% improvement in burn rate), thermal conductivity (improved by 9%), and mechanical performance (33% improvement in modulus) are herein presented. Additionally, it explores the plasticizing effect of the boric acid used in the treatment of TimberFill on the thermal and mechanical properties of the resulting composites, specifically on the thermal expansion coefficient.

Joshy, Joslin [ORNL]↗

Ambient-Pressure Chemical Recycling of Thermoset Carbon Fiber-Reinforced Polymers via Low-Temperature Solvolysis: Techno-Economic and Life-Cycle Assessment

Carbon fiber-reinforced polymer (CFRP) waste is rapidly accumulating, with global volumes projected to exceed 500,000 tons annually by 2050. Virgin carbon fiber production is extremely energy- and cost-intensive, and growing demand has intensified supply-chain vulnerabilities associated with precursor availability and reliance on critical materials. Scalable recycling technologies therefore offer not only environmental benefits but also opportunities to reduce material costs and strengthen domestic composite supply chains. This study introduces an efficient, mild chemical recycling method for epoxy-amine CFRPs, utilizing acetic acid (80 wt%) and zinc acetate (1 wt%) at atmospheric pressure and low temperature (~110 °C). The solvolysis process preserves residual resin (~40%) on recycled carbon fibers (rCF), enhancing interfacial bonding in remanufactured composite materials. Recycled fibers retain ~90% of their mechanical properties, and recovered oligomers show potential for reuse in epoxy formulations, supporting material circularity. Techno-economic analysis (TEA) and life-cycle assessment (LCA) confirm the scalability and sustainability of this process.

Recycling CFRP↗

Hierarchical Reinforcement Learning of a Short-Range Bond-Order Potential for Silica: Analytic Embedding of Coordination with Classical Efficiency

Reinforcement learning (RL) has recently emerged as a data-efficient strategy to parametrize short-range interatomic potentials. Building on our past RL optimization of pairwise silica models, we extend the framework to a bond-order (Tersoff-type) potential that provides an analytic embedding of local coordination through a three-body term. A hierarchical RL workflow combining continuous-action Monte Carlo Tree Search and property-based rewards efficiently explores the 26-dimensional parameter space, sequentially optimizing lattice parameters, densities, angles, and cohesive energies of 21 silica polymorphs. The resulting models, Q-Tersoff and ML-Tersoff, reproduce the energetic ordering of low-energy phases and capture the angular correlations and amorphous structure factors of silica with improved fidelity over pairwise force fields, while remaining orders of magnitude faster than high-dimensional machine-learned potentials. Both models underperform for elastic constants and high-energy frameworks, delineating the limits of the current analytic form. The approach establishes a general and interpretable route to angle-aware, short-range potentials that bridge physics-based and machine-learned descriptions of silicate materials.

36 MATERIALS SCIENCE↗

Ab Initio-Based Bond Order Potential for Arsenene Polymorphs Developed via Hierarchical Reinforcement Learning

Arsenene, a less-explored two-dimensional material, holds the potential for applications in wearable electronics, memory devices, and quantum systems. This study introduces a bond-order potential model with Tersoff formalism, the ML-Tersoff, which leverages multireward hierarchical reinforcement learning (RL), trained on an ab initio data set. This data set covers a spectrum of properties for arsenene polymorphs, enhancing our understanding of its mechanical and thermal behaviors without the complexities of traditional models requiring multiple parameter sets. Our RL strategy utilizes decision trees coupled with a hierarchical reward strategy to accelerate convergence in high-dimensional continuous search spaces. Unlike the Stillinger-Weber approach, which demands separate formalisms for buckled and puckered forms, the ML-Tersoff model concurrently captures multiple properties of the two polymorphs by effectively representing the local environment, thereby avoiding the need for different atomic types. Here, we apply the ML model to understand the mechanical and thermal properties of the arsenene polymorphs and nanostructures. We observe an inverse relationship between the critical strain and temperature in arsenene. Thermal conductivity calculations in nanosheets show good agreement with ab initio data, reflecting a decrease in thermal conductivity attributable to increased anharmonic effects at higher temperatures. We also apply the model to predict the thermal behavior of arsenene nanotubes.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Multiscale Mechanical Characterization of Mineral-Reinforced Wood Cell Walls

Studying the multiscale mechanics of bio-based composites offers unique perspectives on underlying structure–property relations. Cellular materials, such as wood, are highly organized, hierarchical assemblies of load-bearing structural elements that respond to mechanical stimuli at the microscopic, mesoscopic and macroscopic scale. In this study, we modified oak wood with nanocrystalline ferrihydrite, a widespread ferric oxyhydroxide mineral, and characterized the resulting mechanical properties of the composite at various levels of organization. Ferrihydrite nanoparticles were deposited inside the wood cell wall by an in situ chemical reaction, resulting in increased stiffness and hardness of the functionalized secondary cell wall, as evidenced by region-specific nanoindentation tests under an electron microscope. Chemically modified and pristine wood samples were characterized by using atomic force microscopy in the bimodal frequency modulation mode, which produced topographical images from the cellular ultrastructure with high lateral resolution and localized nanomechanical information across distinct cell wall layers. In conclusion, despite mineral reinforcement at the cell wall level, the macroscopic fracture behavior examined through three-point flexural testing remained unchanged upon modification, as cell–cell adhesion could be impaired by harsh chemical conditions.

Cells↗

Electricity use in big area additive manufacturing of fiber-reinforced polymer composites

In recent years, additive manufacturing (AM), especially large-format additive manufacturing (LFAM), has gained momentum in the manufacturing industry. While LFAM offers benefits over conventional manufacturing processes, such as minimizing material waste and providing vast geometric freedom, assessing its sustainability remains challenging due to limited data, particularly on energy consumption. Most existing data pertain to small-scale or desktop AM and are not directly applicable to LFAM. In this study, we conducted real-time measurements of electricity usage for a type of LFAM known as big area additive manufacturing (BAAM), which typically uses fiber-reinforced polymer pellets as feedstock. We collected electricity usage data from fifteen printing jobs over two months in an industrial production setting. These data fill the existing gap and can be reused to enhance the community’s understanding of LFAM electricity usage, support further research, and promote sustainable development in advanced manufacturing technologies.

ecology↗

Decentralised Reinforcement Learning for Dynamic Cyberattack Response in Microgrid Networks

Microgrids rely on communication networks for reliable operation, which makes them inherently vulnerable to cyberattacks. Such attacks can destabilise system dynamics and drive states away from their nominal operating trajectories. Although several physics-informed and machine learning-based strategies have been developed to counter these threats, the rapidly evolving cyber landscape enables adversaries to bypass static defences or rules-based mitigation approaches. This paper proposes a dynamic, online-trained and fully decentralised reinforcement learning (RL)-based cyberattack response framework to protect microgrids from evolving cyberattacks. The proposed framework deploys multiple deep Q-networks (DQNs), each associated with a distributed energy resource (DER), to enable localised and adaptive attack mitigation. In this framework, each DQN processes local voltage and frequency measurements—combined with intrusion detection system (IDS) alerts—as observations and rewards to guide decision-making. Extensive simulation studies demonstrate the robustness of the proposed framework under diverse attack scenarios and varying IDS-induced detection delays. Comparative analysis highlights its superiority over existing static or preexisting rules-based mitigation approaches. Finally, we present an analysis that shows the framework's scalability to real-life microgrids with more interacting agents.

24 POWER TRANSMISSION AND DISTRIBUTION↗

Reinforcement Learning‐Based Adaptation of Grid Following Inverter's Internal Controller to Networked Microgrids' Strengths

The varying topological configurations, generator commitments and dispatches, and dynamic load demand lead to changing system's strengths during the operations of networked microgrids. When the system's strengths significantly change, the fixed control gains at large devices may result in unsatisfactory system performance; this necessitates the tuning of the control gains at large devices to adapt to the changing system's strengths. In this paper, observer-based reinforcement learning (RL) is utilised to automatically tune the proportional-integral (PI) gains of phase lock loop (PLL) controller of grid-following (GFL) inverters to adapt to the changing strengths of microgrids and networked microgrids. The RL agent in this framework augments an observer predicting system's strengths, from which the RL control policy will adjust accordingly to tune the PLL controller's gains towards the system's strengths. Also, to enhance the control performance, the recently introduced Barrier function-based RL framework is leveraged for the design of reward function to prevent the high frequency nadir. An operational 26 kV electric distribution system, which is modelled as networked microgrids, is used to illustrate the need and effectiveness of the proposed RL-tuned control.

frequency response↗

Adaptive Reinforcement Learning (ARL) Control of a Multi-port Resonant Converter in UAV Systems

This study presents an adaptive reinforcement learning (ARL) control framework for a multi-port resonant converter used in hybrid unmanned aerial vehicle (UAV) power systems. The converter integrates high-frequency half-bridge input ports connected to a rectified engine–generator set and a battery energy storage system, along with a semi-bridgeless active rectifier supplying the propulsion load. A deep RL agent is trained to dynamically regulate inter-port phase-shift commands in real time based on flight conditions and load power demand. The ARL controller autonomously identifies phase-shift combinations that maximize conversion efficiency while maintaining stable and coordinated power flow, even under rapidly varying operating scenarios. This data-driven approach eliminates the need for explicit system modeling or extensive manual tuning and enables coordinated control among multiple power ports without inter-port communication. Experimental results validate that the ARL based strategy achieves reliable power sharing and consistently high-efficiency operation across diverse UAV operating conditions.

Asa, Erdem [ORNL] (ORCID:0000000190884812)↗

Reinforcement Learning Control for Buildings Co-Optimizing Energy, Comfort, and Indoor Air Quality: An Annual Assessment

Efficient control of Heating, Ventilation, and Air Conditioning (HVAC) systems is crucial for optimizing energy use and maintaining indoor comfort in buildings. Traditional control methods, such as PID control, cannot handle energy use trade-offs among multiple components in the building energy system at a supervisory level. Reinforcement learning (RL) presents a promising solution, offering adaptive and data-driven control strategies that optimize performance over time. However, RL also faces several challenges, including the conflicts encountered in co-optimizing energy savings, occupant comfort, and indoor air quality, and the requirement for extensive interactions with the environment in training. We proposed a flexible simulation platform that integrates a hybrid model for RL training and designed an RL agent to control the entire central HVAC system, focusing on co-optimizing energy consumption, thermal comfort, and indoor air quality ($\text{CO}_{2}$ and PM2.5 concentrations). Finally, we evaluated the RL agent's performance over an annual cycle. Our findings indicate that the RL agent can effectively manage the HVAC system with 14.7 % energy savings annually and balance multiple objectives, which demonstrates significant potential for improving HVAC system control and sustainability in buildings.

Guo, Fangzhou↗

Deep Multi-Agent Reinforcement Learning for Real-World Signalized Traffic Corridor Control

Signalized traffic control problem has been addressed recently with deep Reinforcement Learning (RL) approaches involving diverse state, action, and reward structures. While significant progress has been noted in the literature, open challenges still remain in the areas of adaptive signal phase timing, coordination in a multi-intersection corridor setting, and consideration of real-world traffic conditions. In the context of deep RL-based problem framing, extensions are needed that enable adaptive signal phase timings in an intersection agent's action space, computationally efficient information sharing among neighboring signalized intersection agents along a corridor, and experimentation in realistic simulation environments. In this paper, we develop a deep Advantage Actor Critic (A2C) multi-agent RL (MARL) approach capturing the research extensions above and apply it within a real-world calibrated Aimsun Next traffic corridor simulation model based on traffic data from the City of Coral Gables, Florida. For a multi-intersection corridor control setting, our numerical simulation experiments with a decentralized A2C MARL algorithm applied at different time periods led to a total average corridor travel delay reduction (expressed in seconds/mile averaged over vehicles) from 4.9% to 19.9% compared to state-of-the-art actuated control.

Shuvo, Salman S. [BATTELLE (PACIFIC NW LAB)]↗

Optimal Management of Grid-Interactive Efficient Buildings via Safe Reinforcement Learning

Reinforcement learning (RL)-based methods have achieved significant success in managing grid-interactive efficient buildings (GEBs). However, RL does not carry intrinsic guarantees of constraint satisfaction, which may lead to severe safety consequences. Besides, in GEB control applications, most existing safe RL approaches rely only on the regularisation parameters in neural networks or penalty of rewards, which often encounter challenges with parameter tuning and lead to catastrophic constraint violations. To provide enforced safety guarantees in controlling GEBs, this paper designs a physics-inspired safe RL method whose decision-making is enhanced through safe interaction with the environment. Different energy resources in GEBs are optimally managed to minimize energy costs and maximize customer comfort. The proposed approach can achieve strict constraint guarantees based on prior knowledge of a set of developed hard steady-state rules. Simulations on the optimal management of GEBs, including heating, ventilation, and air conditioning (HVAC), solar photovoltaics, and energy storage systems, demonstrate the effectiveness of the proposed approach.

Huo, Xiang↗

Adaptive Reinforcement Learning Control for Power Distribution in Multi-Output Resonant Converters

This paper presents an adaptive reinforcement learning (ARL)-based control framework for efficient power distribution in a multi-output resonant converter for UAV applications. The proposed system is based on a high-frequency isolated resonant architecture, where a single energy source supplies multiple propulsion loads through independently controlled output rectifiers, addressing the need for coordinated multi-motor power management. The ARL framework dynamically allocates output power by learning optimal phase-shift control actions under varying load demands and operating conditions. The agent autonomously determines control parameters that maximize conversion efficiency while ensuring accurate power sharing among multiple outputs. In addition, the proposed approach enables adaptive operation without requiring detailed system modeling or manual tuning. Experimental results demonstrate stable and efficient performance over a wide range of operating conditions, confirming the effectiveness and robustness of the learning-based control strategy for multi-output resonant converter system.

Asa, Erdem [ORNL] (ORCID:0000000190884812)↗

Optimal Coordination of Electric Vehicles for Grid Services using Deep Reinforcement Learning

Recent research has shown the effectiveness of reinforcement learning (RL) in coordinating electric vehicles (EVs) with vehicle-to-grid capabilities for grid services. However, many of these studies rely on lookup table and deep Q-network techniques, which can be impractical when dealing with continuous states and actions. In addition, existing RL designs inadequately account for battery aging effects, EV user satisfaction, uncertain departure and arrival time, and trip distance, which may compromise effective coordination. This paper aims to bridge these gaps by developing an innovative deep deterministic policy gradient-based RL framework for optimal coordination of EVs. Case studies were carried out using a test system with 100 EVs, and numerical analysis results showed that the proposed RL framework can effectively coordinate EVs to maximize economic benefits and user satisfaction while ensuring the expected battery lifespan.

Das, Avijit↗

Reinforcement Learning-Based Approach for EMT Automation of Large-Scale PV Plants

In the pursuit of efficient and precise modeling of large-scale power systems, particularly utility-scale photovoltaic (PV) plants, Electromagnetic Transient (EMT) simulations play a crucial role. As utility-scale PV plants increase in size and complexity, traditional computational methods become inadequate, necessitating more advanced techniques. This paper highlights the progressive efforts made to accelerate EMT simulations. A novel continuous reinforcement learning (RL) strategy is explored to automate the differentiation and categorization of stiff and non-stiff differential algebraic equations (DAEs). The use of stiff and non-stiff integration methods applied to relevant parts of the DAEs assists with the speed-up of the simulations. The paper details the data acquisition, development and offline training of the RL model, leading to its validation that demonstrates a high precision in optimizing simulation methods. The proposed RL promises to significantly enhance the efficacy of EMT simulations, offering a robust framework for the future of power system analysis.

Xia, Qianxue↗

A Sequential Model Predictive and Deep Reinforcement Learning-Based Controller for Distribution System Outage Mitigation under Hurricane Events

This paper proposes a proactive outage mitigation framework for power distribution networks to withstand hurricane-induced disruptions. It leverages Model Predictive Control (MPC) to identify safe lines for proactive switching during hurricanes, minimizing the risk of cascading failures and voltage violations. The switching strategies optimized by MPC are sequentially integrated with a Deep Reinforcement Learning agent using the Advantage Actor-Critic algorithm, enabling dynamic line switching to maximize connected buses and minimize voltage violations in real time. Using a probabilistic hurricane model, the framework predicts line failures and adapts to varying conditions to enhance grid resilience. Simulations on the IEEE 123-bus system demonstrate its effectiveness in maintaining high connectivity and minimizing disruptions. Real-time testing with an RTDS confirms the practicality and reliability of the proposed approach.

Selim, Alaa [University of Connecticut]↗