Search NASA⌕ Search

SEARCH · Search NASA

Results for “REINFORCEMENT”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 109 records · Page 6

Safe Reinforcement Learning-Based Transient Stability Control for Islanded Microgrids With Topology Reconfiguration

This paper proposes a safe reinforcement learning (RL)-based transient stability emergency control (TSEC) method for islanded microgrids. RL requires extensive interaction with the environment to learn control strategies, hence, a data-driven approach is used as a substitute for time-consuming time-domain simulation calculations. Deep sigma point processes (DSPP), which is a Gaussian process model, is utilized to predict the normal distribution of transient stability of microgrids and to construct a transient stability chance constraint. Reward-constrained policy optimization (RCPO) can simultaneously achieve objective prediction, policy learning, and constraint cost coefficient update across multiple timescales. RCPO interacts with the DSPP-based microgrid environment through a multi-process parallel manner, greatly increasing the training speed. Case studies on a real islanded microgrid demonstrate that the proposed method can efficiently and quickly obtain the optimal emergency control strategy while adhering to all hard constraints.

14 SOLAR ENERGY↗

Multi-Agent Hierarchical Deep Reinforcement Learning for HVAC Control With Flexible DERs

As electricity consumption in commercial and residential buildings continues to rise, reducing energy costs presents an increasing challenge. Heating, ventilating, and air-conditioning (HVAC) systems, which typically account for 40%-50% of a building's energy use, are prime targets for energy savings. Intelligent control of HVAC temperature through the exploitation of HVAC load flexibility brings significant potential to reduce energy consumption and electricity expenses. The nonlinear models of HVAC systems challenge traditional control methods, while the uncertainty introduced by HVAC load flexibility complicates distributed energy resource (DER) management using conventional optimal dispatch techniques. In response to these challenges, we propose a hierarchical multi-agent deep reinforcement learning (DRL) approach. The lower-level agents focus on balancing comfort and energy conservation, while the upper-level DRL agents optimize the use of DERs to reduce peak demand based on the control outcomes of the HVAC by the lower-level agents. Here, in the upper-level agents, we incorporate a multi-agent structure based on ensemble learning, which acts based on historical and current data without relying on precise load forecasting to address the delayed rewarding issue in DRL. This allows for the effective reduction of energy costs. The proposed method is tested using a real-world microgrid comprising 413 buildings in Southern California, and the results demonstrate that our approach can significantly reduce overall electricity bills while ensuring the comfort of consumers and residents.

24 POWER TRANSMISSION AND DISTRIBUTION↗

A General Overview of Cathodic Protection in Reinforced Concrete

Corrosion of reinforcement steel in concrete is a prevalent issue in infrastructures worldwide, with the direct costs of repair estimated to be $1 trillion annually. Treatment options are complicated by the concrete’s role in the corrosion process. Ordinary, modern-day cement is so alkaline that it maintains a thick passive layer on the steel surface, reducing corrosion to inconsequential rates. However, environmental contaminants in the form of carbon dioxide or chlorides may cause a breakdown in the passive layer, exposing the steel surface to corrosion. Galvanic cells will be free to form due to potential differences along the rebar, and anode sites will create dissolved iron ions that cannot travel far through the still concrete. Thus, rust products will accumulate directly onto the steel/concrete interface. Iron oxides have a much higher volume relative to steel, and even a miniscule amount of rust can produce enough volumetric stress to cause cracking, spalling and delamination of the surrounding concrete, increasing the risk of structural failure. Safeguards and inhibiting technology exist to limit the chances of corrosion initiation by either species, but cracking of the concrete cover and gradual accumulation of contaminants means that corrosion is unavoidable in certain environments and will initiate, given enough time.

36 MATERIALS SCIENCE↗

Technical Report for Bayesian Optimization and Reinforcement Learning for Beam Polarization Increase in the BNL Hadron Injectors

This project developed and evaluated physics-informed Bayesian learning and machine learning (ML)-based optimization methods for improving beam polarization preservation in the BNL hadron injector chain. The work focused on uncertainty-aware digital twin modeling, Bayesian calibration of accelerator simulations using beam measurements, and data-efficient optimization strategies including Bayesian optimization and reinforcement learning. These methods were applied to injector tuning and RF control problems in realistic accelerator settings to support improved operational robustness and readiness for RHIC operations and future Electron–Ion Collider facilities. No subject inventions were disclosed under this award.

43 PARTICLE ACCELERATORS↗

Reinforcement Learning to Enhance Optimal Operation of Resilient Community Energy Systems

This paper presents a novel model-free multi-agent Reinforcement Learning (RL) control method to enhance the resilience of community energy systems in island mode, which coordinates multiple objectives without the necessity of identifying system models that require expert knowledge. Specifically, a community-level coordinator agent is designed to allocate renewable energy resources among different buildings, and multiple building-level agents are developed to optimize load schedules based on limited energy resources and requirements of building loads and occupants’ comfort. In a two-day evaluation, our RL approach demonstrated a similar performance against MPC without requiring system models and formulation of optimization problems as required in MPC.

ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATION↗

Discovering the Most Severe K-Point Failure Based on Reinforcement Learning: Preprint

Smart devices are essential to ensure the stability of the power grid and resilience to intermittent energy production. However, smart devices can also be the target of cyber adversaries that may exploit false data injection attacks (FDIAs) to induce unstable grid conditions. A practical consideration of FDIA mitigation approaches is addressed here: given a finite available budget, for which smart device should cyber-threat mitigation be deployed first? In this work, this question is answered by identifying the so-called most-sensitive devices, i.e., the devices that, if compromised, can let an adversary induce the most serious grid instabilities. The method proposed utilizes an adversarial reinforcement learning (RL) framework to identify the k-mostsensitive smart devices (here, smart inverters). The adversarial agent can tamper with the compromised inverters' active and reactive operating power setup points, with the goal of maximizing voltage deviations. Numerical results show that the proposed RL method finds the optimal attack scenarios for 1-point failure and the near-optimal solution for the 2-point case. Additionally, the proposed RL method achieves an 8.8 speed-up ratio in running time compared to the brute force method for the 2-point case.

97 MATHEMATICS AND COMPUTING↗

Enhancing Autonomous Control of Microreactors Using Multi-Agent Reinforcement Learning

In order for microreactors to be economically competitive, operation costs will need to be minimized through some degree of autonomous control. Previous work has demonstrated the effectiveness of reinforcement learning (RL) for load-following control in a drum-controlled microreactor. This study extends that work by exploring the potential of RL to independently control each of the reactor’s drums. We compare a single-agent RL approach with a multi-agent RL (MARL) framework, testing them for generalization across different load-following power profiles and control timescales, and for robustness in cases of randomly disabled control drums. Since the point kinetics simulation environment used in this study cannot resolve spatial effects, we assume that in the absence of spatially localized disturbances, optimal drum movements should be symmetrical. We demonstrate that single-agent RL is able to achieve accurate performance only when symmetric actions are ignored; otherwise, it fails to train a useful controller. Meanwhile, the MARL framework performs symmetric actions by design and trains a robust, accurate agent, as evidenced by mean absolute errors in power matching of 0.41% for the training power profile, 0.68% for a profile with half the drums disabled, and 0.21% for a profile on a realistic load-following time horizon.

22 - GENERAL STUDIES OF NUCLEAR REACTORS↗

Physics-informed Deep Reinforcement Learning-based Control in Power systems

Incorporating physics information into the deep reinforcement learning (DRL) process is a promising approach for addressing the challenges faced in learning-based control design problems for physical systems. Power grid dynamics, being a physical system, adheres to specific physical laws, constraints, as well as operational and control rules. Therefore, consideration of such physics-based law improves the learning process drastically. In general, traditional grid control schemes rely on rule-based mechanisms that cannot adapt to changing operating conditions. To improve the adaptability and computation time, recent research has seen a surge of DRL-based applications in power grid control. A generic DRL-based control design imposes the system performance requirements through the design of reward functions. In some cases, some of the important physics information is injected through this reward function. However, due to the complex dynamics and large state-action space, learning an optimal DRL policy often becomes challenging. Inspired by the latest developments in general machine learning (ML) research, power system researchers have been investigating more direct ways of incorporating physics knowledge into DRL training. This chapter specifically focuses on these aspects of physics-informed DRL designs in grid control. It discusses the significance, applications, research gaps, and open problems that need to be addressed in future research.

artificial intelligence, machine learning↗

In‐Scribe Silane Bonding for Mechanical Reinforcement of Perovskite Photovoltaic Modules

Despite the extraordinary rise in power conversion efficiency over the last decade, metal halide perovskite (MHP) photovoltaics remain more mechanically fragile than other PV technologies. In this work, the scribe area, created by the monolithic interconnection of thin-film solar cells, is used to extrinsically reinforce the mechanical robustness of packaged MHP solar modules. In contrast to the epoxy-based chemistries often leveraged in the MHP literature, silane-grafted polyolefin encapsulants are designed to form strong covalent bonds to oxide surfaces, specifically to glass and the transparent conductive oxide at the base of the scribe line. Pseudo-modules encapsulated with silane-grafted polyolefin are measured with more than an order-of-magnitude enhancement infracture energy from 0.27 ± 0.01 J·m −2 (no scribes) to 5.97 ± 0.42 J·m −2 (scribes perpendicular to delamination directioncovering ≈2.2% of the module area). The silane-grafted polyolefin retains strong adhesion even after undergoing an accelerated IEC 61215 thermal cycling test consisting of 250 cycles. We find that the in-scribe bonding allows perovskite modules to have adhesion strength comparable to commercial c-Si and CdTe technologies with only 5% reduction in the active module area. This manufacturing-compatible approach offers a practical solution to address the mechanical integrity challenges in MHP solar modules, regardless of cell architecture.

14 SOLAR ENERGY↗

Exploring the holographic entropy cone via reinforcement learning

We develop a reinforcement learning algorithm to study the holographic entropy cone. Given a target entropy vector, our algorithm searches for a graph realization whose min-cut entropies match the target vector. If the target vector does not admit such a graph realization, it must lie outside the cone, in which case the algorithm finds a graph whose corresponding entropy vector most nearly approximates the target and allows us to probe the location of the facets. For the N = 3 cone, we confirm that our algorithm successfully rediscovers monogamy of mutual information beginning with a target vector outside the holographic entropy cone. We then apply the algorithm to the N = 6 cone, analyzing the 6 mystery extreme rays of the subadditivity cone from [1] that satisfy all known holographic entropy inequalities yet lacked graph realizations. We found realizations for 3 of them, proving they are genuine extreme rays of the holographic entropy cone, while providing evidence that the remaining 3 are not realizable, implying unknown holographic inequalities exist for N = 6.

AdS-CFT correspondence↗

Empirical Characterization and Modeling of Cohesive – to – Adhesive Shear Fracture Mode Transition due to Increased Adhesive Layer Thicknesses of Fiber Reinforced Composite Single – Lap Joints

Here, to ensure a strong adhesive bond, most standards and adhesive manufacturers specify a maximum adhesive gap of 1 mm when bonding fiber reinforced composite structures. In manufacturing large components, such as joining two halves of wind turbine blades, meeting this gap tolerance specification is impractical; gaps larger than 10 mm are common in large adhesively bonded composite structures using state-of-the-art manufacturing techniques. Currently, there is a lack of fundamental understanding of the failure mechanics of adhesive gaps larger than 3 mm. To create such understanding, glass fiber - acrylic thermoplastic composite panels bonded using different epoxy adhesives within single-lap joint samples with adhesive thicknesses of 0.1 mm, 0.3 mm, 1 mm, 3 mm, 5 mm, and 10 mm were sheared to failure. A transition from cohesive to adhesive failure was observed to occur about 1 mm to 3 mm joint thicknesses. Plotting the shear stress normalized by the ratio of the joint width to thickness as a function of the joint thickness normalized by the joint length is shown to result in the ability to fit simple empirically derived models of the cohesive-to-adhesive failure transition, regardless of the adhesive. Furthermore, using these normalized variables, all the observed cohesively failed specimens collapse to a single master curve, as do the adhesively failed specimens.

36 MATERIALS SCIENCE↗

Deep reinforcement learning control for co-optimizing energy consumption, thermal comfort, and indoor air quality in an office building

With the recent demand for decarbonization and energy efficiency, advanced HVAC control using Deep Reinforcement Learning (DRL) becomes a promising solution. Due to its flexible structures, DRL has been successful in energy reduction for many HVAC systems. However, only a few researches applied DRL agents to manage the entire central HVAC system and control multiple components in both the water loop and the air loop, owing to its complex system structures. Moreover, those researches have not extended their applications by incorporating the indoor air quality, especially both CO2 and PM2.5concentrations, on top of energy saving and thermal comfort, as achieving those objectives simultaneously can cause multiple control conflicts. What's more, DRL agents are usually trained on the simulation environment before deployment, so another challenge is to develop an accurate but relatively simple simulator. Therefore, we propose a DRL algorithm for a central HVAC system to co-optimize energy consumption, thermal comfort, indoor CO2 level, and indoor PM2.5 level in an office building. To train the controller, we also developed a hybrid simulator that decoupled the complex system into multiple simulation models, which are calibrated separately using laboratory test data. The hybrid simulator combined the dynamics of the HVAC system, the building envelope, as well as moisture, CO2, and particulate matter transfer. Three control algorithms (rule-based, MPC, and DRL) are developed, and their performances are evaluated on the hybrid simulator environment with a realistic scenario (i.e., with stochastic noises). The test results showed that, the DRL controller can save 21.4 % of energy compared to a rule-based controller, and has improved thermal comfort, reduced indoor CO2 concentration. The MPC controller showed an 18.6 % energy saving compared to the DRL controller, mainly due to savings from comfort and indoor air quality boundary violations caused by unmeasured disturbances, and it also highlights computational challenges in real-time control due to non-linear optimization. Finally, we provide the practical considerations for designing and implementing the DRL and MPC controllers based on their respective pros and cons.

Guo, Fangzhou↗

Designing reinforcement learning algorithms for building HVAC control: From experimental observation to simulation comparisons

Advanced supervisory-level control with reinforcement learning (RL) is regarded as a promising solution for HVAC systems to minimize energy consumption while maintaining thermal comfort and indoor air quality. However, most RL applications were conducted in the simulation environment rather than real-world HVAC systems. This paper developed a value-based RL controller termed Deep Q-Network (DQN) for a typical central HVAC system and evaluated its performance in a building test facility. By comparing DQN with a rule-based controller, the study not only demonstrated the cases where DQN could properly maintain indoor comfort but also discussed possible reasons why DQN failed in some other situations. Recognizing the limitations of value-based RL algorithms from the experimental tests, a simulation study was conducted to compare DQN with an alternative RL approach, an actor–critic algorithm termed Deep Deterministic Policy Gradient (DDPG). In scenarios with a relatively large action space, DDPG outperformed DQN by requiring fewer computational resources and achieving better thermal comfort, lower energy consumption, and more stable control actions. The findings suggest that the ability of DDPG to handle continuous control variables more effectively allows for faster convergence in training and more precise control in practice, which enhances the overall efficiency and reliability of the HVAC system.

Guo, Fangzhou↗

Cellulose nanofibrils-based packaging films synergistically reinforced by lignin and tea polyphenols for strawberry preservation

Developing biomass-based food packaging materials could alleviate resource scarcity and environmental pollution. Here, this study successfully fabricated cellulose nanofibrils-based packaging films synergistically reinforced by lignin and tea polyphenols for fruit preservation. Thanks to the abundant carboxyl and hydroxyl groups in modified lignin and tea polyphenol, the composite film delivered the highest tensile strength and Young’s modulus of 230.7 MPa and 11.9 GPa, respectively. The water vapor permeability (WVP) and oxygen permeability (OP) of composite film were as low as 2.53 × 10 –11 gm –1 s –1 Pa –1 and 1.69 × 10 –16 cm m –2 s–1 Pa –1 , respectively, demonstrating excellent barrier performance. Furthermore, the composite film possessed remarkable antibacterial (100 % antibacterial activity against two types of bacteria), antioxidant (DPPH free radical scavenging activity of 90 %), and UV shielding properties (impressive UV shielding rate of 99.0 %). Strawberry preservation experiments showed that these composite films could significantly extend shelf life, owing to the excellent barrier and antimicrobial properties. Our research indicated that these composite films could be promising food packaging material compared with traditional plastic and showed great application potential in fruit preservation.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Reactive extrusion of frontally polymerizing continuous carbon fiber reinforced polymer composites

The manufacturing of carbon fiber-reinforced polymer (CFRP) composites demands rapid and energy-efficient strategies. Frontal polymerization (FP) enables the manufacturing of CFRP using dicyclopentadiene (DCPD) thermoset polymer which meets these requirements. In this work, we introduce reactive extrusion of CFRP (RE-CFRP), where two rollers provide localized heat and pressure to sustain the curing reaction and the consolidation of a continuous carbon fiber tow pre-impregnated with DCPD. We study the effect of the extrusion speed, temperature, and compaction force on the properties of the produced CFRP. Mechanical testing confirms that the resulting fiber volume fraction and the elastic modulus are similar to bulk cured tows. A homogenized thermo-chemical model is developed to study the effect of the process parameters on the polymerization reaction. The process produces hollow woven composite tubes directly via extrusion and in situ curing. Overall, this process offers advantages in curing, tooling, speed, and energy.

36 MATERIALS SCIENCE↗

3D printed carbon fiber reinforced carbon as an energy efficient alternative to graphite for EFAS tooling

As Electric Field Assisted Sintering (EFAS) gains more industrial acceptance and use, it becomes more important to develop more efficient means to implement this technology. To this aim, 3D printed continuous carbon fiber reinforced carbon (CCC) was manufactured and fabricated into tooling for EFAS systems as an alternative to traditional graphite tooling. The impact of fiber orientation on the thermal and electrical properties of the CCC was characterized. Sample material was sintered in Tokai G535 graphite tooling, under common processing conditions and compared with CCC tooling. There was nearly 50 % energy savings compared to graphite while maintaining equivalent sample density and microstructure plus keeping ram temperatures 39 % cooler. This is due to spatial control of generated heat and thermal diffusivity within the molds, by means of fiber orientation anisotropy. Finite element modeling of the tooling design supported the experimental results as well as displays the effect of optimization of this 3D printed CCC material.

36 MATERIALS SCIENCE↗

Physically constrained 3D diffusion for inverse design of fiber-reinforced polymer composite materials

Designing fiber-reinforced polymer composites (FRPCs) with a tailored nonlinear stress-strain response is crucial for applications such as energy absorption in crash structures, flexible robotics, and impact-resistant protective gear. However, the inherent complexities of composite materials and the multitude of parameters involved, render traditional design and optimization methods inadequate for achieving effective inverse design of composites. In this paper, we present an AI-based inverse design framework that effectively and efficiently generates FRPCs with targeted nonlinear stress-strain responses. We introduce a physically constrained diffusion model (PC3D_Diffusion) capable of managing the complexities of composite materials and producing detailed, high-quality designs. We propose a loss-guided, learning-free approach to generate physically feasible microstructure designs by explicitly enforcing physical constraints during the generation process. For training purposes, 1.35 million FRPC samples were created, and their corresponding stress-strain curves were computed using established physics-based computational models. The results show that PC3D_Diffusion consistently generates high-quality designs with tailored mechanical behaviors, while guaranteeing compliance with the physical constraints. PC3D_Diffusion advances FRPC inverse design and may facilitate the inverse design of other 3D materials, offering potential applications in industries reliant on materials with custom mechanical properties.

Xu, Pei [Clemson Univ., SC (United States)]↗

Corrosion characteristics of silicon carbide fiber-reinforced composites in beryllium-bearing molten fluoride salt

Corrosion of SiC-fiber reinforced SiC matrix composites in 2LiF-BeF 2 molten salt was examined following static molten salt exposure for 1000 h at 650 and 750 °C. Composites were fabricated with either Tyranno SA3 or Hi-Nicalon type-S fibers and a chemical vapor infiltration SiC matrix. Both composites showed minimal weight loss after salt exposure and maintained shape. Corrosion was characterized by comprehensive electron microscopy and Raman spectroscopy. In conclusion, based on the results of experimental and thermodynamic analysis, a corrosion mechanism for SiC/SiC is proposed where both the salt and SiC materials play a key role in the phase stability of the composite.

36 MATERIALS SCIENCE↗