Search NASA⌕ Search

SEARCH · Search NASA

Results for “REINFORCEMENT”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 361 records · Page 20

Additive Manufacturing and Characterization of Polylactic Acid (PLA) Composites Containing Metal Reinforcements

Additive manufacturing of polymeric systems using 3D printing has become quite popular recently due to rapid growth and availability of low cost and open source 3D printers. Two widely used 3D printing filaments are based on polylactic acid (PLA) and acrylonitrile butadiene styrene (ABS) systems. PLA is much more environmentally friendly in comparison to ABS since it is made from renewable resources such as corn, sugarcane, and other starches as precursors. Recently, polylactic acid-based metal powder containing composite filaments have emerged which could be utilized for multifunctional applications. The composite filaments have higher density than pure PLA, and the majority of the materials volume is made up of polylactic acid. In order to utilize functionalities of composite filaments, printing behavior and properties of 3-D printed composites need to be characterized and compared with the pure PLA materials. In this study, pure PLA and composite specimens with different metallic reinforcements (Copper, Bronze, Tungsten, Iron, etc) were 3D printed at various layer heights and resulting microstructures and properties were characterized. Differential scanning calorimetry (DSC) and thermogravimetric analysis (TGA) behavior of filaments with different reinforcements were studied. The microscopy results show an increase in porosity between 3-D printed regular PLA and the metal composite PLA samples, which could produce weaker mechanical properties in the metal composite materials. Tensile strength and fracture toughness behavior of specimens as a function of print layer height will be presented.

brass↗

Influence of Aerogel Morphology and Reinforcement Architecture on Gas Convection in Aerogel Composites

A variety of thermal protection applications require lightweight insulation capable of withstanding temperatures well above 900 C. Aerogels offer extremely low-density thermal insulation due to their mesoporous structure, which inhibits both gas convection and solid conduction. Silica aerogel systems are limited to use temperatures of 600-700 C, above which they sinter. Alumina aerogels maintain a porous structure to higher temperatures than silica, before transforming to -alumina and densifying. We have synthesized aluminosilicate aerogels capable of maintaining higher surface areas at temperatures above 1100 C than an all-alumina aerogel using -Boehmite as the aluminum source and tetraethoxysilane (TEOS) as the silicon source. The pore structure of these aerogels varies with thermal exposure temperature and time, as the aluminosilicate undergoes a variety of phase changes to form transition aluminas. Transformation to -alumina is inhibited by incorporation of silica into the alumina lattice. The aerogels are fragile, but can be reinforced using a large variety of ceramic papers, felts or fabrics. The objective of the current study is to characterize the influence of choice of reinforcement and architecture on gas permeability of the aerogel composites in both the as fabricated condition and following thermal exposure, as well as understand the effects of incorporating hydrophobic treatments in the composites.

aerogels↗

Hybrid Fiber Layup and Fiber-Reinforced Polymeric Composites Produced Therefrom

Embodiments of a hybrid fiber layup used to form a fiber-reinforced polymeric composite, and a fiber-reinforced polymeric composite produced therefrom are disclosed. The hybrid fiber layup comprises one or more dry fiber strips and one or more prepreg fiber strips arranged side by side within each layer, wherein the prepreg fiber strips comprise fiber material impregnated with polymer resin and the dry fiber strips comprise fiber material without impregnated polymer resin.

Barnell, Thomas J.↗

Hybrid Fiber Layup and Fiber-Reinforced Polymeric Composites Produced Therefrom

Embodiments of a hybrid fiber layup used to form a fiber-reinforced polymeric composite, and a fiber-reinforced polymeric composite produced therefrom are disclosed. The hybrid fiber layup comprises one or more dry fiber strips and one or more prepreg fiber strips arranged side by side within each layer, wherein the prepreg fiber strips comprise fiber material impregnated with polymer resin and the dry fiber strips comprise fiber material without impregnated polymer resin.

Barnell, Thomas J.↗

Ground Delay Program Analytics with Behavioral Cloning and Inverse Reinforcement Learning

We used historical data to build two types of model that predict Ground Delay Program implementation decisions and also produce insights into how and why those decisions are made. More specifically, we built behavioral cloning and inverse reinforcement learning models that predict hourly Ground Delay Program implementation at Newark Liberty International and San Francisco International airports. Data available to the models include actual and scheduled air traffic metrics and observed and forecasted weather conditions. We found that the random forest behavioral cloning models we developed are substantially better at predicting hourly Ground Delay Program implementation for these airports than the inverse reinforcement learning models we developed. However, all of the models struggle to predict the initialization and cancellation of Ground Delay Programs. We also investigated the structure of the models in order to gain insights into Ground Delay Program implementation decision making. Notably, characteristics of both types of model suggest that GDP implementation decisions are more tactical than strategic: they are made primarily based on conditions now or conditions anticipated in only the next couple of hours.

Bloem, Michael↗

Adaptive Stress Testing: Using Reinforcement Learning to Find Failures in Safety-Critical Systems

Emerging applications in artificial intelligence, such as driverless cars and autonomous aircraft promise to be more efficient, cheaper to operate, and always available. However, ensuring the safety of these systems remains a major challenge to their certification and adoption. These autonomous systems are expected to routinely make safety-critical decisions where failures can have serious consequences including loss of life and property. Testing and validation techniques aim to identify and diagnose potential failures before the system is deployed. However, finding failure scenarios in autonomous systems can be very challenging due to high-dimensional and continuous state spaces, interaction with large environments over many time steps, and the rarity of failures. This talk presents Adaptive Stress Testing (AST), a simulation-based testing framework for finding the most likely path to a failure event of a safety-critical system. The key idea of AST is that stress testing can be formulated as a Partially Observable Markov Decision Process (POMDP), which enables reinforcement learning techniques to be used for finding failure events. Reinforcement learning algorithms can efficiently explore the search space and have been shown to scale to very large systems. We present applications of AST to find failures in various safety-critical systems including the aircraft collision avoidance systems, autonomous cars, and small unmanned aerial vehicles.

autonomous vehicles↗

A Reinforcement Learning Framework for Space Missions in Unknown Environments

A land-and-traverse mission to icy worlds such as Europa and Enceladus is challenging due to lack of prior knowledge regarding the terrain conditions. Previous work [1] showed that rovers with high degrees of freedom (DoF) can achieve robust traversal by leveraging redundant modes for mobility to counter terrain uncertainty (e.g. walking, driving, or inch-worming). This paper presents a generic and scalable reinforcement learning scheme for enabling on-board decision making on rovers to automatically switch between modes of traversal based on online performance feedback. The objective is to maximize energy efficiency, minimize operator input and successfully negotiate unstructured terrain conditions without relying on exhaustive prior knowledge. The proposed methodology is well grounded in the literature on reinforcement learning and has been adapted to address conformance to validation and verification requirements and JPL flight operations history of using per-sol prescribed sequences for a space mission.

Tavallali, Peyman↗

Scheduling Mission Reconfiguration for an Interferometry Synthetic Aperture Radar Using Deep Reinforcement Learning

This paper presents a method to intelligently adapt the baseline of a synthetic aperture radar based on Deep Rein- forcement Learning to help create plans for missions that use formation flight for Earth observation purposes. The main contribution of this paper is the initial results we have found from applying the tool to a toy mission: measuring the ver- tical structure of forests by using a synthetic aperture radar mounted on a formation of 7 satellites orbiting the Earth in a Sun Synchronous Orbit. We have found that with a reward function based on expected science return over time and fuel usage, the Deep Reinforcement Learning planner is able to create plans with positive scientific returns while minimizing fuel usage. We also find that fuel usage and collision avoid- ance planning is better done with traditional methods, as Deep Reinforcement Learning does not converge to optimal solutions.

Viros-i-Martin, Antoni↗

Effects of Debulking on the Fiber Microstructure and Void Distribution in Carbon Fiber Reinforced Plastics

Carbon Fiber Reinforced Plastics (CFRPs) are widely used due to their high stiffness to weight ratios. A common process manufacturers use to increase the strength to weight ratio is debulking. Debulking is the process of compacting a dry fibrous reinforcement prior to resin infusion. This process is meant to decrease the average inter-fiber distance, effectively increasing the fiber volume fraction of the sample. While this process is widely understood macroscopically its effects on fibrous microstructures have not yet been well characterized. The aim of this work is to compare the microstructures of three CFRP laminates, varying only the debulking step in the manufacturing process. High resolution serial sections of all three laminates were taken for analysis. Using these scans, the fiber positions were reconstructed. Statistical descriptors such as local fiber and void volume fractions, fiber orientation, and void distribution and morphology were then generated for each sample. Fiber clusters present within the material were identified and analyzed for each level of debulking applied. Using these descriptors, the effects of debulking on the morphology and organization of the composite microstructure was evaluated.

carbon fiber↗

Micromechanical Design of Carbon Nanotube Ribbon Reinforced Polymer Composite Materials

Lightweight materials are an important component of the design of aerospace structures. Carbon nanotube materials have been considered for this purpose due to the strength and stiffness of individual nanotubes, and the commercial availability of bulk formats such as fibers. These fibers can have a ribbon cross section which results in a different design space for their composites relative to traditional reinforcements which have a round cross section. This work applies brick-and-mortar micromechanical models and classical lamination theory with an inverse approach to investigate the design space of these composites. Using this approach, the influences of fiber geometry and axial and transverse mechanical properties are mapped. Finally, a sensitivity study is performed and the relative impacts of ±10% variations in the constituent material and geometric properties are ranked. Lamina axial moduli were found to range from a maximum of 3x to 1x minimum relative to a target quasi-isotropic laminate modulus depending on the anisotropy and shear modulus of the lamina. The fiber targets depended strongly on the fiber volume fraction in the lamina and the fiber axial modulus target was found to range from 4.8x to 3.2x the quasi-isotropic laminate target. The sensitivity analysis found that the largest driver of performance was the volume fraction, followed by the fiber axial modulus. While bio-based brick-and-mortar composites, such as nacre, can benefit from reinforcement aspect ratios above 10, for carbon nanotube ribbon (or carbon fiber)/polymer composites the sensitivity study indicated that the optimal cross-sectional aspect ratio was relatively smaller, potentially less than three.

Micromechanics↗

Post-Deployment Characterization of Glass Fiber-Reinforced Thermoset and Thermoplastic Composite Tidal Turbine Blades: Preprint

In 2021, the National Renewable Energy Laboratory (NREL) supported Verdant Power with the most successful tidal energy deployment in U.S. history. Three of their tidal turbines were deployed as part of the Roosevelt Island Tidal Energy project. Initially, the three rotors were manufactured from glass fiber-reinforced epoxy composites. Midway through the deployment, one rotor was replaced with one manufactured at NREL. Instead, it was infused with a novel, infusible thermoplastic resin system. Since the deployment, one epoxy rotor and one thermoplastic rotor were returned to NREL for continued materials and manufacturing research. The two rotors underwent full-scale structural testing before being sectioned and cut into specimens for a variety of manufacturing quality tests, thermos-mechanical characterization, and evaluation of material performance in marine environments, to understand the key differences between the fiberglass reinforced epoxy and Elium composites used for the respective rotors. Matrix burnoff tests showed that the Elium blades had a considerably higher fiber volume fraction compared to the epoxy blades (61% vs. 49%). Environmental aging of the specimens showed that the epoxy laminates absorbed more water over the conditioning period, however, it was determined that the Elium laminates had higher diffusion coefficients, so initially absorbed water faster. Finally, one full epoxy blade and one full Elium blade were conditioned at ambient temperatures for up to 11 months, while taking periodic mass measurements. The datasets were extrapolated out to assume a full 20 year operational life span and it was determined that the blades would not reach full saturation during that time span.

composite manufacturing↗

Post-Deployment Characterization of Glass Fiber-Reinforced Thermoset and Thermoplastic Composite Tidal Turbine Blades

In 2021, the National Renewable Energy Laboratory (NREL) supported Verdant Power with the most successful tidal energy deployment in U.S. history. Three of their Gen5d 5 m turbines were deployed as part of the Roosevelt Island Tidal Energy project. Initially, the three deployed rotors were manufactured from glass fiber-reinforced epoxy composites. Midway through the deployment, one rotor was replaced with one manufactured at NREL. The new rotor utilized a novel infusible thermoplastic resin system (Elium from Arkema). Since the deployment, one epoxy rotor and one thermoplastic rotor were returned to NREL for continued materials and manufacturing research. The two rotors underwent full-scale structural testing before being sectioned and cut into specimens for a variety of manufacturing quality tests, thermomechanical characterization, and evaluation of material performance in marine environments to understand the key differences between the fiberglass-reinforced epoxy and Elium composites used for the respective rotors. Matrix burn-off tests showed that the Elium blades had a considerably higher fiber volume fraction compared to the epoxy blades (61% vs. 49%). Environmental aging of the specimens showed that the epoxy laminates absorbed more water over the conditioning period; however, it was determined that the Elium laminates had higher diffusion coefficients, so they initially absorbed water faster. Finally, one full epoxy blade and one full Elium blade were conditioned at ambient temperatures for up to 11 months, while periodic mass measurements were taken. The datasets were extrapolated to assume a full 20-year operational life span, and it was determined that the blades would not reach full saturation during that time span.

composite manufacturing↗

CDRL: Certification-Driven Reinforcement Learning for Neutrino Flavor Model Discovery

Many scientific discovery problems require searching combinatorial hypothesis spaces under complex domain constraints. Reinforcement learning (RL) offers a promising approach, but existing methods rely on scalar rewards that provide limited information about why candidate solutions fail, leading agents to repeatedly explore invalid regions. We introduce Certification-Driven Reinforcement Learning (CDRL), a framework that leverages structured feedback from symbolic reasoning tools. When a candidate violates domain constraints, these tools produce certificates identifying the actions responsible for failure. CDRL converts these certificates into reusable constraints that eliminate classes of invalid solutions and guide exploration toward valid regions. We evaluate CDRL on neutrino flavor model discovery in theoretical particle physics, where the hypothesis space exceeds $10^{26}$ possible models, and compare it with the state-of-the-art RL approach previously used for this task. Across three theory spaces, CDRL achieves up to 1.95$\times$ higher valid model rates and up to 6.33$\times$ higher neutrino model rates while evaluating up to 4$\times$ fewer candidates. We further extract 40 interpretable rules from search trajectories using a post-hoc decision-tree framework and show that reusing them as soft constraints yields gains of up to 2$\times$ in valid model rates and 3$\times$ in neutrino model discovery across all three theory spaces. These results suggest that CDRL uncovers reusable structure in combinatorial search spaces and provides a general framework for scientific model discovery.

Jha, Piyush [Georgia Tech., Atlanta; Georgia Tech]↗

A Physics-Informed Reinforcement Learning Framework for Economic-Thermal Co-Optimization of Crypto Mining Data Centers: Preprint

The rapid expansion of cryptocurrency mining has created a new class of high-density data centers characterized by extreme thermal flux and high sensitivity to volatile economic markets. Traditional thermal management strategies, typically reliant on rule-based control, maintain static setpoints that fail to account for fluctuating electricity prices and cryptocurrency values - factors critical to mining profitability. To address this, we present a physics-informed reinforcement learning (PIRL) framework for economic-thermal co-optimization in crypto mining data centers. This framework consists of a proximal policy optimization (PPO) agent, a virtual testbed powered by high-fidelity physics-based models, and an interactive frontend dashboard. The PPO agent is trained using the virtual testbed and strict hardware safety limits. This physics-informed approach allows the agent to learn a stochastic policy that dynamically balances mining revenue against operational costs by co-optimizing HVAC cooling setpoints and IT computational hashrate. The simulation results demonstrate that the integrated framework achieved an 8.62% increase in net operational profit compared to traditional baseline strategies while strictly adhering to safety-critical temperature constraints (coolant supply temperature < 32 degrees C). This work provides a scalable template for the deployment of reinforcement learning in mission critical facilities where economic volatility and physical safety must be managed simultaneously.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗

Nuclear microreactor transient and load-following control with deep reinforcement learning

The economic feasibility of nuclear microreactors will depend on minimizing operating costs through advancements in autonomous control, especially when these microreactors are operating alongside other types of energy systems (e.g., renewable energy). This study explores the application of deep reinforcement learning (RL) for real-time drum control in microreactors, exploring performance in regard to load-following scenarios. By leveraging a point kinetics model with thermal and xenon feedback, we first establish a baseline using a single-output RL agent, then compare it against a traditional proportional–integral–derivative (PID) controller. This study demonstrates that RL controllers, including both single- and multi-agent RL (MARL) frameworks, can achieve similar or even superior load-following performance as traditional PID control across a range of load-following scenarios. In short transients, the RL agent was able to reduce the tracking error rate in comparison to PID by one half to one third. Over extended 300-minute load-following scenarios in which xenon feedback becomes a dominant factor, PID maintained better accuracy, but RL still remained within a 1% error margin despite being trained only on short-duration scenarios. This highlights RL’s strong ability to generalize and extrapolate to longer, more complex transients, affording substantial reductions in training costs and reduced overfitting. Furthermore, when control was extended to multiple drums, MARL enabled independent drum control as well as maintained reactor symmetry constraints without sacrificing performance---an objective that standard single-agent RL could not learn. We also found that, as increasing levels of Gaussian noise were added to the power measurements, the RL controllers were able to maintain lower error rates than PID, and to do so with at least 10% and upwards of 150% less control effort. These findings illustrate RL's potential for autonomous nuclear reactor control, laying the groundwork for future integration into high-fidelity simulations and experimental validation efforts.

22 - GENERAL STUDIES OF NUCLEAR REACTORS↗

Optimization of the FRIB beam dump: a hybrid genetic algorithm and reinforcement learning approach

The operational envelope of high-power-density systems, such as particle accelerators and advanced nuclear energy systems, is critically constrained by the need to manage extreme thermal loads. To address this, we present a novel hybrid optimization framework combining a genetic algorithm (GA) with a soft actor-critic (SAC) deep reinforcement learning agent. This framework was applied to a practical high-heat-flux problem: redesigning the beam dump at the Facility for Rare Isotope Beams (FRIB) for a power upgrade from 20 kW to 50 kW. The resulting design, validated by three-dimensional conjugate heat transfer simulations, suppresses hazardous hot spots and yields a markedly more uniform temperature distribution. This provides a robust operating margin, increasing the average power-handling capability by 72% relative to the current design, demonstrating the framework’s potential to solve complex thermal management challenges in both accelerator technology and advanced nuclear systems.

Accelerator↗

Visibility-enhanced model-free deep reinforcement learning algorithm for voltage control in realistic distribution systems using smart inverters

Increasing integration of distributed solar photovoltaic (PV) into distribution networks could result in adverse effects on grid operation. Traditional model-based control algorithms require accurate model information that is difficult to acquire and thus are challenging to implement in practice. Here, this paper proposes a surrogate model-enabled grid visibility scheme to empower deep reinforcement learning (DRL) approach for distribution network voltage regulation using PV inverters with minimal system knowledge. In contrast to existing DRL methods, this paper presents and corroborates the adverse impact of missing load information on DRL performance and, based on this finding, proposes a surrogate model methodology to impute load information utilizing observable data. Additionally, a multi-fidelity neural network is utilized to construct the DRL training environment, chosen for its efficient data utilization and enhanced robustness to data uncertainty. The feasibility and effectiveness of the proposed algorithm are assessed by considering DRL testing across varying degrees of observable load information and diverse training environments on a realistic power system.

14 SOLAR ENERGY↗

Development of algorithms for augmenting and replacing conventional process control using reinforcement learning

Here, this work seeks to allow for the online operation and training of model-free reinforcement learning (RL) agents but limit the risk to system equipment and personnel. The parallel implementation of RL alongside more conventional process control (CPC) allows for the RL algorithm to learn from CPC. The past performance of both methods are assessed on a continuous basis allowing for a transition from CPC to RL and, if needed, transitioning back to CPC from RL. This allows for the RL algorithm to slowly and safely assume control of the process without significant degradation in control performance. It is shown that the RL can derive a near optimal policy even when coupled with a suboptimal CPC. It is also demonstrated that the coupled RL-CPC algorithm learns at a faster rate than traditional RL methods of exploration while the algorithm’s performance does not deteriorate below CPC, even when exposed to an unknown operating condition.

30 DIRECT ENERGY CONVERSION↗