Search NASASearch

SEARCH · Search NASA

Results for “RL”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 91 records · Page 5

Explainable physics-based constraints on reinforcement learning for accelerator optimization

We present a reinforcement learning (RL) framework for optimizing particle accelerator experiments that builds explainable physics-based constraints on agent behavior. The goal is to increase transparency and trust by letting users verify that the agent’s decision-making process incorporates suitable physics. Our algorithm uses a learnable surrogate function for physical observables, such as energy, and uses them to fine-tune how actions are chosen. This surrogate can be represented by a neural network or by an interpretable sparse dictionary model. We test our algorithm on a range of particle accelerator optimization environments designed to emulate the Continuous Electron Beam Accelerator Facility at Jefferson Lab. By examining the mathematical form of the learned constraint function, we are able to confirm the agent has learned to use the established physics of each environment. In addition, we find that the introduction of a physics-based surrogate enables our RL algorithms to reliably converge for difficult high-dimensional accelerator optimization environments.

explainability

Integrated Routing and Traffic Signal Control for CAVs via Reinforcement Learning Approach

Incorporating Connected and Automated Vehicles (CAVs) into urban traffic networks presents opportunities and challenges for traffic management systems. This paper aims to develop an integrated routing and traffic signal control system designed explicitly for CAVs, utilizing a Reinforcement Learning (RL) approach. The objective is to enhance traffic flow and improve overall transportation efficiency in the controlled areas. We propose an innovative framework that employs the Deep Reinforcement Learning (DRL) algorithm, especially the Deep Q-network (DQN), to dynamically adjust the number of vehicles in the routes and the duration of traffic signals. Our simulation results demonstrate that a DQN agent successfully optimizes the number of vehicles in the routes and traffic signal timings of traffic signal controllers, eventually reducing total travel time. The study illustrates the potential usage of RL-based systems in managing routing and traffic signals for CAVs, offering a promising opportunity for future urban traffic management strategies.

Park, Jiho [New York University]

Reinforcement Learning-Based Secondary Control Strategy for Voltage and Frequency Regulation in Islanded Inverter-Based Microgrids

This paper presents a reinforcement learning (RL) approach for secondary voltage and frequency control in islanded inverter-based microgrids. The proposed control strategy aims to restore voltage and frequency deviations caused by the primary droop control while ensuring proper power sharing between distributed generators. The RL agent is designed to provide correction signals to the primary control, considering communication delays and system constraints. The effectiveness of the proposed control strategy is validated through simulation results in MATLAB/Simulink environment, demonstrating superior performance in maintaining voltage and frequency within the nominal values.

Rodriguez Martinez, Omar Felipe [University of Pue

Safe Deep Reinforcement Learning for Robust Frequency and Voltage-Constrained Networked Microgrid Restoration

Here, this paper proposes a safe soft actor-critic reinforcement learning (RL) algorithm–based controller for networked microgrid restoration. It formulates the post black-start start as a finite-horizon constrained Markov decision process. The RL agent co-optimizes real and reactive power set-points for both grid-forming and grid-following inverters under explicit voltage and frequency constraints, while enforcing proper power sharing via the Mean Active Power Sharing Index (MPSI) and Mean Reactive Power Sharing Index (MQSI). Numerical results obtained on the IEEE 123-bus distribution system show that the proposed method achieves a mean voltage build-up time of 0.01 s without breaching the 5% sharing-violation budget under various load scenarios, considering MPSI and MQSI indices. These findings demonstrate that the proposed method yields fast and safe black-start schedules without resorting to heuristic penalties.

Selim, Alaa [Dartmouth College, Hanover, NH (Unite

Safe Reinforcement Learning-Based Transient Stability Control for Islanded Microgrids With Topology Reconfiguration

This paper proposes a safe reinforcement learning (RL)-based transient stability emergency control (TSEC) method for islanded microgrids. RL requires extensive interaction with the environment to learn control strategies, hence, a data-driven approach is used as a substitute for time-consuming time-domain simulation calculations. Deep sigma point processes (DSPP), which is a Gaussian process model, is utilized to predict the normal distribution of transient stability of microgrids and to construct a transient stability chance constraint. Reward-constrained policy optimization (RCPO) can simultaneously achieve objective prediction, policy learning, and constraint cost coefficient update across multiple timescales. RCPO interacts with the DSPP-based microgrid environment through a multi-process parallel manner, greatly increasing the training speed. Case studies on a real islanded microgrid demonstrate that the proposed method can efficiently and quickly obtain the optimal emergency control strategy while adhering to all hard constraints.

14 SOLAR ENERGY

Tensorized Interior Radiative Heat Transfer for a Scalable and Calibrated Building Energy Simulator

Building energy simulation is a critical tool for developing and testing advanced control strategies, such as Reinforcement Learning (RL), to provide demand flexibility and affordable energy costs. The recently introduced Smart Buildings Control Suite (sbsim) provides a lightweight, scalable, and data-calibrated simulation environment based on a 2D finite-difference model. However, the initial model primarily focused on conductive and convective heat transfer, neglecting the significant impact of long-wave radiative heat exchange between interior surfaces. This paper presents a significant extension to the sbsim framework by incorporating a physically-grounded model for interior radiative heat transfer. Our primary contribution is the development and integration of a fully tensorized radiative heat transfer module, which preserves the computational efficiency and scalability of the original simulator. This was achieved by developing a pipeline for view factor calculation, including an algorithm to identify directly seeing surfaces within complex floor plans, and formulating the net radiation equations for efficient execution on modern hardware accelerators. We validate the numerical accuracy of our tensorized implementation by comparing its results against a traditional iterative approach, demonstrating identical outcomes. This enhancement increases the physical fidelity of sbsim, enabling more accurate training of RL agents for building energy optimization.

Ham, Sang woo

Explainable and Differentiable Reinforcement Learning for Multi-objective Optimization in Particle Accelerators

Operating particle accelerators involves optimizing multiple goals simultaneously, which can be challenging due to trade-offs among objectives. While evolutionary algorithms like the genetic algorithm (GA) have been used for various Multi-Objective Optimization (MOO) tasks, they are not inherently suited for complex control problems. This talk highlights two variations of Reinforcement Learning (RL) for concurrently optimizing heat load and trip rates at the Continuous Electron Beam Accelerator Facility (CEBAF). The problem involves strict constraints on individual states, actions, and overall energy requirements of the beam. First, this talk highlights how differentiability can be harnessed through a Deep Differentiable Reinforcement Learning (DDRL) approach to address MOO issues within particle accelerators. We examine the DDRL method alongside Model Free Reinforcement Learning (MFRL), GA, and Bayesian Optimization (BO). The performance of these methods is assessed by generating a Pareto-front for two objectives. Our findings indicate that DDRL excels in handling high-dimensional problems more effectively than MFRL, BO, and GA. Next, we will show integration of explainable physics-based constraints into RL algorithms to enhance trans- parency and trust in decision-making processes by enabling users to verify that agents adhere to established physical principles. This surrogate function can be modeled using neural networks or sparse dictionary mod- els. By examining the mathematical form of the learned constraint function, we are able to confirm the agent has learned to use the established physics of each environment provided but the surrogate model. In addi- tion, we find that the introduction of a mathematical functional dictionary based surrogate model enables our reinforcement learning algorithms to reliably converge for difficult high-dimensional accelerator controls environments.

Rajput, Kishansingh [Thomas Jefferson National Acc

Numerical Modeling & Optimization of the iProTech Pitching Inertial Pump (PIP) Wave Energy Converter (WEC) (CRADA Final Report)

This project represents a continuation of the collaboration between iProTech and NLR to simulate, optimize and design the iProTech Pitching Inertial Pump (PIP) device. The objectives of this TEAMER project are twofold: 1. Refining the physical characteristics of the existing iProTech PIP WEC-Sim model to enhance the model’s fidelity and include controllable components. Key model enhancements target the inclusion of Coulomb friction, the introduction of a controllable bypass valve, and the replacement of traditional check valves with advanced motorized ones. 2. Exploring traditional and advanced control algorithms. From traditional methods like latching control to cutting-edge reinforcement learning (RL) algorithms, the goal is to ensure the PIP device's adaptability and optimal performance across a range of ocean conditions. NLR is tasked with augmenting the WEC-Sim model and implementing the control algorithms, culminating in performance comparison analyses. iProTech will update their existing 3D models, advise on model improvements, and determine crucial system metrics. WEC-Sim, developed in MATLAB/SIMULINK with Simscape Multibody, is the main piece of software that will be used in this project. Coupled with the MATLAB RL Toolbox, it offers a robust platform for in-depth simulation and optimization of the iProTech PIP device. Building on previous work to explore the PIP design space and optimize its geometry, mass distribution, center of gravity and other key parameters, this project aims to refine iProTech’s existing numerical models and develop effective control algorithms that can seamlessly integrate into their future hardware testing campaigns.

16 TIDAL AND WAVE POWER

Optimization and stabilization of Fermilab Booster using hybrid Bayesian/RL framework

PIPII project will raise Fermilab Booster intensity and ramp rate. Beam losses will limit average power and are hard to simulate. Presently, Booster uses operator-guided empirical tuning. This task is challenging due to high dimensionality, multiple objectives, critical safety constraints, and drifts. We developed a synergistic suite of Bayesian optimization (BO) and reinforcement learning (RL) tools to optimize and stabilize beam losses. First, active learning was used to build a rough model. Data was collected parasitically using two novel safety constraint types – nonlinear input space restrictions (based on optics model), and uncertainty constraints (to stop bad steps/beam aborts). We then applied online multi-objective BO with scalarized objectives and fitting to improve/rebalance losses, increasing safety margins by 25%. Using BO model as a safety veto, we tried several on/off-policy RL agents for long term stabilization; SAC had best performance. We found that adding contextual (state) information further improved performance, eventually integrating key knobs like linac phase and temperature into the parameter space. Long term testing is ongoing to enable operational use.

Kuklev, Nikita [Fermilab]

Adaptive Cybersecurity for Distributed Energy Resources (AdCyDER): Online Reinforcement Learning with Stackelberg-Optimized Defenses — Pipeline Architecture, Evaluation Methodology, and Findings from a Synthetic-Data Evaluation

This report documents the design and evaluation of an integrated online-learning pipeline developed within the AdCyDER project for Distributed Energy Resource (DER) cybersecurity. The pipeline couples a Reinforcement Learning (RL) attack classifier — which produces an attack-type probability distribution — with a Stackelberg game-theoretic (GT) defense selector that consumes those distributions alongside SME-encoded priors over (defense, attack) effectiveness pairings and perdefense costs to choose grid-health-preserving defenses. The objective is not attack classification per se but production of distributions that drive effective defense selection through the Stackelberg layer, learned from delayed grid-health feedback rather than labeled attack data. AdCyDER as a whole is broader than the work presented here; this report covers the specific RL/GT loop integration and its evaluation. We present the integrated pipeline (SCADA telemetry with Fronius inverter physics, Suricata IDS, time-windowed aggregation, per-facility LSTM classifier, Stackelberg optimizer, OpenC2 actuators), an experimental campaign of 28 eight-hour iterations across three baseline modes, and a pipeline-ordered diagnostic protocol. The protocol identifies two distinct failure modes within the loop: paired supervised ceilings on the same features establish that the deployed online RL classifier (macro F1 ≈ 0.07) sits at least 4.7× below a same-architecture supervised LSTM (≈ 0.34) and 10–11× below a linear feature-signal ceiling (≈ 0.70–0.79 depending on per-facility isolation), localizing the dominant failure to the training procedure; and the reward signal driving online updates carries weak directional coupling with classifier correctness in the methodology-expected direction (multi-lens convergent: top-decile P(true) records produce more frequent state changes and slightly larger improvements, top-vs-bot Cohen’s 𝑑 ≈ −0.19), but at effect magnitudes too small to drive gradient-based learning at the campaign sample size. The original learning hypothesis is not supported by the data. The primary contributions are the diagnostic methodology — proposed as a transferable falsification protocol for online RL/GT defense pipelines learning from delayed environmental reward — and the open, reproducible experimental infrastructure. We outline reward reformulation as the highest-priority aspirational next step given the underpowered-but-aligned Q6 reading, with hardware-in-the-loop evaluation as the broadest scope-expansion option.

Blakely, Benjamin [Argonne National Laboratory (AN

ARM SGP PBLH and MLH datasets from Raman lidar and Doppler lidar

The planetary boundary layer (PBL) plays a critical role in the atmosphere by transferring heat, moisture, and momentum. The warm PBL has a distinct diurnal cycle including the daytime convective mixing layer (ML) and nighttime residual layer developments. Thus, simultaneous determinations of PBL height (PBLH) and ML height (MLH) are necessary for studying PBL characterization and processes. Here, new approaches are developed to provide reliable PBLH and MLH estimates to characterize warm PBL evolution. The approaches use Raman lidar (RL) water vapor mixing ratio (WVMR) and Doppler lidar (DL) vertical velocity measurements at the Southern Great Plains (SGP) atmospheric observatory, which was established by the Atmospheric Radiation Measurement (ARM) User Facility. Compared to widely used lidar aerosol measurements for PBLH, WVMR is a better tracer for PBL vertical mixing. For PBLH, the approach classifies PBL water vapor structures into a few general patterns, then uses a slope method and dynamic threshold method to determine PBLH. For MLH, wavelet analysis is used to reconstruct 2D variance from DL vertical wind velocity measurements according to the turbulence eddy size to minimize the impacts of gravity wave and eddy size on variance calculations; then, a dynamic threshold method is used to determine MLH. Remotely-sensed PBLHs and MLHs are compared with radiosonde measurements based on the Richardson number method. Good agreements between them confirm that the proposed new algorithms are reliable for PBLH and MLH characterization. The algorithms are applied to warm-season RL and ML measurements at the SGP site for five years to study warm-season PBL structure and processes. The weekly composited diurnal evolutions of PBLHs and MLHs in a warm climate were provided to illustrate diurnal and seasonal PBL evolutions. This reliable data set of PBLH and MLH values will be valuable for studying PBL processes, model evolution, and PBL parameterization improvements. The MLH dataset includes the MLH in values of km above ground level. The PBLH dataset includes the PBLH in values of km above ground level, along with a flag ("situation_PBLH") to determine the state of the PBL (1 = Cloudy Condition, 2 = Stable Layer, 3 = Multi-layer WVMR structure, 4 = Well-Mixed PBL, 5 = A de-coupled layer, 6 = Other).

mixing layer height

Photographic Study of Liquid-Oxygen Boiling and Gas Injection in the Injector of a Chugging Rocket Engine

High-speed motion pictures were taken of conditions in the injector liquid-oxygen cavity of an RL-10 rocket engine during throttled engine operation. Photographs were taken during operation of the engine in the chugging region as the helium gas was injected to stabilize combustion, during operation at rated thrust, and during transition into chugging conditions as the gas injection was discontinued. Results of the investigation indicate that, during chugging rocket operation of the RL-10 engine, a high population of fairly large bubbles formed and collapsed within the liquid-oxygen cavity at the same frequency as the chamber pressure oscillations. When gaseous helium was injected into the liquid-oxygen cavity, a fog rapidly spread over the entire field of view, and the system immediately became stable. The injection of gaseous helium at rated conditions produced a very slight increase in engine performance but not enough to produce a net gain in a typical mission payload with the extra equipment needed. The inherent low-frequency system instability associated with the fuel system at low thrust levels was reduced by injecting either gaseous helium or hydrogen. Complete stabilization was achieved in some cases, and a reduction in the severity of the oscillations in others. This was apparently due to the ·anchoring of the phase change front to the location of the gas injection.

Conrad, E. William

Design study of RL10 derivatives. Volume 2: Engine design characteristics

The design characteristics of the RL-10 rocket engine are discussed. The results from critical elements evaluation, baseline engine design, parametric and special study tasks are presented. Critical element evaluation established the feasibility of various engine features such as tank head idle, pumped idle, autogenous tank pressurization, and two-phase pumping. Three baseline engines, derived from the RL-10 were conceptually designed. Parametric life and performance data were generated. Special studies were conducted to establish the impact on the engine design of environment, safety, interchangeability, and maintenance.

Adams, A.

Orbit transfer vehicle engine study. Volume 2: Technical report

The orbit transfer vehicle (OTV) engine study provided parametric performance, engine programmatic, and cost data on the complete propulsive spectrum that is available for a variety of high energy, space maneuvering missions. Candidate OTV engines from the near term RL 10 (and its derivatives) to advanced high performance expander and staged combustion cycle engines were examined. The RL 10/RL 10 derivative performance, cost and schedule data were updated and provisions defined which would be necessary to accommodate extended low thrust operation. Parametric performance, weight, envelope, and cost data were generated for advanced expander and staged combustion OTV engine concepts. A prepoint design study was conducted to optimize thrust chamber geometry and cooling, engine cycle variations, and controls for an advanced expander engine. Operation at low thrust was defined for the advanced expander engine and the feasibility and design impact of kitting was investigated. An analysis of crew safety and mission reliability was conducted for both the staged combustion and advanced expander OTV engine candidates.

Source record

Lunar Transportation Facilities and Operations Study, option 2

During the Option 2 period of the Lunar Transportation Facilities and Operations Study (LTFOS), a joint McDonnell Douglas Space Systems Company Kennedy Space Center (MDSSC-KSC) and National Aeronautics and Space Administration Kennedy Space Center (NASA-KSC) Study team conducted a comparison of the functional testing of the RL-10 and Space Shuttle Main Engine, a quick-look impact assessment of the Synthesis Group Report, and a detailed assessment of the Synthesis Group Report. The results of these KSC LTFOS team efforts are included. The most recent study task effort was a detailed assessment of the Synthesis Group Report. The assessment was conducted to determine the impact on planetary launch and landing facilities and operations. The result of that effort is a report entitled 'Analysis of the Synthesis Group Report, its Architectures and their Impacts on PSS Launch and Landing Operations' and is contained in Appendix A. The report is structured in a briefing format with facing pages as opposed to a narrative style. A quick-look assessment of the Synthesis Group Report was conducted to determine the impact of implementing the recommendations of the Synthesis Group on KSC launch facilities and operations. The data was documented in a presentation format as requested by Kennedy Space Center Technology and Advanced Projects Office and is included in Appendix B. Appendix C is a white paper on the comparison of the functional testing of the RL-10 and Space Shuttle Main Engine. The comparison was undertaken to provide insight regarding common test requirements that would be applicable to Lunar and Mars Excursion Vehicles (LEV and MEV).

Source record

Propagation of Electrical Excitation in a Ring of Cardiac Cells: A Computer Simulation Study

The propagation of electrical excitation in a ring of cells described by the Noble, Beeler-Reuter (BR), Luo-Rudy I (LR I), and third-order simplified (TOS) mathematical models is studied using computer simulation. For each of the models it is shown that after transition from steady-state circulation to quasi-periodicity achieved by shortening the ring length (RL), the action potential duration (APD) restitution curve becomes a double-valued function and is located below the original ( that of an isolated cell) APD restitution curve. The distributions of APD and diastolic interval (DI) along a ring for the entire range of RL corresponding to quasi-periodic oscillations remain periodic with the period slightly different from two RLs. The 'S' shape of the original APD restitution curve determines the appearance of the second steady-state circulation region for short RLs. For all the models and the wide variety of their original APD restitution curves, no transition from quasi-periodicity to chaos was observed.

Kogan, B. Y.

Modeling Controller Tasks for Safety Analysis

As control systems become more complex, the use of automated control has increased. At the same time, the role of the human operator has changed from primary system controller to supervisor or monitor. Safe design of the human computer interaction becomes more difficult. In this paper, we present a visual task modeling language that can be used by system designers to model human-computer interactions. The visual models can be translated into SpecTRM-RL, a blackbox specification language for modeling the automated portion of the control system. The SpecTRM-RL suite of analysis tools allow the designer to perform formal and informal safety analyses on the task model in isolation or integrated with the rest of the modeled system.

Brown, Molly

Institutional Memory Preservation at NASA Glenn Research Center

In this era of downsizing and deficit reduction, the preservation of institutional memory is a widespread concern for U.S. companies and governmental agencies. The National Aeronautical and Space Administration faces the pending retirement of many of the agency's long-term, senior engineers. NASA has a marvelous long-term history of success, but the agency faces a recurring problem caused by the loss of these engineers' unique knowledge and perspectives on NASA's role in aeronautics and space exploration. The current work describes a knowledge elicitation effort aimed at demonstrating the feasibility of preserving the more personal, heuristic knowledge accumulated over the years by NASA engineers, as contrasted with the "textbook" knowledge of launch vehicles. Work on this project was performed at NASA Glenn Research Center and elsewhere, and focused on launch vehicle systems integration. The initial effort was directed toward an historic view of the Centaur upper stage which is powered by two RL-10 engines. Various experts were consulted, employing a variety of knowledge elicitation techniques, regarding the Centaur and RL-10. Their knowledge is represented in searchable Web-based multimedia presentations. This paper discusses the various approaches to knowledge elicitation and knowledge representation employed, and assesses successes and challenges in trying to perform large-scale knowledge preservation of institutional memory. It is anticipated that strategies for knowledge elicitation and representation that have been developed in this grant will be utilized to elicit knowledge in a variety of domains including the complex heuristics that underly use of simulation software packages such as that being explored in the Expert System Architecture for Rocket Engine Numerical Simulators.

Coffey, J.