Search NASA⌕ Search

SEARCH · Search NASA

Results for “Reinforcement Learning”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 217 records · Page 12

Quantum Circuit Partitioning for Scalable Noise-Aware Quantum Circuit Re-Synthesis

Re-synthesis techniques are utilized to optimize the quantum circuit. To enable scalable re-synthesis a divide-and-conquer approach is adopted that partitions the circuit into smaller blocks, which are optimized independently. Several algorithms have been proposed to minimize the block number while maximizing the gate count of each block. However, they vary in their performance and may not yield the highest output fidelity. We propose a reinforcement learning-based quantum circuit partitioning framework that incorporates the physical properties of the quantum hardware to maximize the output fidelity post-quantum circuit optimization. To accelerate the training, we also propose a noise injection method that enables on-the-fly optimization in the reinforcement learning environment, independent of the adopted optimization/re-synthesis method at the block level. We evaluate our approach compared to different partitioning techniques using various quantum benchmarks executed on IBM Q Hanoi quantum computer.

Charrwi, Mohammad Walid↗

Learning-Based Building Flexibility Estimation and Control to Improve Microgrid Economics and Resilience: Preprint

This paper proposes a learning-based building flexibility estimation and control framework to improve system economics and resilience. A data-driven building load flexibility model consisting of weather forecasting and estimating load consumption is proposed to quantify building heating, ventilation, and air conditioning (HVAC) load flexibility. A reinforcement learning-based microgrid controller is proposed to dispatch distributed generators, distributed energy resources, and build HVAC loads while taking flexibility information as one of the inputs. Simulation analysis is conducted on the model of a real microgrid in California. The effectiveness of the proposed learning-based building flexibility estimation and control in reducing microgrid energy costs and improving the sustainability of critical loads is demonstrated.

building load flexibility↗

Integrating Crack Detection and Pipe Shape Optimization for Enhanced Sewage System Durability

Crack detection in underground reinforced concrete pipes has been essential in determining the state of stormwater infrastructure. Detection models have been implemented for detecting cracks and other defects in pipes using CCTV footage for stormwater drainage systems. In addition, Finite element models have been used to determine optimum shapes and pipe thickness for different boundary conditions such as header pipes in power plants. The concept of shape optimization emerges as a crucial factor in power plant design and operation, with the potential to maximize performance while minimizing the use of materials. Shape optimization not only enhances efficiency but also contributes to reducing the environmental footprint. This paper discusses the integration of both topics by using the cracks detected in underground pipes as boundary conditions for shape optimization of the pipes. A machine learning model has been developed which uses limited data for training and outlines the location of detected cracks. A shape optimization methodology is proposed in which ANSYS modules are used to analyze fluid flow and then optimize the shape of the pipe. The crack detection model developed has been applied to a crack detected in lab setting and machine learning model used has an accuracy of 98% using a random forest algorithm.

20 FOSSIL-FUELED POWER PLANTS↗

Solving high-dimensional partial integral differential equations: The finite expression method

Partial integro-differential equations (PIDEs) have broad applications in the sciences, from electro-magnetism to options pricing. Here, in this paper, we introduce a new finite expression method (FEX) to solve PIDEs. This approach builds upon the original FEX and its inherent advantages with new advances: 1) A novel method of parameter grouping is proposed to reduce the number of coefficients in high-dimensional function approximation; 2) A Taylor series approximation method is implemented to significantly improve the computational efficiency and accuracy of the evaluation of the integral terms of PIDEs. The new FEX based method, denoted FEX-PG to indicate the addition of the parameter grouping (PG) step to the algorithm, provides both high accuracy and interpretable numerical solutions, with the outcome being an explicit equation that facilitates intuitive understanding of the underlying solution structures. These features are often absent in traditional methods, such as finite element methods (FEM) and finite difference methods, as well as in deep learning-based approaches. To benchmark our method against recent advances, we apply the new FEX-PG to solve benchmark PIDEs in the literature. In high-dimensional settings, FEX-PG exhibits strong and robust performance, achieving relative errors on the order of single precision machine epsilon, significantly outperforming existing approaches based on neural networks.

Combinatorial optimization↗

Towards intelligent emergency control for large-scale power systems: Convergence of learning, physics, computing and control

Here, this paper has delved into the pressing need for intelligent emergency control in large-scale power systems, which are experiencing significant transformations and are operating closer to their limits with more uncertainties. Learning-based control methods are promising and have shown effectiveness for intelligent power system control. However, when they are applied to large-scale power systems, there are multifaceted challenges such as scalability, adaptiveness, and security posed by the complex power system landscape, which demand comprehensive solutions. The paper first proposes and instantiates a convergence framework for integrating power systems physics, machine learning, advanced computing, and grid control to realize intelligent grid control at a large scale. Our developed methods and platform based on the convergence framework have been applied to a large (more than 3000 buses) Texas power system, and tested with 56 000 scenarios. Our work achieved a 26% reduction in load shedding on average and outperformed existing rule-based control in 99.7% of the test scenarios. The results demonstrated the potential of the proposed convergence framework and DRL-based intelligent control for the future grid.

24 POWER TRANSMISSION AND DISTRIBUTION↗

Foveal vision reduces neural resources in agent-based game learning

Efficient processing of information is crucial for the optimization of neural resources in both biological and artificial visual systems. In this paper, we study the efficiency that may be obtained via the use of a fovea. Using biologically-motivated agents, we study visual information processing, learning, and decision making in a controlled artificial environment, namely the Atari Pong video game. We compare the resources necessary to play Pong between agents with and without a fovea. Our study shows that a fovea can significantly reduce the neural resources, in the form of number of neurons, number of synapses, and number of computations, while at the same time maintaining performance at playing Pong. To our knowledge, this is the first study in which an agent must simultaneously optimize its visual system, along with its decision making and action generation capabilities. That is, the visual system is integral to a complete agent.

60 APPLIED LIFE SCIENCES↗

Multi-agent voltage control in distribution systems using GAN-DRL-based approach

Active distribution grids can experience voltage fluctuations and violations due to the high penetration of variable distributed energy resources (DERs). These problems might occur because of the uncertain and variable generation natures of these resources, especially solar photovoltaic resources, during panel shadowing scenarios. Volt-VAR control (VVC) is an efficient method that controls the reactive power set-points of the inverters to regulate the voltage of distribution grids. Although several VVC approaches have been proposed recently, the performance of these approaches degrades significantly if behind-the-meter solar generation data are unobservable/missing. Therefore, it is necessary to impute missing/unobservable PV data accurately to be utilized in VVC approaches. Further, this paper proposes a model-free, data-driven, centrally trained, and decentrally executed multi-agent deep reinforcement learning-based VVC architecture to regulate the voltage of distribution networks. A generative adversarial network (GAN) is incorporated to impute the unobservable PV data accurately, which improves the performance of the proposed control architecture. The proposed multi-agent-soft-actor–critic algorithm (MASAC)-based VVC technique utilizes the actual PV dataset as well as the imputed dataset from the GAN framework to learn the optimal coordinated control policy for controlling the optimal reactive power set-points of PV inverters. The effectiveness of the proposed approach is analyzed on a modified IEEE 34-bus test case with added PV inverters. The results are compared and analyzed with a base case model with no VVC and VVC with a local droop control approach, genetic algorithm optimization, and a centralized soft actor–critic-based approach. Moreover, the performance of the proposed approach is compared with that of a multi-agent VVC framework without using the PV generation data and load information as the system state. The results illustrate that the proposed method with more state input improves the voltage profile and reduces the power loss of the network across various loading and PV generation scenarios.

14 SOLAR ENERGY↗

Agentic AI and the Cyber Arms Race

Here, in this article, we examine the implications for cyberwarfare and global politics as agentic artificial intelligence becomes more powerful and enables the broad proliferation of capabilities only available to the most well-resourced actors today.

Cybersecurity↗

Decision-making based on Markov decision process in integrated artificial reasoning framework—Part I: Theory

This paper presents a decision-making framework based on an integrated artificial reasoning framework and Markov decision process (MDP). The integrated artificial reasoning framework provides a physics-based approach that converts system information into state transition models, and the analysis result will be represented by the transition probabilities that can be used with an MDP to find a traceable and explainable optimal pathway. A dynamic Bayesian network (DBN) is well suited for representing the structure of an MDP. The causality information among process variables (or among subsystems) is mathematically represented in a DBN by the conditional probabilities of the node’s states provided different probabilities of the parent node’s states. To define node states in a physically understandable manner, we used multilevel flow modeling (MFM). An MFM follows the fundamental energy and mass conservation laws and supports the selection of process variables that represent the system of interest so that causal relations among process variables are properly captured. An MFM-based DBN supports developing state transition models in an MDP to capture the effect of process variables of system having physical relations. The operators of the target system can capture stochastic system dynamics as multiple subsystem state transitions based on their physical relations and uncertainties coming from component degradation or random failures. We analyzed a simplified exemplary system to illustrate an optimal operational policy using the suggested approach.

Markov decision process↗

Dynamic Model Development of a Wind Power Plant Using Neural Net Method to Forecast Wind Power Output (CRADA Final Report)

This project is intended to model wind power plant based on monitored data at the wind power plant. This project will promote the university research in Renewable Energy area and trains the future highly qualified engineers. The dynamic model will be based on neural net model with the input from the two met towers (12 inputs), and the number of turbines in operation (one input). The overall input will be 13 inputs to drive the simulations. The output power at the point of interconnection will be used to tune the neural net weight coefficients. Two neural net concepts will be investigated (the back propagation neural net and the dynamic recurrent neural net with feedback).

17 WIND ENERGY↗

NEXT Generation Energy Technologies for Connected and Automated On-Road Vehicles (NEXTCAR Phase I & II)

The Ohio State University’s ARPA-E NEXTCAR project was a multi-phase, multi-year research, development, and demonstration program focused on improving the energy efficiency of connected and automated vehicles (CAVs). The team developed and validated advanced vehicle motion and powertrain control algorithms that coordinate propulsion and automation systems to optimize energy use. Key technologies included Dynamic Skip Fire engine control, predictive eco-driving functions such as Eco-Approach and Departure (Eco-AND) and Eco-Adaptive Cruise Control (Eco-ACC), and powertrain-agnostic optimization frameworks for hybrid, plug-in hybrid, and battery electric vehicles. The project successfully demonstrated up to 30% energy-efficiency improvement during real-world testing at the Transportation Research Center and the American Center for Mobility. The outcomes provide a foundation for scalable, cost-effective deployment of energy-optimized CAV technologies across the automotive industry.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗

Reinforcement expectation in the honeybee ( Apis mellifera ): Can downshifts in reinforcement show conditioned inhibition?

When animals learn the association of a conditioned stimulus (CS) with an unconditioned stimulus (US), later presentation of the CS invokes a representation of the US. When the expected US fails to occur, theoretical accounts predict that conditioned inhibition can accrue to any other stimuli that are associated with this change in the US. Empirical work with mammals has confirmed the existence of conditioned inhibition. But the way it is manifested, the conditions that produce it, and determining whether it is the opposite of excitatory conditioning are important considerations. Invertebrates can make valuable contributions to this literature because of the well-established conditioning protocols and access to the central nervous system (CNS) for studying neural underpinnings of behavior. Nevertheless, although conditioned inhibition has been reported, it has yet to be thoroughly investigated in invertebrates. Here, we evaluate the role of the US in producing conditioned inhibition by using proboscis extension response conditioning of the honeybee (Apis mellifera). Specifically, using variations of a “feature-negative” experimental design, we use downshifts in US intensity relative to US intensity used during initial excitatory conditioning to show that an odorant in an odor–odor mixture can become a conditioned inhibitor. We argue that some alternative interpretations to conditioned inhibition are unlikely. However, we show variation across individuals in how strongly they show conditioned inhibition, with some individuals possibly revealing a different means of learning about changes in reinforcement. We discuss how the resolution of these differences is needed to fully understand whether and how conditioned inhibition is manifested in the honeybee, and whether it can be extended to investigate how it is encoded in the CNS. It is also important for extension to other insect models. In particular, work like this will be important as more is revealed of the complexity of the insect brain from connectome projects.

60 APPLIED LIFE SCIENCES↗

Integrating multi-modal remote sensing, deep learning, and attention mechanisms for yield prediction in plant breeding experiments

In both plant breeding and crop management, interpretability plays a crucial role in instilling trust in AI-driven approaches and enabling the provision of actionable insights. The primary objective of this research is to explore and evaluate the potential contributions of deep learning network architectures that employ stacked LSTM for end-of-season maize grain yield prediction. A secondary aim is to expand the capabilities of these networks by adapting them to better accommodate and leverage the multi-modality properties of remote sensing data. In this study, a multi-modal deep learning architecture that assimilates inputs from heterogeneous data streams, including high-resolution hyperspectral imagery, LiDAR point clouds, and environmental data, is proposed to forecast maize crop yields. The architecture includes attention mechanisms that assign varying levels of importance to different modalities and temporal features that, reflect the dynamics of plant growth and environmental interactions. The interpretability of the attention weights is investigated in multi-modal networks that seek to both improve predictions and attribute crop yield outcomes to genetic and environmental variables. This approach also contributes to increased interpretability of the model's predictions. The temporal attention weight distributions highlighted relevant factors and critical growth stages that contribute to the predictions. The results of this study affirm that the attention weights are consistent with recognized biological growth stages, thereby substantiating the network's capability to learn biologically interpretable features. Accuracies of the model's predictions of yield ranged from 0.82-0.93 R 2 ref in this genetics-focused study, further highlighting the potential of attention-based models. Further, this research facilitates understanding of how multi-modality remote sensing aligns with the physiological stages of maize. The proposed architecture shows promise in improving predictions and offering interpretable insights into the factors affecting maize crop yields, while demonstrating the impact of data collection by different modalities through the growing season. By identifying relevant factors and critical growth stages, the model's attention weights provide valuable information that can be used in both plant breeding and crop management. The consistency of attention weights with biological growth stages reinforces the potential of deep learning networks in agricultural applications, particularly in leveraging remote sensing data for yield prediction. To the best of our knowledge, this is the first study that investigates the use of hyperspectral and LiDAR UAV time series data for explaining/interpreting plant growth stages within deep learning networks and forecasting plot-level maize grain yield using late fusion modalities with attention mechanisms.

59 BASIC BIOLOGICAL SCIENCES↗

Physics-Informed Neural Network (PINN) Prediction of Mixed Mass-Heat-Crystallization Limited Methane Hydrate Formation and Dissociation in Micro-Confinement

The creation and use of Physics-Informed Neural Networks (PINNs) for simulating the dynamics of methane hydrate formation and dissociation will be presented. The PINN framework's main benefit is its capacity to impose physical consistency with only a partial comprehension of the governing equations. This makes the algorithm especially useful for systems with little experimental evidence or a lack of theoretical knowledge. A strong basis for forecasting methane hydrate behavior over the verified operating ranges of 30.0-80.9 bar pressure and 1.0-4.0 K sub-cooling conditions is provided by the combination of conductive heat transfer equations and mixed mass-transfer–crystallization kinetics. PINNs were more accurate at predicting the mixed mass-heat-crystallization limited kinetics than conventional Artificial Neural Networks (ANNs), demonstrating remarkable predictive accuracy for methane hydrate production over the ANN model. The efficiency of incorporating physical limitations from first principles into machine learning frameworks for methane hydrate crystallizations is reinforced by these findings. For hydrate-related applications in energy generation, carbon sequestration, and climate modelling, our study establishes PINNs as a computational tool that is both scalable and efficient. The proven capacity to close the gap between conventional physics-based simulations and solely data-driven models creates new opportunities for expedited hydrate research and practical applications.

Hartman, Ryan L [NYU Tandon School of Engineering]↗

Predicting Mechanical Properties from Microstructure Images in Fiber-Reinforced Polymers Using Convolutional Neural Networks

Evaluating the mechanical response of fiber-reinforced composites can be extremely time-consuming and expensive. Machine learning (ML) techniques offer a means for faster predictions via models trained on existing input–output pairs and have exhibited success in composite research. This paper explores a fully convolutional neural network modified from StressNet, which was originally used for linear elastic materials, and extended here for a non-linear finite element (FE) simulation to predict the stress field in 2D slices of segmented tomography images of a fiber-reinforced polymer specimen. The network was trained and evaluated on data generated from the FE simulations of the exact microstructure. The testing results show that the trained network accurately captures the characteristics of the stress distribution, especially on fibers, solely from the segmented microstructure images. The trained model can make predictions within seconds in a single forward pass on an ordinary laptop, given the input microstructure, compared to 92.5 h to run the full FE simulation on a high-performance computing cluster. These results show promise in using ML techniques to conduct fast structural analysis for fiber-reinforced composites and suggest a corollary that the trained model can be used to identify the location of potential damage sites in fiber-reinforced polymers.

Sun, Yixuan (ORCID:0000000311093380)↗

CRCNS US-France Research Proposal: Collaborative Research: Encoding reward expectation in Drosophilia

The fruit fly Drosophila melanogaster has been a valuable model for investigating the genetic and neural bases that underlie learning and memory. Early and most current studies use basic behavior conditioning protocols to study learning in controlled laboratory settings. More recently, the ability to transgenically manipulate many of the brain neurons in the fruit fly with exquisite specificity, and the recent knowledge of the synaptic ‘connectome’ of the fruit fly brain, makes these animals almost unique as a comprehensive model for studies of learning, memory and motivated behavior. In fact, the connectome has revealed many types of new connections that had until now been overlooked. Within this context, the thesis of this proposal is that studies of learning and memory will be greatly enhanced by using more sophisticated means for evaluating memory representations, such as have been developed in vertebrates, and combining those studies with information from the connectome guided by computational modelling. We propose to push beyond the boundaries of existing conditioning protocols for fruit flies to investigate more complex memory representations. In particular, we will investigate the function of reinforcement pathways in relation to the absence of expected reinforcement. More specifically, we propose a series of experiments designed to investigate the memory representations in fruit flies when an expected consequence of a Conditioned Stimulus (CS) fails to occur. Although studies have evaluated how this failure can establish extinction memory for the CS, our studies will go beyond studying extinction. Specifically, we predict that in Drosophila when a CS is associated with a failed expectation of an appetitive food reinforcement it will acquire aversive value, and vice versa for a failed expectation of an aversive reinforcer. We combine these studies with manipulations of reinforcement pathways in the CNS inspired from the connectome, iteratively knitted in with established computational models. Intellectual Merit: The concept of reinforcement expectation and incentive contrast have been influential in the development of studies of associative learning in mammals. These questions are particularly challenging to answer in vertebrates because they require exquisite cellular, temporal, and genetic specificity of experimental manipulations. The recent development of work with identified neurons and their connectomes makes the larval and adult fly brains ripe as models for pushing our understanding of neural bases for these higher- order conditioning phenomena. Broader Impacts: Public health: These analyses and the conceptual framework of prediction error processing underlying them have a profound impact on our understanding of reinforcement-related behavior in humans, including monetary rewards and the mnemonic consequences of traumatic experiences, and for pathologies of the dopamine reinforcement system. Educational: This project will provide interdisciplinary training for postdoctoral researchers, Ph.D. and undergraduate students. The PIs will act as co-supervisors or mentors of students working in the different labs via face-to-face and internet-based technologies. We will also work with ASU’s award-winning Ask- A-Biologist program. This is an online science program designed to enrich the learning experiences of students of all ages and to provide classroom material for use by K-12 teachers. We will develop an extension of a game developed under a prior NSF award, and the new game will include modules to teach K-12 students about how insects learn. We will also integrate into the AAB site a program developed by a collaborator (B Gerber) at the Leibniz Institut für Neurobiologie, Magdeburg, and now in use in schools in Germany, to teach K-12 students how to train animals using the fruit fly larval learning paradigm. Underrepresented groups: All PIs will work with their university offices of Academic Diversity and Equal Opportunity for reaching underrepresented students.

59 BASIC BIOLOGICAL SCIENCES↗