Search NASA⌕ Search

SEARCH · Search NASA

Results for “multiple agents”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 235 records · Page 13

Collaborative Resource Allocation

Collaborative Resource Allocation Networking Environment (CRANE) Version 0.5 is a prototype created to prove the newest concept of using a distributed environment to schedule Deep Space Network (DSN) antenna times in a collaborative fashion. This program is for all space-flight and terrestrial science project users and DSN schedulers to perform scheduling activities and conflict resolution, both synchronously and asynchronously. Project schedulers can, for the first time, participate directly in scheduling their tracking times into the official DSN schedule, and negotiate directly with other projects in an integrated scheduling system. A master schedule covers long-range, mid-range, near-real-time, and real-time scheduling time frames all in one, rather than the current method of separate functions that are supported by different processes and tools. CRANE also provides private workspaces (both dynamic and static), data sharing, scenario management, user control, rapid messaging (based on Java Message Service), data/time synchronization, workflow management, notification (including emails), conflict checking, and a linkage to a schedule generation engine. The data structure with corresponding database design combines object trees with multiple associated mortal instances and relational database to provide unprecedented traceability and simplify the existing DSN XML schedule representation. These technologies are used to provide traceability, schedule negotiation, conflict resolution, and load forecasting from real-time operations to long-range loading analysis up to 20 years in the future. CRANE includes a database, a stored procedure layer, an agent-based middle tier, a Web service wrapper, a Windows Integrated Analysis Environment (IAE), a Java application, and a Web page interface.

Wang, Yeou-Fang↗

Software for Partly Automated Recognition of Targets

The Feature Analyst is a computer program for assisted (partially automated) recognition of targets in images. This program was developed to accelerate the processing of high-resolution satellite image data for incorporation into geographic information systems (GIS). This program creates an advanced user interface that embeds proprietary machine-learning algorithms in commercial image-processing and GIS software. A human analyst provides samples of target features from multiple sets of data, then the software develops a data-fusion model that automatically extracts the remaining features from selected sets of data. The program thus leverages the natural ability of humans to recognize objects in complex scenes, without requiring the user to explain the human visual recognition process by means of lengthy software. Two major subprograms are the reactive agent and the thinking agent. The reactive agent strives to quickly learn the user s tendencies while the user is selecting targets and to increase the user s productivity by immediately suggesting the next set of pixels that the user may wish to select. The thinking agent utilizes all available resources, taking as much time as needed, to produce the most accurate autonomous feature-extraction model possible.

Opitz, David↗

Collectives for Multiple Resource Job Scheduling Across Heterogeneous Servers

Efficient management of large-scale, distributed data storage and processing systems is a major challenge for many computational applications. Many of these systems are characterized by multi-resource tasks processed across a heterogeneous network. Conventional approaches, such as load balancing, work well for centralized, single resource problems, but breakdown in the more general case. In addition, most approaches are often based on heuristics which do not directly attempt to optimize the world utility. In this paper, we propose an agent based control system using the theory of collectives. We configure the servers of our network with agents who make local job scheduling decisions. These decisions are based on local goals which are constructed to be aligned with the objective of optimizing the overall efficiency of the system. We demonstrate that multi-agent systems in which all the agents attempt to optimize the same global utility function (team game) only marginally outperform conventional load balancing. On the other hand, agents configured using collectives outperform both team games and load balancing (by up to four times for the latter), despite their distributed nature and their limited access to information.

Tumer, K.↗

Cloning the promoter for transforming growth factor-beta type III receptor. Basal and conditional expression in fetal rat osteoblasts

Transforming growth factor-beta binds to three high affinity cell surface molecules that directly or indirectly regulate its biological effects. The type III receptor (TRIII) is a proteoglycan that lacks significant intracellular signaling or enzymatic motifs but may facilitate transforming growth factor-beta binding to other receptors, stabilize multimeric receptor complexes, or segregate growth factor from activating receptors. Because various agents or events that regulate osteoblast function rapidly modulate TRIII expression, we cloned the 5' region of the rat TRIII gene to assess possible control elements. DNA fragments from this region directed high reporter gene expression in osteoblasts. Sequencing showed no consensus TATA or CCAAT boxes, whereas several nuclear factors binding sequences within the 3' region of the promoter co-mapped with multiple transcription initiation sites, DNase I footprints, gel mobility shift analysis, or loss of activity by deletion or mutation. An upstream enhancer was evident 5' proximal to nucleotide -979, and a silencer region occurred between nucleotides -2014 and -2194. Glucocorticoid sensitivity mapped between nucleotides -687 and -253, whereas bone morphogenetic protein 2 sensitivity co-mapped within the silencer region. Thus, the TRIII promoter contains cooperative basal elements and dispersed growth factor- and hormone-sensitive regulatory regions that can control TRIII expression by osteoblasts.

Non-NASA Center↗

Thermal Interface Evaluation of Heat Transfer from a Pumped Loop to Titanium-Water Thermosyphons

Titanium-water thermosyphons are being considered for use in the heat rejection system for lunar outpost fission surface power. Key to their use is heat transfer between a closed loop heat source and the heat pipe evaporators. This work describes laboratory testing of several interfaces that were evaluated for their thermal performance characteristics, in the temperature range of 350 to 400 K, utilizing a water closed loop heat source and multiple thermosyphon evaporator geometries. A gas gap calorimeter was used to measure heat flow at steady state. Thermocouples in the closed loop heat source and on the evaporator were used to measure thermal conductance. The interfaces were in two generic categories, those immersed in the water closed loop heat source and those clamped to the water closed loop heat source with differing thermal conductive agents. In general, immersed evaporators showed better overall performance than their clamped counterparts. Selected clamped evaporator geometries offered promise.

Jaworske, Donald A.↗

Electric Vehicle Charging Demand in the Chicago Metropolitan Area through 2030

This report outlines the collaborative efforts between Argonne National Laboratory and Exelon in advancing the Agent-Based Transportation Energy Analysis Model (ATEAM). Aligning with ComEd’s beneficial electrification plan, this study developed eight scenarios to access the temporal and spatial distribution of charging load and demand stemming from the widespread adoption of battery electric vehicles (BEV) adoption, augmented public charging infrastructure deployment, and increased multi-unit dwelling (MUD) charging availability. Enhancements to the ATEAM model encompassed the simulation of multiple days of travel behavior, estimation of total public charging infrastructure needs, user interface refinements, and output tracking at both vehicle and charging station levels. The total electricity consumption for residential and public charging to support over 800,000 BEVs in Chicago in 2029 is projected at approximately 10.2 GWh. Enhanced MUD home charging accessibility (70%) amplifies the home charging load in the study area by 1.5% compared to the baseline scenario (10%). The widespread adoption of BEVs reduces peak charging loads, owing to their inclusion across households with diverse income levels, thus fostering a more dispersed charging activity pattern. However, widespread BEV adoption increases the peak home charging load in areas with lower median household incomes, reflecting a higher BEV concentration in these locales and, subsequently, heightened peak charging demands. In the Widespread BEV adoption scenario, fewer census tracts exhibit elevated peak loads for combined home and public charging, indicating a more even distribution of charging demand across the study area. Predominantly, peak loads for combined charging—both home and public— occur between 2 p.m. and 10 p.m. across all scenarios, encompassing the majority of census tracts.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗

A dual-input nonlinear system analysis of autonomic modulation of heart rate

Linear analyses of fluctuations in heart rate and other hemodynamic variables have been used to elucidate cardiovascular regulatory mechanisms. The role of nonlinear contributions to fluctuations in hemodynamic variables has not been fully explored. This paper presents a nonlinear system analysis of the effect of fluctuations in instantaneous lung volume (ILV) and arterial blood pressure (ABP) on heart rate (HR) fluctuations. To successfully employ a nonlinear analysis based on the Laguerre expansion technique (LET), we introduce an efficient procedure for broadening the spectral content of the ILV and ABP inputs to the model by adding white noise. Results from computer simulations demonstrate the effectiveness of broadening the spectral band of input signals to obtain consistent and stable kernel estimates with the use of the LET. Without broadening the band of the ILV and ABP inputs, the LET did not provide stable kernel estimates. Moreover, we extend the LET to the case of multiple inputs in order to accommodate the analysis of the combined effect of ILV and ABP effect on heart rate. Analyzes of data based on the second-order Volterra-Wiener model reveal an important contribution of the second-order kernels to the description of the effect of lung volume and arterial blood pressure on heart rate. Furthermore, physiological effects of the autonomic blocking agents propranolol and atropine on changes in the first- and second-order kernels are also discussed.

NASA Discipline Regulatory Physiology↗

Autonomy Architectures for a Constellation of Spacecraft

Until the past few years, missions typically involved fairly large expensive spacecraft. Such missions have primarily favored using older proven technologies over more recently developed ones, and humans controlled spacecraft by manually generating detailed command sequences with low-level tools and then transmitting the sequences for subsequent execution on a spacecraft controller. This approach toward controlling a spacecraft has worked spectacularly on previous missions, but it has limitations deriving from communications restrictions - scheduling time to communicate with a particular spacecraft involves competing with other projects due to the limited number of deep space network antennae. This implies that a spacecraft can spend a long time just waiting whenever a command sequence fails. This is one reason why the New Millennium program has an objective to migrate parts of mission control tasks onboard a spacecraft to reduce wait time by making spacecraft more robust. The migrated software is called a "remote agent" and has 4 components: a mission manager to generate the high level goals, a planner/scheduler to turn goals into activities while reasoning about future expected situations, an executive/diagnostics engine to initiate and maintain activities while interpreting sensed events by reasoning about past and present situations, and a conventional real-time subsystem to interface with the spacecraft to implement an activity's primitive actions. In addition to needing remote planning and execution for isolated spacecraft, a trend toward multiple-spacecraft missions points to the need for remote distributed planning and execution. The past few years have seen missions with growing numbers of probes. Pathfinder has its rover (Sojourner), Cassini has its lander (Huygens), and the New Millenium Deep Space 3 (DS3) proposal involves a constellation of 3 spacecraft for interferometric mapping. This trend is expected to continue to progressively larger fleets. For example, one mission proposed to succeed DS3 would have 18 spacecraft flying in formation in order to detect earth-sized planets orbiting other stars. A proposed magnetospheric constellation would involve 5 to 500 spacecraft in Earth orbit to measure global phenomena within the magnetosphere. This work describes and compares three autonomy architectures for a system that continuously plans to control a fleet of spacecraft using collective mission goals instead of goals or command sequences for each spacecraft. A fleet of self-commanding spacecraft would autonomously coordinate itself to satisfy high level science and engineering goals in a changing partially-understood environment making feasible the operation of tens or even a hundred spacecraft (such as for interferometry or plasma physics missions). The easiest way to adapt autonomous spacecraft research to controlling constellations involves treating the constellation as a single spacecraft. Here one spacecraft directly controls the others as if they were connected. The controlling "master" spacecraft performs all autonomy reasoning, and the slaves only have real-time subsystems to execute the master's commands and transmit local telemetry/observations. The executive/diagnostics module starts actions and the master's real-time subsystem controls the action either locally or remotely through a slave. While the master/slave approach benefits from conceptual simplicity, it relies on an assumption that the master spacecraft's executive can continuously monitor the slaves' real-time subsystems, and this relies on high-bandwidth highly-reliable communications. Since unintended results occur fairly rarely, one way to relax the bandwidth requirements involves only monitoring unexpected events in spacecraft. Unfortunately, this disables the ability to monitor for unexpected events between spacecraft and leads to a host of coordination problems among the slaves. Also, failures in the communications system can result in losing slaves. The other two architectures improve robustness while reducing communications by progressively distributing more of the other three remote agent components across the constellation. In a teamwork architecture, all spacecraft have executives and real-time subsystems - only the leader has the planner/scheduler and mission manager. Finally, distributing all remote agent components leads to a peer-to-peer approach toward constellation control.

Barrett, Anthony↗

Simulating nationwide coupled disease and fear spread in an agent-based model

Human cognitive responses, behavioral responses, and disease dynamics co-evolve over the course of any disease outbreak, and can result in complex feedbacks. We present a dynamic agent-based model that explicitly couples the spread of disease with the spread of fear surrounding the disease, implemented within the EpiCast simulation framework. EpiCast models transmission within a realistic synthetic population, capturing individual-level interactions. In our model, fear propagates through both in-person contact and broadcast media, prompting individuals to adopt protective behaviors that reduce disease spread. In order to better understand these coupled dynamics, we create and compare a range of compartmental models to ensure that introducing additional disease states does not prevent the emergence of multiple waves in these simpler models. Additionally, we compare a range of behavioral scenarios within EpiCast, varying the level and intensity of fear and behavior change. Our results show that the addition of asymptomatic, exposed, and pre-symptomatic disease states can impact both the rate at which an outbreak progresses and its overall trajectory in compartmental models. In EpiCast, the combination of non-local fear spread via broadcasters and strong behavioral responses by fearful individuals generally leads to multiple epidemic waves, an outcome that occurs only within a narrow parameter range when fear spreads purely through local contact. Accounting for the coupled spread of fear and disease is critical for understanding disease dynamics and designing timely, targeted responses to emerging infectious threats.

60 APPLIED LIFE SCIENCES↗

ISS Operations Cost Reductions Through Automation of Real-Time Planning Tasks

In 2008 the Johnson Space Center s Mission Operations Directorate (MOD) management team challenged their organization to find ways to reduce the costs of International Space station (ISS) console operations in the Mission Control Center (MCC). Each MOD organization was asked to identify projects that would help them attain a goal of a 30% reduction in operating costs by 2012. The MOD Operations and Planning organization responded to this challenge by launching several software automation projects that would allow them to greatly improve ISS console operations and reduce staffing and operating costs. These projects to date have allowed the MOD Operations organization to remove one full time (7 x 24 x 365) ISS console position in 2010; with the plan of eliminating two full time ISS console support positions by 2012. This will account for an overall 10 EP reduction in staffing for the Operations and Planning organization. These automation projects focused on utilizing software to automate many administrative and often repetitive tasks involved with processing ISS planning and daily operations information. This information was exchanged between the ground flight control teams in Houston and around the globe, as well as with the ISS astronaut crew. These tasks ranged from managing mission plan changes from around the globe, to uploading and downloading information to and from the ISS crew, to even more complex tasks that required multiple decision points to process the data, track approvals and deliver it to the correct recipient across network and security boundaries. The software solutions leveraged several different technologies including customized web applications and implementation of industry standard web services architecture between several planning tools; as well as a engaging a previously research level technology (TRL 2-3) developed by Ames Research Center (ARC) that utilized an intelligent agent based system to manage and automate file traffic flow, archiving f data, and generating console logs. This technology called OCAMS (OCA (Orbital Communication System) Management System), is now considered TRL level 9 and is in daily use in the Mission Control Center in support of ISS operations. These solutions have not only allowed for improved efficiency on console; but since many of the previously manual data transfers are now automated, many of the human error prone steps have been removed, and the quality of the planning products has improved tremendously. This has also allowed our Planning Flight Controllers more time to focus on the abstract areas of the job, (like the complexities of planning a mission for 6 international crew members with a global planning team), instead of being burdened with the administrative tasks that took significant time each console shift to process. The resulting automation solutions have allowed the Operations and Planning organization to realize significant cost savings for the ISS program through 2020 and many of these solutions could be a viable

Hall, Timothy A.↗

Infectious Disease risks associated with exposure to stressful environments

Multiple environmental factors asociated with space flight can increase the risk of infectious illness among crewmembers thereby adversely affecting crew health and mission success. Host defences can be impaired by multiple physiological and psychological stressors including: sleep deprivation, disrupted circadian rhythms, separation from family, perceived danger, radiation exposure, and possibly also by the direct and indirect effects of microgravity. Relevant human immunological data from isolated or stressful environments including spaceflight will be reviewed. Long-duration missions should include reliable hardware which supports sophisticated immunodiagnostic capabilities. Future advances in immunology and molecular biology will continue to provide therapeutic agents and biologic response modifiers which should effectively and selectively restore immune function which has been depressed by exposure to environmental stressors.

Meehan, Ichard T.↗

Swing Contract-Based Valuation for Distributed Energy Resources in Transactive Energy Systems: A Reinforcement Learning Approach

With the proliferation of distributed energy resources (DERs) and power grids with high fractions of renewable energy, market constructs are evolving to allow DERs to participate in multiple possible markets, at different levels of grid hierarchy. The effective participation of DERs in market environments is aided by swing contract-based pricing mechanisms, whereby DERs have a two-part compensation structure – one for their reservation/commitment and another for performancedriven ex-post payment for their actual mobilization during dispatch. In this paper, we propose a reinforcement learningbased (Q-learning) approach that allows a rational DER agent to select the market it wants to participate in within a composite market environment where individual markets are coordinated by possibly different actors. The proposed Q-learning framework aids DERs in their self-valuation by implicitly maximizing their own payoff through market participation, assuming a swing contract-based compensation structure. We complement our work through simulation-based investigations where factors affecting the DER decision making process, such as parametric uncertainties in market (and grid) environments, are studied.

Naqvi, Syed Ahsan Raza↗

Adaptive Reinforcement Learning (ARL) Control of a Multi-port Resonant Converter in UAV Systems

This study presents an adaptive reinforcement learning (ARL) control framework for a multi-port resonant converter used in hybrid unmanned aerial vehicle (UAV) power systems. The converter integrates high-frequency half-bridge input ports connected to a rectified engine–generator set and a battery energy storage system, along with a semi-bridgeless active rectifier supplying the propulsion load. A deep RL agent is trained to dynamically regulate inter-port phase-shift commands in real time based on flight conditions and load power demand. The ARL controller autonomously identifies phase-shift combinations that maximize conversion efficiency while maintaining stable and coordinated power flow, even under rapidly varying operating scenarios. This data-driven approach eliminates the need for explicit system modeling or extensive manual tuning and enables coordinated control among multiple power ports without inter-port communication. Experimental results validate that the ARL based strategy achieves reliable power sharing and consistently high-efficiency operation across diverse UAV operating conditions.

Asa, Erdem [ORNL] (ORCID:0000000190884812)↗

Fine-Water-Mist Multiple-Orientation-Discharge Fire Extinguisher

A fine-water-mist fire-suppression device has been designed so that it can be discharged uniformly in any orientation via a high-pressure gas propellant. Standard fire extinguishers used while slightly tilted or on their side will not discharge all of their contents. Thanks to the new design, this extinguisher can be used in multiple environments such as aboard low-gravity spacecraft, airplanes, and aboard vehicles that may become overturned prior to or during a fire emergency. Research in recent years has shown that fine water mist can be an effective alternative to Halons now banned from manufacture. Currently, NASA uses carbon dioxide for fire suppression on the International Space Station (ISS) and Halon chemical extinguishers on the space shuttle. While each of these agents is effective, they have drawbacks. The toxicity of carbon dioxide requires that the crew don breathing apparatus when the extinguishers are deployed on the ISS, and Halon use in future spacecraft has been eliminated because of international protocols on substances that destroy atmospheric ozone. A major advantage to the new system on occupied spacecraft is that the discharged system is locally rechargeable. Since the only fluids used are water and nitrogen, the system can be recharged from stores of both carried aboard the ISS or spacecraft. The only support requirement would be a pump to fill the water and a compressor to pressurize the nitrogen propellant gas. This system uses a gaseous agent to pressurize the storage container as well as to assist in the generation of the fine water mist. The portable fire extinguisher hardware works like a standard fire extinguisher with a single storage container for the agents (water and nitrogen), a control valve assembly for manual actuation, and a discharge nozzle. The design implemented in the proof-of-concept experiment successfully extinguished both open fires and fires in baffled enclosures.

Butz, James R.↗

Localization of Ad-Hoc Lunar Constellations in Communication Failure Modes for Distributed Spacecraft Autonomy

As Lunar missions increase in complexity, inspired by NASA’s Artemis Program, they will require reliable and sufficient Position, Navigation, and Timing (PNT) capability to support the upcoming Lunar users. The navigation service should also be compatible with the smaller platforms, like CubeSats, being sent by the public and private sectors. A non-dedicated, ad-hoc Lunar navigation constellation can provide PNT services on-demand using the non-dedicated swarm assets. Swarm members cooperatively and autonomously localize themselves with minimal interaction from Earth, freeing up valuable bandwidth and ground segment resources. The autonomous localization of Lunar constellations utilizes neighbor two-way intersatellite link (ISL) measurements in a distributed extended Kalman filter (DEKF) system to minimize operating costs. Because the decentralized Lunar PNT system relies on relay communication amongst the agents, network failures or loss of assets among ad-hoc Lunar constellations may impact localization performance. This study presents an evaluation of localization performance under increasing levels of network degradation. A simulation of an ad-hoc Lunar PNT swarm is augmented to include system faults and the impacts of intermittent and permanent failures on localization performance are evaluated. We investigate three potential causes of network degradation: single spacecraft loss, multiple spacecraft loss, and antenna failure. The numerical assessments from the simulation show that the LPNT system under study, based on an autonomous decentralized concept of operation, is highly robust and resilient to communication failures. Minor faults, such as single spacecraft loss, solar interference, technical malfunctions, message delays, and antenna outages, have minimal impact on state estimation, with only a 4.47% and 3.75% degradation in median position error for assets and a representative ground user, respectively, compared to an ideal communication scenario. However, major faults, such as hardware failures or meteor strikes leading to the loss of multiple spacecrafts, are more concerning. The permanent loss of three spacecraft results in a more severe performance degradation, with median position error increasing by 23.3% for assets and 11.7% for a representative ground user, despite the Lunar PNT system remaining functional.

Yeji Kim↗

Piezoelectric impedance-based high-accuracy damage identification using sparsity conscious multi-objective optimization inverse analysis

Two elements are essential in structural health monitoring utilizing dynamic responses: response measurement with high-frequency contents, i.e., small characteristic wavelengths, that can adequately reflect damage features, and effective inverse identification analysis that is however oftentimes under-determined. The advancement of smart structure integration has led to active interrogation through frequency-sweeping piezoelectric impedance measurement at high frequency range. In this research we develop a multi-objective optimization formulation for the identification of damage location and severity utilizing piezoelectric impedance. While one optimization objective is to match the response measurement with finite element model prediction in the damage parametric space, the other is the number of locations of damage, i.e., the sparsity of damage index as the solution vector, since damage usually occurs within a small number of locations. This multi-objective formulation fits well the under-determined nature of damage identification, as it naturally provides multiple solutions as basis for further elucidation. The challenge remaining is how to find a small solution set that can include the actual damage scenario. Here we develop a novel inverse identification framework utilizing the intelligent swarm optimizer which possesses flexibility for enhancement. We first embed a sparsity enforcement process into the population generation of the optimizer, which yields a solution repository intrinsically possessing sparsity. We then apply reinforcement learning so the agents can adaptively opt for local strategies with the aim of enriching the searching patterns to diversify the solutions. Through the incorporation of a Q-table, searching toward more promising directions will be rewarded. Our case analyses employing experimental data indicate that this sparsity-conscious multi-objective particle swarm optimization technique can lead to a small solution set which generally encompasses the true damage scenario. This effectively solves the structural damage identification problem with piezoelectric impedance measurement.

Yang Zhang↗

Analytic solutions for single and multiple cylinders of gravitating polytropes in magnetostatic equilibrium

Exact analytic solutions for the static equilibrium of a gravitating plasma polytrope in the presence of magnetic fields are presented. The means of generating various equilibrium configurations to illustrate directly the complex physical relationships between pressure, magnetic fields, and gravity in self-gravitating systems is demonstrated. One of the solutions is used to model interstellar clouds suspended by magnetic fields against the galactic gravity such as may be formed by the Parker (1966) instability. It is concluded that the pinching effect of closed loops of magnetic fields in the clouds may be a dominant agent in further collapsing the clouds following their formation.

Lerche, I.↗

Design and optimization of a modular hydrogen-based integrated energy system to maximize revenue via nuclear-renewable sources

Here, this paper demonstrates a novel modular distributed framework that uses optimal energy-dispatching strategies to enable greater flexibility and profitability in nuclear-renewable integrated energy systems (NR-IES). Hydrogen is used as a commodity in this framework since its production can improve grid stability and system operational flexibility, decarbonize heavy industry, and create an additional revenue stream for electricity generators, particularly nuclear power plants with high operational expenses. The proposed solution addresses the challenges associated with merging multiple software and services from various domains by using functional mock-up units (FMU) to co-simulate diverse subsystems designed in various platforms. The tightly coupled integrated energy system (IES) is optimized to maximize revenue by utilizing the deep reinforcement learning (DRL) technique to make smart dispatching decisions based on variable electricity prices and the availability of renewable energy. Proximal policy optimization (PPO) algorithm is used in training and testing the DRL agent. Over a period of 120 days, the proposed hydrogen-based IES framework showed about 10% revenue boost compared to a non-hydrogen generating baseline IES while also providing an easily-adoptable framework which can help to improve the flexibility of future generation nuclear power plants.

29 ENERGY PLANNING, POLICY, AND ECONOMY↗