Search NASA⌕ Search

SEARCH · Search NASA

Results for “sequential decision making”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

Correlated Trajectory Uncertainty for Adaptive Sequential Decision Making

One of the great challenges with decision making tasks on real world systems is the fact that data is sparse and acquiring additional data is expensive. In these cases, it is often crucial to make a model of the environment to assist in making decisions. At the same time, limited data means that learned models are erroneous, making it just as important to equip the model with good predictive uncertainties. In the context of learning sequential decision making policies, these uncertainties can prove useful for informing which data to collect for the greatest improvement in policy performance \citep{mehta2021experimental, mehta2022exploration} or informing the policy about unsure regions of state and action space to avoid during test time \citep{yu2020mopo}. Additionally, assuming that realistic samples of the environment can be drawn, an adaptable policy can be trained that attempts to make optimal decisions for any given possible instance of the environment \citep{ghosh2022offline, chen2021offline}. In this work, we examine the so-called ``probabilistic neural network'' (PNN) model that is ubiquitous in model-based reinforcement learning (MBRL) works. We argue that while PNN models may have good marginal uncertainties, they form a distribution of non-smooth transition functions. Not only are these samples unrealistic and may hamper adaptability, but we also assert that this leads to poor uncertainty estimates when predicting multiple step trajectory estimates. To address this issue, we propose a simple sampling method that can be implemented on top of pre-existing models.We evaluate our sampling technique on a number of environments, including a realistic nuclear fusion task, and find that, not only do smooth transition function samples produce more calibrated uncertainties, but they also lead to better downstream performance for an adaptive policy.

Offline Reinforcement Learning↗

Scalable Bayesian optimization with randomized prior networks

Several fundamental problems in science and engineering consist of global optimization tasks involving unknown high-dimensional (black-box) functions that map a set of controllable variables to the outcomes of an expensive experiment. Bayesian Optimization (BO) techniques are known to be effective in tackling global optimization problems using a relatively small number objective function evaluations, but their performance suffers when dealing with high-dimensional outputs. To overcome the major challenge of dimensionality, here we propose a deep learning framework for BO and sequential decision making based on bootstrapped ensembles of neural architectures with randomized priors. Using appropriate architecture choices, we show that the proposed framework can approximate functional relationships between design variables and quantities of interest, even in cases where the latter take values in high-dimensional vector spaces or even infinite-dimensional function spaces. In the context of BO, we augmented the proposed probabilistic surrogates with re-parameterized Monte Carlo approximations of multiple-point (parallel) acquisition functions, as well as methodological extensions for accommodating black-box constraints and multi-fidelity information sources. We test the proposed framework against state-of-the-art methods for BO and demonstrate superior performance across several challenging tasks with high-dimensional outputs, including a constrained multi-fidelity optimization task involving shape optimization of rotor blades in turbo-machinery.

97 MATHEMATICS AND COMPUTING↗

Thermal Reservoir Networks for Modularly Expandable Thermal Microgrids

The Department of Defense (DoD) faces the substantial challenge of cost-effectively retrofitting one to two installations per month, each comprising approximately 1,000 buildings, to improve resilience, reduce energy consumption, and enhance energy supply security. Achieving these objectives requires optimal system selection and effective risk mitigation during system integration. To address this need, we introduce Platform-Based Design (PBD), a structured, hierarchical methodology adapted from other industrial sectors to the domain of energy system retrofits. We demonstrate the effectiveness of PBD through a techno-economic feasibility study comparing geothermal-coupled thermal energy networks (TENs) with conventional energy systems for heating, cooling, and powering 17 buildings at Joint Base Andrews (JBA) in Maryland. Our analysis illustrates that the PBD approach enables rigorous, data-driven, sequential decision making, resulting in a family of Pareto-optimal systems, among which the TEN emerged as the most promising solution. The selected TEN design integrates geothermal borefields, heat recovery heat pumps, photovoltaic (PV) arrays, and battery storage. Compared to the baseline system – gas heating combined with air-source chillers – the proposed TEN reduces annual imported energy by 74% and peak electricity demand by 45%, achieves a levelized cost of energy of $\$0.210$/kWh, and substantially enhances resilience. Life-cycle costs increase by approximately 6%, and initial investment costs are about 2.5 times higher than the baseline. However, if central plant infrastructure, district loops, and utility-scale PV and battery systems are privately funded and operated, the initial investment would fall below the baseline system cost. Critical to achieving these significant performance improvements were detailed nonlinear dynamic simulations coupling geothermal heat transfer, energy system operation, and realistic feedback control logic. These simulations identified essential design modifications and control strategy refinements that substantially reduced energy use, peak demand, and compressor shortcycling, thereby improving durability and reliability—issues that would have been significantly more expensive to resolve during operation. Additionally, the verification step highlighted sensitivities to key design parameters that could reduce initial investment by approximately $\$2$ million and reduce annual life-cycle costs more than $\$300,000$. We recommend adopting the PBD methodology for future feasibility studies and TEN pilot projects to gain valuable operational experience. Furthermore, we recommend that DoD invest in transferring and scaling the PBD methodology to other installations. This entails developing standardized computational frameworks and component libraries as well as training industry in conducting PBD. Such investments would enable rapid, robust, reliable, and cost-effective retrofits, supporting DoD’s ambitious energy system modernization goals.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗

A control-inspired approach for energy transition planning under uncertainty

As the global carbon footprint continues to grow, many countries are implementing carbon emission reduction policies which have incentivized the expansion of low-carbon and renewable technologies. However, the speed and scale of deployment falls short of that needed to meet climate goals. Energy system models serve as key tools for guiding investment decisions and helping policymakers evaluate the effects of various policies on the development of an energy system. This study focuses on the energy system of the United States and builds upon prior work by incorporating more geographic granularity to account for the trade of commodities and addresses transmission congestion through electricity price adjustments. Furthermore, real-world characteristics, such as delays in constructing new liquid fuel production and electricity generation facilities, are integrated using a sequential decision-making approach that better reflects how decisions can be updated as uncertainties unfold. Results demonstrate that stochastic programming combined with sequential decision-making produces energy transition pathways that are robust to multiple uncertain futures. Additionally, considering real-world characteristics significantly impacts the deployment of renewable technologies and the ability to meet carbon emission reduction goals while also reliably meeting demand. These findings highlight the importance of accounting for uncertainty and real-world characteristics to avoid overly optimistic projections in energy system planning.

energy systems↗

Energy efficiency in industrial drying: A hybrid ultrasonic system with a novel dynamic optimization framework

Drying processes are among the most energy-consuming operations in industrial and manufacturing settings, demanding strategic selection, design, and control for enhanced efficiency. Advancing drying technologies is critical for improving sustainability, lowering energy use, reducing carbon emissions, and minimizing waste. This study explores two innovative strategies aimed at transforming drying processes into sustainable, low-carbon systems by reducing energy consumption, minimizing waste, and maintaining a strong emphasis on preserving product quality. The first strategy showcases a sub-pilot scale hybrid ultrasonic-convective dryer for agrifood products. This technology, powered by electricity (process electrification), integrates non-thermal ultrasonic dehydration with convective heating and is presented as a sustainable and energy-efficient solution that enhances eco-friendly practices. The second strategy involves introducing and implementing a novel, multiobjective, mixed integer dynamic optimization technique to determine the optimal time-dependent process parameter values for the drying operation. This optimization technique yields operating conditions that are piecewise constant in time aiming to maximize the energy efficiency of the hybrid ultrasonic-convective dryer while ensuring strict adherence to product quality constraints. By adopting the hybrid ultrasonic-convective dryer, a notable 35% improvement in energy efficiency was achieved compared to conventional hot-air drying systems for drying apple slices. The proposed optimization framework further enhanced energy efficiency by nearly 14% over the most efficient process on the identical testbed, under static operating conditions. The reported enhancements have been experimentally validated. Regarding drying time (thereby improving production yield), the developed hybrid ultrasonic-convective dryer demonstrates as much as a 41% reduction in total processing time, which is further optimized by an additional 10% using our proposed optimization framework. The research outcomes have profound implications for the design and operation of drying systems, encompassing crucial aspects such as process electrification, cost-effectiveness, energy savings, time efficiency, product yield, product quality, and process automation.

Dynamic optimization↗

Deep Reinforcement Learning for Distribution System Operations: A Tutorial and Survey

Here, the rapid evolution of modern electric power distribution systems into complex networks of interconnected active devices, distributed generation (DG), and storage poses increasing difficulties for system operators. The large-scale integration of distributed energy resources (DERs) and the rapid exchange of measurement data via communication networks present major opportunities for advancing grid operations but also introduce greater uncertainty, higher data dimensionality, more complex network and device models, and challenging control and optimization problems. Deep reinforcement learning (DRL) algorithms are promising in addressing these challenges. However, they have not been effectively adapted for power systems applications, requiring extensive customization for implementation and evaluation. This has resulted in reproducibility challenges and a steep learning curve for researchers new to applying DRL algorithms to the power systems domain. To bridge these gaps, this tutorial aims to serve as a valuable resource for researchers interested in exploring learning-based algorithms to operate active power distribution networks. Specifically, this work presents a generalized process for translating sequential decision-making problems in power distribution systems into Markov decision process (MDP) formulations, illustrated through concrete grid service examples. Additionally, we introduce a simple environment design strategy to develop and evaluate example DRL algorithms for distribution system applications, complete with an included code repository to guide users through environment construction.

24 POWER TRANSMISSION AND DISTRIBUTION↗

Inverse Reinforcement Learning based Bayesian Goal Inference Method for Early Nuclear Proliferation Detection

Traditional methods for detection of nuclear proliferation indicators are usually applied after nuclear proliferation has already occurred. There is a need to advance these methods to perform early detection of nuclear proliferation indicators. In this project, we formulated an early detection problem as a sequential, decision-making, goal inference problem based on research publications of authors, to determine whether it is possible to infer whether an author will publish on a research activity before it has occurred. To develop and test our approach, we selected a civil nuclear activity for our case study. We constructed a state-action-state transition graph from publications of authors associated with the activity and the co-authors of their publications, using titles, abstracts, and author publication sequences. We then used inverse reinforcement learning to model the goal-directed behavior of authors in trajectories that terminate at selected goal states. Using a Bayesian formulation, we computed the probability that authors would reach each selected state from partially observed trajectories of their state transitions in their research topic space. The state with the highest probability was selected as the most probable goal state. Based on our results, we found that 60% of the times we can infer the correct goal state early; sometimes the inference is either delayed, or multiple states could be inferred as goal states. Overall, our results show that it is possible to perform early detection of research activities of authors in a nuclear technology area. Further research is necessary to establish a more accurate understanding of how topic modeling, topic space grid discretization, and the extent of overlap among trajectories of different goal states, affect the goal inference results. The methods developed in this work may be used to enhance data-driven methods for early detection of nuclear proliferation indicators.

97 MATHEMATICS AND COMPUTING↗

Early Inference of Nuclear Technology-Directed Research Activities of Authors from Scientific Publications

Nuclear research articles can provide information about early nuclear proliferation indicators such as influential research entities and technology capability levels of a country, but detection of nuclear activities typically occurs after they have started. We investigate the extent to which nuclear research articles can be used to infer whether a research entity will acquire or develop a nuclear technology before it happens. Early detection of nuclear proliferation or technology development indicators from data is challenging due to partial observability, sparse and unlabeled information, and confounding signals from multiple concurrent activities. This paper presents the early detection problem as a sequential decision-making, goal inference problem, where the objective is to characterize and predict an individual’s, organization’s, or a country’s intent (unobserved goal-directed behavior) towards developing a nuclear capability from partially observed sequences of their research publications, using inverse reinforcement learning and Bayesian goal inference methods. A computational framework is presented, and its application demonstrated using 29,196 Scopus records for a case study related to a civil nuclear capability. The case study results serve as a proof-of-concept demonstration for inference of technology-directed research activity of authors who publish in the nuclear domain. The inference method, combined with advanced computing, may be used to assess and monitor activities pertaining to early developmental stages of a nuclear technology or capability, which in turn can help to identify and prioritize activities with nuclear proliferation potential for further investigation.

98 NUCLEAR DISARMAMENT, SAFEGUARDS, AND PHYSICAL P↗

Learning to Trigger: Reinforcement Learning at the Large Hadron Collider

High-throughput scientific facilities such as the Large Hadron Collider depend on real-time event filtering (\textit{triggering}) under tight constraints on bandwidth, latency, and storage. In practice, trigger menus are largely static and hand-tuned and can become suboptimal as detector conditions, pileup, and background composition drift over time. We cast online threshold tuning as a sequential decision-making problem: a reinforcement learning agent ingests streaming summaries of recent rates and signal-sensitive features and updates trigger thresholds to maximize signal efficiency while tracking a target background rate within a tolerance band. We adapt Group-Filtered Policy Optimization (GFPO) to streaming control and introduce two variants (GFPO-F, GFPO-FR) that enforce background rate feasibility during training. On a benchmark that emulates realistic collider operation, we study two representative triggers: a total transverse energy ($H_{T}$) trigger sensitive to pileup variation, and an anomaly-detection (AD) trigger based on reconstruction loss for rare or non-standard signatures. On Monte Carlo streams, our agent increases the fraction of in-tolerance time intervals by 48% ($H_T$) and 28% (AD), with a cumulative gain of up to 2% in signal efficiency on those in-tolerance intervals. Transferring from simulation to \emph{real} collision data (CMS Run 283408), the same agent, without fine-tuning, achieves a 56% ($H_T$) and 28% (AD) in-tolerance improvement over baselines, with further signal-efficiency gain on both triggers. To our knowledge, this is the \emph{first} demonstration of RL-based trigger control on real Large Hadron Collider collision data. Code is available at https://github.com/Zixind/GFPO_LHC (see repo for details).

Ding, Zixin [Chicago U.]↗

Safe Exploration Reinforcement Learning for Load Restoration using Invalid Action Masking

This paper addresses the load restoration problem after a power outage event. Our primary proposed methodology uses a multi-agent reinforcement learning method to make the optimal sequential decisions on picking up critical loads. Typically, a negative reward is provided to discourage the agents from selecting decisions that violate physical constraints during the restoration process. However, the main disadvantage of this approach is its difficulty in applying it to large-scale systems due to the curse of dimensionality. This paper introduces the invalid action masking technique to overcome this limitation. The features of this technique include zero physical constraint violations, reduced training time, and stabilization of the explo- ration process. Simulation results are performed in IEEE 13-node and IEEE 123-node systems showing the better performance of the proposed algorithm in comparison to the conventional approaches both in terms of restored power and learning curve.

reinforcement learning, blackstart, artificial int↗

Optimal CO 2 storage management considering safety constraints in multi-stakeholder multi-site GCS projects: A Markov game perspective

Geological carbon storage (GCS) projects could involve a diverse array of stakeholders or players from public, private, and regulatory sectors, each with different objectives and responsibilities. Given the complexity, scale, and long-term nature of GCS operations, determining whether individual stakeholders can independently optimize their interests — or whether collaborative coalition agreements are needed — remains a central question for effective GCS project planning and management. To access large, high-quality storage resources, future GCS deployment may increasingly occur in geologically connected sites, where shared geological features such as pressure space and reservoir pore capacity can lead to competitive behavior among stakeholders. In this work, we propose a paradigm based on Markov games to quantitatively investigate how different coalition structures affect the goals of stakeholders. We frame this multi-stakeholder multi-site problem as a multi-agent reinforcement learning problem with safety constraints. Our approach enables agents to learn optimal strategies while complying with safety regulations. We present an example where multiple operators are injecting CO 2 into their respective project areas in a geologically connected basin. To address the high computational cost of repeated simulations of high fidelity models, a previously developed surrogate model based on the Embed-to-Control (E2C) framework is employed. Our results demonstrate the effectiveness of the proposed framework in addressing optimal management of CO 2 storage when multiple stakeholders with different objectives and goals are involved.

58 GEOSCIENCES↗

Automobile and Technology Lifecycle-Based Assignment (ATLAS) v2.0.12

ATLAS is a comprehensive vehicle transaction and technology adoption microsimulator. ATLAS evolves the fleet mix of individual households by simulating the transaction (vehicle addition, disposal, and replacement) and choice (vehicle type, vintage, powertrain, and tenure) decisions in response to the co-evolving demographics, land use, and vehicle technology simulations. Different from the existing vehicle models that are either static or aggregated (e.g. stock model), ATLAS is fully disaggregated and dynamic following a sequential and circumstantial decision-making trajectory. This fine-grained approach not only enhances the realism of the simulation but also provides a nuanced understanding of the dynamics inherent in vehicle fleet evolution. ATLAS outputs are fully compatible with subsequent agent-based transportation modeling system and can enable distributional effect analysis regarding the fleet turnover among heterogeneous populations. ATLAS expands the typical new sale focused vehicle choice modeling to including used vehicle transactions that are of increasing interests to understanding the vehicle adoption behavior among lower income households.

Jin, Ling↗

Multicriteria-Based Selection of Microalgae Biorefineries: Biomass Composition Defines the Most Suitable Product Portfolios

The supply of microalgae-derived biofuels and bioproducts will be vital in a global-scale bioeconomy. As the diversity of microalgae strains and their compositional plasticity may yield a wide array of products of market interest, the design of effective biorefineries is a central aspect in the path to making microalgae-based products available in the market. In this way, the conversion of microalgae biomass should be planned to employ mature technologies, while maximizing economic and environmental benefits obtained from a diverse product portfolio. This study applied sequential, hybrid Multicriteria Decision Analyses (MCDA) to aid the decision-making process of outlining the most suitable biorefining pathways for specific compositional profiles. For this, the methodology simultaneously considered multiple technical, economic, and environmental criteria, such as market aspects for the main algae-derived products, potential reduction in greenhouse gas (GHG) emissions provided by the biobased alternatives, and technology readiness level of conversion routes, among others. The framework was tested using productivity and compositional data for 13 high-productivity summer strains cultivated under varying nutrient availability. The analysis pointed to compositional profile being a key driver in defining the core biorefining strategy of algae biomass, with lipid-to-carbohydrate ratios higher than roughly 1 warranting the preferential processing of lipids into hydrocarbon fuels with the remainder of the algae biomass compounds being sent to higher-value applications, such as carboxylic acids and renewable thermoplastic substitutes. An in-depth process simulation, techno-economic assessment (TEA), and life-cycle analysis (LCA) effort was carried out as a closing step to corroborate the results stemming from the proposed framework. This study validates the use of MCDAs as a screening method prior to implementing more time-consuming analysis techniques and makes a compelling case for this approach as a product selection tool on a wide range of algae species, thus helping establish species-agnostic (but composition-driven) biorefineries.

biofuels↗

A systematic multicriteria-based approach to support product portfolio selection in microalgae biorefineries

Here this work proposes and applies a sequential approach of objective methods to aid the decision-making process for the deployment of microalgae biorefineries. The strategy combines Multicriteria Decision Analysis (MCDA) and weight assignment methods to simultaneously consider technical, economic, and environmental criteria to (1) outrank the best bioproduct options from different biomass fractions present in microalgae biomass at different ratios (namely carbohydrates, lipids, and protein) and (2) define the most suitable biorefining pathways associated with specific pairings of microalgae strains and cultivation conditions. The first part of the assessment identified succinic acid, acrylic acid, and citric acid as the top-ranked bioproducts from carbohydrates, polyurethane from lipids, and thermoplastic extrusion co-feed from protein. The second step of the analysis determined that, when production of a hydrocarbon fuel is desired, the compositional profile of a strain is paramount in defining the biorefining setup that should be pursued. In summary, microalgae lipids should be sent to the production of hydrocarbon fuels if the ratio between neutral lipids and fermentable carbohydrates is higher than roughly 1, with carbohydrates and protein being converted to the higher-value products noted above. Finally, this result was corroborated through process simulations, which indicated superior economic and environmental metrics when strains are paired with suitable conversion pathways identified through MCDA based on their compositional profiles. The outcomes of this work provide clear, objective, guidelines for establishing the best biorefining approach for a large suite of biochemical compositions as a screening method prior to employing detailed process simulations alongside rigorous techno-economic and life-cycle assessments.

09 BIOMASS FUELS↗

A game-theoretic approach to nuclear fuel cycle transition analysis under uncertainty

We present a novel methodology for optimizing nuclear fuel cycle transitions that incorporates a game-theoretic approach and captures interactions among multiple decision makers. The methodology is demonstrated using a two-person sequential game with uncertainty, where the two players represent a policy maker and an electric utility company, though the method generalizes to any number and type of individual decision making entities. Coupled with a sophisticated nuclear fuel cycle simulator, rich transition scenarios may be analyzed to identify robust transition strategies. These strategies explicitly treat uncertainties using a stochastic programming approach, devising optimal near-term hedging strategies that simultaneously consider all possible states of the world, maintaining flexibility to allow for intelligent recourse decisions once uncertainties are resolved. In the demonstration game, reactor technology and fuel cycle scheme adopted by the electric utility are shown to depend on both the policy maker’s decisions and the distributions over uncertain technological and economic outcomes.

11 NUCLEAR FUEL CYCLE AND FUEL MATERIALS↗

Connected and Learning Based Optimal Freight Management for Efficiency

The management of the future heterogenous fleet is a complex decision-making problem. The heterogenous fleet is emerging as decarbonization technologies are deployed by fleets toward lowering the freight operation emissions in Medium and Heavy-duty vehicles. Traditionally, in fleets characterized by a homogeneous Diesel Internal Combustion Engine (ICE) powertrain, the process of fleet planning and operational optimization unfolds sequentially without the necessity to account for powertrain and vehicle-specific characteristics during dispatch decisions. Fleets with trucks less than 5 years old tend to maintain stable vehicle efficiency with minimal operational reliability risks for fleet managers. However, the landscape changes with the incorporation of emerging powertrain technologies, which lack extensive operational data and service experiences. This includes technologies like hybrid, Electric, Fuel Cell, or alternative fuel ICE. Operational decisions for fleets featuring heterogeneous powertrain technologies and facing limited access to alternative fueling and charging stations become intricate, requiring careful consideration and optimization at each dispatch. The difference in efficiency characteristics of emerging technologies, their range limitations, and the restricted availability of charging/alternative fueling infrastructure, coupled with sensitivity to driving conditions (e.g., EV range reduction in low temperatures) and their impact on component aging (such as batteries), become pivotal factors influencing the reliable and efficient freight transportation. To make the path toward low emission freight transportation efficient and reliable, an AI-assisted fleet management software is developed in this project to help fleet managers in optimizing both adoption of emerging powertrain decarbonization, connected and automated technologies and also operating the fleet after such technologies are deployed as schematically. Freight transportation requirements are different depending on the cargos to be shipped, customer requirements and regions of operations. This further highlights the need for software and digital solutions to tailor deployment and operation of emerging powertrain, connectivity, and automation technologies toward the specific fleet operation requirements. The fleet management optimizer was also integrated with a model of the fleet to simulate the operation of the fleet over 1 year of the baseline fleet operation (250,000+ shipments) indicating the significance of day-to-day variations on emissions and energy consumption of a freight transportation fleet. The results demonstrate ≥20% improvement in freight efficiency in terms of WTW CO2 per ton-mile of cargo shipments while all fleet operation constraints are enforced, and the cost (CapEx and OpEx) is minimized.

33 ADVANCED PROPULSION SYSTEMS↗

AIF for Vis (Active Inference for simulating human interpretation of data visualization) [SWR-26-084]

AIF for Vis contains the Active Inference models and analysis scripts used to study a simple visualization-interpretation task: estimating the average value of two bars in a bar chart. The work is a proof of concept for translating hypothesized cognitive strategies into executable, inspectable process models. We implement two idealized strategies inspired by dual-process accounts of visualization-aided decision making: *Fast model: a compressed, heuristic strategy that estimates the visual midpoint of the two bars and maintains a single belief over their average. *Slow model: a sequential, analytic strategy that estimates the two bar heights separately and maintains them in working memory before computing an average. Both models use a common Active-Inference-inspired framework for sequential perception, belief updating, action selection, and reporting. Their different internal representations produce distinct predicted vulnerabilities: *the Fast model is more susceptible to tick-salience bias; *the Slow model is more susceptible to working-memory decay. The repository includes the model implementations, scripts used for the experiments reported in the paper, precomputed trial-level results, and plotting scripts.

Goldwyn, Harrison [National Laboratory of the Rock↗

Active learning path-dependent properties using a cloud-based materials acceleration platform

Solid state materials are central to many modern technologies in which a given material may be exposed to a variety of environments. The material properties often vary with the sequence of environments in an irreversible manner, resulting in a quintessential path-dependency in experimental observables. While sequential learning techniques have been effectively deployed for accelerating learning of state properties of materials, they often use a consistent environment path in all experiments. To elevate such techniques for making optimal decisions in experimental investigations of path-dependent properties, we introduce an iterated expected information gain acquisition function that optimizes over entire experimental trajectories. This approach is implemented within a cloud-based Materials Acceleration Platform architecture utilizing an event-driven stateful broker coupled with remote HELAO (Hierarchical Experimental Laboratory Automation and Orchestration) instances and an AI science manager. The platform's efficacy was demonstrated through a case study optimizing multi-step spectro-electrochemical experiments to identify optically stable potential windows in (Co–Ni–Sb)O z metal oxides. The system successfully integrated AI-driven experiment design, remote laboratory automation, and cloud-based data infrastructure, validating the platform's capability for managing complex, adaptive, path-dependent workflows in materials discovery.

Guevarra, Dan [California Institute of Technology ↗