Search NASA⌕ Search

SEARCH · Search NASA

Results for “sequential decision making”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

Correlated Trajectory Uncertainty for Adaptive Sequential Decision Making

One of the great challenges with decision making tasks on real world systems is the fact that data is sparse and acquiring additional data is expensive. In these cases, it is often crucial to make a model of the environment to assist in making decisions. At the same time, limited data means that learned models are erroneous, making it just as important to equip the model with good predictive uncertainties. In the context of learning sequential decision making policies, these uncertainties can prove useful for informing which data to collect for the greatest improvement in policy performance \citep{mehta2021experimental, mehta2022exploration} or informing the policy about unsure regions of state and action space to avoid during test time \citep{yu2020mopo}. Additionally, assuming that realistic samples of the environment can be drawn, an adaptable policy can be trained that attempts to make optimal decisions for any given possible instance of the environment \citep{ghosh2022offline, chen2021offline}. In this work, we examine the so-called ``probabilistic neural network'' (PNN) model that is ubiquitous in model-based reinforcement learning (MBRL) works. We argue that while PNN models may have good marginal uncertainties, they form a distribution of non-smooth transition functions. Not only are these samples unrealistic and may hamper adaptability, but we also assert that this leads to poor uncertainty estimates when predicting multiple step trajectory estimates. To address this issue, we propose a simple sampling method that can be implemented on top of pre-existing models.We evaluate our sampling technique on a number of environments, including a realistic nuclear fusion task, and find that, not only do smooth transition function samples produce more calibrated uncertainties, but they also lead to better downstream performance for an adaptive policy.

Offline Reinforcement Learning↗

Sequential decision making and stochastic networks

To solve the problems inherent in working with sequential decision processes, it is proposed to (1) utilize concepts of dominance through bounding in the decision processes (DP) formalism to reduce the amount of computing required. This advocates the marrying of DP recursion and Branch-and-Bound methodology; and (2) relax the requirement of strict optimality in the search over the state space, and be content with a tolerable error.

Elmaghraby, S.↗

Scalable Bayesian optimization with randomized prior networks

Several fundamental problems in science and engineering consist of global optimization tasks involving unknown high-dimensional (black-box) functions that map a set of controllable variables to the outcomes of an expensive experiment. Bayesian Optimization (BO) techniques are known to be effective in tackling global optimization problems using a relatively small number objective function evaluations, but their performance suffers when dealing with high-dimensional outputs. To overcome the major challenge of dimensionality, here we propose a deep learning framework for BO and sequential decision making based on bootstrapped ensembles of neural architectures with randomized priors. Using appropriate architecture choices, we show that the proposed framework can approximate functional relationships between design variables and quantities of interest, even in cases where the latter take values in high-dimensional vector spaces or even infinite-dimensional function spaces. In the context of BO, we augmented the proposed probabilistic surrogates with re-parameterized Monte Carlo approximations of multiple-point (parallel) acquisition functions, as well as methodological extensions for accommodating black-box constraints and multi-fidelity information sources. We test the proposed framework against state-of-the-art methods for BO and demonstrate superior performance across several challenging tasks with high-dimensional outputs, including a constrained multi-fidelity optimization task involving shape optimization of rotor blades in turbo-machinery.

97 MATHEMATICS AND COMPUTING↗

Human-Computer Interaction with Medical Decisions Support Systems

Decision Support Systems (DSSs) have been available to medical diagnosticians for some time, yet their acceptance and use have not increased with advances in technology and availability of DSS tools. Medical DSSs will be necessary on future long duration space missions, because access to medical resources and personnel will be limited. Human-Computer Interaction (HCI) experts at NASA's Human Factors and Ergonomics Laboratory (HFEL) have been working toward understanding how humans use DSSs, with the goal of being able to identify and solve the problems associated with these systems. Work to date consists of identification of HCI research areas, development of a decision making model, and completion of two experiments dealing with 'anchoring'. Anchoring is a phenomenon in which the decision maker latches on to a starting point and does not make sufficient adjustments when new data are presented. HFEL personnel have replicated a well-known anchoring experiment and have investigated the effects of user level of knowledge. Future work includes further experimentation on level of knowledge, confidence in the source of information and sequential decision making.

Adolf, Jurine A.↗

Thermal Reservoir Networks for Modularly Expandable Thermal Microgrids

The Department of Defense (DoD) faces the substantial challenge of cost-effectively retrofitting one to two installations per month, each comprising approximately 1,000 buildings, to improve resilience, reduce energy consumption, and enhance energy supply security. Achieving these objectives requires optimal system selection and effective risk mitigation during system integration. To address this need, we introduce Platform-Based Design (PBD), a structured, hierarchical methodology adapted from other industrial sectors to the domain of energy system retrofits. We demonstrate the effectiveness of PBD through a techno-economic feasibility study comparing geothermal-coupled thermal energy networks (TENs) with conventional energy systems for heating, cooling, and powering 17 buildings at Joint Base Andrews (JBA) in Maryland. Our analysis illustrates that the PBD approach enables rigorous, data-driven, sequential decision making, resulting in a family of Pareto-optimal systems, among which the TEN emerged as the most promising solution. The selected TEN design integrates geothermal borefields, heat recovery heat pumps, photovoltaic (PV) arrays, and battery storage. Compared to the baseline system – gas heating combined with air-source chillers – the proposed TEN reduces annual imported energy by 74% and peak electricity demand by 45%, achieves a levelized cost of energy of $\$0.210$/kWh, and substantially enhances resilience. Life-cycle costs increase by approximately 6%, and initial investment costs are about 2.5 times higher than the baseline. However, if central plant infrastructure, district loops, and utility-scale PV and battery systems are privately funded and operated, the initial investment would fall below the baseline system cost. Critical to achieving these significant performance improvements were detailed nonlinear dynamic simulations coupling geothermal heat transfer, energy system operation, and realistic feedback control logic. These simulations identified essential design modifications and control strategy refinements that substantially reduced energy use, peak demand, and compressor shortcycling, thereby improving durability and reliability—issues that would have been significantly more expensive to resolve during operation. Additionally, the verification step highlighted sensitivities to key design parameters that could reduce initial investment by approximately $\$2$ million and reduce annual life-cycle costs more than $\$300,000$. We recommend adopting the PBD methodology for future feasibility studies and TEN pilot projects to gain valuable operational experience. Furthermore, we recommend that DoD invest in transferring and scaling the PBD methodology to other installations. This entails developing standardized computational frameworks and component libraries as well as training industry in conducting PBD. Such investments would enable rapid, robust, reliable, and cost-effective retrofits, supporting DoD’s ambitious energy system modernization goals.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗

Collaborative, Sequential and Isolated Decisions in Design

The Massachusetts Institute of Technology (MIT) Commission on Industrial Productivity, in their report Made in America, found that six recurring weaknesses were hampering American manufacturing industries. The two weaknesses most relevant to product development were 1) technological weakness in development and production, and 2) failures in cooperation. The remedies to these weaknesses are considered the essential twin pillars of CE: 1) improved development process, and 2) closer cooperation. In the MIT report, it is recognized that total cooperation among teams in a CE environment is rare in American industry, while the majority of the design research in mathematically modeling CE has assumed total cooperation. In this paper, we present mathematical constructs, based on game theoretic principles, to model degrees of collaboration characterized by approximate cooperation, sequential decision making and isolation. The design of a pressure vessel and a passenger aircraft are included as illustrative examples.

Lewis, Kemper↗

Transition-Independent Decentralized Markov Decision Processes

There has been substantial progress with formal models for sequential decision making by individual agents using the Markov decision process (MDP). However, similar treatment of multi-agent systems is lacking. A recent complexity result, showing that solving decentralized MDPs is NEXP-hard, provides a partial explanation. To overcome this complexity barrier, we identify a general class of transition-independent decentralized MDPs that is widely applicable. The class consists of independent collaborating agents that are tied up by a global reward function that depends on both of their histories. We present a novel algorithm for solving this class of problems and examine its properties. The result is the first effective technique to solve optimally a class of decentralized MDPs. This lays the foundation for further work in this area on both exact and approximate solutions.

Becker, Raphen↗

A control-inspired approach for energy transition planning under uncertainty

As the global carbon footprint continues to grow, many countries are implementing carbon emission reduction policies which have incentivized the expansion of low-carbon and renewable technologies. However, the speed and scale of deployment falls short of that needed to meet climate goals. Energy system models serve as key tools for guiding investment decisions and helping policymakers evaluate the effects of various policies on the development of an energy system. This study focuses on the energy system of the United States and builds upon prior work by incorporating more geographic granularity to account for the trade of commodities and addresses transmission congestion through electricity price adjustments. Furthermore, real-world characteristics, such as delays in constructing new liquid fuel production and electricity generation facilities, are integrated using a sequential decision-making approach that better reflects how decisions can be updated as uncertainties unfold. Results demonstrate that stochastic programming combined with sequential decision-making produces energy transition pathways that are robust to multiple uncertain futures. Additionally, considering real-world characteristics significantly impacts the deployment of renewable technologies and the ability to meet carbon emission reduction goals while also reliably meeting demand. These findings highlight the importance of accounting for uncertainty and real-world characteristics to avoid overly optimistic projections in energy system planning.

energy systems↗

A Risk-Constrained Multi-Stage Decision Making Approach to the Architectural Analysis of Mars Missions

This paper presents a novel risk-constrained multi-stage decision making approach to the architectural analysis of planetary rover missions. In particular, focusing on a 2018 Mars rover concept, which was considered as part of a potential Mars Sample Return campaign, we model the entry, descent, and landing (EDL) phase and the rover traverse phase as four sequential decision-making stages. The problem is to find a sequence of divert and driving maneuvers so that the rover drive is minimized and the probability of a mission failure (e.g., due to a failed landing) is below a user specified bound. By solving this problem for several different values of the model parameters (e.g., divert authority), this approach enables rigorous, accurate and systematic trade-offs for the EDL system vs. the mobility system, and, more in general, cross-domain trade-offs for the different phases of a space mission. The overall optimization problem can be seen as a chance-constrained dynamic programming problem, with the additional complexity that 1) in some stages the disturbances do not have any probabilistic characterization, and 2) the state space is extremely large (i.e, hundreds of millions of states for trade-offs with high-resolution Martian maps). To this purpose, we solve the problem by performing an unconventional combination of average and minimax cost analysis and by leveraging high efficient computation tools from the image processing community. Preliminary trade-off results are presented.

entry, descent, and landing (EDL)↗

Tradeoffs When Considering Deep Reinforcement Learning for Contingency Management in Advanced Air Mobility

Air transportation is undergoing a rapid evolution globally with the introduction of Advanced Air Mobility (AAM) and with it comes novel challenges and opportunities for transforming aviation. As AAM operations introduce increasing heterogeneity in vehicle capabilities and density, increased levels of automation are likely necessary to achieve operational safety and efficiency goals. This paper focuses on one example where increased automation has been suggested. Autonomous operations will need contingency management systems that can monitor evolving risk across a span of interrelated (or interdependent) hazards and, if necessary, execute appropriate control interventions via supervised or automated decision making. Accommodating this complex environment may require automated functions (autonomy) that apply artificial intelligence (AI) techniques that can adapt and respond to a quickly changing environment. This paper explores the use of Deep Reinforcement Learning (DRL) which has shown promising performance in complex and high-dimensional environments where the objective can be constructed as a sequential decision-making problem. An extension of a prior formulation of the contingency management problem as a Markov Decision Process (MDP) is presented and uses a DRL framework to train agents that mitigate hazards present in the simulation environment. A comparison of these learning-based agents and classical techniques is presented in terms of their performance, verification difficulties, and development process.

machine learningautonomous systems; flight simulat↗

Tradeoffs When Considering Deep Reinforcement Learning for Contingency Management in Advanced Air Mobility

Air transportation is undergoing a rapid evolution globally with the introduction of Advanced Air Mobility (AAM) and with it comes novel challenges and opportunities for transforming aviation. As AAM operations introduce increasing heterogeneity in vehicle capabilities and density, increased levels of automation are likely necessary to achieve operational safety and efficiency goals. This paper focuses on one example where increased automation has been suggested. Autonomous operations will need contingency management systems that can monitor evolving risk across a span of interrelated (or interdependent) hazards and, if necessary, execute appropriate control interventions via supervised or automated decision making. Accommodating this complex environment may require automated functions (autonomy) that apply artificial intelligence (AI) techniques that can adapt and respond to a quickly changing environment. This paper explores the use of Deep Reinforcement Learning (DRL) which has shown promising performance in complex and high-dimensional environments where the objective can be constructed as a sequential decision-making problem. An extension of a prior formulation of the contingency management problem as a Markov Decision Process (MDP) is presented and uses a DRL framework to train agents that mitigate hazards present in the simulation environment. A comparison of these learning-based agents and classical techniques is presented in terms of their performance, verification difficulties, and development process.

machine learning↗

Energy efficiency in industrial drying: A hybrid ultrasonic system with a novel dynamic optimization framework

Drying processes are among the most energy-consuming operations in industrial and manufacturing settings, demanding strategic selection, design, and control for enhanced efficiency. Advancing drying technologies is critical for improving sustainability, lowering energy use, reducing carbon emissions, and minimizing waste. This study explores two innovative strategies aimed at transforming drying processes into sustainable, low-carbon systems by reducing energy consumption, minimizing waste, and maintaining a strong emphasis on preserving product quality. The first strategy showcases a sub-pilot scale hybrid ultrasonic-convective dryer for agrifood products. This technology, powered by electricity (process electrification), integrates non-thermal ultrasonic dehydration with convective heating and is presented as a sustainable and energy-efficient solution that enhances eco-friendly practices. The second strategy involves introducing and implementing a novel, multiobjective, mixed integer dynamic optimization technique to determine the optimal time-dependent process parameter values for the drying operation. This optimization technique yields operating conditions that are piecewise constant in time aiming to maximize the energy efficiency of the hybrid ultrasonic-convective dryer while ensuring strict adherence to product quality constraints. By adopting the hybrid ultrasonic-convective dryer, a notable 35% improvement in energy efficiency was achieved compared to conventional hot-air drying systems for drying apple slices. The proposed optimization framework further enhanced energy efficiency by nearly 14% over the most efficient process on the identical testbed, under static operating conditions. The reported enhancements have been experimentally validated. Regarding drying time (thereby improving production yield), the developed hybrid ultrasonic-convective dryer demonstrates as much as a 41% reduction in total processing time, which is further optimized by an additional 10% using our proposed optimization framework. The research outcomes have profound implications for the design and operation of drying systems, encompassing crucial aspects such as process electrification, cost-effectiveness, energy savings, time efficiency, product yield, product quality, and process automation.

Dynamic optimization↗

Deep Reinforcement Learning for Distribution System Operations: A Tutorial and Survey

Here, the rapid evolution of modern electric power distribution systems into complex networks of interconnected active devices, distributed generation (DG), and storage poses increasing difficulties for system operators. The large-scale integration of distributed energy resources (DERs) and the rapid exchange of measurement data via communication networks present major opportunities for advancing grid operations but also introduce greater uncertainty, higher data dimensionality, more complex network and device models, and challenging control and optimization problems. Deep reinforcement learning (DRL) algorithms are promising in addressing these challenges. However, they have not been effectively adapted for power systems applications, requiring extensive customization for implementation and evaluation. This has resulted in reproducibility challenges and a steep learning curve for researchers new to applying DRL algorithms to the power systems domain. To bridge these gaps, this tutorial aims to serve as a valuable resource for researchers interested in exploring learning-based algorithms to operate active power distribution networks. Specifically, this work presents a generalized process for translating sequential decision-making problems in power distribution systems into Markov decision process (MDP) formulations, illustrated through concrete grid service examples. Additionally, we introduce a simple environment design strategy to develop and evaluate example DRL algorithms for distribution system applications, complete with an included code repository to guide users through environment construction.

24 POWER TRANSMISSION AND DISTRIBUTION↗

Early Inference of Nuclear Technology-Directed Research Activities of Authors from Scientific Publications

Nuclear research articles can provide information about early nuclear proliferation indicators such as influential research entities and technology capability levels of a country, but detection of nuclear activities typically occurs after they have started. We investigate the extent to which nuclear research articles can be used to infer whether a research entity will acquire or develop a nuclear technology before it happens. Early detection of nuclear proliferation or technology development indicators from data is challenging due to partial observability, sparse and unlabeled information, and confounding signals from multiple concurrent activities. This paper presents the early detection problem as a sequential decision-making, goal inference problem, where the objective is to characterize and predict an individual’s, organization’s, or a country’s intent (unobserved goal-directed behavior) towards developing a nuclear capability from partially observed sequences of their research publications, using inverse reinforcement learning and Bayesian goal inference methods. A computational framework is presented, and its application demonstrated using 29,196 Scopus records for a case study related to a civil nuclear capability. The case study results serve as a proof-of-concept demonstration for inference of technology-directed research activity of authors who publish in the nuclear domain. The inference method, combined with advanced computing, may be used to assess and monitor activities pertaining to early developmental stages of a nuclear technology or capability, which in turn can help to identify and prioritize activities with nuclear proliferation potential for further investigation.

98 NUCLEAR DISARMAMENT, SAFEGUARDS, AND PHYSICAL P↗

Risk-Constrained Dynamic Programming for Optimal Mars Entry, Descent, and Landing

A chance-constrained dynamic programming algorithm was developed that is capable of making optimal sequential decisions within a user-specified risk bound. This work handles stochastic uncertainties over multiple stages in the CEMAT (Combined EDL-Mobility Analyses Tool) framework. It was demonstrated by a simulation of Mars entry, descent, and landing (EDL) using real landscape data obtained from the Mars Reconnaissance Orbiter. Although standard dynamic programming (DP) provides a general framework for optimal sequential decisionmaking under uncertainty, it typically achieves risk aversion by imposing an arbitrary penalty on failure states. Such a penalty-based approach cannot explicitly bound the probability of mission failure. A key idea behind the new approach is called risk allocation, which decomposes a joint chance constraint into a set of individual chance constraints and distributes risk over them. The joint chance constraint was reformulated into a constraint on an expectation over a sum of an indicator function, which can be incorporated into the cost function by dualizing the optimization problem. As a result, the chance-constraint optimization problem can be turned into an unconstrained optimization over a Lagrangian, which can be solved efficiently using a standard DP approach.

Ono, Masahiro↗

Adaptive Stress Testing of Trajectory Predictions in Flight Management Systems

To find failure events and their likelihoods in flight-critical systems, we investigate the use of an advanced black-box stress testing approach called adaptive stress testing. We analyze a trajectory predictor from a developmental commercial flight management system which takes as input a collection of lateral waypoints and en-route environmental conditions. Our aim is to search for failure events relating to inconsistencies in the predicted lateral trajectories. The intention of this work is to find likely failures and report them back to the developers so they can address and potentially resolve shortcomings of the system before deployment. To improve search performance, this work extends the adaptive stress testing formulation to be applied more generally to sequential decision-making problems with episodic reward by collecting the state transitions during the search and evaluating at the end of the simulated rollout. We use a modified Monte Carlo tree search algorithm with progressive widening as our adversarial reinforcement learner. The performance is compared to direct Monte Carlo simulations and to the cross-entropy method as an alternative importance sampling baseline. The goal is to find potential problems otherwise not found by traditional requirements-based testing. Results indicate that our adaptive stress testing approach finds more failures and finds failures with higher likelihood relative to the baseline approaches.

adaptive stress testing↗

Learning to Trigger: Reinforcement Learning at the Large Hadron Collider

High-throughput scientific facilities such as the Large Hadron Collider depend on real-time event filtering (\textit{triggering}) under tight constraints on bandwidth, latency, and storage. In practice, trigger menus are largely static and hand-tuned and can become suboptimal as detector conditions, pileup, and background composition drift over time. We cast online threshold tuning as a sequential decision-making problem: a reinforcement learning agent ingests streaming summaries of recent rates and signal-sensitive features and updates trigger thresholds to maximize signal efficiency while tracking a target background rate within a tolerance band. We adapt Group-Filtered Policy Optimization (GFPO) to streaming control and introduce two variants (GFPO-F, GFPO-FR) that enforce background rate feasibility during training. On a benchmark that emulates realistic collider operation, we study two representative triggers: a total transverse energy ($H_{T}$) trigger sensitive to pileup variation, and an anomaly-detection (AD) trigger based on reconstruction loss for rare or non-standard signatures. On Monte Carlo streams, our agent increases the fraction of in-tolerance time intervals by 48% ($H_T$) and 28% (AD), with a cumulative gain of up to 2% in signal efficiency on those in-tolerance intervals. Transferring from simulation to \emph{real} collision data (CMS Run 283408), the same agent, without fine-tuning, achieves a 56% ($H_T$) and 28% (AD) in-tolerance improvement over baselines, with further signal-efficiency gain on both triggers. To our knowledge, this is the \emph{first} demonstration of RL-based trigger control on real Large Hadron Collider collision data. Code is available at https://github.com/Zixind/GFPO_LHC (see repo for details).

Ding, Zixin [Chicago U.]↗