Search NASASearch

SEARCH · Search NASA

Results for “online optimization”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

Autonomous online optimization of a closed-circuit reverse osmosis system

As freshwater becomes increasingly scarce, many industrial and municipal water utilities look at premise-scale water treatment and reuse to meet water demand. Closed-circuit reverse osmosis (CCRO) has been proposed as a promising process design to do so. This sequencing batch process enables operation at higher brine salinity levels by means of a recycle flow. Optimal operation requires that the maximum salinity level at the membrane surface represents an optimal trade-off between brine disposal costs and energy efficiency. This maximum salinity level may change over time as the feed water composition changes and electricity markets fluctuate. In this article, we present the results of the experimental evaluation of an automatic technique for continuous online optimization, known as extremum seeking control. This technique has a long history in the process control community but has received little traction so far in the water industry. We modify this technique to enable its use for online optimization of CCRO, specifically to account for its sequential batch operation. We challenge the optimization schemes through several experimental tests and illustrate the advantages and drawbacks of extremum-seeking control.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI

Performance Optimization of the IOTA Duoplasmatron Proton Source

We present results from online optimization studies of a duoplasmatron ion source designed to produce 50~keV protons for acceleration to 2.5~MeV and subsequent injection into the Integrable Optics Test Accelerator (IOTA) at Fermilab. Using a Bayesian exploration technique, we developed multi-parameter models of the source s proton current and employed these models to optimize its performance. Depending on the spectrometer configuration used to isolate the proton beam and the chosen optimization objective, we identified three candidate operating points, achieving normalized 50~\% emittances between 0.57 and 1.3~\textmu m and a maximum proton current of $14.5 \pm 0.6$~mA.

Banerjee, Nilanjan [Fermilab] (ORCID:0000000344660

Lila: Optimal Dispatching in Probabilistic Temporal Networks using Monte Carlo Tree Search

Executing a Probabilistic Simple Temporal Network (PSTN) amounts at scheduling, i.e. \textit{dispatch}, a set of events under time uncertainty. This constitutes a NP-hard online optimization problem. The right execution time must be dynamically assigned to each event of the PSTN such that the temporal constraints are met, whereas activity durations are progressively observed as the execution unfolds. We propose a dispatching algorithm based on Monte Carlo Tree Search, called Lila, with the following characteristics: (i) it is an anytime algorithm, both offline and online, proven asymptotically optimal; (ii) it returns the current probability of success, either before or at any moment during operations; (iii) it handles any possible continuous or discrete, even non-parametric, probability distributions, as well as inter-dependencies between random variables, exogenous and endogenous uncertainty; and (iv) can be easily extended to handle probabilistic external events, PSTNs with resources, PSTNs with cutoff times and precondition chains, etc. Lila is universal in the sense that it can handle any dispatching protocol, simply by specifying it to the algorithm. It has the unlimited flexibility offered by the simulation paradigm, whilst it asymptotically converges to optimal decisions and/or robustness approximations.

Chien, Steve A.

Optimal Control Prediction Method for Control Allocation

This paper proposes a novel prediction method for online optimal control allocation that extends the volume of moments achievable with the Moore-Penrose generalized inverse to the entire Attainable Moment Set. This method formulates the control allocation problem using selected basis vectors and associated gains which reduces the optimization problem dimensions and provides physical insight into the resulting optimal solutions. The proposed algorithm finds the entire family of unique optimal control solutions along the desired moment vector from the origin to the boundary of the Attainable Moment Set. Numerical results for the Moore-Penrose prediction method show that the unique minimal controls obtained yield the desired moment with near machine precision accuracy while maintaining control effectors within specified position limits. This method has been fully validated against the unique solution obtained on the boundary of the Attainable Moment Set using the Durham Direct Allocation method. Minimal control solutions obtained for moments in the interior of the Attainable Moment Set, similarly yield the desired moment to near machine precision while providing control solutions that are smaller (i.e. 2-norm) than solutions found with traditional control allocation algorithms (e.g. interior point methods) applied to the minimal control problem. Numerical simulations using a Matlab® autocoded executable (MEX) for the representative real world problem of 3-moments with 20 individual control effectors and prescribed control position limits show a mean computation speed of approximately 125 Hz which is sufficient to enable real-time flight allocation.

Acheson, Michael J.

Stable Machine‐Learning Parameterization of Subgrid Processes in a Comprehensive Atmospheric Model Learned From Embedded Convection‐Permitting Simulations

Modern climate projections often suffer from inadequate spatial and temporal resolution due to computational limitations, resulting in inaccurate representations of sub-grid processes. A promising technique to address this is the multiscale modeling framework (MMF), which embeds a kilometer-resolution cloud-resolving model (CRM) within each atmospheric column of a host climate model to replace traditional convection and cloud parameterizations. Machine learning offers a unique opportunity to make MMF more accessible by emulating the embedded CRM and reducing its substantial computational cost. Although many studies have demonstrated proof-of-concept success of achieving stable hybrid simulations, it remains a challenge to achieve near operational-level success with real geography and comprehensive variable emulation that includes, for example, explicit cloud condensate coupling. In this study, we present a stable hybrid model capable of integrating for at least 5 years with near operational-level complexity, including coarse-grid geography, seasonality, explicit cloud condensate and wind predictions, and land coupling. Our model demonstrates skillful online performance, achieving a 5-year zonal mean tropospheric temperature bias within 2 K, water vapor bias within 1 g/kg, and a precipitation root mean square error of 0.96 mm/day. Key factors contributing to our online performance include an expressive U-Net architecture and physical thermodynamic constraints for microphysics. With microphysical constraints mitigating unrealistic cloud formation, our work is the first to demonstrate realistic multi-year cloud condensate climatology under the MMF framework. Despite these advances, online diagnostics reveal persistent biases in certain regions, highlighting the need for innovative strategies to further optimize online performance.

Hu, Zeyuan [NVIDIA Corporation, Santa Clara, CA (U

Online regularization of Poincaré map of storage rings with Shannon entropy

A measurable chaos indicator is used as the online optimization objective in tuning a complicated nonlinear system—the National Synchrotron Light Source-II storage ring. Through analyzing the Shannon entropy in measured Poincaré maps, not only can the commonly used nonlinear characterizations be extracted, but more importantly, the chaos can be quantified and then used for an online regularization of these maps. The method itself is general and applicable to other tunable nonlinear systems as well. Published by the American Physical Society 2025

36 MATERIALS SCIENCE

The poststall nonlinear dynamics and control of an F-18: A preliminary investigation

The successful high angle of attack (HAOA) operation of fighter aircraft will necessarily require the introduction of a new onboard control methodology that address the nonlinearity of the system when flown at the stall/poststall limits of the craft's flight envelope. As a precursor to this task, a researcher endeavored to familarize himself with the dynamics of one specific aircraft, the F-18, when it is flown at HAOA. This was accomplished by conducting a number of real time flight sorties using the NASA-Langley Research Center's F-18 simulator, which was operated with a pilot in the loop. In addition to developing a first hand familarity with the aircraft's dynamic characteristic at HAOA, work was also performed to identify the input/output operational footprint of the F-18's control surfaces. This investigator proposes to employ the nonlinear models of the plant identified this summer in a subsequent research effort that will make it possible to fly the F-18 effectively at poststall angles of attack. The controller design used there will rely on a new technique proposed by this investigator that provides for the automatic generation of online optimal control solutions for nonlinear dynamic systems.

Patten, William N.

Online multi-objective Bayesian optimization of injection efficiency and beam lifetime with skew quadrupoles at NSLS-II

At NSLS-II, the vertical emittance of electron beam is typically blown up to ~30 pm with a coupling wave to increase beam lifetime during user operation. As more and more insertion devices are added to the storage ring, injection efficiency to the ring drops noticeably in certain machine states, apparently due to degraded dynamic apertures. To help alleviate this issue, we have recently performed online multi-objective Bayesian optimization to increase injection efficiency while maintaining beam lifetime, by adjusting the strengths of 15 skew quadrupoles in non-dispersive sections. We report the results of this optimization effort.

Hidaka, Yoshiteru [Brookhaven]

Stochastic Optimization to Find Optimum Beginning-of-Life Core Configuration of Stable Salt Reactor with Online Refueling

A stochastic optimization method has been developed to find an optimum equilibrium cycle core configuration of the waste-burning stable salt reactor, which is a fast-spectrum molten salt reactor with frequent online refueling. An optimum core configuration was determined with the goal of minimizing radial power peaking. Because of the vast number of potential candidate core configurations, stochastic optimization was applied based on simulated annealing and an additional acceleration method, which screened out unpromising core configurations. It has been demonstrated that the developed stochastic optimization method successfully finds the optimal core configuration regardless of the initial guess and outperforms the gradient descent approach. In addition, it has been observed that the use of a so-called out-in core configuration as the initial guess speeds up convergence of the iterative solution more than five times. Based on the searched optimum equilibrium cycle core configuration, new beginning-of-life (BOL) core configurations have been developed. In conclusion, the new BOL core configurations will be used in developing optimum refueling strategies.

Moltex static salt reactor

Adaptive Online Model Update Algorithm for Predictive Control in Networked Systems

In this article, we introduce an adaptive on-line model update algorithm designed for predictive control applications in networked systems, particularly focusing on power distribution systems. Unlike traditional methods that depend on historical data for offline model identification, our approach utilizes real-time data for continuous model updates. This method integrates seamlessly with existing online control and optimization algorithms and provides timely updates in response to real-time changes. This methodology offers significant advantages, including a reduction in the communication network bandwidth requirements by minimizing the data exchanged at each iteration and enabling the model to adapt after disturbances. Furthermore, our algorithm is tailored for non-linear convex models, enhancing its applicability to practical scenarios. The efficacy of the proposed method is validated through a numerical study, demonstrating improved control performance using a synthetic IEEE test case.

data-driven model predictive control

Software For Multivariate Bayesian Classification

PHD general-purpose classifier computer program. Uses Bayesian methods to classify vectors of real numbers, based on combination of statistical techniques that include multivariate density estimation, Parzen density kernels, and EM (Expectation Maximization) algorithm. By means of simple graphical interface, user trains classifier to recognize two or more classes of data and then use it to identify new data. Written in ANSI C for Unix systems and optimized for online classification applications. Embedded in another program, or runs by itself using simple graphical-user-interface. Online help files makes program easy to use.

Saul, Ronald

A Learn-To-Fly Approach for Adaptively Tuning Flight Control Systems

A method is presented for adaptively tuning feedback control gains in a ight control sys- tem to achieve desired closed-loop performance. The method combines efficient parameter estimation for identifying closed-loop dynamics models, with online nonlinear optimization for sequentially perturbing and updating control gains to improve performance. Prior in- formation on stability and control derivatives is not needed, nor is any knowledge about the control system architecture. Following convergence, the optimized control gains (with uncertainties), the open-loop dynamics model, and the closed-loop dynamics model are available. The method is demonstrated for tuning a longitudinal stability augmentation system using a realistic nonlinear ight dynamics simulation of the NASA FASER airplane. Convergence was attained using five piloted maneuvers that spanned approximately one minute of ight test time. Although demonstrated for a relatively simple case, the method is general and can be applied to other aircraft, axes, performance metrics, and control systems.

Grauer, Jared A.

A Learn-to-Fly Approach for Adaptively Tuning Flight Control Systems

A method is presented for adaptively tuning feedback control gains in a flight control system to achieve desired closed-loop performance. The method combines efficient aircraft parameter estimation, for identifying closed-loop dynamics models, with online nonlinear optimization, for sequentially perturbing and updating control gains to improve performance. Access to flight measurements and the control gains is required, but no prior information about the aircraft dynamics or the flight control architecture is needed. After the procedure, the optimized control gains (with uncertainties), the open-loop dynamics model, and the closed-loop dynamics model are available. The method is demonstrated for tuning a longitudinal stability augmentation system using a realistic nonlinear flight dynamics simulation of a subscale airplane. Convergence was attained using five maneuvers that spanned approximately one minute of flight test time. Although shown for a relatively simple case, the method is general and can be applied to other aircraft, axes of motion, performance metrics, and control system designs.

Adaptive tuning

Shock-Stationary Application of Pseudoshock Models During High-Amplitude Combustion-Driven Unsteadiness

The isolator pseudo-shock provides necessary compression within a dual-mode scramjet engine and buffers the engine system against unstart. Quasi-1D flux-conserved models are the state-of-the-art reduced-order model for optimization and online control of dual-mode scramjet engines. The stability and efficacy of this modeling approach is evaluated against data from a combustor-driven direct-connect experiment. The experiment exhibited strong combustor-driven unsteadiness that produced upstream propagating weak shocks into the isolator, interacting with the pseudo-shock. While this configuration resulted in unsteadiness that is atypical of standard operation, the experiment provided an opportunity to evaluate the modeling techniques in highly transient states. Such transients could occur during maneuvering or result from unexpected combustor events. A flexible quasi-1D formulation, the Fievet flux-conserved model is fit using Bayesian inference in the laboratory and shock-stationary reference frames. Model performance is analyzed using the Bayesian posteriors and model evaluations over the measured shock train speed range. It is concluded that to produce consistent isolator pressure profile estimates in this unsteady environment, the model must be implemented in a shock-stationary reference frame. Implementing this conclusion in model-based engine controllers may reduce needed unstart safety margins and increase maximum performance.

Bayesian

Adaptive Cybersecurity for Distributed Energy Resources (AdCyDER): Online Reinforcement Learning with Stackelberg-Optimized Defenses — Pipeline Architecture, Evaluation Methodology, and Findings from a Synthetic-Data Evaluation

This report documents the design and evaluation of an integrated online-learning pipeline developed within the AdCyDER project for Distributed Energy Resource (DER) cybersecurity. The pipeline couples a Reinforcement Learning (RL) attack classifier — which produces an attack-type probability distribution — with a Stackelberg game-theoretic (GT) defense selector that consumes those distributions alongside SME-encoded priors over (defense, attack) effectiveness pairings and perdefense costs to choose grid-health-preserving defenses. The objective is not attack classification per se but production of distributions that drive effective defense selection through the Stackelberg layer, learned from delayed grid-health feedback rather than labeled attack data. AdCyDER as a whole is broader than the work presented here; this report covers the specific RL/GT loop integration and its evaluation. We present the integrated pipeline (SCADA telemetry with Fronius inverter physics, Suricata IDS, time-windowed aggregation, per-facility LSTM classifier, Stackelberg optimizer, OpenC2 actuators), an experimental campaign of 28 eight-hour iterations across three baseline modes, and a pipeline-ordered diagnostic protocol. The protocol identifies two distinct failure modes within the loop: paired supervised ceilings on the same features establish that the deployed online RL classifier (macro F1 ≈ 0.07) sits at least 4.7× below a same-architecture supervised LSTM (≈ 0.34) and 10–11× below a linear feature-signal ceiling (≈ 0.70–0.79 depending on per-facility isolation), localizing the dominant failure to the training procedure; and the reward signal driving online updates carries weak directional coupling with classifier correctness in the methodology-expected direction (multi-lens convergent: top-decile P(true) records produce more frequent state changes and slightly larger improvements, top-vs-bot Cohen’s 𝑑 ≈ −0.19), but at effect magnitudes too small to drive gradient-based learning at the campaign sample size. The original learning hypothesis is not supported by the data. The primary contributions are the diagnostic methodology — proposed as a transferable falsification protocol for online RL/GT defense pipelines learning from delayed environmental reward — and the open, reproducible experimental infrastructure. We outline reward reformulation as the highest-priority aspirational next step given the underpowered-but-aligned Q6 reading, with hardware-in-the-loop evaluation as the broadest scope-expansion option.

Blakely, Benjamin [Argonne National Laboratory (AN

Progress of 2.05 μm Fiber Laser Development for CO2 DIAL Measurements of Martian CO2 and Pressure

The Decadal Survey by the US National Academy of Sciences in the United States and NASA’s Science Mission Directorate (SMD) Science Plan both require global CO2 observations for the Martian atmospheric pressure, dynamics and CO2 variations. However, there are significant observational gaps in polar regions and during nighttime for global air pressure and CO2 on Mars. Therefore, recently we proposed a new concept of Martian differential absorption lidar (DIAL) operating in the 2.05 μm CO2 absorption band for global, including poles, atmospheric CO2 and pressure observations day and night. Based on the concept, we are awarded to develop 2.05 μm fiber lasers by the NASA Planetary Instrument Concepts for the Advancement of Solar System Observations (PICASSO) Program. The laser is designed to be an all-fiber master oscillator and power amplifier (MOPA) system with laser out put of ~3 mJ at a repetition frequency of 2 kHz. The primary master oscillator (PMO) is locked to the center (2.0504280 μm) of the selected CO2 absorption line. The frequency of a second MO (SMO) is locked to that of PMO. This SMO frequency is adjustable and switched between the online (2.05044156 μm) and offline (2.05050812 μm) wavelengths. The online wavelength is optimally selected so that the CO2 absorption optical depth (AOD) is ~1.1 at 3 km and the measurement in the low Martian atmosphere has largest signal-to-noise ratio. The online wavelength can also be adjusted to a line slope location where AOD is larger to observe atmospheric pressure at higher altitudes. We will present more details about this project and instrument development.

CO2

Development of 2.05 µm Fiber Lasers for CO2 DIAL Lidar Measuring Martian CO2 and Pressure

The Decadal Survey by the National Academy of Sciences in the United States and NASA’s Science Mission Directorate (SMD) Science Plan both require global CO 2 observations for the Martian atmospheric pressure, dynamics and chemistry. However, there are significant observational gaps in polar regions and during nighttime for global air pressure and CO 2 on Mars. Therefore, recently we proposed a new concept of Martian differential absorption lidar (DIAL) operating in the 2.05 µm CO 2 absorption band for global, including poles, atmospheric CO 2 and pressure observations day and night [1]. Based on the concept, we are awarded to develop 2.05 µm fiber lasers by the NASA Planetary Instrument Concepts for the Advancement of Solar System Observations (PICASSO) Program. The laser is designed to be an all-fiber master oscillator and power amplifier (MOPA) system with laser output of ~3 mJ at a repetition frequency of 2 kHz. The primary master oscillator (PMO) is locked to the center (2.0504280 µm) of the selected CO 2 absorption line. The frequency of a second MO (SMO) is locked to that of PMO. This SMO frequency is adjustable and switched between the online (2.05044156 µm) and offline (2.05050812 µm) wavelengths. The online wavelength is optimally selected so that the CO 2 absorption optical depth (AOD) is ~1.1 at 3 km and the measurement in the low Martian atmosphere has largest signal-to-noise ratio. The online wavelength can also be adjusted to a line slope location where AOD is larger to observe atmospheric pressure at higher altitudes. We will present more detail about this project and instrument development at the conference.

Zhaoyan Liu

Risk-Aware Reinforcement Learning Framework for User-Centric O-RAN

The evolution of Open Radio Access Networks (O-RAN) presents an opportunity to enhance network performance by enabling dynamic orchestration of configuration and optimization parameters (COPs) through online learning methods. However, leveraging this potential requires overcoming the limitations of traditional cell-centric RAN architectures, which lack the necessary flexibility. On the other hand, despite their recent popularity, the practical deployment of online learning frameworks, such as Deep Reinforcement Learning (DRL)-based COP optimization solutions, remains limited due to their risk of deteriorating network performance during the exploration phase. In this article, we propose and analyze a novel risk-aware DRL framework for user-centric RAN (UC-RAN), which offers both the architectural flexibility and COP optimization to exploit this flexibility. We investigate and identify UC-RAN COPs that can be optimized via a soft actor-critic algorithm implementable as an O-RAN application (rApp) to jointly maximize latency satisfaction, reliability satisfaction, area spectral efficiency, and energy efficiency. We use the offline learning on UC-RAN to reliably accelerate DRL training, thus minimizing the risk of DRL deteriorating cellular network performance. Results show that our proposed solution approaches near-optimal performance in just a few hundred iterations with a decrease in risk score by a factor of ten.

6G and beyond