Search NASA⌕ Search

SEARCH · Search NASA

Results for “algorithm timings”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 145 records · Page 8

TorchBraid: High-Performance Layer-Parallel Training of Deep Neural Networks with MPI and GPU Acceleration

TorchBraid is a high-performance implementation of layer-parallel training for deep neural networks (DNNs) supporting MPI-based parallelism and GPU acceleration. Layer-parallel training has been developed to overcome the serialization inherent in forward and backward propagation of DNNs that limits utilization of computational resources in the strong scaling limit. To achieve this, TorchBraid integrates the PyTorch neural network framework with the state-of-the-art XBraid time-parallel library. Furthermore, this article presents the use and performance of TorchBraid, in addition to solutions for overcoming the algorithmic challenges inherent in combining automatic differentiation with layer-parallel. Results are presented with and without GPU acceleration for the Tiny ImageNet and MNIST image classification data sets, as well as recurrent neural networks. Overall, TorchBraid enables fast training of DNNs, both in a strong and weak scaling context. In addition to the TorchBraid software, several new advances in applying layer-parallel algorithms are detailed. Integration of layer-parallel with data-parallel algorithms is presented for the first time, showing the computational advantages of the combination. Standard deep learning techniques, like batch-normalization, are developed for layer-parallel training. Finally, a new approach combining layer-parallel with spatial coarsening in order to accelerate training for 3D image classification shows roughly a 10× speedup over serial execution.

Layer-parallel↗

Optimizing Traffic Signal Control to Enhance Transportation Efficiency and Maximize Pedestrian Benefits in the Road Network

Increasing urban mobility requirements demand efficient transportation system strategies for both vehicular and pedestrian movement. This study enhances the Decentralized Graph-based Multi-Agent Reinforcement Learning (DGMARL) approach, originally tailored for vehicular traffic signal timing, to incorporate pedestrian traffic dynamics. The improved algorithm considers crucial metrics such as Eco_PI, assesses vehicle fuel consumption by factoring in stops and delays, and addresses pedestrian waiting time, crucial for system efficiency while acknowledging driver waiting time impact. Utilizing Digital Twin simulation along the MLK Smart Corridor in Chattanooga, Tennessee, the algorithm's performance is compared for various pedestrian control scenarios. To evaluate the effectiveness of DGMARL, this study compared DGMARL-enabled signal management with automated pedestrian traffic detection and an actuated signal management system (real-word baseline) with pedestrian recall, which predetermingly enforces a pedestrian phase every cycle. Findings indicate substantial improvements with DGMARL, showing a 28.29% enhancement in vehicle Eco_PI, a 60.55 % reduction in pedestrian waiting time, and a 55.74% decrease in driver stop delay, on average, compared to the baseline actuated signal timing plan.

Kumarasamy, Vijayalakshmi K [The University of Ten↗

Control simulations of many-body quantum systems by a synergism of discrete real-time learning and optimal control theory

We present a self-consistent algorithm for optimal control simulations of many-body quantum systems. The algorithm features a two-step synergism that combines discrete real-time machine learning (DRTL) with Quantum Optimal Control Theory (QOCT) using the time-dependent Schrödinger equation. Specifically, in step (1), DRTL is employed to identify a compact working space (i.e., the important portion of the Hilbert space) for the time evolution of the many-body quantum system in the presence of a control field (i.e., the initial or previously updated field), and in step (2), QOCT utilizes the DRTL-determined working space to find a newly updated control field for a chosen objective. Steps 1 and 2 are iterated until a self-consistent control objective value is reached such that the resulting optimal control field yields the same targeted objective value when the corresponding working space is systematically enlarged. Furthermore, to demonstrate this two-step self-consistent DRTL-QOCT synergistic algorithm, we perform optimal control simulations of strongly interacting 1D as well as 2D Heisenberg spin systems. In both scenarios, only a single spin (at the left end site for 1D and the upper left corner site for 2D) is driven by the time-dependent control fields to create an excitation at the opposite site as the target. It is found that, starting from all spin-down zero excitation states, the synergistic method is able to identify working spaces and convergence of the desired controlled dynamics with just a few iterations of the overall algorithm. In the cases studied, the dimensionality of the working space scales only quasi-linearly with the number of spins.

Artificial neural networks↗

Online randomized interpolative decomposition with a posteriori error estimator for temporal PDE data reduction

Traditional low-rank approximation is a powerful tool for compressing large data matrices that arise in simulations of partial differential equations (PDEs), but suffers from high computational cost and requires several passes over the PDE data. The compressed data may also lack interpretability thus making it difficult to identify feature patterns from the original data. Here, to address these issues, we present an online randomized algorithm to compute the interpolative decomposition (ID) of large-scale data matrices in situ. Compared to previous randomized IDs that used the QR decomposition to determine the column basis, we adopt a streaming ridge leverage score-based column subset selection algorithm that dynamically selects proper basis columns from the data and thus avoids an extra pass over the data to compute the coefficient matrix of the ID. In particular, we adopt a single-pass error estimator based on the non-adaptive Hutch++ algorithm to provide real-time error approximation for determining the best coefficients. As a result, our approach only needs a single pass over the original data and thus is suitable for large and high-dimensional matrices stored outside of core memory or generated in PDE simulations. A strategy to improve the accuracy of the reconstructed data gradient, when desired, within the ID framework is also presented. We provide numerical experiments on turbulent channel flow and ignition simulations, and on the NSTX Gas Puff Image dataset, comparing our algorithm with the offline ID algorithm to demonstrate its utility in real-world applications.

Column subset selection↗

Exponentially Reduced Circuit Depths Using Trotter Error Mitigation

Product formulas are a popular class of digital quantum simulation algorithms due to their conceptual simplicity, low overhead, and performance, which often exceeds theoretical expectations. Recently, Richardson extrapolation and polynomial interpolation have been proposed to mitigate the Trotter error incurred by the use of these formulas. This work provides a rigorous, general analysis of these techniques for computing time-evolved observables, simplifying the interpolation algorithm in the process, and shows that extrapolation generically improves the performance of product formulas for this task. We demonstrate that, to achieve error 𝜖 in a simulation of time 𝑇 using a 𝑝 ⁢th-order product formula with extrapolation, circuit depths of 𝑂⁡(𝑇 1+1/𝑝 ⁢polylog (1/𝜖)) are sufficient—an exponential improvement in the precision over product formulas alone. Furthermore, we prove that these algorithms achieve commutator scaling, and improve the 𝑇 complexity for the interpolation algorithm. By relaxing the requirement of performing exact Chebyshev interpolation, our simplified algorithm eliminates the need for fractional implementations of Trotter steps, reducing computational overhead. Finally, we show these techniques can be combined with the classical shadows method to estimate many time-evolved local observables. Taken together, our findings provide the strongest evidence yet for the utility of Trotter error-mitigation techniques in algorithmic applications.

quantum algorithms & computation↗

An interregional optimization approach for time series aggregation in continent-scale electricity system models

Modeling electric power systems with high shares of weather-dependent resources requires tradeoffs between temporal, spatial, and operational resolution. Many studies perform time series aggregation using clustering algorithms to reduce the temporal dimension, but when modeling continent-scale electricity systems that are large enough to contain multiple independent weather systems, this approach requires large numbers of representative periods to minimize errors in regional wind and solar capacity factors. Here, a new optimization-based approach for representative period selection and weighting is introduced that minimizes regional errors in average renewable capacity factors and electricity demand. The method delivers higher regional fidelity with fewer representative periods than alternative clustering methods when applied to wind, solar, and demand profiles for the contiguous United States. When representative periods are selected from multiple weather years, the optimized method reproduces regional averages with lower error than a complete 365-day time series from any single weather year. The method identifies only representative (as opposed to outlying) periods but can be combined with an iterative "stress period" identification approach to guide efficient decision-making considering both average and high-risk weather conditions.

24 POWER TRANSMISSION AND DISTRIBUTION↗

An Innovative Energy Management System for Microgrids with Multiple Grid-Forming Inverters

As increasingly more grid-forming (GFM) inverter-based resources replace traditional fossil-fueled synchronous generators as the GFM sources in microgrids, the existing microgrid energy management systems (EMS) need to be updated to control and coordinate multiple GFM inverters that consider system control objectives under different microgrid connection states. For each state, we formulate an optimization problem and apply a real-time feedback-based control algorithm; altogether, the control algorithms seamlessly connect the states into a generic microgrid EMS that controls the nodal voltages and frequencies, becomes a virtual power plant (VPP) when connected to the main grid, and coordinates power sharing responsibility among GFM sources when islanded. We showcase the EMS on a real-world simulation of a microgrid under the different states to demonstrate its operational effectiveness.

energy management systems↗

Development of High-Granularity Dual-Readout Calorimetry with psec Timing

Dual-readout and particle flow algorithm (PFA) are technologies proposed for precise jet energy measurement in future colliders. While PFA requires highly granular calorimeters, dual-readout has mainly been used with fiber-based calorimeters that do not have highly segmented capabilities. It is still non-trivial to combine these two technologies in one calorimeter system because of the use of fibers in most of the dualreadout calorimeters, which is not compatible with the high granularity requirement of PFA technologies. The aim of this study is to develop a novel calorimetry that combines dual-readout and PFA by adopting a highly segmented tile-based configuration. This paper compares the improvement of energy resolution using dual-readout approach across several configurations of highly granular hadron calorimeters through simulation. Results indicate that setups with fine sampling and close placement of scintillators and Cherenkov detectors improve dual-readout performance.

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND ↗

Model-free stabilization via Extremum Seeking using a cost neural estimator

In this paper, a fully model-free architecture for vertical stabilization of thermonuclear plasmas in tokamak experimental reactors is presented. For the first time, an Extremum Seeking control algorithm is combined with neural networks to estimate the Lyapunov function to be minimized, resulting in a fully data-driven control architecture. The performance of different neural networks are compared. Specifically, Multilayer Perceptrons and Extreme Learning Machines are considered. The proposed architecture is tested in simulation to show that it can counteract relevant plasma disturbances, resulting in a significant improvement in terms of the achievable operative space compared to the Extremum Seeking algorithm, which still relies on model-based cost estimator.

42 ENGINEERING↗

Simple algorithm for polarized parton evolution

We present an algorithm to include the correlation between the production and decay planes of gluons in a parton-shower simulation. The technique is based on identifying the charge currents responsible for the creation and annihilation of the vector field. It is applicable in both the hard-collinear and the soft wide-angle region. As a function of the number of particles, the algorithm scales linear in computing time and memory. We demonstrate agreement with fixed-order perturbative calculations in the relevant kinematical limits, and present a new observable that can be used to probe correlations beyond current-current interactions.

Höche, Stefan [Fermi National Accelerator Laborato↗

Data-Driven Analysis of Multipactor Dynamics via Dynamic Mode Decomposition

Multipactor effect is a performance-limiting kinetic plasma effect that can occur in high-power microwave and radio frequency (RF) devices. Multipactor effect is of special concern in vacuum or near-vacuum conditions such as those in particle accelerators and spaceborne devices. In this work, we present a data-driven reduced-order model (ROM) based on dynamic mode decomposition (DMD) for modeling of multipactor effects. We study multipactor effects and the resulting nonlinear harmonic generation by processing high-fidelity data generated from electromagnetic particle-in-cell (EMPIC) simulations using the DMD algorithm. We also investigate time-delay embedding extensions of DMD with improved generalizability and accuracy for modeling the electron plasma current density behavior. Here, the results show that DMD provides valuable insights into multipactor phenomena by extracting relevant modal spatiotemporal patterns and frequencies. In addition, DMD offers the potential to time extrapolate EMPIC simulations at a minimal cost, thereby reducing overall simulation time.

43 PARTICLE ACCELERATORS↗

CGSim: A Simulation Framework for Large Scale Distributed Computing Environment

Large-scale distributed computing infrastructures such as the Worldwide LHC Computing Grid (WLCG) require comprehensive simulation tools for evaluating performance, testing new algorithms, and optimizing resource allocation strategies. However, existing simulators suffer from limited scalability, hardwired algorithms, lack of real-time monitoring, and inability to generate datasets suitable for modern machine learning approaches. We present CGSim, a simulation framework for large-scale distributed computing environments that addresses these limitations. Built upon the validated SimGrid simulation framework, CGSim provides high-level abstractions for modeling heterogeneous grid environments while maintaining accuracy and scalability. Key features include a modular plugin mechanism for testing custom workflow scheduling and data movement policies, interactive real-time visualization dashboards, and automatic generation of event-level datasets suitable for AI-assisted performance modeling. We demonstrate CGSim’s capabilities through a comprehensive evaluation using production ATLAS PanDA workloads, showing significant calibration accuracy improvements across WLCG computing sites. Scalability experiments show near-linear scaling for multi-site simulations, with distributed workloads achieving 6 × better performance compared to single-site execution. The framework enables researchers to simulate WLCG-scale infrastructures with hundreds of sites and thousands of concurrent jobs within practical time budget constraints on commodity hardware.

Vatsavai, Sairam Sri [Brookhaven National Laborato↗

Kinetic Model of HoxEFU reduction by NADH [SWR-26-087]

This repository is used to release code generated for manuscripts on the Photosynthetic Energy Transduction core program. This code simulates the reduction of HoxEFU by NADH. The electro transfer rate constants for the simulation are specified in the .csv files. The two .csv files correspond tot he two models described in Dawson et al. Cell. Rep. Phys. Sci. 2026. The code utilizes a chemical master equation, a set of differential equations, defining the time evolution of the oxidation and reduction kinetics of NAD+, NADH, a FMN flavin, and a set of iron sulfur clusters. The kinetics of HoxEFU reduction by NADH are evaluated by numerical integration of the chemical master equation using a variable-time-step Runge-Kutta algorithm.

Dahl, Peter [National Laboratory of the Rockies (N↗

Development of Automated Atom Probe Tomography capability to study the influence of applied voltage and laser power on the final apparent composition of the analyzed specimen

This study presents the development and implementation of an autonomous Bayesian optimization (BO) framework for controlling and optimizing experimental parameters in Atom Probe Tomography (APT). Using commercial silicon needle samples as a benchmark system, we demonstrate that BO can efficiently navigate the complex parameter space of voltage and laser power to achieve target charge state ratios (specifically Si + /(Si + +Si 2+ )) with minimal experimental evaluations. Our implementation integrates Gaussian Process modeling with the CAMECA atom probe control framework, enabling autonomous adjustment of experimental conditions in real-time. Results show that the algorithm successfully converges to target ratios under different scenarios: maintaining a reference ratio, increasing the ratio (favoring Si 1+ ), and decreasing the ratio (favoring Si 2+ ). The system adapts to specimen evolution during analysis, compensating for changes in apex geometry while maintaining optimization targets. This work establishes a proof of concept for AI-driven optimization in APT, addressing the traditional challenges of manual parameter tuning and paving the way for applications to more complex materials where compositional accuracy is critical.

36 MATERIALS SCIENCE↗

Digital Twin Framework for PIP-II Linac: AI-Driven Multi-Scale Modeling from Ion Source to 800 MeV

The PIP-II linac will enable >1.2 MW beam power for DUNE, requiring unprecedented operational reliability across its warm front-end (RFQ, MEBT) and five distinct SRF sections operating at 162.5/325/650 MHz. We present a comprehensive digital twin framework uniquely combining a fully differentiable fast beam transport code with neural network surrogates trained on high-fidelity PIC simulations, capturing space charge and nonlinear dynamics beyond traditional envelope codes while achieving 10⁴ speedup at <1% accuracy. End-to-end differentiability enables gradient-based optimization across 500+ parameters simultaneously previously impossible with conventional tools while the model incorporates static/dynamic errors and serves as a virtual commissioning platform for diverse hardware integration. The framework facilitates reinforcement learning for pulsed/CW mode transitions, predictive maintenance through anomaly detection, and autonomous tuning algorithm development with real-time execution capability. Validation against physics simulations shows excellent agreement for the front-end, with initial results demonstrating potential for 30% commissioning time reduction and proactive fault mitigation, providing a scalable blueprint for operating next-generation high-intensity accelerators.

Pathak, Abhishek [Fermilab] (ORCID:000000021704208↗

Neutrino Physics with Deep Learning on NOvA

The NOvA experiment has made both νμ \nu_\mu disappearance and νe \nu_e appearance measurements in Fermilab's NuMI beam, and is working on cross section measurements using near detector data. At the core of NOvA's measurements is the use of deep learning algorithms for identification and reconstruction of the neutrino flavor and energy. These algorithms, used for the first time on NOvA in 2016, yielded large improvements in selection efficiency, and will be applied to our first anti-neutrino results to be released this year. Presented here is the extension of our deep learning efforts for identification of neutrino signal events, final state identification, single particle tagging, and reconstruction using instance segmentation techniques. We will describe the new implementations of modified Convolutional Neural Networks for anti-neutrino events, single particles and their performance for analysis final states selection, standard candle measurements, and reconstruction.

Psihas, Fernanda [Indiana U.]↗

An Innovative Energy Management System for Microgrids with Multiple Grid-Forming Inverters: Preprint

As more and more grid-forming (GFM) inverter-based resources replace traditional fossil-fueled synchronous generators as the GFM sources in microgrids, the existing microgrid energy management systems (EMSs) need to be updated to control and coordinate multiple GFM inverters that consider system control objectives under different microgrid connection states. For each state, we formulate an optimization problem and apply a real-time feedback-based control algorithm; all together, the control algorithms seamlessly connect the states into a generic EMS. We showcase the EMS on a real-world simulation of a microgrid under the different states to demonstrate its operational effectiveness.

energy management system↗

An Innovative Energy Management System for Microgrids with Multiple Grid-Forming Inverters

As increasingly more grid-forming (GFM) inverter-based resources replace traditional fossil-fueled synchronous generators as the GFM sources in microgrids, the existing microgrid energy management systems (EMS) need to be updated to control and coordinate multiple GFM inverters that consider system control objectives under different microgrid connection states.For each state, we formulate an optimization problem and apply a real-time feedback-based control algorithm; altogether, the control algorithms seamlessly connect the states into a generic microgrid EMS that controls the nodal voltages and frequencies, becomes a virtual power plant (VPP) when connected to the main grid, and coordinates power sharing responsibility among GFM sources when islanded. We showcase the EMS on a real-world simulation of a microgrid under the different states to demonstrate its operational effectiveness.

energy management system↗