Search NASA⌕ Search

SEARCH · Search NASA

Results for “algorithm timings”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 433 records · Page 24

Lanczos Algorithm, the Transfer Matrix, and the Signal-to-Noise Problem

This Letter introduces a method for determining the energy spectrum of lattice quantum chromodynamics by applying the Lanczos algorithm to the transfer matrix and using a bootstrap generalization of the Cullum-Willoughby method to filter out spurious eigenvalues. Proof-of-principle analyses of the simple harmonic oscillator and the lattice quantum chromodynamics proton mass demonstrate that this method provides faster ground-state convergence than the “effective mass,” which is related to the power-iteration algorithm. Lanczos provides more accurate energy estimates than multistate fits to correlation functions with small imaginary times while achieving comparable statistical precision. Two-sided error bounds are computed for Lanczos results and guarantee that excited-state effects cannot shift Lanczos results far outside their statistical uncertainties.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

Bayesian optimization algorithms for accelerator physics

Accelerator physics relies on numerical algorithms to solve optimization problems in online accelerator control and tasks such as experimental design and model calibration in simulations. The effectiveness of optimization algorithms in discovering ideal solutions for complex challenges with limited resources often determines the problem complexity these methods can address. The accelerator physics community has recognized the advantages of Bayesian optimization algorithms, which leverage statistical surrogate models of objective functions to effectively address complex optimization challenges, especially in the presence of noise during accelerator operation and in resource-intensive physics simulations. In this review article, we offer a conceptual overview of applying Bayesian optimization techniques toward solving optimization problems in accelerator physics. We begin by providing a straightforward explanation of the essential components that make up Bayesian optimization techniques. We then give an overview of current and previous work applying and modifying these techniques to solve accelerator physics challenges. Finally, we explore practical implementation strategies for Bayesian optimization algorithms to maximize their performance, enabling users to effectively address complex optimization challenges in real-time beam control and accelerator design. Published by the American Physical Society 2024

43 PARTICLE ACCELERATORS↗

A Computational Framework to design 3D stiffness gradient acoustic metamaterials for impedance matching

Acoustic waves play a crucial role in various applications, including medical imaging, non-destructive testing, and sonar systems. One of the significant challenges in these applications is impedance matching, which is essential for minimizing reflections and maximizing the transfer of acoustic energy between different media. Acoustic metamaterials offer a promising solution to this challenge. In addition to impedance control, gradient stiffness can enhance structural efficiency and enable spatial control of wave propagation, making it a valuable feature in acoustic metamaterial design. In this pa- per, we present our developed computational method to design 3D stiffness gradient acoustic metamaterials for impedance matching. The key steps in our approach include generating initial designs using a periodic covariance function to provide unit cells that are both periodic on the boundaries and randomly formed inside the unit cell. Furthermore, we integrated manufacturing constraints into the design process, ensuring that the structures are interconnected for fabrication. We propose two computational optimization algorithms: GenUnit, based on a non-dominated sorting genetic algorithm (NSGA-II), and MLMatch, which leverages differentiable machine learning. The two approaches are not separate contributions but complementary com- ponents of a unified framework. GenUnit requires no training data and directly interfaces with physics-based simulations, making it highly accurate but slower for large-scale exploration. In contrast, MLMatch is data-hungry during training but, once trained, enables near-instantaneous inference and broad design-space coverage. Together, they form a hybrid strategy: ML- Match rapidly explores the global design space, and GenUnit provides local refinement with high-fidelity accuracy. This balance between training cost, inference time, and precision is the motivation for including both methods in the same study. We applied this dual-algorithm framework to generate two metallic-based metamaterial designs that match the acoustic impedance of water while exhibiting a controlled gradient in stiffness (from stiff to soft). The stiffness gradient is particularly advantageous in applications where one side of the structure must interface with soft or sensitive surfaces, such as human tissue or delicate components. Here, this work paves the way for improved materials in various acoustic applications, particularly in ultrasound devices, by providing better impedance.

Metamaterial↗

Initial position optimization in molecular dynamics simulations for a Coulomb system

A new algorithm for molecular dynamics (MD) simulations is developed to optimize plasma particle distributions at given initial temperatures. By combining velocity scaling and reassignment, the method effectively eliminates the initial rise and oscillation in temperatures observed with randomly distributed positions. These rises and oscillations are undesired numerical artifacts observed in conventional plasma MD simulations, arising from unoptimized particle positions. The algorithm demonstrates temperature relaxation without initial rises or oscillations, as well as precise flow velocity relaxation, enabling accurate measurement of relaxation times. The code is accelerated using graphics processing units for parallel processing, enhancing the study of plasma dynamics. The proposed method for distributing physically valid particles in MD simulations enables accurate studies of intrinsic collision processes in plasmas, including the dynamics of strongly coupled plasmas, plasma–wave interactions, and transport phenomena in magnetized plasmas. The paper concludes with a discussion of potential applications and future enhancements to the algorithm.

Jo, Jawon (ORCID:0009000924193285)↗

Toroid Signal Processing Container with Redis Integration

This report presents the research, design, and implementation of a Toroid Signal Processing Container with Redis Integration, conducted during my internship assignment. The project focused on developing a modular signal processing sys- tem capable of performing baseline correction, droop compensation, and real- time data communication for beam diagnostic signals. Using Redis as a message broker and configuration manager, the system ensures modularity, scalability, and efficient inter-process communication. Signal correction algorithms, such as Asymmetric Least Squares (ALS) baseline smoothing and a numerical droop correction approach, were implemented and tested using Linac beam data. The results confirm that this system improves signal fidelity and supports real-time analysis requirements, offering valuable contributions to the field of beam instrumentation and data acquisition systems.

Bowers, Elliot [Cabrillo Coll.]↗

Toroid Signal Processing Container with Redis Integration

This report presents the research, design, and implementation of a Toroid Signal Processing Container with Redis Integration, conducted during my internship assignment. The project focused on developing a modular signal processing sys- tem capable of performing baseline correction, droop compensation, and real- time data communication for beam diagnostic signals. Using Redis as a message broker and configuration manager, the system ensures modularity, scalability, and efficient inter-process communication. Signal correction algorithms, such as Asymmetric Least Squares (ALS) baseline smoothing and a numerical droop correction approach, were implemented and tested using Linac beam data. The results confirm that this system improves signal fidelity and supports real-time analysis requirements, offering valuable contributions to the field of beam in- strumentation and data acquisition systems.

Bowers, Elliot [Cabrillo Coll.; Fermilab]↗

Distributed Stochastic Optimization of a Neural Representation Network for Time-Space Tomography Reconstruction

4D time-space reconstruction of dynamic events or deforming objects using X-ray computed tomography (CT) is an important inverse problem in non-destructive evaluation. Conventional back-projection based reconstruction methods assume that the object remains static for the duration of several tens or hundreds of X-ray projection measurement images (reconstruction of consecutive limited-angle CT scans). However, this is an unrealistic assumption for many in-situ experiments that causes spurious artifacts and inaccurate morphological reconstructions of the object. To solve this problem, we propose to perform a 4D time-space reconstruction using a distributed implicit neural representation (DINR) network that is trained using a novel distributed stochastic training algorithm. Our DINR network learns to reconstruct the object at its output by iterative optimization of its network parameters such that the measured projection images best match the output of the CT forward measurement model. Here, we use a forward measurement model that is a function of the DINR outputs at a sparsely sampled set of continuous valued 4D object coordinates. Unlike previous neural representation architectures that forward and back propagate through dense voxel grids that sample the object's entire time-space coordinates, we only propagate through the DINR at a small subset of object coordinates in each iteration resulting in an order-of-magnitude reduction in memory and compute for training. DINR leverages distributed computation across several compute nodes and GPUs to produce high-fidelity 4D time-space reconstructions. We use both simulated parallel-beam and experimental cone-beam X-ray CT datasets to demonstrate the superior performance of our approach.

36 MATERIALS SCIENCE↗

Block encoding bosons by signal processing

Block Encoding (BE) is a crucial subroutine in many modern quantum algorithms, including those with near-optimal scaling for simulating quantum many-body systems, which often rely on Quantum Signal Processing (QSP). Currently, the primary methods for constructing BEs are the Linear Combination of Unitaries (LCU) and the sparse oracle approach. In this work, we demonstrate that QSP-based techniques, such as Quantum Singular Value Transformation (QSVT) and Quantum Eigenvalue Transformation for Unitary Matrices (QETU), can themselves be efficiently utilized for BE implementation. Specifically, we present several examples of using QSVT and QETU algorithms, along with their combinations, to block encode Hamiltonians for lattice bosons, an essential ingredient in simulations of high-energy physics. We also introduce a straightforward approach to BE based on the exact implementation of Linear Operators Via Exponentiation and LCU (LOVE-LCU). We find that, while using QSVT for BE results in the best asymptotic gate count scaling with the number of qubits per site, LOVE-LCU outperforms all other methods for operators acting on up to qubits, highlighting the importance of concrete circuit constructions over mere comparisons of asymptotic scalings. Using LOVE-LCU to implement the BE, we simulate the time evolution of single-site and two-site systems in the lattice theory using the Generalized QSP algorithm and compare the gate counts to those required for Trotter simulation.

Kane, Christopher F↗

Efficient derivative computation for unsteady fatigue-constrained nonlinear aero-structural wind turbine blade optimization

Gradient-based optimization offers significant efficiency advantages for wind turbine blade design, but its application has often been limited by the cost and accuracy of finite-difference derivative calculations, especially when fatigue constraints are considered. In this work, we systematically compare and evaluate four differentiation techniques, namely algorithmic differentiation, implicit differentiation, sparsity exploitation, and parallelization, to determine their effectiveness in computing accurate gradients through time-domain aero-structural simulations. By integrating these techniques with unsteady nonlinear aerodynamic and structural models, we develop software designed for accurate gradient computation. We show that combining these techniques addresses memory and runtime challenges associated with long simulations required by design load cases. Specifically, the most effective combination reduces derivative computation wall time by over an order of magnitude compared to finite differencing while maintaining superior accuracy. We demonstrate this approach in a proof-of-concept aero-structural optimization of a wind turbine blade that improves the cost of energy by 12.78 %. This comparative study establishes a viable approach for fatigue-aware blade design that balances computational efficiency with modeling accuracy.

17 WIND ENERGY↗

Special Nuclear Material Mass Estimates from Neutron Singles Count Rate [Poster]

The objective of this research was to create an algorithm to provide an estimate of special nuclear material (SNM) mass using only neutron count rate data from a Radioisotope Identification Device (RIID), rather than using time-correlated data from a neutron multiplicity counter. To meet this objective neutron count rate measurements of a 252 Cf source were taken at varying distances with an ORTEC Detective X, FLIR Identifinder 2, and an ORTEC RADEAGLET-R. An algorithm was created to estimate mass of SNM utilizing the singles rate equation and the measured absolute efficiency curves.

FLIR↗

Aerial drone fleet deployment optimization with endogenous battery replacements for direct delivery of time-sensitive products

Aerial drones offer a distinct potential to reduce the delivery time and energy consumption for the delivery of time-sensitive and small products. However, there is still a need in the relevant industry to understand the performance of drone-based delivery under different business needs and drone operating conditions. We studied a drone deployment optimization problem for direct delivery of time-sensitive products with release dates to customers maintaining a specified time window. This paper presents a new mixed-integer programming model, new valid inequalities, a new greedy heuristic algorithm, and a Genetic algorithm to help business owners optimally schedule and route their drone fleet minimizing the required fleet size, the required number of additional batteries, and total energy consumption. A realistic feature of the optimization method is that instead of replacing the drone battery after each return to the depot, it keeps track of the remaining energy in the drone battery and decides on battery replacements accounting for the drone routing and the user-specified minimum required battery energy. Numerical results based on real data from drone flight tests and prepared food delivery industry provide insights into the effect of different practical drone operating parameters on the required fleet size, the required number of battery replacements, and energy consumption. Here, results demonstrate that the proposed heuristic algorithm substantially outperforms the accelerated CPLEX in runtime while sacrificing the solution quality by a small amount. Additionally, results show that using a mixed fleet of hexacopter and quadcopter drones reduces the total energy consumption by 48.52% compared to using a homogeneous fleet of only hexacopters.

Drone energy consumption↗

Advanced Shuttle Strategies for Parallel QCCD Architectures

Trapped ions (TIs) are at the forefront of quantum computing implementation, offering unparalleled coherence, fidelity, and connectivity. However, the scalability of TI systems is hampered by the limited capacity of individual ion traps, necessitating intricate ion shuttling for advanced computational tasks. The quantum charge-coupled device (QCCD) framework has emerged as a promising solution, facilitating ion mobility for universal quantum computation. Current QCCD architectures predominantly feature a linear topology, which is increasingly recognized as inefficient for complex quantum operations. Anticipating the shift toward more efficacious designs, this article introduces an innovative quantum scheduling strategy optimized for parallel QCCD topologies. Our strategy proposes a probabilistic formula for ion movement, alongside ingenious methods for local layer generation and layer compression, yielding a significant reduction in ion shuttle times. Through simulations, we demonstrate that our strategy not only substantially outstrips the linear model but also exhibits better performance over other parallel strategies that employ greedy algorithms. This is achieved through our nuanced resolution of complexities, such as traffic blocks and trap capacity limitations. The consequent reduction in shuttle operations leads to lower energy consumption and an enhancement in the quantum computer's fidelity, ultimately accelerating program execution times.

43 PARTICLE ACCELERATORS↗

Statistically elucidated responses from low-signal contrast mechanisms in ultrafast electron microscopy

The emergence of ultrafast electron microscopy (UEM) has enabled the discovery of strongly correlated dynamic mechanisms, including electron–phonon coupling, structural phase transitions, thermal transport, and electromagnetic deflection. Most UEM systems operate stroboscopically, meaning that the technique is susceptible to artifacts, mistakes, and misinterpretation of the data due to extensive experimental effort. In contrast to the ultrafast designation, data acquisition is extraordinarily slow because the electron beam has significantly reduced signal compared to traditional transmission electron microscopy due to pulsing the electron beam. Consequently, the sample may drift, tilt, or undergo irreversible structural changes that are independent of the time-resolved dynamics throughout the experimental time frame. Furthermore, these datasets require significant user interpretation that can be problematic when proper controls are not implemented thoroughly. Here, we demonstrate a new algorithm designed to separate ultrafast structural dynamics from long-term artifacts using a LiNbO 3 sample experiencing electrically driven surface acoustic wave propagation. Additionally, we provide examples of the impact of user bias when analyzing the data and provide a methodology, which enables the extraction of time-resolved responses when the image signal is extraordinarily low. Overall, the goal of this publication is to provide methods that validate the experimental results and reduce researcher biases during UEM data interpretation.

75 CONDENSED MATTER PHYSICS, SUPERCONDUCTIVITY AND↗

Automated Hybrid Variance Reduction on Advanced Architectures in the Shift Monte Carlo Code

Monte Carlo transport methods are the most accurate schemes for solving problems with complex energy and spatial features, but they come with a high computational cost. Although hybrid methods have enabled the use of Monte Carlo transport for a large class of problems, they still require significant computing resources. Modern multicore CPUs with large numbers of compute cores and graphical processing units (GPUs) provide opportunities to optimize the memory and run-time costs of hybrid Monte Carlo methods. This paper documents the development and analysis of three Monte Carlo transport algorithms that support hybrid transport using the consistent adjoint-driven importance sampling (CADIS) and forward-weighted CADIS methods in the Shift Monte Carlo code: history-based transport using static and dynamic threading on multicore CPUs and event-based transport enabling weight window tracking on GPUs. The results are shown for two challenging hybrid problems on the Frontier supercomputer at the Oak Ridge Leadership Computing Facility. The results show that all three methods yield good performance and enable solutions of difficult fixed-source transport problems in less than 2 min on 20 nodes of Frontier. Dynamic threading was observed to give up to 20% better scaling behavior than static threading. Moreover, the AMD Instinct 250X GPU was found to give 9 to 11 times greater throughput per graphics compute die than the best CPU performance. In conclusion, additional opportunities for optimization of hybrid transport on GPUs are discussed.

Denovo↗

Chaos as the cause of randomness in tearing mode onset times in DIII-D ITER baseline scenario plasmas

The first evidence of widespread chaotic plasma dynamics relevant to the onset of disruptive tearing modes is found in DIII-D tokamak discharges. This evidence is obtained by calculating positive maximal Lyapunov exponents from Mirnov signals using the Rosenstein algorithm, implying the chaotic growth and/or rotation of magnetohydrodynamic modes in DIII-D [Luxon, Nucl. Fusion 42, 614 (2002)] plasmas. Such chaotic dynamics offer an explanation for the random tearing mode onset times previously observed in ITER baseline scenario discharges. Further, the Lyapunov timescales associated with this chaos could inhibit reliable prediction of tearing mode onset beyond the angular momentum confinement timescale; this constraint should inform the design of plasma control systems on ITER.

Chaotic dynamics↗

Scalable Circuit Cutting and Scheduling in a Resource-constrained and Distributed Quantum System

Despite quantum computing's rapid development, current systems remain limited in practical applications due to their limited qubit count and quality. Various technologies, such as superconducting, trapped ions, and neutral atom quantum computing technologies are progressing towards a fault tolerant era, however they all face a diverse set of challenges in scalability and control. Recent efforts have focused on multi-node quantum systems that connect multiple smaller quantum devices to execute larger circuits. Future demonstrations hope to use quantum channels to couple systems, however current demonstrations can leverage classical communication with circuit cutting techniques. This involves cutting large circuits into smaller subcircuits and reconstructing them post-execution. However, existing cutting methods are hindered by lengthy search times as the number of qubits and gates increases. Additionally, they often fail to effectively utilize the resources of various worker configurations in a multi-node system. To address these challenges, we introduce FitCut, a novel approach that transforms quantum circuits into weighted graphs and utilizes a community-based, bottom-up approach to cut circuits according to resource constraints, e.g., qubit counts, on each worker. FitCut also includes a scheduling algorithm that optimizes resource utilization across workers. Implemented with Qiskit and evaluated extensively, FitCut significantly outperforms the Qiskit Circuit Knitting Toolbox, reducing time costs by factors ranging from 3 to 2000 and improving resource utilization rates by up to 3.88 times on the worker side, achieving a system-wide improvement of 2.86 times.

Kan, Shuwen [Fordham University]↗

Monitoring and modeling hydrologic conditions in Ukraine for hydropower generation

Study region: The Dnieper and Dniester Rivers of Ukraine. Study focus: The ongoing conflict in Ukraine has caused disruptions to electricity generation, of which hydroelectric sources contribute approximately 9 % to the country’s needs. With the takeover of the Zaporizhzhia nuclear power plant by enemy forces, the loss of the Kakhovka hydroelectric dam, and the future impacts of the conflict on electricity generation unclear, it may be valuable for the Ukrainian government to better understand how it could leverage hydroelectric power sources in the near future. Unfortunately, measurements of river discharge throughout Ukraine ceased data collection in the late 1980’s to early 1990’s. To address this data gap, we developed a protocol that combined satellite-based time-series measurements of river width at seven locations throughout Ukraine from 2013 to 2023 with reanalysis data, climate-model predictions, and hydrologic models to both provide a means of monitoring a proxy for near-real-time discharge and also predict near-term (i.e., 2023–2030) hydrologic patterns for the region. New hydrological insights for the region: We ran new algorithms on 144 WorldView-2 and WorldView-3 satellite images to map rivers and extract width, one of which was validated against river gauge data located along the same river but in a neighboring country. Hydrologic models using two climate scenarios found minimal change in annual discharge at all sites, but magnitude and timing of peak discharge showed a moderate trend. The results suggest that hydropower is underutilized in Ukraine.

13 HYDRO ENERGY↗

Variable rate neural compression for sparse detector data

Particle colliders produce data at extraordinary rates, posing major challenges for transmission and storage. High-throughput compression algorithms are therefore essential. In the sPHENIX experiment taking data at the Relativistic Heavy Ion Collider, a time projection chamber records three-dimensional (3D) particle trajectories that are highly sparse, making conventional learning-free lossy compression ineffective. Convolutional neural networks have surpassed traditional methods in compression ratio and accuracy. However, they fail to exploit sparsity for efficiency. To address these gaps, we present BCAE-VS, a bicephalous convolutional autoencoder with variable compression ratio for sparse data, which adapts compression to input complexity through key-point identification and sparse convolution. BCAE-VS achieves higher accuracy and compression ratios than prior neural approaches while being orders of magnitude smaller. Moreover, its throughput increases with sparsity—a property not observed in other methods. Although it was developed for collider experiments, BCAE-VS readily extends to other sparse data domains, such as light detection and ranging (LiDAR) sensing and 3D microscopy.

73 NUCLEAR PHYSICS AND RADIATION PHYSICS↗