Search NASA⌕ Search

SEARCH · Search NASA

Results for “efficient”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 361 records · Page 20

Sequential Kalman tuning of the t -preconditioned Crank-Nicolson algorithm: efficient, adaptive and gradient-free inference for Bayesian inverse problems

Ensemble Kalman Inversion (EKI) has been proposed as an efficient method for the approximate solution of Bayesian inverse problems with expensive forward models. However, when applied to the Bayesian inverse problem EKI is only exact in the regime of Gaussian target measures and linear forward models. Here, in this work we propose embedding EKI and Flow Annealed Kalman Inversion, its normalizing flow (NF) preconditioned variant, within a Bayesian annealing scheme as part of an adaptive implementation of the t-preconditioned Crank-Nicolson (tpCN) sampler. The tpCN sampler differs from standard pCN in that its proposal is reversible with respect to the multivariate t-distribution. The more flexible tail behaviour allows for better adaptation to sampling from non-Gaussian targets. Within our Sequential Kalman Tuning (SKT) adaptation scheme, EKI is used to initialize and precondition the tpCN sampler for each annealed target. The subsequent tpCN iterations ensure particles are correctly distributed according to each annealed target, avoiding the accumulation of errors that would otherwise impact EKI. We demonstrate the performance of SKT for tpCN on three challenging numerical benchmarks, showing significant improvements in the rate of convergence compared to adaptation within standard SMC with importance weighted resampling at each temperature level, and compared to similar adaptive implementations of standard pCN. The SKT scheme applied to tpCN offers an efficient, practical solution for solving the Bayesian inverse problem when gradients of the forward model are not available. Code implementing the SKT schemes for tpCN is available at https://github.com/RichardGrumitt/KalmanMC.

97 MATHEMATICS AND COMPUTING↗

Enhancing nitrogen fixation efficiency in glow-like discharge by reducing cathode-fall voltage

In plasma nitrogen fixation devices, discharge electrodes are crucial yet susceptible to oxidation and corrosion due to plasma’s high temperatures and oxygen content, which could alter discharge modes. This research evaluates the impact of different electrode materials, including iron, chromium, nickel, copper, and 304 stainless steel, on nitrogen fixation efficiency in glow-like discharges driven by high-voltage DC power. Notably, iron and 304 stainless steel cathodes undergo a mode transition at increased currents, evident from plasma color shifts and significant voltage reductions. Fourier transform infrared spectroscopy analyses reveal that such mode changes minimally affect nitrogen oxide production rates, leading to a notable decrease in energy consumption for nitrogen fixation by up to 40%. OES and SEM-EDS measurements suggest that iron oxide, with its higher secondary electron emission, replaces metal as the cathode material, facilitating mode transitions and maintaining discharge current at lower voltages. Further, this voltage change is largely attributed to the cathode voltage drop, highlighting the minimal role of the cathode fall region in NO x synthesis. These findings underscore the potential for improving plasma nitrogen fixation energy efficiency by choosing suitable cathode materials to lower the cathode-fall voltage.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

Light in the dark forest. Part I. An efficient optimal estimator for 3D Lyman-alpha forest power spectrum

The highly anisotropic nature of the Lyman-alpha (Lyα) forest data introduces a complex survey window function that complicates the measurement of the three-dimensional power spectrum ( P 3D ). In this paper, we present the first fully optimal estimator for P 3D , which exactly deconvolves the survey window function and marginalizes contaminated modes that distort the power spectrum. Our approach adapts optimal estimator techniques developed for the 2D cosmic microwave background data to the 3D case. To achieve computational feasibility, we employ the conjugate gradient method and implement the P 3 M formalism to handle large-scale and small-scale operations separately and efficiently. We validate our estimator using Monte Carlo mocks and Gaussian simulations, demonstrating its accuracy and computational efficiency. We confirm that mode marginalization eliminates distortions arising from quasar continuum errors and delivers robust power spectrum estimation, though it also inflates errors at large scales. This first implementation works in the flat-sky case; we discuss the remaining steps needed to generalize it to the curved-sky case. This formalism offers a foundation for the Lyα forest P 3D measurements and a new path toward cosmological constraints from the Lyα forest data.

Lyman alpha forest↗

Simultaneous enhancement of tritium burn efficiency and fusion power with low-tritium spin-polarized fuel

This study demonstrates that using spin-polarized deuterium-tritium (D-T) fuel with more deuterium than tritium can increase tritium burn efficiency (TBE) by at least an order of magnitude without compromising fusion power output, compared to unpolarized fuel. Although previous studies show that a low tritium fraction can enhance TBE, this strategy resulted in reduced fusion power density. The surprising improvement in TBE at fixed power reported here is due to the TBE increasing nonlinearly with decreasing tritium fraction but the fusion power density increasing roughly linearly with D-T cross section. A study is performed for an ARC-like tokamak producing 481 MW of fusion power with unpolarized 53:47 D-T fuel, finding the minimum startup tritium inventory (I startup,min ) is 0.69 kg. By spin-polarizing half of the fuel and using a 60:40 D-T mix, I startup,min is reduced to 0.08 kg, and fully spin-polarizing the fuel with a 63:37 D-T mix further reduces I startup,min to 0.03 kg. Some ARC-like scenarios are predicted to achieve plasma ignition with relatively modest spin polarization. These findings indicate that, with advancements in helium divertor pumping efficiency, TBE values of approximately 10%–40% could be achieved using low-tritium-fraction and spin-polarized fuel with minimal power loss. This would dramatically lower tritium startup inventory requirements and reduce the amount of on-site tritium. More generally than just for spin-polarized fuels, increased plasma performance can be used to increase TBE. This strongly motivates the development of spin-polarized fuels and low-tritium-fraction operation for burning plasmas.

11 NUCLEAR FUEL CYCLE AND FUEL MATERIALS↗

Investigating resource-efficient neutron/gamma classification ML models targeting eFPGAs

There has been considerable interest and resulting progress in implementing machine learning (ML) models in hardware over the last several years from the particle and nuclear physics communities. A big driver has been the release of the Python package, hls4ml, which has enabled porting models specified and trained using Python ML libraries to register transfer level (RTL) code. So far, the primary end targets have been commercial field-programmable gate arrays (FPGAs) or synthesized custom blocks on application specific integrated circuits (ASICs). However, recent developments in open-source embedded FPGA (eFPGA) frameworks now provide an alternate, more flexible pathway for implementing ML models in hardware. These customized eFPGA fabrics can be integrated as part of an overall chip design. In general, the decision between a fully custom, eFPGA, or commercial FPGA ML implementation will depend on the details of the end-use application. In this work, we explored the parameter space for eFPGA implementations of fully-connected neural network (fcNN) and boosted decision tree (BDT) models using the task of neutron/gamma classification with a specific focus on resource efficiency. We used data collected using an AmBe sealed source incident on Stilbene, which was optically coupled to an OnSemi J-series silicon photomultiplier (SiPM) to generate training and test data for this study. We investigated relevant input features and the effects of bit-resolution and sampling rate as well as trade-offs in hyperparameters for both ML architectures while tracking total resource usage. The performance metric used to track model performance was the calculated neutron efficiency at a gamma leakage of 10 -3 . The results of the study will be used to aid the specification of an eFPGA fabric, which will be integrated as part of a test chip.

47 OTHER INSTRUMENTATION↗

Stokes-dependent droplet collection efficiency on a NACA 0012 airfoil from droplet-informed simulations with statistical overloading

Accurate modelling of ice accretion on aircraft wings requires analysing droplet impingement on the surface to optimize the design of ice-protection systems. We perform Euler–Lagrange simulations of a droplet-laden flow impinging on a NACA 0012 airfoil. Our study includes water droplets with eight discrete sizes ranging from 1 to 160 microns. We vary the free-stream velocity of the incoming airflow in the range 60 ≤ U ≤ 240 m s −1 and the chord length of the airfoil in the range 0.5 ≤ c ≤ 2 m. Due to the dilute nature of supercooled clouds, one-way coupling is used in the simulations. The effects of droplet breakup and collision are also neglected. To reduce the computational cost, we employ statistical overloading of droplets, allowing us to simulate millions of impinging droplets in a time span on the order of milliseconds. Our results show that the droplet collection efficiency, which measures the likelihood of droplet impingement on the airfoil surface, increases with droplet size and free-stream velocity but decreases with airfoil size. We demonstrate that collection efficiency, impingement velocity and impingement angle are primarily dictated by a single non-dimensional parameter, the droplet Stokes number. We also identify a critical stagnation-streamline Stokes number below which impingements do not occur and use it to estimate the minimum droplet size for impingement. In addition, we observe droplet behaviour to become Stokes number independent at large values of the Stokes number. This article is part of the theme issue ‘Heat and mass transfer in frost and ice’.

Science & Technology - Other Topics↗

Efficient Measurement-Driven Eigenenergy Estimation with Classical Shadows

Quantum algorithms exploiting real-time evolution under a target Hamiltonian have demonstrated remarkable efficiency in extracting key spectral information. However, the broader potential of these methods, particularly beyond ground-state calculations, is underexplored. In this work, we introduce the framework of multiobservable dynamic mode decomposition (MODMD), which combines the observable dynamic mode decomposition (DMD), a measurement-driven eigensolver tailored for near-term implementation, with classical shadow tomography. MODMD leverages random scrambling in the classical shadow technique to construct, with exponentially reduced resource requirements, a signal subspace that encodes rich spectral information. Notably, we replace typical Hadamard-test circuits with a protocol designed to predict low-rank observables, thereby broadening the use of classical shadow tomography for predicting many low-rank observables. We establish theoretical guarantees on the spectral approximation from MODMD, taking into account distinct sources of error. In the ideal case, we prove that the spectral error scales as exp (−Δ⁢𝐸⁢𝑡 max ), where Δ⁢𝐸 is the Hamiltonian spectral gap and 𝑡 max is the maximal simulation time. This analysis provides a rigorous justification of the rapid convergence observed across simulations. To demonstrate the utility of our framework, we consider its application to fundamental tasks, such as determining the low-lying, i.e., ground or excited, energies of representative many-body systems. Our work paves the path for efficient designs of measurement-driven algorithms on near-term and early fault-tolerant quantum devices.

quantum algorithms & computation↗

Efficient Preparation of Dicke States

Here, we present an algorithm utilizing midcircuit measurement and feedback that prepares Dicke states with polylogarithmically many ancillae and polylogarithmic depth. Our algorithm uses only global midcircuit projective measurements and adaptively chosen global rotations. This improves over prior work that was only efficient for Dicke states of low weight or was not efficient in both depth and width. Our algorithm can also naturally be implemented in a cavity QED context using logarithmic time, zero ancillae, and atom-photon coupling scaling with the square root of the system size.

cavity methods↗

Efficient Simulation of Logical Magic State Preparation Protocols

Developing space- and time-efficient logical magic state preparation (MSP) protocols will likely be an essential step toward building a large-scale fault-tolerant quantum computer. Motivated by this need, we introduce a scalable method for simulating logical MSP protocols under the standard circuit-level noise model. When applied to protocols based on code-switching, magic state cultivation, and magic state distillation, our method yields a complexity polynomial in (i) the number of qubits and (ii) the nonstabilizerness, e.g., stabilizer rank or Pauli rank, of the target encoded magic state. The efficiency of our simulation method is rooted in a curious fact: every circuit-level Pauli error in these protocols propagates to a Clifford error at the end. This property is satisfied by a large family of protocols, including those that repeatedly measure a transversal Clifford that squares to a Pauli. We provide a proof-of-principle numerical simulation that prepares a magic state using such logical Clifford measurements. Our work enables practical simulation of logical MSP protocols without resorting to approximations or resource-intensive state-vector simulations.

71 CLASSICAL AND QUANTUM MECHANICS, GENERAL PHYSIC↗

Efficient many-jet event generation with flow matching

We apply for the first time, to the best of our knowledge, the flow matching method to the problem of phase-space sampling for event generation in high-energy collider physics. By training the model to remap the random numbers used to generate the momenta and helicities of the scattering matrix elements as implemented in the portable partonic event generator pepper, we find substantial efficiency improvements in the studied processes. We focus our study on the highest final-state multiplicities in Drell-Yan and top-antitop pair production used in simulated samples for the Large Hadron Collider, which computationally are the most relevant ones. We find that the unweighting efficiencies improve by factors of 184 and 25, respectively, when compared to the standard approach of using a vegas-based optimization. We also compare continuous normalizing flows trained with flow matching against the previously studied normalizing flows based on coupling layers and find that the former leads to better results, faster training and a better scaling behavior across the studied multiplicity range, while the latter evaluate faster. When combining the advantages of both methods using the regflow approach, we find parton-level unweighted event generation walltime gains of about a factor of 10 at the highest final-state multiplicities.

Bothmann, E. [CERN; Gottingen U.] (ORCID:000000016↗

Efficient simulation of low-temperature physics in one-dimensional gapless systems

Here, we discuss the computational efficiency of the finite-temperature simulation with minimally entangled typical thermal states (METTS). To argue that METTS can be efficiently represented as matrix product states, we present an analytic upper bound for the average entanglement Rényi entropy of METTS for a Rényi index 0 < q ≤ 1. In particular, for one-dimensional (1D) gapless systems described by conformal field theories, the upper bound scales as O⁡(cN 0 ⁢log⁡β) where c is the central charge and N is the system size. Furthermore, we numerically find that the average Rényi entropy exhibits a universal behavior characterized by the central charge and is roughly given by half of the analytic upper bound. Based on these results, we show that METTS can provide a speedup compared to employing the purification method to analyze thermal equilibrium states at low temperatures in 1D gapless systems.

71 CLASSICAL AND QUANTUM MECHANICS, GENERAL PHYSIC↗

Efficient state preparation for the Schwinger model with a theta term

We present a comparison of different quantum state preparation algorithms and their overall efficiency for the Schwinger model with a theta term. While adiabatic state preparation is proved to be effective, in practice it leads to large gate counts to prepare the ground state. The quantum approximate optimization algorithm (QAOA) provides excellent results while keeping the counts small by design, at the cost of an expensive classical minimization process. We introduce a “blocked” modification of the Schwinger Hamiltonian to be used in the QAOA that further decreases the length of the algorithms as the size of the problem is increased. The rodeo algorithm (RA) provides a powerful tool to efficiently prepare any eigenstate of the Hamiltonian, as long as its overlap with the initial guess is large enough. We obtain the best results when combining the blocked QAOA ansatz and the RA, as this provides an excellent initial state with a relatively short algorithm without the need to perform any classical steps for large problem sizes. Published by the American Physical Society 2025

Bazavov, Alexei (ORCID:0000000321411901)↗

Startup regime of high-efficiency tapering-enhanced FEL oscillator

In this paper, we present a design of a high-efficiency high-gain free-electron laser oscillator based on the use of a strongly tapered undulator for extracting energy from high-brightness electron beams. We provide an analytical model of the setup followed by numerical simulations for lasing at the wavelength of 13.5 nm. We discuss the optimization of the system in steady state and the conditions necessary for the pass-per-pass buildup of the power from shot noise level. We propose the use of fast phase shifters as a way to accelerate the buildup. Finally, we present time-dependent simulations of the oscillator and discuss the role of spectral filtering. The optimized working point yields a total energy conversion efficiency from the electron beam to output radiation above 1% at the wavelength of 13.5 nm.

Beam dynamics↗

Hardware-Efficient Quantum Phase Estimation via Local Control

Quantum phase estimation plays a central role in quantum simulation as it enables the study of spectral properties of many-body quantum systems. Most variants of the phase estimation algorithm require the application of the global unitary evolution conditioned on the state of one or more auxiliary qubits, posing a significant challenge for current quantum devices. In this work, we present an approach to quantum phase estimation that uses only locally controlled operations, resulting in a significantly reduced circuit depth. At the heart of our approach are efficient routines to measure the complex phase of the expectation value of the time-evolution operator, the so-called Loschmidt echo, for both circuit dynamics and Hamiltonian dynamics. By tracking changes in the phase during the dynamics, the routines trade circuit depth for increased sampling cost and classical postprocessing. Our approach does not rely on reference states and is applicable to any efficiently preparable state, regardless of its correlations. We provide a comprehensive analysis of the sample complexity and illustrate the results with numerical simulations. Our methods offer a practical pathway for measuring spectral properties in large many-body quantum systems using current quantum devices.

Schiffer, Benjamin F. [Max Planck Institute of Qua↗

Efficient Unitary Designs from Random Sums and Permutations

A unitary k-design is an ensemble of unitaries that matches the first k moments of the Haar measure. In this work, we provide two efficient constructions of k-designs on n-qubits using new random matrix theory techniques. Our first construction is based on exponentiating sums of random i.i.d. Hermitian matrices and uses O(k2n2)-many gates. In the spirit of central limit theorems, we show that this random sum approximates the Gaussian Unitary Ensemble (GUE). We then show that the product of just two exponentiated GUE matrices is already approximately Haar random. Our second construction is based on products of exponentiated sums of random permutations and uses Õ(k poly (n)) many gates. The k dependence is optimal (up to polylogarithmic factors) and is inherited from the efficiency of existing k-wise independent permutations. Furthermore, replacing random permutations with quantum-secure pseudorandom permutations (PRPs), we also obtain a pseudorandom unitary (PRU) ensemble that is secure under nonadaptive queries. A central feature of both proofs is a new connection between the polynomial method in quantum query complexity and the large-dimension (N) expansion in random matrix theory. In particular, the first construction uses the polynomial method to control high moments of certain random matrix ensembles without requiring delicate Weingarten calculations. In doing so, we define and solve a moment problem on the unit circle, asking whether a finite number of equally weighted points can reproduce a given set of moments. In our second construction, the key step is to exhibit an orthonormal basis for irreducible representations of the partition algebra that has a low-degree large-N expansion. This allows us to show that the distinguishing probability is a low-degree rational polynomial of the dimension N.

algebra↗

A Parallel Alternative for Energy-Efficient Neural Network Training and Inferencing

Energy efficiency of training and inferencing with large neural network models is a critical challenge facing the future of sustainable large-scale machine learning workloads. This paper introduces an alternative strategy, called phantom parallelism, to minimize the net energy consumption of traditional tensor (model) parallelism, the most energy-inefficient component of large neural network training. The approach is presented in the context of feed-forward network architectures as a preliminary, but comprehensive, proof-of-principle study of the proposed methodology. We derive new forward and backward propagation operators for phantom parallelism, implement them as custom autograd operations within an end-to-end phantom parallel training pipeline and compare its parallel performance and energy-efficiency against those of conventional tensor parallel training pipelines. Formal analyses that predict lower bandwidth and FLOP counts are presented with supporting empirical results on up to 256 GPUs that corroborate these gains. Experiments are shown to deliver ∼50% reduction in the energy consumed to train FFNs using the proposed phantom parallel approach when compared with conventional tensor parallel methods. Additionally, the proposed approach is shown to train smaller phantom models to the same model loss on smaller GPU counts as larger tensor parallel models on larger GPU counts offering the possibility for even greater energy savings.

Seal, Sudip [ORNL] (ORCID:0000000332330656)↗

On the Origin of Holes During Polarization Reset in Floating Body Ferroelectric FETs Towards Improving Switching Efficiency

In this article, we performed a comprehensive combined experimental and modeling study on the polarization reset mechanisms of floating body (i.e., channel) ferroelectric FETs, an important class of device with growing interests due to added functionalities and improved reliabilities. Using fully-depleted silicon-on-insulator (FDSOI) FeFET as a classical example, we demonstrate that: 1) without hole generation mechanisms, floating body FeFETs during reset is simply a capacitor divider, with negligible ferroelectric voltage drop for switching; ii) Band-to-band-tunneling (BTBT) around gate-to-S/D overlap even with zero drain bias generates holes to facilitate the reset in FDSOI FeFET, though at a slower speed and hold the reset state; iii) With scaling, S/D inner fringe field can enable fast reset, thus offering a potential efficiency boost approach; iv) a compact FDSOI FeFET model is developed that can capture the BTBT effect and reproduce the observed behaviors; v) the reset mechanism is also validated in a NAND string composed of FDSOI FeFETs, demonstrating its relevant applications. These insights show the strategies in improving reset efficiency, i.e., enhanced BTBT and inner fringe field.

42 ENGINEERING↗

Picasso: Memory-Efficient Graph Coloring Using Palettes With Applications in Quantum Computing

A coloring of a graph is an assignment of colors to vertices such that no two neighboring vertices have the same color. The need for memory-efficient coloring algorithms is motivated by their application in computing clique partitions of graphs arising in quantum computations where the objective is to map a large set of Pauli strings into a compact set of unitaries. We present Picasso, a randomized memory-efficient iterative parallel graph coloring algorithm with theoretical sublinear space guarantees under practical assumptions. The parameters of our algorithm provide a trade-off between coloring quality and resource consumption. To assist the user, we also propose a machine learning model to predict the coloring algorithm’s parameters considering these trade-offs. We provide a sequential and a parallel implementation of the proposed algorithm. We perform an experimental evaluation on a 64-core AMD CPU equipped with 512 GB of memory and an Nvidia A100 GPU with 40GB of memory. For a small dataset where existing coloring algorithms can be executed within the 512 GB memory budget, we show up to 68× memory savings. On massive datasets we demonstrate that GPU-accelerated Picasso can process inputs with 49.5× more Pauli strings (vertex set in our graph) and 2,478× more edges than state-of-the-art parallel approaches.

artificial intelligence, quantum computing↗