Search NASA⌕ Search

SEARCH · Search NASA

Results for “Numerical Methods”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 217 records · Page 12

A Game-Theoretic Quantum Algorithm for Solving Magic Squares

Variational quantum algorithms (VQAs) offer a promising near-term approach to finding optimal quantum strategies for playing non-local games. These games test quantum correlations beyond classical limits and enable entanglement verification. In this work, we present a variational framework for the Magic Square Game (MSG), a two-player non-local game with perfect quantum advantage. We construct a value Hamiltonian that encodes the game’s parity and consistency constraints, then optimize parameterize quantum circuits to minimize this cost. Our approach build on the stabilizer formalism, leverages commutation structure for circuit design, and is hardware-efficient. Compared to existing work, our contribution emphasizes algebraic structure an interpretability. We validate our method through numerical experiments and outline generalizations to larger games.

Chehade, Sarah [ORNL]↗

Preliminary Study on Fine-Grained Power and Energy Measurements on Grace Hopper GH200 with Open-Source Performance Tools

The increasing adoption of tightly integrated, heterogeneous architectures, combined with the slowdown of Moore’s law, has made application power and energy-driven optimizations critical to efficiently use high-performance computing systems. This paper introduces a newly developed open-source toolkit that seamlessly integrates the Linux real-time hardware monitoring program hwmon with the Performance Application Programming Interface and the Score-P performance measurement system, thereby enabling fine-grained power and energy measurements for high-performance computing applications. Our primary target platform is the Wombat test bed, which is a system based on the NVIDIA GH200 superchip. The toolkit can capture transient power peaks with high temporal resolution (50 ms) and, thanks to Score-P integration, can map power metrics to specific code regions, thereby providing actionable information on power-intensive operations and inefficiencies. The toolkit also provides a holistic view of both the power and the energy consumption of the entire GH200 superchip by covering all major components: the Grace CPU, the Hopper GPU, and the I/O subsystem. Experiments that use Locally Self-consistent Multiple Scattering, which is an application for first-principles calculations of materials developed at Oak Ridge National Laboratory, have demonstrated the tool’s ability to identify transient power spikes and uncover opportunities for energy-aware optimizations. Additionally, we introduce a Python-based utility for converting Open Trace Format 2 traces to Parquet format, thus enabling advanced data analysis for numerical integration methods applied to power data for accurate energy profiling.

Hernandez Mendoza, Oscar [ORNL] (ORCID:00000002538↗

Stability-preserving Lossy Compression for Large-scale Partial Differential Equations

Checkpoint/Restart (C/R) strategies are vital for fault tolerance in PDE-based scientific simulations, yet traditional checkpointing incurs significant I/O overhead. Lossy compression offers a scalable solution by reducing checkpoint data size, but conventional methods often lack control over physical invariants (e.g., energy), leading to instability such as oscillations or divergence in Partial Differential Equations (PDE) systems. This paper introduces a stability-preserving compression approach tailored for PDE simulations by explicitly controlling kinetic and potential energy perturbations to ensure stable restarts. Extensive experiments conducted across diverse PDE configurations demonstrate that our method maintains numerical stability with minimal error magnification—even across multiple checkpoint-restart cycles—outperforming state-of-the-art lossy compressors. Parallel evaluations on the Frontier supercomputer show up to 8.4× improvement in checkpoint write performance and 6.3× in read performance, while maintaining relative L2 errors ∼ 2e-6 throughout continued simulation. These results provide practical guidance for balancing compression accuracy, stability, and computational efficiency in large-scale PDE applications.

Gong, Qian [ORNL] (ORCID:0000000235704142)↗

HygroThermFEM v1.0

HygroThermFEM is a Finite Element Method-based numerical calculation engine for solving 2-D heat and moisture transfer problems. This numerical engine is used in the THERM software tool, and its primary purpose is for the analysis of building envelopes (e.g., windows, walls, roofs, foundations, etc.). However, the engine can also be used for any heat and moisture transfer problems that require solving fundamental 2-D energy and mass transfer equations. Fluid flow solutions (Navier-Stokes momentum equations) are not included, but the correlations for various convection heat transfer situations are provided, including the translation of complex cavity geometries into those for which correlations are applicable. The calculation engine is written in C++ and includes an API for connecting to third-party tools.

Vidanovic, Dragan [Lawrence Berkeley National Labo↗

Portable, heterogeneous ensemble workflows at scale using libEnsemble

libEnsemble is a Python-based toolkit for running dynamic ensembles, developed as part of the DOE Exascale Computing Project. The toolkit utilizes a unique generator–simulator–allocator paradigm, where generators produce input for simulators, simulators evaluate those inputs, and allocators decide whether and when a simulator or generator should be called. The generator steers the ensemble based on simulation results. Generators may, for example, apply methods for numerical optimization, machine learning, or statistical calibration. libEnsemble communicates between a manager and workers. Flexibility is provided through multiple manager–worker communication substrates each of which has different benefits. These include Python’s multiprocessing, mpi4py, and TCP. Multisite ensembles are supported using Balsam or Globus Compute. We overview the unique characteristics of libEnsemble as well as current and potential interoperability with other packages in the workflow ecosystem. We highlight libEnsemble’s dynamic resource features: libEnsemble can detect system resources, such as available nodes, cores, and GPUs, and assign these in a portable way. These features allow users to specify the number of processors and GPUs required for each simulation; and resources will be automatically assigned on a wide range of systems, including Frontier, Aurora, and Perlmutter. Such ensembles can include multiple simulation types, some using GPUs and others using only CPUs, sharing nodes for maximum efficiency. We also describe the benefits of libEnsemble’s generator–simulator coupling, which easily exposes to the user the ability to cancel, and portably kill, running simulations based on models that are updated with intermediate simulation output. We demonstrate libEnsemble’s capabilities, scalability, and scientific impact via a Gaussian process surrogate training problem for the longitudinal density profile at the exit of a plasma accelerator stage. In conclusion, the study uses gpCAM for the surrogate model and employs either Wake-T or WarpX simulations, highlighting efficient use of resources that can easily extend to exascale.

Dynamic ensembles↗

Precision Computations in Strongly Coupled Conformal Field Theories (Final Technical Report)

Conformal Field Theories (CFTs) are quantum field theories that are invariant under the conformal symmetry group (which includes translations and rotations, but also local rescalings of spacetime). They are building blocks of general quantum field theories, and appear in many areas of physics, including statistical physics, condensed matter physics, particle physics, and quantum gravity. Because of their extra symmetries, the mathematical structure of CFTs is tightly constrained, and this leads to the idea of the ``conformal bootstrap," which is to use these mathematical structures to constrain, and in some cases determine, CFT observables. A new numerical implementation of the conformal bootstrap idea appeared in 2008 with the work of Rattazzi, Rychkov, Tonni, and Vichi. Their observation was that certain bootstrap constraints (conformal symmetry and unitarity) could be combined to yield a convex optimization problem that constraints CFT data. By solving this convex optimization problem on a computer, one could obtain bounds on observables like critical exponents and operator product expansion (OPE) coefficients. Over the course of this award, the PI has improved numerical bootstrap techniques by optimizing known algorithms and finding new ones for performing the required convex optimization computations. The PI has applied these techniques to compute high-precision observables in several important strongly-coupled systems. The PI has also explored both analytical and numerical bootstrap methods for constraining the space of low energy effective field theories of quantum gravity, and developed new analytical techniques for CFT and QFT more broadly.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

On the minimum number of radiation field parameters to specify gas cooling and heating functions

Fast and accurate approximations of gas cooling and heating functions are needed for hydrodynamic galaxy simulations. We use machine learning to analyze atomic gas cooling and heating functions in the presence of a generalized incident local radiation field computed by Cloudy. We characterize the radiation field through binned radiation field intensities instead of the photoionization rates used in our previous work. We find a set of 6 energy bins whose intensities exhibit relatively low correlation. We use these bins as features to train machine learning models to predict Cloudy cooling and heating functions at fixed metallicity. We compare the relative SHapley Additive exPlanation (SHAP) value importance of the features. From the SHAP analysis, we identify a feature subset of 3 energy bins (0.5-1, 1-4, and 13-16Ry) with the largest importance and train additional models on this subset. We compare the mean squared errors and distribution of errors on both the entire training data table and a randomly selected 20% test set withheld from model training. The machine learning models trained with 3 and 6 bins, as well as 3 and 4 photoionization rates, have comparable accuracy everywhere, with errors ≳10 times smaller than for the interpolation table of Gnedin and Hollon (2012). We conclude that 3 energy bins (or 3 analogous photoionization rates: molecular hydrogen photodissociation, neutral hydrogen HI, and fully ionized carbon CVI) are sufficient to characterize the dependence of the gas cooling and heating functions on our assumed incident radiation field model.

79 ASTRONOMY AND ASTROPHYSICS↗

Do We Know How to Model Reionization?

I compare the power spectra of the radiation fields from two recent sets of fully-coupled simulations that model cosmic reionization: “Cosmic Reionization On Computers” (CROC) and “Thesan”. While both simulations have similar power spectra of the radiation sources, the power spectra of the photoionization rate are significantly different at the same values of cosmic time or the same values of the mean neutral hydrogen fraction. However, the power spectra of the photoionization rate can be matched at large scales for the two simulations when the matching snapshots are allowed to vary independently. I.e., on large scales, the clustering of the radiation field in two simulations evolves similarly, but the exact timing of this evolution is different in different simulations and is not parameterized by an easily interpretable physical quantity like the mean neutral fraction or the mean free path. On small scales, large differences are present and remain partially unexplained. Both CROC and Thesan use the Variable Eddington Tensor approximation for modeling radiative transfer, but adopt different closure relations (optically thin OTVET versus M1). The role of this key difference is tested by using smaller simulations with a new cosmological simulation code that implements both closure relations in a controlled environment (the same hydro, cooling, and gravity solvers and the star formation recipe). In these controlled tests, both the M1 closure and the OTVET ansatz follow the expected behavior from a simple analytical approximation, demonstrating that the differences in the 2-point function of the radiation field induced by the choice of the Eddington tensor are not dominant.

79 ASTRONOMY AND ASTROPHYSICS↗

A geometric framework for momentum-based optimizers for low-rank training

Low-rank pre-training and fine-tuning have recently emerged as promising techniques for reducing the computational and storage costs of large neural networks. Training low-rank parameterizations typically relies on conventional optimizers such as heavy ball momentum methods or Adam. In this work, we identify and analyze potential difficulties that these training methods encounter when used to train low-rank parameterizations of weights. In particular, we show that classical momentum methods can struggle to converge to a local optimum due to the geometry of the underlying optimization landscape. To address this, we introduce novel training strategies derived from dynamical low-rank approximation, which explicitly account for the underlying geometric structure. Our approach leverages and combines tools from dynamical low-rank approximation and momentum-based optimization to design optimizers that respect the intrinsic geometry of the parameter space. We validate our methods through numerical experiments, demonstrating faster convergence, and stronger validation metrics at given parameter budgets.

Schotthoefer, Steffen [ORNL] (ORCID:00000002156965↗

Performance of high-order Godunov-type methods in simulations of astrophysical low Mach number flows

High-order Godunov methods for gas dynamics have become a standard tool for simulating different classes of astrophysical flows. Their accuracy is mostly determined by the spatial interpolant used to reconstruct the pair of Riemann states at cell interfaces and by the Riemann solver that computes the interface fluxes. In most Godunov-type methods, these two steps can be treated independently, so that many different schemes can in principle be built from the same numerical framework. Because astrophysical simulations often test out the limits of what is feasible with the computational resources available, it is essential to find the scheme that produces the numerical solution with the desired accuracy at the lowest computational cost. However, establishing the best combination of numerical options in a Godunov-type method to be used for simulating a complex hydrodynamic problem is a nontrivial task. In fact, formally more accurate schemes do not always outperform simpler and more diffusive methods, especially if sharp gradients are present in the flow. For this work, we used our fully compressible Seven-League Hydro (SLH) code to test the accuracy of six reconstruction methods and three approximate Riemann solvers on two- and three-dimensional (2D and 3D) problems involving subsonic flows only. We considered Mach numbers in the range from 10 −3 to 10 −1 , which are characteristic of many stellar and geophysical flows. In particular, we considered a well-posed, 2D, Kelvin–Helmholtz instability problem and a 3D turbulent convection zone that excites internal gravity waves in an overlying stable layer. Although the different combinations of numerical methods converge to the same solution with increasing grid resolution for most of the quantities analyzed here, we find that (i) there is a spread of almost four orders of magnitude in computational cost per fixed accuracy between the methods tested in this study, with the most performant method being a combination of a low-dissipation Riemann solver and a sextic reconstruction scheme; (ii) the low-dissipation solver always outperforms conventional Riemann solvers on a fixed grid when the reconstruction scheme is kept the same; (iii) in simulations of turbulent flows, increasing the order of spatial reconstruction reduces the characteristic dissipation length scale achieved on a given grid even if the overall scheme is only second order accurate; (iv) reconstruction methods based on slope-limiting techniques tend to generate artificial, high-frequency acoustic waves during the evolution of the flow; and (v) unlimited reconstruction methods introduce oscillations in the thermal stratification near the convective boundary, where the entropy gradient is steep.

79 ASTRONOMY AND ASTROPHYSICS↗

Conservative projection-based data-driven model order reduction of a fluid-kinetic spectral solver

Kinetic simulations are computationally intensive due to six-dimensional phase space discretization. Many kinetic spectral solvers use the asymmetrically weighted Hermite expansion due to its conservation and fluid-kinetic coupling properties, i.e., the lower-order Hermite moments capture and describe the macroscopic fluid dynamics, and higher-order Hermite moments describe the microscopic kinetic dynamics. We leverage this structure by developing a parametric data-driven reduced-order model based on the proper orthogonal decomposition, which projects the higher-order kinetic moments while retaining the fluid moments intact. We demonstrate analytically and numerically that the method ensures local and global mass, momentum, and energy conservation. The numerical results show that the proposed method effectively replicates the high-dimensional spectral simulations at a fraction of the computational cost and memory, as validated on the weak Landau damping and two-stream instability benchmark problems.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

Explicit Monotone Stable Super-Time-stepping Methods for Finite Time Singularities

We explore a novel way to numerically resolve the scaling behavior of finite-time singularities in solutions of nonlinear parabolic PDEs. The Runge–Kutta–Legendre (RKL) and Runge–Kutta–Gegenbauer (RKG) super-time-stepping methods were originally developed for nonlinear complex physics problems with diffusion. These are multistage single step second-order, forward-in-time methods with no implicit solves. The advantage is that the time-step size for stability scales with stage number 𝑠 as $\mathcal{O}$⁡(𝑠 2 ). Many interesting nonlinear PDEs have finite-time singularities, and the presence of diffusion often limits one to using implicit or semi-implicit time-step methods for stability constraints. Finite-time singularities are particularly challenging due to the large range of scales that one desires to resolve, often with adaptive spatial grids and adaptive time steps. Here, in this study, we show two examples of nonlinear PDEs for which the self-similar singularity structure has time and space scales that are resolvable using the RKL and RKG methods, without forcing even smaller time steps. Compared to commonly used implicit numerical methods, we achieve a significantly smaller run time while maintaining comparable accuracy. We also prove numerical monotonicity for both the RKL and RKG methods under their linear stability conditions for the constant coefficient heat equation, in the case of infinite domain and periodic boundary condition, leading to a theoretical guarantee of the superiority of the RKL and RKG methods over traditional super-time-stepping methods, such as the Runge-Kutta-Chebyshev and the orthogonal Runge-Kutta-Chebyshev methods. Code can be found at https://github.com/ZT220501/SRK-Singularity.

97 MATHEMATICS AND COMPUTING↗

Numerical integration in the virtual element method with the scaled boundary cubature scheme

Abstract The virtual element method (VEM) is a stabilized Galerkin method on meshes that consist of arbitrary (convex and nonconvex) polygonal and polyhedral elements. A crucial ingredient in the implementation of low‐ and high‐order VEM is the numerical integration of monomials and nonpolynomial functions over such elements. In this article, we apply the recently proposed scaled boundary cubature (SBC) scheme to compute the weak form integrals in various virtual element formulations over polygonal and polyhedral meshes. In doing so, we demonstrate the flexibility of the approach and the accuracy that it delivers on a broad suite of boundary‐value problems in 2D and 3D over polytopes with affine faces as well as on elements with curved boundaries. In addition, the use of the SBC scheme is exemplified in an enriched Poisson formulation of the VEM in which weakly singular functions are required to be integrated. This study establishes the SBC method as a simple, accurate and efficient integration scheme for use in the VEM.

Chin, Eric B.↗

Steady-state properties of multi-orbital systems using quantum Monte Carlo

A precise dynamical characterization of quantum impurity models with multiple interacting orbitals is challenging. In quantum Monte Carlo methods, this is embodied by sign problems. A dynamical sign problem makes it exponentially difficult to simulate long times. A multi-orbital sign problem generally results in a prohibitive computational cost for systems with multiple impurity degrees of freedom even in static equilibrium calculations. Here, we present a numerically exact inchworm method that simultaneously alleviates both sign problems, enabling simulation of multi-orbital systems directly in the equilibrium or nonequilibrium steady-state. The method combines ideas from the recently developed steady-state inchworm Monte Carlo framework [Erpenbeck et al., Phys. Rev. Lett. 130, 186301 (2023)] with other ideas from the equilibrium multi-orbital inchworm algorithm [Eidelstein et al., Phys. Rev. Lett. 124, 206405 (2020)]. We verify our method by comparison with analytical limits and numerical results from previous methods.

Chemistry↗

Numerical analysis of a time discretized method for nonlinear filtering problem with Lévy process observations

Abstract In this paper, we consider a nonlinear filtering model with observations driven by correlated Wiener processes and point processes. We first derive a Zakai equation whose solution is an unnormalized probability density function of the filter solution. Then, we apply a splitting-up technique to decompose the Zakai equation into three stochastic differential equations, based on which we construct a splitting-up approximate solution and prove its half-order convergence. Furthermore, we apply a finite difference method to construct a time semi-discrete approximate solution to the splitting-up system and prove its half-order convergence to the exact solution of the Zakai equation. Finally, we present some numerical experiments to demonstrate the theoretical analysis.

Mathematics↗

Effect of parallel flow on resonant layer responses in high beta plasmas

Abstract Resonant layers in a tokamak respond to non-axisymmetric magnetic perturbations by amplifying the mode amplitude and balancing the plasma rotation through magnetic reconnection and force balance, respectively. This resonant response can be characterized by local layer parameters and especially by a single quantity in the linear regime, the so-called inner-layer Δ. The computation of Δ under two-fluid drift-MHD formalism has been progressed by reducing the order of the system in the phase space, where the shielding current is approximated as being only carried by electrons, a posteriori . In this study, we relax the approximation and compute Δ accounted for by the parallel flow associated with the ion shielding current. The posteriori is numerically verified in great agreement with the original SLAYER developed in a previous paper (J.-K. Park 2022 Phys. Plasmas 29 072506). Extending the resonant layer response theory to high β plasmas, our research findings answer two important questions: how the parallel flow influences the resonant layer response and why the parallel flow effect appears in high β plasmas. The complicated plasma compression in high β regime allows the parallel flow response to give rise to the ion shielding current, which not only shifts the zero-crossing condition of the ExB flow but also enhances the field penetration threshold. Technically, the Riccati matrix transformation method is adapted to handle the numerical stiffness due to the increased order of the system. The high fidelity of this numerical method makes use of further extension of the model to higher-order systems to take other physical phenomena into account. This work is envisaged to predict the resonant layer response under high β fusion reactor conditions.

Lee, Yeongsun (ORCID:000000034474416X)↗

Mixing by internal gravity waves in stars: assessing numerical simulations against theory

ABSTRACT Here we present a study of radial chemical mixing in non-rotating massive main-sequence stars driven by internal gravity waves (IGWs), based on multidimensional hydrodynamical simulations with the fully compressible code MUSIC. We examine two proposed mechanisms of material mixing in stars by IGWs that are commonly quoted, relating to thermal diffusion and sub-wavelength shearing. Thermal diffusion provides a non-restorative effect to the waves, leaving material displaced from its previous equilibrium, while shearing arising within the waves drives weak localized flows, mixing the fluid there. Using IGW spectra from the simulations, we evaluate theoretical predictions of mixing rates due to these mechanisms. We show, for $20\, \mathrm{M}_\odot$ main-sequence stars, that neither of these mechanisms are likely to create mixing sufficient to correct inaccuracies in current stellar evolution models. Furthermore, we compare these predictions to results obtained from Lagrangian tracer particles, following a method recently used for global simulations of stellar interiors to measure mixing by IGWs in their radiative zones. We demonstrate that tracer particle methods face significant numerical challenges in measuring the small diffusion coefficients predicted by the aforementioned theories, for which they are prone to yielding artificially enhanced coefficients. Diffusion coefficients based on such methods are currently used with stellar evolution codes for asteroseismic studies, but should be viewed with caution. Finally, in a case where tracer particles do not suffer from numerical artefacts, we suggest that a diffusion model is not suitable for time-scales typically considered by 2D numerical simulations.

79 ASTRONOMY AND ASTROPHYSICS↗