Leveraging Approximation Theory for Efficient Scientific Machine Learning
Explore the source record for details and available documents.
SEARCH · Search NASA
Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.
Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.
Explore the source record for details and available documents.
Molecular simulations of biological and physical phenomena generally involve sampling complicated, rough energy landscapes characterized by multiple local minima. In this work, we introduce a new family of methods for advanced sampling that draw inspiration from functional representations used in machine learning and approximation theory. As shown here, such representations are particularly well suited for learning free energies using artificial neural networks. As a system evolves through phase space, the proposed methods gradually build a model for the free energy as a function of one or more collective variables, from both the frequency of visits to distinct states and generalized force estimates corresponding to such states. Implementation of the methods is relatively simple and, more importantly, for the representative examples considered in this work, they provide computational efficiency gains of up to several orders of magnitude over other widely used simulation techniques.
The neutron survival probability (and related quantities including probabilities of extinction and initiation) is a central element of the broader stochastic theory of neutron populations and finds application in fields including reactor start-up, analysis of reactor power bursts and criticality accidents, and safeguards. In a full neutron transport formulation, the equation governing the single-neutron survival probability is a backward or adjoint-like integro-partial differential equation with the added complexity of being highly nonlinear. Analogous formulations of this equation exist in the context of many approximate theories of neutron transport, with the point kinetics formulation having received significant theoretical attention since the 1940s. This work continues this tradition by providing a novel analysis of the single-neutron survival probability equation using the tools of boundary layer theory. The analysis reveals that the “fully dynamic” solution of the single-neutron survival probability equation—and some key probability distributions derived from it—may be cast as a singular perturbation around the underlying quasi-static single-neutron probability of initiation. In this perturbation solution, the expansion parameter is the ratio of the neutron generation time to a macroscopic time scale characterizing the overall system evolution; this interpretation illuminates some of the fundamental structural aspects of neutron survival phenomena.
We analyzed the hydroxyl librational signatures of five structurally related aluminum (oxy)hydroxides, using inelastic neutron scattering (INS) and plane-wave lattice dynamics simulations. A clear trend across these aluminum-containing phases illustrates the relationship between hydrogen bonding, local atomic structure, and the spectral location and profile of the librational bands. The INS spectra have been compared to previous optical spectroscopy and computational studies, highlighting the complementary nature of the INS technique. Taking into account other structurally or chemically related material analogs, we have identified a correlation between a blueshift (to higher energy) of the upper librational band edge and the geometry of the hydrogen bond interactions, mirroring (with opposite correlation) the well-known redshift in the intramolecular O–H stretching energy with increasing hydrogen bond strength. For hydroxyl groups that do not participate in hydrogen bonding effectively, the bending librations occur at lower energies and hybridize with metal–oxygen lattice modes. Standard density functional theory approximations, including dispersion corrections, struggle to correctly predict vibrational frequencies of motions dominated by H but perform well for metal–oxygen modes, allowing us to make detailed mode assignments in several cases, including a demonstration of how layer-to-layer disorder in boehmite hydrogen bond orientations is reflected in the sharp but minor low energy peaks (at ∼70–80 meV) of the INS spectrum.
Sustained advances in the mathematics of modeling and simulation have resulted in the capability today for routine simulation of a number of large scale complex DOE-relevant systems. As remarkable as this capability for solving the so-called forward problem is, it is typically only the first step-an inner loop within an outer loop that explores the simulation model's parameter space and decision space to characterize uncertainty in the model's predictions, learn unknown model parameters from data, design the most informative experiments, determine optimal control strategies, and create optimal designs. Broadly, what unifies all of these outer loop problems is that they are, in one form or another, optimization problems over parameter/control/design space that are constrained by complex uncertain models. To fully realize the power of scientific simulation as a basis for scientific discovery, technological innovation, and rational decision-making, it is imperative to move beyond simulation to tackle the outer loop of optimization for learning from data, experimental design, and control with complex uncertain models. When the models under consideration are large-scale and complex, and when the optimization variable and uncertain parameter spaces are high (or infinite) dimensional, this constitutes a grand challenge of the highest order, and is intractable with conventional methods. To overcome these challenges, the AEOLUS Center was established to develop a unified mathematical, computational, and statistical framework for (1) Learning predictive models from complex data via Bayesian inference and optimization, and (2) Optimizing experiments, processes, and designs using the resulting uncertain models. These problems are intractable with conventional methods, for several reasons: (1) The simulation problems that govern the inner loops of the optimization problems are expensive to execute (due to severe nonlinearity, heterogeneity, multiphysics/multiscale coupling); (2) The optimization variable and uncertain parameter spaces are high dimensional, often stemming from discretizations of infinite dimensional fields such as initial conditions, sources, or material properties. We argue that the key to overcoming these challenges is to develop new mathematical, computational, and statistical methods that exploit the structure of the Bayesian inference and optimization problems mediated by their underlying complex uncertain models. This structure includes the regularity, sparsity, geometry, low intrinsic dimensionality, and multifidelity nature of the maps from uncertain parameter/optimization variable spaces to the specific objectives targeted: Bayesian inference, optimal experimental design, and optimal control design. Black box methods developed as generic tools are incapable of exploiting this structure. To be successful, we must create, integrate, and cross-fertilize ideas across multiple areas of applied math--including approximation theory, Bayesian inference, data science, experimental design, information theory, machine learning, model reduction, optimal control theory, parallel algorithms, PDE-constrained optimization, randomized algorithms, stochastic optimization, and uncertainty quantification--all while exploiting the structure of the problems at hand. With this goal in mind, we have marshaled a team of leading authorities in these areas. While the methods we develop will be broadly applicable across a wide spectrum of DOE problems in which experiments inform models and the systems those models describe must be optimized under uncertainty, we have chosen a specific area, advanced manufacturing and materials, to drive our work. AMM is characterized by complex models across multiple scales, and is a rich source of challenging problems in inference, experimental design, and optimal control, requiring multifaceted and integrated advances in applied mathematics. As such, AMM serves as an excellent vehicle to motivate and demonstrate the advances in applied mathematics developed by our center.
Security-constrained unit commitment (SCUC) is a key component in power system operations. When AC power flow constraints are considered in the SCUC model (AC-SCUC), the problem becomes extremely difficult due to its discrete and non-convex nature, as described in “Grid Optimization Competition Challenge 3 Problem Formulation (GOCC)”. There are four main challenges: (i) Discrete decisions regarding unit online/offline status and start-up/shut-down procedures for every single unit. The number of discrete decision variables increases considerably when a system integrates multiple generators; (ii) Configuration-based combined-cycle formulations, and multi-commodity models that include ramping products, spin/non-spin products, and regulation up/down products. The combined-cycle units introduce additional discrete decision variables and auxiliary service products further complicate the model by connecting multi-commodity products’ continuous and discrete variables; (iii) SCUC models with AC power flow constraints are far more complex due to massive bilinear terms in the large-scale nonlinear power balance equations. The nonlinear power balance equations are further complicated by the discrete step control variables of shunts; (iv) N − 1 contingency analysis. The size of the model increases linearly with the number of contingencies considered, greatly increasing the size of the optimization model. Accordingly, there is an emergent need to develop a robust algorithm capable of deriving a high-quality solution in a short time and passing through contingency tests simultaneously. In this project, we explore innovative techniques to address this challenging problem by integrating advanced polyhedral theory, approximation methods, relaxation strategies, decomposition techniques, and parallel computing. Each technique approaches the problem from a different perspective, leveraging its specific strengths to tackle distinct challenges. Each individual method has demonstrated its effectiveness in the PI’s previous research. Their integration is expected to significantly reduce the computational time required to solve the proposed complex problem. Successful completion of this project has the potential to transform the industry by enhancing optimization solvers capable of handling large-scale day-ahead energy market clearing models within strict time constraints, while incorporating AC power flow constraints. This advancement will lead to reduced overall generation costs and, consequently, increased social welfare.
The electron temperature of photoionized plasmas characterizes the thermalization of photoelectrons, impacts the charge-state distribution, emissivity, and opacity through atomic recombination processes, and is needed to perform detailed comparisons with theory predictions. We discuss temperature measurements in laboratory photoionized plasmas and a comparison with model calculations done with several theory approximations and codes. These include a radiation-hydrodynamics simulation and two nonequilibrium heating and ionization models that tracked the evolution of the internal energy of the electrons. Furthermore, for the same physics model and X-ray flux time history, calculations were performed assuming steady-state or time-dependent conditions. The time history of steady-state results correlates with that of the X-ray flux, while that of the time-dependent cases does not, and it is qualitatively and quantitatively different from the steady-state case. Steady-state results significantly overestimated temperature measurements, while time-dependent results produced better approximations, which suggests the importance of transient effects in the experiment and also the need for time-resolved measurements.
Low-Rank Adaptation (LoRA) has become a widely used method for parameter-efficient fine-tuning of large-scale, pre-trained neural networks. However, LoRA and its extensions face several challenges, including the need for rank adaptivity, robustness, and computational efficiency during the fine-tuning process. We introduce GeoLoRA, a novel approach that addresses these limitations by leveraging dynamical low-rank approximation theory. GeoLoRA requires only a single backpropagation pass over the small-rank adapters, significantly reducing computational cost as compared to similar dynamical low-rank training methods and making it faster than popular baselines such as AdaLoRA. This allows GeoLoRA to efficiently adapt the allocated parameter budget across the model, achieving smaller low-rank adapters compared to heuristic methods like AdaLoRA and LoRA, while maintaining critical convergence, descent, and error-bound theoretical guarantees. The resulting method is not only more efficient but also more robust to varying hyperparameter settings. We demonstrate the effectiveness of GeoLoRA on several state-of-the-art benchmarks, showing that it outperforms existing methods in both accuracy and computational efficiency.
We study hydrodynamic theories with approximate symmetries in the recently developed effective action approach on the Schwinger-Keldysh contour. We employ the method of spurious symmetry transformation for small explicit symmetry-breaking parameters to systematically constrain symmetry-breaking effects in the nonequilibrium effective action for hydrodynamics. We apply our method to the hydrodynamic theory of chiral symmetry in quantum chromodynamics at finite temperature and density and its explicit breaking by quark masses. We show that the spurious symmetry and the Kubo-Martin-Schwinger relation dictate that the Ward-Takahashi identity for the axial symmetry, i.e., the partial conservation of axial vector current (PCAC) relation, contains a relaxational term proportional to the axial chemical potential, whose kinetic coefficient is at least of the second order in the quark mass. In the phase where the chiral symmetry is spontaneously broken, and the pseudo-Nambu-Goldstone pions appear as hydrodynamic variables, this relaxation effect is subleading compared to the conventional pion mass term in the PCAC relation, which is of the first order in the quark mass. On the other hand, in the chiral symmetry restored phase, we show that our relaxation term, which is of the second order in the quark mass, becomes the leading contribution to the axial charge relaxation. Therefore, the leading axial charge relaxation mechanism is parametrically different in the quark mass across a chiral phase transition.
Tungsten (W) and tungsten alloys are being considered as leading candidates for structural and functional materials in future fusion energy devices. The most attractive properties of tungsten for the design of magnetic and inertial fusion energy reactors are its high melting point, high thermal conductivity, low sputtering yield, and low long-term disposal radioactive footprint. Despite these relevant features, there is a lack of understanding of how the structural and mechanical properties of W-based alloys are affected by the temperature in fusion power plants. In this work, we present a study on the thermo-mechanical properties of five W-based plasma-facing materials. First-principles density functional theory (DFT) calculations are combined with the quasi-harmonic approximation (QHA) theory to investigate the electronic, structural, mechanical, and thermal properties of these W-based alloys as a function of temperature. The coefficient of thermal expansion, temperature-dependent elastic constants, and several elastic parameters, including bulk and Young’s modulus, are calculated. Our work advances the understanding of the structural and thermo-mechanical behavior of W-based materials, thus providing insights into the design and selection of candidate plasma-facing materials in fusion energy devices.
The random phase approximation (RPA), a method for treating electron correlation, has been shown to be superior to standard density functional theory (DFT) approximations in numerous cases. However, the RPA’s computational cost is substantially higher than that of DFT, particularly restricting its application to extended surfaces. The recently introduced embedded RPA (emb-RPA) approach [Wei et al., J. Chem. Phys. 159(19), 194108 (2023)] reduces this computational cost by approximately two orders of magnitude. While previous applications of emb-RPA focused on non-spin-polarized systems, here we extend the approach to ferromagnetic ones. Unlike other embedded correlated wavefunction methods, such as embedded complete active space self-consistent field theory, emb-RPA is advantageous for spin-polarized systems because the RPA is compatible with unrestricted DFT solutions, which are eigenfunctions of the spin angular momentum operator S z but not the total spin-squared operator S 2 . By applying emb-RPA with specific magnetization constraints, we achieved a speedup of two to three orders of magnitude (one order when accounting for the one-time embedding potential optimization cost) with only small errors (∼50 meV) compared to full periodic RPA. Moreover, emb-RPA significantly reduces the over-binding errors of DFT approximations. In conclusion, we anticipate that the acceleration enabled by the spin-polarized emb-RPA approach will broaden the applicability of RPA to magnetic materials.
The application of Glauber theory has been playing an increasingly important role with the study of unstable or exotic nuclei. Its adaptation to medium and high-energy nucleus-nucleus collisions is severely limited because one has to evaluate the matrix elements of multiple-scattering operators. The extraction of physical observables has been done using ‘approximate’ Glauber theory whose validity is hard to evaluate. Here, we perform a full calculation of the matrix elements using Monte Carlo integration and analyze the elastic differential cross sections and the total reaction cross sections for p+¹²C, ⁴,⁶He+¹²C, and ¹²C+¹²C collisions. We use the variational Monte Carlo wave functions for ⁴,⁶He and ¹²C obtained by using realistic two- and three-nucleon potentials. We demonstrate the performance of the Glauber-theory calculations by comparing with available experimental data. We further discuss the accuracy of the conventional approximate methods in the light of the cumulant expansion for Glauber’s phase-shift function.
We propose a new way to perform path integrals in quantum mechanics by using a quantum version of Hamilton-Jacobi (HJ) theory. In classical mechanics, Hamilton-Jacobi theory is a powerful formalism, however, its utility is not explored in quantum theory beyond approximation schemes. The canonical transformation enables one to set the new Hamiltonian to constant or zero, but keeps the information about solution in Hamilton’s characteristic function. To benefit from this in quantum theory, one must work with a formulation in which classical Hamiltonian is used. This uniquely points to phase space path integral. However, the main variable in HJ formalism is energy, not time. Thus, we are led to consider the Fourier transform of the path integral, the spectral path integral Z ˜ ( E ) . The evaluation of path integrals reduces to determining the quantum Hamilton characteristic functions (which can be achieved via an asymptotic analysis) and a discrete sum over the quantum period lattice, generalizing Gutzwiller’s sum. Published by the American Physical Society 2025
Large-scale eigenvalue problems arise in various fields of science and engineering and demand computationally efficient solutions. In this study, we investigate the subspace approximation for parametric linear eigenvalue problems, aiming to mitigate the computational burden associated with high-fidelity systems. Furthermore, we provide general error estimates under non-simple eigenvalue conditions, establishing some theoretical foundations for understanding the convergence behavior of subspace approximations. Numerical examples, including problems with one-dimensional to three-dimensional spatial domain and one-dimensional to two-dimensional parameter domain, are presented to demonstrate the efficacy of reduced basis method in handling parametric variations in boundary conditions and coefficient fields to achieve significant computational savings while maintaining high accuracy, making them promising tools for practical applications in large-scale eigenvalue computations.
The accurate description of non-covalent interactions is critical for understanding the structure, dynamics, and eventual function of biomolecules. The adenine dimer serves as a benchmark system for computational methods due to its role in nucleic acid structures and its rich conformational landscape. In this study, we employ benchmark diffusion quantum Monte Carlo (DMC) methods to investigate the relative energies and role of electron correlation on a set of adenine dimer conformations generated via a search of the potential energy landscape using the global optimizer algorithm. Relative DMC energies are compared against a wide range of density functional theory (DFT) approximation results. We find that although most of the DFT functionals perform well for low-energy structures, their accuracy varies significantly for higher-energy conformations, including stacked and T-shaped structures. A large fraction of the variation is due to the treatment of the van der Waals interaction. BLYP, B3LYP, and PBE0 significantly improve with added D4 dispersion, while the recent r2SCAN-D4 and ωB97M-V functionals show the least scatter and closest agreement with the DMC. These findings highlight the delicate nature of these interactions in biomolecular systems and provide guidance for simulations of their structure and dynamics and for the development of machine learned interatomic potentials.
We propose a Coefficient-to-Basis Network (C2BNet), a novel framework for solving inverse problems within the operator learning paradigm. C2BNet efficiently adapts to different discretizations through fine-tuning, using a pre-trained model to significantly reduce computational cost while maintaining high accuracy. Unlike traditional approaches that require retraining from scratch for new discretizations, our method enables seamless adaptation without sacrificing predictive performance. Furthermore, we establish theoretical approximation and generalization error bounds for C2BNet by exploiting low-dimensional structures in the underlying datasets. Our analysis demonstrates that C2BNet adapts to low-dimensional structures without relying on explicit encoding mechanisms, highlighting its robustness and efficiency. To validate our theoretical findings, we conducted extensive numerical experiments that showcase the superior performance of C2BNet on several inverse problems. The results confirm that C2BNet effectively balances computational efficiency and accuracy, making it a promising tool to solve inverse problems in scientific computing and engineering applications.
Neural scaling laws play a pivotal role in the performance of deep neural networks and have been observed in a wide range of tasks. However, a complete theoretical framework for understanding these scaling laws remains underdeveloped. In this paper, we explore the neural scaling laws for deep operator networks, which involve learning mappings between function spaces, with a focus on the Chen and Chen style architecture. These approaches, which include the popular Deep Operator Network (DeepONet), approximate the output functions using a linear combination of learnable basis functions and coefficients that depend on the input functions. We establish a theoretical framework to quantify the neural scaling laws by analyzing its approximation and generalization errors. We articulate the relationship between the approximation and generalization errors of deep operator networks and key factors such as network model size and training data size. Moreover, we address cases where input functions exhibit low-dimensional structures, allowing us to derive tighter error bounds. These results also hold for deep ReLU networks and other similar structures. Our results offer a partial explanation of the neural scaling laws in operator learning and provide a theoretical foundation for their applications.
We present a real-space method for computing the random phase approximation (RPA) correlation energy within Kohn–Sham density functional theory, leveraging the low-rank nature of the frequency-dependent density response operator. In particular, we employ a cubic-scaling formalism based on density functional perturbation theory that circumvents the calculation of the response function matrix, instead relying on the ability to compute its product with a vector through the solution of the associated Sternheimer linear systems. We develop a large-scale parallel implementation of this formalism using the subspace iteration method in conjunction with the spectral quadrature method while employing the Kronecker product-based method for the application of the Coulomb operator and the conjugate orthogonal conjugate gradient method for the solution of the linear systems. We demonstrate convergence with respect to key parameters and verify the method’s accuracy by comparing with plane-wave results. We show that the framework achieves good strong scaling to many thousands of processors, reducing the time to solution for a lithium hydride system with 128 electrons to around 150 s on 4608 processors.