Search NASA⌕ Search

SEARCH · Search NASA

Results for “Matrix-Free”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

Matrix-Free High-Performance Saddle-Point Solvers for High-Order Problems in \(\boldsymbol{H}(\operatorname{\textbf{div}})\)

Here, this work describes the development of matrix-free GPU-accelerated solvers for high-order finite element problems in H(div). The solvers are applicable to grad-div and Darcy problems in saddle-point formulation, and have applications in radiation diffusion and porous media flow problems, among others. Using the interpolation–histopolation basis, efficient matrix-free preconditioners can be constructed for the (1, 1)-block and Schur complement of the block system. With these approximations, block-preconditioned MINRES converges in a number of iterations that is independent of the mesh size and polynomial degree. The approximate Schur complement takes the form of an M-matrix graph Laplacian and therefore can be well-preconditioned by highly scalable algebraic multigrid methods. High-performance GPU-accelerated algorithms for all components of the solution algorithm are developed, discussed, and benchmarked. Numerical results are presented on a number of challenging test cases, including the “crooked pipe” grad-div problem, the SPE10 reservoir modeling benchmark problem, and a nonlinear radiation diffusion test case.

97 MATHEMATICS AND COMPUTING↗

Multi-material ALE remap with interface sharpening using high-order matrix-free finite element methods

The arbitrary Lagrangian-Eulerian (ALE) technique involves remapping field quantities from a Lagrangian mesh to an optimized mesh in a conservative, accurate and bounds-preserving manner. For methods based on arbitrary order finite elements, as described in a reference, material volume fractions are advected in pseudo-time using flux-corrected transport (FCT) without any form of interface reconstruction. In practice, this can lead to excessive propagation of small volume fractions throughout the domain. In addition, this method requires assembly of a global advection matrix to compute the bounds-preserving low-order FCT solution. In this work, we introduce a new approach for ALE remap using a high-order matrix-free technique which incorporates a flux modification to sharpen material interfaces in a conservative manner. Our approach begins with computing a bounds-preserving low-order solution to the ALE remap equations at the element level. We then compute a sharp interface solution (not guaranteed to be bounds-preserving) which comes from solving an augmented version of the ALE remap equations with a conservative flux modification which acts to sharpen material volume fractions based on their gradients and transport directions. Using the sharp interface solution, we make global corrections to the bounds-preserving solution while maintaining preservation of bounds. By blending with the sharpened solution at the global level we are able to globally conserve mass without hindering the remap pseudo-time step. This new interface-aware ALE remap method is based entirely on partial assembly techniques where globally assembled matrix operators are no longer needed, resulting in a globally matrix-free FCT method for multi-material, multi-field ALE remap with high performance on GPU architectures. We present results of our new remap method on 1D, 2D and 3D benchmarks and describe the algorithmic tailoring for GPU architectures that was developed.

Vargas, Arturo [Lawrence Livermore National Labora↗

Fast and Scalable FFT-Based GPU-Accelerated Algorithms for Block-Triangular Toeplitz Matrices with Application to Linear Inverse Problems Governed by Autonomous Dynamical Systems

In this work, we present an efficient and scalable algorithm for performing matrix-vector multiplications (matvecs) for block Toeplitz matrices. Such matrices, which are shift-invariant with respect to their blocks, arise in the context of solving inverse problems governed by autonomous systems, and time-invariant systems in particular. In this article, we consider inverse problems that infer unknown parameters from observational data of a linear time-invariant dynamical system given in the form of partial differential equations (PDEs). Matrix-free Newton-conjugate-gradient methods are often the gold standard for solving these inverse problems, but they require numerous actions of the Hessian on a vector. Matrix-free adjoint-based Hessian matvecs require solution of a pair of linearized forward/adjoint PDE solves per Hessian action, which may be prohibitive for large-scale inverse problems. Time invariance of the forward PDE problem leads to a block Toeplitz structure of the discretized parameter-to-observable (p2o) map defining the mapping from inputs (parameters) to outputs (observables) of the PDEs. This block Toeplitz structure enables us to exploit two key properties: (1) compact storage of the p2o map and its adjoint, and (2) efficient fast Fourier transform–based Hessian matvecs. The proposed algorithm is mapped onto large multi-GPU clusters and achieves more than 80% of peak bandwidth on NVIDIA A100 GPUs. Excellent weak scaling is shown for up to 48 A100 GPUs. For the targeted problems, the implementation executes Hessian matvecs within fractions of a second, which is orders of magnitude faster than can be achieved by conventional matrix-free Hessian matvecs via forward/adjoint PDE solves.

97 MATHEMATICS AND COMPUTING↗

ALESQP: An Augmented Lagrangian Equality-Constrained SQP Method for Optimization with General Constraints

Here we present a new algorithm for infinite-dimensional optimization with general constraints, called ALESQP. In short, ALESQP is an augmented Lagrangian method that penalizes inequality constraints and solves equality-constrained nonlinear optimization subproblems at every iteration. The subproblems are solved using a matrix-free trust-region sequential quadratic programming (SQP) method that takes advantage of iterative, i.e., inexact linear solvers, and is suitable for large-scale applications. A key feature of ALESQP is a constraint decomposition strategy that allows it to exploit problem-specific variable scalings and inner products. We analyze convergence of ALESQP under different assumptions. We show that strong accumulation points are stationary. Consequently, in finite dimensions ALESQP converges to a stationary point. In infinite dimensions we establish that weak accumulation points are feasible in many practical situations. Under additional assumptions we show that weak accumulation points are stationary. We present several infinite-dimensional examples where ALESQP shows remarkable discretization-independent performance in all of its iterative components, requiring a modest number of iterations to meet constraint tolerances at the level of machine precision. Also, we demonstrate a fully matrix-free solution of an infinite-dimensional problem with nonlinear inequality constraints.

97 MATHEMATICS AND COMPUTING↗

Solution of the Schrödinger equation for quasi-one-dimensional materials using helical waves

We formulate and implement a spectral method for solving the Schrödinger equation, as it applies to quasi-one-dimensional materials and structures. This allows for computation of the electronic structure of important technological materials such as nanotubes (of arbitrary chirality), nanowires, nanoribbons, chiral nanoassemblies, nanosprings and nanocoils, in an accurate, efficient and systematic manner. Our work is motivated by the observation that one of the most successful methods for carrying out electronic structure calculations of bulk/crystalline systems — the plane-wave method — is a spectral method based on eigenfunction expansion. Our scheme avoids computationally onerous approximations involving periodic supercells often employed in conventional plane-wave calculations of quasi-one-dimensional materials, and also overcomes several limitations of other discretization strategies, e.g., those based on finite differences and atomic orbitals. The basis functions in our method — called helical waves (or twisted waves) — are eigenfunctions of the Laplacian with symmetry adapted boundary conditions, and are expressible in terms of plane waves and Bessel functions in helical coordinates. We describe the setup of fast transforms to carry out discretization of the governing equations using our basis set, and the use of matrix-free iterative diagonalization to obtain the electronic eigenstates. Miscellaneous computational details, including the choice of eigensolvers, use of a preconditioning scheme, evaluation of oscillatory radial integrals and the imposition of a kinetic energy cutoff are discussed. We have implemented these strategies into a computational package called HelicES (Helical Electronic Structure). We demonstrate the utility of our method in carrying out systematic electronic structure calculations of various quasi-one-dimensional materials through numerous examples involving nanotubes, nanoribbons and nanowires. We also explore the convergence properties of our method, and assess its accuracy and computational efficiency by comparison against reference finite difference, transfer matrix method and plane-wave results. We anticipate that our method will find applications in computational nanomechanics and multiscale modeling, for carrying out transport calculations of interest to the field of semiconductor devices, and for the discovery of novel chiral phases of matter that are of relevance to the burgeoning quantum hardware industry.

75 CONDENSED MATTER PHYSICS, SUPERCONDUCTIVITY AND↗

Designing the Protocols for Programmable Ammonia Catalysis

Programmable catalysis can provide a more energy-efficient and cost-effective route to enhancing commercial ammonia production, a key process in the advancement of renewable energy technologies and the manufacture of fertilizers and basic chemicals. This work explores the computational discovery of optimal forcing protocols to drive such dynamic catalysis models. By employing matrix-free time-stepper methods, coupled with an optimization approach, that integrates Bayesian optimization with a Bayesian continuation strategy to efficiently discover the periodic steady states of such periodically forced systems, we enable the discovery of complex optimal catalyst strain waveforms, while ensuring robust solver convergence. We demonstrate the flexibility of our approach to discover optimized forcing protocols under varying physical constraints on strain modulation or other catalyst operating parameters. We show that these can have a temporal structure more complex than simple step functions. In order to detect undesirable catalytic loops that may correlate with overall reduced performance, we perform a study using graph-theoretical analysis to investigate the dynamics of catalytic kinetic networks formed.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

A combinatorial droplet microfluidic device integrated with mass spectrometry for enzyme screening

Mass spectrometry (MS) enables detection of different chemical species with a very high specificity; however, it can be limited by its throughput. Integrating MS with microfluidics has a tremendous potential to improve throughput and accelerate biochemical research. In this work, we introduce Drop-NIMS, a combination of a passive droplet loading microfluidic device and a matrix-free MS laser desorption ionization technique called nanostructure-initiator mass spectrometry (NIMS). This platform combines different droplets at random to generate a combinatorial library of enzymatic reactions that are deposited directly on the NIMS surface without requiring additional sample handling. The enzyme reaction products are then detected with MS. Drop-NIMS was used to rapidly screen enzymatic reactions containing low (on the order of nL) volumes of glycoside reactants and glycoside hydrolase enzymes per reaction. MS “barcodes” (small compounds with unique masses) were added to the droplets to identify different combinations of substrates and enzymes created by the device. We assigned xylanase activities to several putative glycoside hydrolases, making them relevant to food and biofuel industrial applications. Overall, Drop-NIMS is simple to fabricate, assemble, and operate and it has potential to be used with many other small molecule metabolites.

47 OTHER INSTRUMENTATION↗

Scalable learning of potentials to predict time-dependent Hartree–Fock dynamics

We propose a framework to learn the time-dependent Hartree–Fock (TDHF) inter-electronic potential of a molecule from its electron density dynamics. Although the entire TDHF Hamiltonian, including the inter-electronic potential, can be computed from first principles, we use this problem as a testbed to develop strategies that can be applied to learn a priori unknown terms that arise in other methods/approaches to quantum dynamics, e.g., emerging problems such as learning exchange–correlation potentials for time-dependent density functional theory. We develop, train, and test three models of the TDHF inter-electronic potential, each parameterized by a four-index tensor of size up to 60 × 60 × 60 × 60. Two of the models preserve Hermitian symmetry, while one model preserves an eight-fold permutation symmetry that implies Hermitian symmetry. Across seven different molecular systems, we find that accounting for the deeper eight-fold symmetry leads to the best-performing model across three metrics: training efficiency, test set predictive power, and direct comparison of true and learned inter-electronic potentials. All three models, when trained on ensembles of field-free trajectories, generate accurate electron dynamics predictions even in a field-on regime that lies outside the training set. To enable our models to scale to large molecular systems, we derive expressions for Jacobian-vector products that enable iterative, matrix-free training.

97 MATHEMATICS AND COMPUTING↗

Sparsity-Independent Lyapunov Exponent in the Sachdev-Ye-Kitaev Model

The saturation of a recently proposed universal bound on the Lyapunov exponent has been conjectured to signal the existence of a gravity dual. This saturation occurs in the low-temperature limit of the dense Sachdev-Ye-Kitaev (SYK) model, N Majorana fermions with q body ( q > 2 ) infinite-range interactions. We calculate certain out-of-time-order correlators (OTOCs) for N ≤ 64 fermions for a highly sparse SYK model and find no significant dependence of the Lyapunov exponent on sparsity up to near the percolation limit where the Hamiltonian breaks up into blocks. This provides strong support to the saturation of the Lyapunov exponent in the low-temperature limit of the sparse SYK. A key ingredient to reaching N = 64 is the development of a novel quantum spin model simulation library that implements highly optimized matrix-free Krylov subspace methods on graphical processing units. This leads to a significantly lower simulation time as well as vastly reduced memory usage over previous approaches, while using modest computational resources. Strong sparsity-driven statistical fluctuations require both the use of a much larger number of disorder realizations with respect to the dense limit and a careful finite size scaling analysis. The saturation of the bound in the sparse SYK points to the existence of a gravity analog that would enlarge substantially the number of field theories with this feature. Published by the American Physical Society 2024

Physics↗

Energy Exascale Computational Fluid Dynamics Simulations With the Spectral Element Method

Development and application of the open-source GPU-based fluid-thermal simulation code, NekRS, are described. Time advancement is based on an efficient kth-order accurate timesplit formulation coupled with scalable iterative solvers. Spatial discretization is based on the high-order spectral element method (SEM), which affords the use of fast, low-memory, matrix-free operator evaluation. Further, recent developments include support for nonconforming meshes using overset grids and for GPU-based Lagrangian particle tracking. Results of large-eddy simulations of atmospheric boundary layers for wind-energy applications as well as extensive nuclear energy applications are presented.

42 ENGINEERING↗

Low-Order Preconditioning for the High-Order Finite Element de Rham Complex

Here, we present a unified framework for constructing spectrally equivalent low-order-refined discretizations for the high-order finite element de Rham complex. This theory covers diffusion problems in H 1 , H(curl), and H(div) and is based on combining a low-order discretization posed on a refined mesh with a high-order basis for Nédélec and Raviart–Thomas elements that makes use of the concept of polynomial histopolation (polynomial fitting using prescribed mean values over certain regions). This spectral equivalence, coupled with algebraic multigrid methods constructed using the low-order discretization, results in highly scalable matrix-free preconditioners for high-order finite element problems in the full de Rham complex. Additionally, a new lowest-order (piecewise constant) preconditioner is developed for high-order interior penalty discontinuous Galerkin (DG) discretizations, for which spectral equivalence results and convergence proofs for algebraic multigrid methods are provided. In all cases, the spectral equivalence results are independent of polynomial degree and mesh size; for DG methods, they are also independent of the penalty parameter. These new solvers are flexible and easy to use; any “black-box” preconditioner for low-order problems can be used to create an effective and efficient preconditioner for the corresponding high-order problem. A number of numerical experiments are presented, based on an implementation in the finite element library MFEM. A range of challenging three-dimensional problems are used to corroborate the theoretical properties and demonstrate the flexibility and scalability of the method.

97 MATHEMATICS AND COMPUTING↗

End-to-end GPU acceleration of low-order-refined preconditioning for high-order finite element discretizations

In this article, we present algorithms and implementations for the end-to-end GPU acceleration of matrix-free low-order-refined preconditioning of high-order finite element problems. The methods described here allow for the construction of effective preconditioners for high-order problems with optimal memory usage and computational complexity. The preconditioners are based on the construction of a spectrally equivalent low-order discretization on a refined mesh, which is then amenable to, for example, algebraic multigrid preconditioning. The constants of equivalence are independent of mesh size and polynomial degree. For vector finite element problems in H(curl) and H(div) (e.g., for electromagnetic or radiation diffusion problems), a specially constructed interpolation–histopolation basis is used to ensure fast convergence. Detailed performance studies are carried out to analyze the efficiency of the GPU algorithms. The kernel throughput of each of the main algorithmic components is measured, and the strong and weak parallel scalability of the methods is demonstrated. The different relative weighting and significance of the algorithmic components on GPUs and CPUs is discussed. Results on problems involving adaptively refined nonconforming meshes are shown, and the use of the preconditioners on a large-scale magnetic diffusion problem using all spaces of the finite element de Rham complex is illustrated.

97 MATHEMATICS AND COMPUTING↗

Advancing attenuation estimation through integration of the Hessian in multiparameter viscoacoustic full-waveform inversion

Accurate seismic attenuation models of subsurface structures not only enhance subsequent migration processes by improving fidelity, resolution, and facilitating amplitude-compliant angle gather generation but also provide valuable constraints on subsurface physical properties. Leveraging full-wavefield information, multiparameter viscoacoustic full-waveform inversion ( Q-FWI) simultaneously estimates seismic velocity and attenuation ( Q) models. However, a major challenge in Q-FWI is the contamination of crosstalk artifacts, where inaccuracies in the velocity model are mistakenly mapped to the inverted attenuation model. While incorporating the Hessian is expected to mitigate these artifacts, the explicit implementation is prohibitively expensive due to its formidable computational cost. In this study, we formulate and develop a Q-FWI algorithm via the Newton-conjugate gradient (CG) framework, where the search direction at each iteration is determined through an internal CG loop. In particular, the Hessian is integrated into each CG step in a matrix-free fashion using the second-order adjoint-state method. We find through synthetic experiments that our Newton-CG Q-FWI significantly mitigates crosstalk artifacts compared with the limited-memory Broyden-Fletcher-Goldfarb-Shanno method and the CG method, albeit with a notable computational cost. In the discussion of several key implementation details, we also determine the significance of the approximate Gauss-Newton Hessian, the second-order adjoint-state method, and the two-stage inversion strategy.

Geochemistry & Geophysics↗

Neutron (and other Particle) Transport at LANL: An Overview [Presentation]

For decades, Los Alamos National Laboratory has been at the forefront of neutron transport methods research and code development. One such code is PARTISN, the LANL parallel time-dependent discrete ordinate neutron transport code. In this presentation, we describe the various research efforts currently underway by the PARTISN and other code teams. Some examples of current research are a block automated mesh refinement scheme, the application of tensor trains to the discretized neutron transport equation, and GPU code porting. The block automated mesh refinement scheme uses cross section information to refine and coarsen the solution mesh to improve time to solution and reduce memory. The tensor train approach expresses discretized transport operators as tensor products of vectors and matrices to compress the size of linear systems being solved by transport codes. Rather than relying on matrix-free methods such as the transport sweep, we have access to an operator that can be inverted, reshaped, or manipulated algebraically. Finally, we describe how PARTISN is used, what problems we are looking to solve, and what the future holds for neutron transport at LANL. In addition to this, we briefly describe the various research efforts in other particle transport teams using both deterministic and Monte Carlo methods. In the presentation, we list possible opportunities for collaboration between the laboratory and faculty and students.

73 NUCLEAR PHYSICS AND RADIATION PHYSICS↗

Batched Sparse Linear Algebra (Final Report for Subcontract B648960)

This report finalizes design specifications for developing batched kernels for small tensor operations for unassembled matrix-free iterative solvers, batched solvers for partially assembled operators, and batched solvers with support for various sparse formats. The outcome of the project milestones is a set of interfaces to Batched Sparse LA solvers running on hardware accelerators for use in ECP Libraries and Applications. It is part of the development of sparse batched kernels, solvers/preconditioners as well as creating interoperability in xSDK libraries with sparse and dense batched functions to benefit ECP applications. The participants included representatives from ECP libraries (not limited to the xSDK project), applications, and vendors (AMD, Intel, and NVIDIA). Batched sparse linear algebra solvers form the new frontier for algorithmic development and performance engineering. Many applications (ECP and non-ECP alike) require simultaneous solutions of small linear systems of equations that are structurally sparse. To move towards high hardware utilization, it is important to provide these applications with appropriate interfaces to efficient batched sparse solvers running on modern hardware accelerators. We present interface designs in use by HPC software libraries supporting batched sparse linear algebra and the development of sparse batched kernel codes for solvers and preconditioners. We also address the potential interoperability opportunities to keep the software portable between the major hardware accelerators from AMD, Intel, and NVIDIA. The presented interface specifications includes batched band, sparse iterative, and sparse direct solvers. This report summarizes progress in Kokkos Kernels and the xSDK libraries MAGMA, Ginkgo, hypre, SUNDIALS, and SuperLU_dist.

97 MATHEMATICS AND COMPUTING↗

LDRD Abbreviated report: High-Order General-Discrete-Ordinates Method Enabling Efficient Deterministic Transport in Hydrodynamic Simulations

Deterministic transport simulations for national-security and energy applications often operate in high-dimensional phase-space, where accuracy and cost both become major challenges. A common numerical artifact in such problems is the “ray-effect,” which appears as unphysical streaks. Beyond misinterpretation, these artifacts can contaminate tightly coupled physics, such as fluid dynamics, radiation-hydrodynamics, and laser-plasma interactions, eroding the predictive capability of entire multiphysics workflows. Our objective was to make high-dimension studies practical on modern hardware while mitigating the ray-effect without relying on prohibitively expensive sampling approaches such as Monte Carlo methods. We developed the Generic Discretization Library (GenDiL), a Graphics Processing Unit (GPU)-first framework that uses high-order Discontinuous Galerkin (DG) methods and matrix-free algorithms to reduce memory usage and improve computational efficiency, critical for phase-space simulations. GenDiL supports phase-space adaptivity in both mesh size and polynomial order (hp-adaptivity) to place resolution only where it is needed. A central capability is Local Dimensional Refinement (LDR), which couples lower-dimension continuum models to higher-dimension kinetic models through stable and conservative interfaces, so that high-fidelity physics is applied only in regions where it is essential. Building on the GenDiL framework, we developed the General SN (GSN) family of algorithms as a true generalization of the polar SN approach (discrete ordinates, often denoted SN). Rather than tying discrete ordinates to a specific polar change of coordinates, GSN formulates transport on an arbitrary change of coordinates chosen to reduce ray-effect. We studied two complementary variants: an analytic variant, where the coordinate map is prescribed in advance by a closed-form function; and a data-driven variant, where a quantity of interest, such as the net flux, guides the coordinate system. GenDiL provides the library infrastructure for efficient GPU execution, but the GSN concept is algorithmic and independent of any one library. Across representative high-dimension tests, including non-symmetric solutions, both variants delivered strong ray-effect mitigation at practical cost, moving four- to six-dimensional analysis toward repeatable, routine studies.

73 NUCLEAR PHYSICS AND RADIATION PHYSICS↗

Equation-Free Coarse Control of Distributed Parameter Systems via Local Neural Operators

The control of high-dimensional distributed parameter systems (DPS) remains a challenge when explicit coarse-grained equations are unavailable. Classical equation-free (EF) approaches rely on fine-scale simulators treated as black-box timesteppers. However, repeated simulations for steady-state computation, linearization, and control design are often computationally prohibitive, or the microscopic timestepper may not even be available, leaving us with data as the only resource. We propose a data-driven alternative that uses local neural operators, trained on spatiotemporal microscopic/mesoscopic data, to obtain efficient short-time solution operators. These surrogates are employed within Krylov subspace methods to compute coarse steady and unsteady-states, while also providing Jacobian information in a matrix-free manner. Krylov-Arnoldi iterations then approximate the dominant eigenspectrum, yielding reduced models that capture the open-loop slow dynamics without explicit Jacobian assembly. Both discrete-time Linear Quadratic Regulator (dLQR) and pole-placement (PP) controllers are based on this reduced system and lifted back to the full nonlinear dynamics, thereby closing the feedback loop.

93B52, 93C20, 47N70, 65J15, 65M32, 68T07, 68T20, 6↗