Search NASA⌕ Search

SEARCH · Search NASA

Results for “tensor processing units”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

A high-resolution large-eddy simulation framework for wildland fire predictions using TensorFlow

Background: Wildfires are becoming more severe, so we need improved tools to predict them over a wide range of conditions and scales. One approach towards this goal entails the use of coupled fire/atmosphere modelling tools. Although significant progress has been made in advancing their physical fidelity, existing tools have not taken full advantage of emerging programming paradigms and computing architectures to enable high-resolution wildfire simulations. Aims: The aim of this study was to present a new framework that enables landscape-scale wildfire simulations with physical representation of combustion at an affordable cost. Methods: We developed a coupled fire/atmosphere simulation framework using TensorFlow, which enables efficient and scalable computations on Tensor Processing Units. Key Results: Simulation results for a prescribed fire were compared with experimental data. Predicted fire behavior and statistical analysis for fire spread rate, scar area, and intermittency showed overall reasonable agreement. Scalability analysis was performed, showing close to linear scaling. Conclusions: While mesh refinement was shown to have less impact on global quantities, such as fire scar area and spread rate, it benefits predictions of intermittent fire behavior, buoyancy-driven dynamics, and small-scale turbulent motion. Implications: This new simulation framework is efficient in capturing both global quantities and unsteady dynamics of wildfires at high spatial resolutions.

54 ENVIRONMENTAL SCIENCES↗

Mixed-Precision S/DGEMM Using the TF32 and TF64 Frameworks on Low-Precision AI Tensor Cores

Using NVIDIA graphics processing units (GPUs) equipped with Tensor Cores has enabled the significant acceleration of general matrix multiplication (GEMM) for applications in machine learning (ML) and artificial intelligence (AI) and in high-performance computing (HPC) generally. The use of such power-efficient, specialized accelerators can provide a performance increase between 8 × and 20 ×, albeit with a loss in precision. However, a high level of precision is required in many large scientific and HPC applications, and computing in single or double precision is still necessary for many of these applications to maintain accuracy. Fortunately, mixed-precision methods can be employed to maintain a higher level of numerical precision while also taking advantage of the performance increases from computing with lower-precision AI cores. With this in mind, we extend the state of the art by using NVIDIA’s new TF32 framework. This new framework not only burdens some constraints of the previous frameworks, such as costly 32 16-bit castings but also provides an equivalent precision and performance by using a much simpler approach. We also propose a new framework called TF64 that attempts double-precision arithmetic with low-precision Tensor Cores. Although this framework does not exist yet, we validated the correctness of this idea and achieved an equivalent of 64-bit precision on 32-bit hardware.

Valero Lara, Pedro↗

Coupled cluster theory on modern heterogeneous supercomputers

This study examines the computational challenges in elucidating intricate chemical systems, particularly through ab-initio methodologies. This work highlights the Divide-Expand-Consolidate (DEC) approach for coupled cluster (CC) theory—a linear-scaling, massively parallel framework—as a viable solution. Detailed scrutiny of the DEC framework reveals its extensive applicability for large chemical systems, yet it also acknowledges inherent limitations. To mitigate these constraints, the cluster perturbation theory is presented as an effective remedy. Attention is then directed towards the CPS (D-3) model, explicitly derived from a CC singles parent and a doubles auxiliary excitation space, for computing excitation energies. The reviewed new algorithms for the CPS (D-3) method efficiently capitalize on multiple nodes and graphical processing units, expediting heavy tensor contractions. As a result, CPS (D-3) emerges as a scalable, rapid, and precise solution for computing molecular properties in large molecular systems, marking it an efficient contender to conventional CC models.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Massively parallel GPU enabled third-order cluster perturbation excitation energies for cost-effective large scale excitation energy calculations

We present here a massively parallel implementation of the recently developed CPS(D-3) excitation energy model that is based on cluster perturbation theory. The new algorithm extends the one developed in Baudin et al. [J. Chem. Phys., 150, 134110 (2019)] to leverage multiple nodes and utilize graphical processing units for the acceleration of heavy tensor contractions. Furthermore, we show that the extended algorithm scales efficiently with increasing amounts of computational resources and that the developed code enables CPS(D-3) excitation energy calculations on large molecular systems with a low time-to-solution. More specifically, calculations on systems with over 100 atoms and 1000 basis functions are possible in a few hours of wall clock time. This establishes CPS(D-3) excitation energies as a computationally efficient alternative to those obtained from the coupled-cluster singles and doubles model.

38 RADIATION CHEMISTRY, RADIOCHEMISTRY, AND NUCLEA↗

Adaptive variational quantum dynamics simulations with compressed circuits and fewer measurements

The adaptive variational quantum dynamics simulation (AVQDS) method performs real-time evolution of quantum states using automatically generated parametrized quantum circuits that often contain substantially fewer gates than Trotter circuits. Here we report an improved version of the method, which we call AVQDS(T), by porting the tiling efficient trial circuits with rotations implemented simultaneously technique. The algorithm adaptively adds layers of disjoint unitary gates to the ansatz circuit so as to keep the McLachlan distance, a measure of the accuracy of the variational dynamics, below a fixed threshold. Here we perform benchmark noiseless AVQDS(T) simulations of quench dynamics in local spin models and compare with an alternative adaptive variational approach on quantum resource requirement. Quantum dynamical simulations implementing realistic noise channels are also reported. Finally, we propose a way to substantially alleviate the measurement overhead of AVQDS(T) while maintaining high accuracy by synergistically integrating quantum circuit calculations on quantum processing units with classical calculations using, e.g., tensor networks to evaluate the quantum geometric tensor. We showcase that this approach enables AVQDS(T) to deliver more accurate results than simulations using a fixed ansatz of comparable final depth for a significant time duration with fewer quantum resources.

75 CONDENSED MATTER PHYSICS, SUPERCONDUCTIVITY AND↗

GPU-acceleration of tensor renormalization with PyTorch using CUDA

We show that numerical computations based on tensor renormalization group (TRG) methods can be significantly accelerated with PyTorch on graphics processing units (GPUs) by leveraging NVIDIA's Compute Unified Device Architecture (CUDA). Here we find improvement in the runtime and its scaling with bond dimension for two-dimensional systems. Our results establish that the utilization of GPU resources is essential for future precision computations with TRG.

97 MATHEMATICS AND COMPUTING↗

GPAW: An open Python package for electronic structure calculations

We review the GPAW open-source Python package for electronic structure calculations. GPAW is based on the projector-augmented wave method and can solve the self-consistent density functional theory (DFT) equations using three different wave-function representations, namely real-space grids, plane waves, and numerical atomic orbitals. The three representations are complementary and mutually independent and can be connected by transformations via the real-space grid. This multi-basis feature renders GPAW highly versatile and unique among similar codes. By virtue of its modular structure, the GPAW code constitutes an ideal platform for the implementation of new features and methodologies. Moreover, it is well integrated with the Atomic Simulation Environment (ASE), providing a flexible and dynamic user interface. In addition to ground-state DFT calculations, GPAW supports many-body GW band structures, optical excitations from the Bethe–Salpeter Equation, variational calculations of excited states in molecules and solids via direct optimization, and real-time propagation of the Kohn–Sham equations within time-dependent DFT. A range of more advanced methods to describe magnetic excitations and non-collinear magnetism in solids are also now available. In addition, GPAW can calculate non-linear optical tensors of solids, charged crystal point defects, and much more. Recently, support for graphics processing unit (GPU) acceleration has been achieved with minor modifications to the GPAW code thanks to the CuPy library. We end the review with an outlook, describing some future plans for GPAW.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

TAMM: Tensor algebra for many-body methods

Tensor algebra operations such as contractions in computational chemistry consume a significant fraction of the computing time on large-scale computing platforms. The widespread use of tensor contractions between large multi-dimensional tensors in describing electronic structure theory has motivated the development of multiple tensor algebra frameworks targeting heterogeneous computing platforms. In this paper, we present Tensor Algebra for Many-body Methods (TAMM), a framework for productive and performance-portable development of scalable computational chemistry methods. TAMM decouples the specification of the computation from the execution of these operations on available high-performance computing systems. With this design choice, the scientific application developers (domain scientists) can focus on the algorithmic requirements using the tensor algebra interface provided by TAMM, whereas high-performance computing developers can direct their attention to various optimizations on the underlying constructs, such as efficient data distribution, optimized scheduling algorithms, and efficient use of intra-node resources (e.g., graphics processing units). The modular structure of TAMM allows it to support different hardware architectures and incorporate new algorithmic advances. We describe the TAMM framework and our approach to the sustainable development of scalable ground- and excited-state electronic structure methods. We present case studies highlighting the ease of use, including the performance and productivity gains compared to other frameworks.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Accurate numerical simulations of open quantum systems using spectral tensor trains

Decoherence between qubits is a major bottleneck in quantum computations. Decoherence results from intrinsic quantum and thermal fluctuations as well as noise in the external fields that perform the measurement and preparation processes. With prescribed colored noise spectra for intrinsic and extrinsic noise, we present a numerical method, Quantum Accelerated Stochastic Propagator Evaluation (Q-ASPEN), to solve the time-dependent noise-averaged reduced density matrix in the presence of intrinsic and extrinsic noise. Q-ASPEN is arbitrarily accurate and can be applied to provide estimates for the resources needed to error-correct quantum computations. We employ spectral tensor trains, which combine the advantages of tensor networks and pseudospectral methods, as a variational ansatz to the quantum relaxation problem and optimize the ansatz using methods typically used to train neural networks. Here, the spectral tensor trains in Q-ASPEN make accurate calculations with tens of quantum levels feasible. We present benchmarks for Q-ASPEN on the spin-boson model in the presence of intrinsic noise and on a quantum chain of up to 32 sites in the presence of extrinsic noise. In our benchmark, the memory cost of Q-ASPEN scales as a low-order polynomial in the size of the system once the number of system states surpasses the number of basis functions used in the spectral expansion.

71 CLASSICAL AND QUANTUM MECHANICS, GENERAL PHYSIC↗

A Flexible Forwarding Scheme to Improve Latency-Bound Irregular P2P Communication in MPI

We propose an algorithm to efficiently perform latency-bound communication scenarios that consist of many small messages. In these parallel scenarios, processes typically pass around a lot of small-sized messages of a few KBs of size. Performing communication operations with P2P MPI routines or collective MPI routines (including neighborhood collectives) in such scenarios may not always yield the optimal results and may not resolve the latency bottleneck. To this end, we develop a regular structure called virtual process topology (VPT) on which the messages can be communicated in a structured and controlled manner. Using parameters of this topology, one can tune the rate of aggression in tackling the latency costs. We demonstrate that our communication algorithm is preferable to MPI P2P and collective routines for latency-bound communication and it can easily be adapted only by replacing calls to MPI routines in a parallel application. We show how to adapt existing topology-aware mapping heuristics to address the volume overhead due to communicating messages on the VPT. Moreover, we propose a novel swap-based mapping heuristic to address this overhead by optimizing the maximum volume handled by a process. Experiments on synthetic communication graphs as well as real-world applications such as parallel Canonical Polyadic sparse tensor decomposition and parallel sparse matrix-dense matrix multiplication show that our approach is a powerful way of overcoming the bottlenecks posed by sparse and latency-bound irregular communication.

communication algorithm↗

Quantitative kinetic rules for plastic strain-induced α - ω phase transformation in Zr under high pressure

Plastic strain-induced phase transformations (PTs) and chemical reactions under high pressure are broadly spread in modern technologies, friction and wear, geophysics, and astrogeology. However, because of very heterogeneous fields of plastic strain $E$ p and stress σ tensors and volume fraction c of phases in a sample compressed in a diamond anvil cell (DAC) and impossibility of measurements of σ and $E$ p , there are no strict kinetic equations for them. Here, we develop a kinetic model, finite element method (FEM) approach, and combined FEM-experimental approaches to determine all fields in strongly plastically predeformed Zr compressed in DAC, and specific kinetic equation for α-ω PT consistent with experimental data for the entire sample. Since all fields in the sample are very heterogeneous, data are obtained for numerous complex 7D paths in the space of 3 components of the plastic strain tensor and 4 components of the stress tensor. Kinetic equation depends on accumulated plastic strain (instead of time) and pressure and is independent of plastic strain and deviatoric stress tensors, i.e., it can be applied for various above processes. Our results initiate kinetic studies of strain-induced PTs and provide efforts toward more comprehensive understanding of material behavior in extreme conditions.

36 MATERIALS SCIENCE↗

Tree tensor network hierarchical equations of motion based on time-dependent variational principle for efficient open quantum dynamics in structured thermal environments

In this work, we introduce an efficient method, TTN-HEOM, for exactly calculating the open quantum dynamics for driven quantum systems interacting with highly structured bosonic baths by combining the tree tensor network (TTN) decomposition scheme with the bexcitonic generalization of the numerically exact hierarchical equations of motion (HEOM). The method yields a series of quantum master equations for all core tensors in the TTN that efficiently and accurately capture the open quantum dynamics for non-Markovian environments to all orders in the system–bath interaction. These master equations are constructed based on the time-dependent Dirac–Frenkel variational principle, which isolates the optimal dynamics for the core tensors given the TTN ansatz. The dynamics converges to the HEOM when increasing the rank of the core tensors, a limit in which the TTN ansatz becomes exact. We introduce TENSO, tensor equations for non-Markovian structured open systems, as a general-purpose Python code to propagate the TTN-HEOM dynamics. We implement three general propagators for the coupled master equations: two fixed-rank methods that require a constant memory footprint during the dynamics and one adaptive-rank method with a variable memory footprint controlled by the target level of computational error. We exemplify the utility of these methods by simulating a two-level system coupled to a structured bath containing one Drude–Lorentz component and eight Brownian oscillators, which is beyond what can presently be computed using the standard HEOM. Our results show that the TTN-HEOM is capable of simulating both dephasing and relaxation dynamics of driven quantum systems interacting with structured baths, even those of chemical complexity, with an affordable computational cost.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Kernel fusion in atomistic spin dynamics simulations on Nvidia GPUs using tensor core

In atomistic spin dynamics simulations, the time cost of constructing the space- and time-displaced pair correlation function in real space increases quadratically as the number of spins N, leading to significant computational effort. The GEMM subroutine can be adopted to accelerate the calculation of the dynamical spin-spin correlation function, but the computational cost of simulating large spin systems (>40000 spins) on CPUs remains expensive. In this work, we perform the simulation on the graphics processing unit (GPU), a hardware solution widely used as an accelerator for scientific computing and deep learning. Here we show that GPUs can accelerate the simulation up to 25-fold compared to multi-core CPUs when using the GEMM subroutine on both. To hide memory latency, we fuse the element-wise operation into the GEMM kernel using CUTLASS that can improve the performance by 26% ~ 33% compared to implementation based on cuBLAS. Furthermore, we perform the on-the-fly calculation in the epilogue of the GEMM subroutine to avoid saving intermediate results on global memory, which makes the large-scale atomistic spin dynamics simulation feasible and affordable.

97 MATHEMATICS AND COMPUTING↗

Impact of tensor forces on quasifission product yield distributions

Quantum shell effects are crucial for the stability and structure of atomic nuclei and play a key role in the discovery of superheavy elements. Furthermore, during nuclear collisions, the dynamical evolution of these shell effects disrupts the equilibration process necessary for forming a compound nucleus, leading to the breakup of the initial composite as a result of quasifission. As such, quasifission reactions hinder the formation of a superheavy element.

73 NUCLEAR PHYSICS AND RADIATION PHYSICS↗

The 10 September 2025 M w 4.1 Earthquake in Northeastern Utah, United States: An Archetypal Continental Mantle Event

The 10 September 2025 M w 4.1 earthquake in northeastern Utah, United States, had a focal depth 68 km beneath sea level, which is ∼20–25 km greater than estimates of local crustal thickness, making it a rare example of a continental mantle earthquake (CME). The focal depth is well resolved from arrival-time inversion (nearest station ∼13 km away) and moment tensor inversion of regional waveforms. Similar to other CMEs in the Intermountain West, there were no obvious aftershocks or foreshocks, and the waveforms were enriched in high-frequency energy. Spectral modeling gives a stress drop of ∼80 MPa and a radiation efficiency of ∼0.08, albeit with large uncertainties. The high stress drop and low radiation efficiency are consistent with a dissipative source process such as thermal runaway. Also similar to previous Intermountain West CMEs, the event occurred along the boundary of the Archean Wyoming craton, where pressure–temperature conditions favor ductile deformation. We hypothesize that edge-driven or regional-scale mantle convection produces increased strain rates near the craton boundary that make either conventional brittle failure or thermal runaway feasible at relatively high pressure–temperature conditions. High conductivity inferred around the edge of the craton may suggest that fluids also contribute to CME occurrence.

Koper, Keith D. [Univ. of Utah, Salt Lake City, UT↗

The Poisson tensor completion non-parametric differential entropy estimator

We introduce the Poisson tensor completion (PTC) estimator, a non-parametric differential entropy estimator. The PTC estimator leverages inter-sample relationships to compute a low-rank Poisson tensor decomposition of the frequency histogram. Our crucial observation is that the histogram bins are an instance of a space partitioning of counts and thus can be identified with a spatial Poisson process. The Poisson tensor decomposition leads to a completion of the intensity measure over all bins—including those containing few to no samples—and leads to our proposed PTC differential entropy estimator. A Poisson tensor decomposition models the underlying distribution of the count data and guarantees non-negative estimated values and so can be safely used directly in entropy estimation. Our estimator is the first tensor-based estimator that exploits the underlying spatial Poisson process related to the histogram explicitly when estimating the probability density with low-rank tensor decompositions for the purpose of tensor completion. Furthermore, we demonstrate that our PTC estimator is a substantial improvement over standard histogram-based estimators for sub-Gaussian probability distributions because of the concentration of norm phenomenon.

42 ENGINEERING↗

Architectures and random properties of symplectic quantum circuits

Parametrized and random unitary (or orthogonal) n-qubit circuits play a central role in quantum information. As such, one could naturally assume that circuits implementing symplectic transformations would attract similar attention. However, this is not the case, as $\mathbb{SP}(d/2)$—the group of d × d unitary symplectic matrices—has thus far been overlooked. In this work, we aim at starting to fill this gap. We begin by presenting a universal set of generators $\mathcal{G}$ for the symplectic algebra $\mathfrak{sp}(d/2)$, consisting of one- and two-qubit Pauli operators acting on neighboring sites in a one-dimensional lattice. Here, we uncover two critical differences between such set, and equivalent ones for unitary and orthogonal circuits. Namely, we find that the operators in $\mathcal{G}$ cannot generate arbitrary local symplectic unitaries and that they are not translationally invariant. We then review the Schur–Weyl duality between the symplectic group and the Brauer algebra, and use tools from Weingarten calculus to prove that Pauli measurements at the output of Haar random symplectic circuits can converge to Gaussian processes. As a by-product, such analysis provides us with concentration bounds for Pauli measurements in circuits that form t-designs over $\mathbb{SP}(d/2)$. To finish, we present tensor-network tools to analyze shallow random symplectic circuits, and we use these to numerically show that computational-basis measurements anti-concentrate at logarithmic depth.

71 CLASSICAL AND QUANTUM MECHANICS, GENERAL PHYSIC↗

Computing Nonequilibrium Responses with Score-Shifted Stochastic Differential Equations

Using equilibrium fluctuations to understand the response of a physical system to an externally imposed perturbation is the basis for linear response theory, which is widely used to interpret experiments and shed light on microscopic dynamics. For nonequilibrium systems, perturbations cannot be interpreted simply by monitoring fluctuations in a conjugate observable and general response results rely on path ensemble averaging. Furthermore, these techniques do not apply to perturbations that affect the diffusion tensor in a stochastic system. Here, we introduce an “effective” physical process that represents the diffusion perturbed dynamics and enables accurate calculations of responses to a change in the diffusion. Interestingly, the effective dynamics contain an additional drift involving the instantaneous “score” of the system, and we leverage score matching algorithms to carry out nonequilibrium response calculations on systems for which the exact stationary distribution is unknown.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗