Search NASA⌕ Search

SEARCH · Search NASA

Results for “tensor computations”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

Improving Runtime Performance of Tensor Computations using Rust From Python

In this work, we investigate improving the runtime performance of key computational kernels in the Python Tensor Toolbox (pyttb), a package for analyzing tensor data across a wide variety of applications. Recent runtime performance improvements have been demonstrated using Rust, a compiled language, from Python via extension modules leveraging the Python C API—e.g., web applications, data parsing, data validation, etc. Using this same approach, we study the runtime performance of key tensor kernels of increasing complexity, from simple kernels involving sums of products over data accessed through single and nested loops to more advanced tensor multiplication kernels that are key in low-rank tensor decomposition and tensor regression algorithms. In numerical experiments involving synthetically generated tensor data of various sizes and these tensor kernels, we demonstrate consistent improvements in runtime performance when using Rust from Python over 1) using Python alone, 2) using Python and the Numba just-in-time Python compiler (for loop-based kernels), and 3) using the NumPy Python package for scientific computing (for pyttb kernels).

97 MATHEMATICS AND COMPUTING↗

Communication Lower Bounds and Optimal Algorithms for Multiple Tensor-Times-Matrix Computation

Multiple tensor-times-matrix (Multi-TTM) is a key computation in algorithms for computing and operating with the Tucker tensor decomposition, which is frequently used in multidimensional data analysis. Here, we establish communication lower bounds that determine how much data movement is required (under mild conditions) to perform the Multi-TTM computation in parallel. The crux of the proof relies on analytically solving a constrained, nonlinear optimization problem. We also present a parallel algorithm to perform this computation that organizes the processors into a logical grid with twice as many modes as the input tensor. We show that, with correct choices of grid dimensions, the communication cost of the algorithm attains the lower bounds and is therefore communication optimal. Finally, we show that our algorithm can significantly reduce communication compared to the straightforward approach of expressing the computation as a sequence of tensor-times-matrix operations when the input and output tensors vary greatly in size.

HBL-inequalities↗

Computing Sparse Tensor Decompositions via Chapel and C++/MPI Interoperability without Intermediate I/O

We extend an existing approach for efficient use of shared mapped memory across Chapel and C++ for graph data stored as 1-D arrays to sparse tensor data stored using a combination of 2-D and 1-D arrays. We describe the specific extensions that provide use of shared mapped memory tensor data for a particular C++ tensor decomposition tool called GentenMPI. We then demonstrate our approach on several real-world datasets, providing timing results that illustrate minimal overhead incurred using this approach. Finally, we extend our work to improve memory usage and provide convenient random access to sparse shared mapped memory tensor elements in Chapel, while still being capable of leveraging high performance implementations of tensor algorithms in C++.

97 MATHEMATICS AND COMPUTING↗

torch-einshard v1.0

torch-einshard is a Python library for describing local and distributed PyTorch tensor computations with compact, einsum-like notation. Its expressions name logical axes, specify how they are sharded across a PyTorch DeviceMesh, and represent partial reductions. The library automatically performs contractions, permutations, reshaping, splitting, gathering, reduction, reduce-scatter, and repartitioning while preserving autograd. Additional features include sharding-aware FFTs, tensor rolls, halo exchange, sliding windows, 1D–3D convolutions, uneven-shard handling, parameter initialization and gradient management, and cost-based execution planning. It is designed for scientific machine learning and large-model workloads, including tensor-, sequence-, and spatial-parallel MLPs, attention, convolutions, and spectral operations. Compared with manually combining torch.einsum and distributed collectives, torch-einshard expresses both the mathematical operation and data placement in one readable formula. This reduces boilerplate and synchronization errors, keeps forward and backward communication consistent, and allows the library to select optimized collective strategies without changing model code.

Morozov, Dmitriy [Lawrence Berkeley National Labor↗

Bootstrapping the 3d Ising stress tensor

We compute observables of the critical 3d Ising model to high precision by applying the numerical conformal bootstrap to mixed correlators of the leading scalar operators σ and ϵ, and the stress tensor T μν . We obtain new precise determinations of scaling dimensions (∆ σ , ∆ ϵ ) = (0.518148806(24), 1.41262528(29)) as well as OPE coefficients involving σ, ϵ, and T μν . We also describe several improvements made along the way to algorithms and software tools for the numerical bootstrap.

Conformal and W Symmetry↗

Georgia Tech Accelerated, Compressed, and Regularized Compute of Kinetic-based PDEs (Final Report)

This report summarizes the collaborative effort between Lawrence Livermore National Laboratory and Georgia Tech to enhance the BoBa library for tensor train computation in PDE solvers, with a target on kinetic equations and their continuum limits. We aimed to reduce computational cost and memory usage by replacing traditional array-based computations with tensor trains. We examined the compressibility of time-evolving solutions to the Euler equations with discontinuities. We also explored using the first invsicid and linear regularization of the compressible flow equations via the information geometric regularization (IGR). We explored this in a tensor train formulation. To identify that inverse terms in the IGR equations pose problems for tensor train formulations and investigate efficient methods for batched inversion of tensor trains.

97 MATHEMATICS AND COMPUTING↗

Real-time estimators for scattering observables: A full account of finite-volume errors for quantum simulation

The real-time correlators of quantum field theories can be directly probed through new approaches to simulation, such as quantum computing and tensor networks. This provides a new framework for computing scattering observables in lattice formulations of strongly interacting theories, such as lattice quantum chromodynamics. In this paper, we prove that the proposal of real-time estimators of scattering observables is universally applicable to all scattering observables of gapped quantum field theories. All finite-volume errors are exponentially suppressed, and the rate of this suppression is controlled by the regulator considered, namely, a displacement of the spectrum of the theory into the complex plane. A partial restoration of Lorentz symmetry by averaging over different boosts gives an additional suppression of finite volume errors. Our results also apply to the simulation of wave packet scattering, where a similar averaging is performed to construct the wave packets that regulate the finite volume effects. This result represents a necessary key step toward determining a broad class of scattering observables via quantum computing that are currently inaccessible via classical computing. Such observables are relevant for various applications, including hadron spectroscopy, hadron structure, and precision tests of the Standard Model. We also comment on potential applications of our results to traditional computational schemes.

Burbano, Ivan M. [University of California, Berkel↗

Sign Problem in Tensor-Network Contraction

We investigate how the computational difficulty of contracting tensor networks depends on the sign structure of the tensor entries. Using results from computational complexity, we observe that the approximate contraction of tensor networks with only positive entries has lower computational complexity as compared to tensor networks with general real or complex entries. This raises the question of how this transition in computational complexity manifests itself in the hardness of different tensor-network-contraction schemes. We pursue this question by studying random tensor networks with varying bias toward positive entries. First, we consider contraction via Monte Carlo sampling and find that the transition from hard to easy occurs when the tensor entries become predominantly positive; this can be understood as a tensor-network manifestation of the well-known negative-sign problem in quantum Monte Carlo. Second, we analyze the commonly used contraction based on boundary tensor networks. The performance of this scheme is governed by the number of correlations in contiguous parts of the tensor network (which by analogy can be thought of as entanglement). Remarkably, we find that the transition from hard to easy—i.e., from a volume-law to a boundary-law scaling of entanglement—already occurs for a slight bias of the tensor entries toward a positive mean, scaling inversely with the bond dimension D , and thus the problem becomes easy the earlier the larger D occurs. This is in contrast both to expectations and to the behavior found in Monte Carlo contraction, where the hardness at fixed bias increases with the bond dimension. To provide insight into this early breakdown of computational hardness and the accompanying entanglement transition, we construct an effective classical statistical-mechanical model that predicts a transition at a bias of the tensor entries of 1 / D , confirming our observations. We conclude by investigating the computational difficulty of computing expectation values of tensor-network wave functions (projected entangled-pair states, PEPSs) and find that in this setting, the complexity of entanglement-based contraction always remains low. We explain this by providing a local transformation that maps PEPS expectation values to a positive-valued tensor network. This not only provides insight into the origin of the observed boundary-law entanglement scaling but also suggests new approaches toward PEPS contraction based on positive decompositions. Published by the American Physical Society 2025

Chen, Jielun (ORCID:0000000178411545)↗

An Incremental Tensor Train Decomposition Algorithm

We present a new algorithm for incrementally updating the tensor train decomposition of a stream of tensor data. This new algorithm, called the tensor train incremental core expansion (TT-ICE) improves upon the current state-of-the-art algorithms for compressing in tensor train format by developing a new adaptive approach that incurs significantly slower rank growth and guarantees compression accuracy. This capability is achieved by limiting the number of new vectors appended to the TT-cores of an existing accumulation tensor after each data increment. These vectors represent directions orthogonal to the span of existing cores and are limited to those needed to represent a newly arrived tensor to a target accuracy. We provide two versions of the algorithm: TT-ICE and TT-ICE accelerated with heuristics (TT-ICE*). Here, we provide a proof of correctness for TT-ICE and empirically demonstrate the performance of the algorithms in compressing large-scale video and scientific simulation datasets. Compared to existing approaches that also use rank adaptation, TT-ICE* achieves 57× higher compression and up to 95% reduction in computational time.

97 MATHEMATICS AND COMPUTING↗

Inclusive reactions from finite Minkowski spacetime correlation functions

The need to determine scattering amplitudes of few-hadron systems for arbitrary kinematics expands a broad set of subfields of modern-day nuclear and hadronic physics. In this work, we expand upon previous explorations on the use of real-time methods, like quantum computing or tensor networks, to determine few-body scattering amplitudes. Such calculations must be performed in a finite Minkowski spacetime, where scattering amplitudes are not well defined. Our previous work presented a conjecture of a systematically improvable estimator for scattering amplitudes constructed from finite-volume correlation functions. Here we provide further evidence that the prescription works for larger kinematic regions than previously explored as well as a broader class of scattering amplitudes. Finally, we devise a new method for estimating the order of magnitude of the error associated with finite time separations needed for such calculations. In units of the lightest mass of the theory, we find that to constrain amplitudes using real-time methods within O ( 10 % ) , the spacetime volumes must satisfy m L ∼ O ( 10 – 10 2 ) ) and m T ∼ O ( 10 2 – 10 4 ) . Published by the American Physical Society 2024

73 NUCLEAR PHYSICS AND RADIATION PHYSICS↗

GPU-acceleration of tensor renormalization with PyTorch using CUDA

We show that numerical computations based on tensor renormalization group (TRG) methods can be significantly accelerated with PyTorch on graphics processing units (GPUs) by leveraging NVIDIA's Compute Unified Device Architecture (CUDA). Here we find improvement in the runtime and its scaling with bond dimension for two-dimensional systems. Our results establish that the utilization of GPU resources is essential for future precision computations with TRG.

97 MATHEMATICS AND COMPUTING↗

Multinuclear Solid-State NMR and NMR Crystallography of Solid Forms of Creatine and Creatinine

Creatine is a performance-enhancing supplement with two widely available commercial solid forms, namely, creatine monohydrate (creatine·H 2 O) and creatine HCl, the latter of which does not have a reported crystal structure. Moreover, commercial formulations of creatine may contain creatinine, an undesired impurity phase resulting from the self-cyclization of creatine during manufacturing. Therefore, reliable methods for characterizing the different solid forms of creatine and detecting the presence of creatinine are essential. Herein, we address these challenges using 13 C, 15 N, and 35 Cl solid-state NMR (SSNMR) spectroscopy to obtain distinct spectral fingerprints for creatine·H 2 O and creatine HCl, along with creatinine and creatinine HCl. The acquisition of these SSNMR spectra offers a robust approach for both the rapid characterization of each solid form and the detection of the impurity phases. Additionally, quadrupolar NMR crystallography-guided crystal structure prediction (QNMRX-CSP) was applied for the de novo crystal structure determination of creatine HCl, which was validated by the subsequently determined single-crystal X-ray diffraction (SCXRD) structure. Finally, to investigate the relationship between NMR parameters and structural features, 13 C and 15 N chemical shifts and 35 Cl electric field gradient (EFG) tensors were computed from geometry-optimized structures of the four solid forms by using dispersion-corrected DFT-D2* methods. Finally, this integrative approach offers a powerful framework for advancing the structural understanding and quality control of creatine-based supplements and next-generation formulations, as well as a wide range of other solid pharmaceuticals and nutraceuticals.

NMR↗

Second-harmonic generation tensors from high-throughput density-functional perturbation theory

Optical materials play a key role in enabling modern optoelectronic technologies in a wide variety of domains such as the medical or the energy sector. Among them, nonlinear optical crystals are of primary importance to achieve a broader range of electromagnetic waves in the devices. However, numerous and contradicting requirements significantly limit the discovery of new potential candidates, which, in turn, hinders the technological development. In the present work, the static nonlinear susceptibility and dielectric tensor are computed via density-functional perturbation theory for a set of 579 inorganic semiconductors. The computational methodology is discussed and the provided database is described with respect to both its data distribution and its format. Several comparisons with both experimental and ab initio results from literature allow to confirm the reliability of our data. The aim of this work is to provide a relevant dataset to foster the identification of promising nonlinear optical crystals in order to motivate their subsequent experimental investigation.

75 CONDENSED MATTER PHYSICS, SUPERCONDUCTIVITY AND↗

Phase field dislocation dynamics formulation coupled with Fourier based micromechanics solver and its application to grain boundary–dislocation interactions

A new phase field dislocation dynamics (PFDD) formulation for homogeneous and heterogeneous materials is presented, which couples micromechanical solvers and the time-dependent Ginzburg–Landau equation. The strain fields are obtained from the micromechanical solver by solving the Lippmann–Schwinger equation and then used to define energy terms to model the evolution of the dislocations. Grain boundary (GB)–dislocation interactions are studied using the coupled PFDD formulation and by describing GBs as inclusions. GB energy and stiffness tensors are computed from molecular statics simulations, and a newly proposed lattice energy term that is dependent on the GB energy is considered in the calculations. Interaction of a screw dislocation with minimum energy and metastable states of low and high angle ⟨110⟩ symmetric tilt grain boundaries are studied. We show good agreement between predictions from our PFDD formulation and molecular dynamics simulations of grain boundary–dislocation interactions.

36 MATERIALS SCIENCE↗

A high-resolution large-eddy simulation framework for wildland fire predictions using TensorFlow

Background: Wildfires are becoming more severe, so we need improved tools to predict them over a wide range of conditions and scales. One approach towards this goal entails the use of coupled fire/atmosphere modelling tools. Although significant progress has been made in advancing their physical fidelity, existing tools have not taken full advantage of emerging programming paradigms and computing architectures to enable high-resolution wildfire simulations. Aims: The aim of this study was to present a new framework that enables landscape-scale wildfire simulations with physical representation of combustion at an affordable cost. Methods: We developed a coupled fire/atmosphere simulation framework using TensorFlow, which enables efficient and scalable computations on Tensor Processing Units. Key Results: Simulation results for a prescribed fire were compared with experimental data. Predicted fire behavior and statistical analysis for fire spread rate, scar area, and intermittency showed overall reasonable agreement. Scalability analysis was performed, showing close to linear scaling. Conclusions: While mesh refinement was shown to have less impact on global quantities, such as fire scar area and spread rate, it benefits predictions of intermittent fire behavior, buoyancy-driven dynamics, and small-scale turbulent motion. Implications: This new simulation framework is efficient in capturing both global quantities and unsteady dynamics of wildfires at high spatial resolutions.

54 ENVIRONMENTAL SCIENCES↗

First simultaneous global QCD analysis of dihadron fragmentation functions and transversity parton distribution functions

We perform a comprehensive study within quantum chromodynamics (QCD) of dihadron observables in electron-positron annihilation, semi-inclusive deep-inelastic scattering, and proton-proton collisions, including recent cross section data from Belle and azimuthal asymmetries from STAR. We extract simultaneously for the first time π + π − dihadron fragmentation functions (DiFFs) and the nucleon transversity distributions for up and down quarks as well as antiquarks. For the transversity distributions we impose their small- x asymptotic behavior and the Soffer bound. In addition, we utilize a new definition of DiFFs that has a number density interpretation to then calculate expectation values for the dihadron invariant mass and momentum fraction. Furthermore, we investigate the compatibility of our transversity results with those from single-hadron fragmentation (from a transverse momentum dependent/collinear twist-3 framework) and the nucleon tensor charges computed in lattice QCD. We find a universal nature to all of this available information. Future measurements of dihadron production can significantly further this research, especially, as we show, those that are sensitive to the region of large parton momentum fractions. Published by the American Physical Society 2024

73 NUCLEAR PHYSICS AND RADIATION PHYSICS↗

Quantum tensor network algorithms for evaluation of spectral functions on quantum computers

We investigate quantum algorithms derived from tensor networks to simulate the static and dynamic properties of quantum many-body systems. Using a sequentially prepared quantum circuit representation of a matrix product state (MPS) that we call a quantum tensor network (QTN), we demonstrate algorithms to prepare ground and excited states on a quantum computer and apply them to molecular nanomagnets (MNMs) as a paradigmatic example. In this setting, we develop two approaches for extracting the spectral correlation functions measured in neutron-scattering experiments: (a) a generalization of the SWAP test for computing wave function overlaps and, (b) a generalization of the notion of matrix product operators to the QTN setting which generates a linear combination of unitaries. The latter method is discussed in detail for translationally invariant spin-half systems, where it is shown to reduce the qubit resource requirements compared with the SWAP method and may be generalized to other systems. We demonstrate the versatility of our approaches by simulating spin-1/2 and spin-3/2 MNMs, with the latter being an experimentally relevant model of a Cr$^{3+}_{8}$ ring. Here, our approach has qubit requirements that are independent of the number of constituents of the many-body system and scale only logarithmically with the bond dimension of the MPS representation, making them appealing for implementation on near-term quantum hardware with mid-circuit measurement and reset.

Neutron scattering↗