Search NASA⌕ Search

SEARCH · Search NASA

Results for “Tensor”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 577 records · Page 32

Dispatch Manager for NEML2 Constitutive Model Calculations Embedded in MOOSE

This report describes the extended capabilities of the NEML2 constitutive modeling library, including a flexible and efficient work dispatching system designed to leverage both CPU and GPU resources. This enhancement addresses one of the primary computational challenges in large-scale simulations: the ability to distribute and execute batches of material model evaluations across heterogeneous computing devices. The new dispatch system introduces a modular set of dispatcher and scheduler classes that coordinate the flow of data and execution between devices. The dispatcher is responsible for efficiently packaging work, managing device-specific memory operations, and synchronizing results. This modularity allows for extensibility, making it straightforward to integrate additional computing backends in the future. From an implementation standpoint, the dispatcher system interfaces seamlessly with NEML2's existing models. They handle device-aware tensor operations, optimize memory transfers, and support asynchronous execution when applicable. This design ensures that batches of material points can be evaluated concurrently, substantially improving throughput compared to previous single-device or serial implementations. These improvements not only enhance the raw performance of NEML2 but also improve its usability in multiscale and high-fidelity simulations, where the simultaneous evaluation of large material point batches is critical. Benchmarks included in the report demonstrate the system’s scalability, highlighting its effectiveness when leveraging modern GPU architectures.

22 GENERAL STUDIES OF NUCLEAR REACTORS↗

Hybrid learning techniques for scientific data reduction with performance guarantees

The research initiatives supported by the U.S. Department of Energy (DOE) Grant DE-SC0022265 are fundamentally aimed at pioneering advanced machine learning (ML) techniques for scientific data compression within high-performance computing (HPC) environments. This comprehensive body of work addresses the critical challenge posed by the exponential growth of data generated by scientific simulations in domains such as fusion energy, climate modeling, and computational fluid dynamics (CFD). A core objective is to develop compression algorithms that achieve substantial data reduction—often by orders of magnitude—while rigorously ensuring the fidelity of both the primary data (PD) and scientifically crucial derived quantities of interest (QoI). The methodologies deployed under this grant integrate sophisticated deep learning architectures, prominently featuring autoencoders, advanced generative models like conditional diffusion, and hybrid learning techniques. Key innovations include the development of Guaranteed Autoencoders (GAE) and the Guaranteed Conditional Diffusion with Tensor Correction (GCDTC) framework, which provide explicit, instance-level error bounds on reconstructed data. Furthermore, specialized strategies such as nonlinear constraint satisfaction are employed to preserve the integrity of QoI, a vital requirement for the trustworthiness of downstream scientific analyses. This research also focuses on the design and implementation of scalable, GPU-accelerated software pipelines that seamlessly integrate into existing HPC workflows, ensuring both computational efficiency and practical applicability. The CAESAR framework, for example, unifies foundation and generative models to create an adaptive and efficient compression solution for spatio-temporal scientific data. Collectively, these efforts represent a significant advancement in mitigating the scientific data deluge, enabling more effective data management, accelerated scientific discovery, and optimized utilization of HPC resources.

97 MATHEMATICS AND COMPUTING↗

Final report- UFL - RAPIDS2: A SciDAC Institute for Computer Science, Data, and Artificial Intelligence

The research initiatives supported by the U.S. Department of Energy (DOE) Grant DE-SC0022265 are fundamentally aimed at pioneering advanced machine learning (ML) techniques for scientific data compression within high-performance computing (HPC) environments. This comprehensive body of work addresses the critical challenge posed by the exponential growth of data generated by scientific simulations in domains such as fusion energy, climate modeling, and computational fluid dynamics (CFD). A core objective is to develop compression algorithms that achieve substantial data reduction—often by orders of magnitude—while rigorously ensuring the fidelity of both the primary data (PD) and scientifically crucial derived quantities of interest (QoI). The methodologies deployed under this grant integrate sophisticated deep learning architectures, prominently featuring autoencoders, advanced generative models like conditional diffusion, and hybrid learning techniques. Key innovations include the development of Guaranteed Autoencoders (GAE) and the Guaranteed Conditional Diffusion with Tensor Correction (GCDTC) framework, which provide explicit, instance-level error bounds on reconstructed data. Furthermore, specialized strategies such as nonlinear constraint satisfaction are employed to preserve the integrity of QoI, a vital requirement for the trustworthiness of downstream scientific analyses. This research also focuses on the design and implementation of scalable, GPU-accelerated software pipelines that seamlessly integrate into existing HPC workflows, ensuring both computational efficiency and practical applicability. The CAESAR framework, for example, unifies foundation and generative models to create an adaptive and efficient compression solution for spatio-temporal scientific data. Collectively, these efforts represent a significant advancement in mitigating the scientific data deluge, enabling more effective data management, accelerated scientific discovery, and optimized utilization of HPC resources.

97 MATHEMATICS AND COMPUTING↗

Enabling Efficient Sparse Computations using Linear Algebra Aware Compilers

This project developed the LAPIS compiler framework, built on the Multilevel Intermediate Representation (MLIR), to optimize sparse linear algebra operations and support performance portability across diverse architectures. The main innovation of LAPIS is the Kokkos dialect, which allows for lowering codes from a high productivity language to different architectures in an elegant way. The dialect also allows the conversion of lower-level MLIR code to C++ Kokkos code, facilitating the integration of scientific machine learning (SciML) models into applications. To extend LAPIS for distributed memory architectures, a new partition dialect was created to manage the distribution of sparse tensors and express communication patterns for sparse linear algebra operations. This dialect also supports the distributed execution of operators and includes algorithmic optimizations to minimize communication to improve performance. The project also demonstrates that MLIR can enable effective linear algebra-level optimizations, improving performance on different GPUs for both sparse and dense linear algebra kernels. Key applications of LAPIS include sparse linear algebra and graph kernels, TenSQL, a relational database management solution built on GraphBLAS, and the development of subgraph isomorphism and monomorphism kernels, showcasing performance portability. In summary, the LAPIS framework supports productivity, performance, portability, and distributed memory execution, while also enabling linear algebra-level optimizations that are challenging in traditional programming languages, with successful applications ranging from simple sparse linear algebra to complex graph kernels.

97 MATHEMATICS AND COMPUTING↗

Quantum-Inspired Bayesian Sampling for Uncertainty Quantification and Machine Learning (Final Technical Report)

With increasing simulation and measurement data, machine learning and artificial intelligence have been widely used in computational decision-making of complex engineering systems. The resulting tools, such as uncertainty quantification solvers, reinforcement learning, and physics-informed machine learning, have achieved great success in critical DOE tasks such as material discovery and design, energy system modeling and control, and numerical weather and climate prediction. A core topic in scientific machine learning and artificial intelligence is Bayesian inference: given an observed data set, people want to estimate the posterior distribution of a (possibly large) number of hidden parameters. Due to the flexibility and weak assumptions, Bayesian sampling has been the mainstream Bayesian inference solvers despite the rapid progress of approximate Bayesian inference. Classical Bayesian sampling methods such as Markov-chain Monte Carlo suffer from a low-acceptance rate due to the random walk nature, therefore state-of-the-art techniques use Hamiltonian Monte Carlo and its variants to efficiently draw posterior samples in a high dimension. The key idea of Hamiltonian Monte Carlo and its variants is to simulate the Hamiltonian dynamics of a classical particle with a fixed mass, and their performance significantly degrades when the posterior distribution is highly spiky or has multiple modes. Leveraging the idea of quantum physics, this project has investigated new theory, algorithms and applications of Bayesian inference (especially Bayesian sampling). The main results include: (1) novel quantum-inspired Bayesian sampling methods that can lead to better accuracy for challenging multi-modal or spiky distributions, (2) more scalable machine learning framework leveraging tensor-compressed Bayesian inference, and (3) Bayesian and sampling approaches for verifying the robustness of continuous and binary neural networks.

97 MATHEMATICS AND COMPUTING↗

Frequency-Selectable Laser Source (FLS) for Cosmic Microwave Background Experiments

Cosmic Microwave Background (CMB) experiments measure remnant radiation from the early universe and use that data to determine fundamental properties of the universe. We can constrain key parameters, such as \textit{r}, the cosmic tensor-to-scalar ratio, and $N_\text{eff}$, the effective number of relativistic species, by analyzing the CMB power spectra. Improving our measurements of the CMB requires improving our instrument systematics, one of the most important of which is detector bandpass. Current experiments use a Fourier Transform Spectrometer (FTS) to measure bandpass. However, the FTS is systematics limited, and cannot achieve the accuracy needed to make improved CMB measurements. For this reason, we are developing a new instrument, the Frequency-Selectable Laser Source (FLS) to decrease the uncertainty in bandpass by an order of magnitude. In this paper, we describe work completed to support the version 2 upgrade to the FLS. Using ray-tracing software, we modeled the FLS optics to set physical tolerances for the new design. We also discuss the laser calibration, and future work to be completed in further development of the FLS upgrade.

Rosen-Turits, Gabriel M.↗

Technical Report on Adjoint Waveform Tomography of East Asia for Improved Waveform Prediction

We present a preliminary version of the East Asia Tomography (EAT) model, an adjoint waveform tomography model of East and Southeast Asia. We used SPiRaL (Simmons et al., 2021) as our starting model and source parameters for 250 earthquakes from the Global Centroid Moment Tensor catalogue (Ekström et al., 2012). Over 198 iterations, we have iterated the EAT model down to a minimum period of 35 seconds. We plan on continuing our iteration technique down to 30 seconds period before updating our misfit function to use a normalized cross correlation-based misfit functions (e.g., Tao et al., 2018) to better constrain Earth structure. We hope to iterate the current extent of the model to 25 seconds minimum period before iterating to shorter periods for a smaller subregion of the full model.

58 GEOSCIENCES↗

Improving the Performance of NEML2 with Modern Graph Compilation Backends

NEML2 vectorizes constitutive-model evaluation for large-scale multiphysics simulation, using PyTorch as its tensor backend so that a batch of material-point updates runs on CPU or GPU through a single implementation. In the two prior reports in this series it was a C++-native library, deployed through TorchScript tracing and just-in-time (JIT) compilation; it has since been rewritten from the ground up into a Python-native library deployed through Ahead-of-Time Inductor (AOTInductor), a modern PyTorch graph-compilation backend. The rewrite is driven by a persistent tension, not a language preference: NEML2 composes constitutive models at runtime from a registry of small, independently-authored pieces, and that flexibility is difficult to reconcile with the compile-time knowledge an efficient GPU kernel needs. This report documents the rewrite and the investment that accompanied it: the AOTInductor export pipeline that turns a Python-authored model into a portable, Python-free compiled artifact loadable from pure C++; the eager and compiled runtimes and the new implicit solver layer built on them; a head-to-head benchmark of legacy JIT against AOTInductor; the physics-model catalog and its worked examples; the developer tooling; and the corresponding overhaul of MOOSE’s NEML2 integration that lets MOOSE consume it. A central objective is to examine whether modern PyTorch graph-compilation backends are effective for MOOSE GPU integration. The benchmark answers directly: AOTInductor outperforms legacy JIT on every GPU scenario measured, by 1.0–4.5×. Modern graph-compilation backends are effective for MOOSE GPU integration, and AOTInductor specifically – not compilation in the abstract – is why.

Hu, Gary (Tianchen) [Argonne National Laboratory (↗

The Future Polarized Target Program at Jefferson Lab

Polarized targets have played a crucial role in Jefferson Lab's exploration of nuclear structure over the past four decades. The three original experimental halls have seen 19 separate installations of polarized solid or gas targets for use in the particle physics scattering experiments, and this trend will continue in the next decade. Five polarized target systems are in preparation for use at JLab in the coming years. Hall B will see the use of two solid polarized targets, one longitudinally polarized to the beam, the other transversely, as well as a novel 3He gas polarized target. In Hall C, new experiments will augment the tensor polarization in dynamically polarized solids. Plans are under development to bring a polarized solid target to Jefferson Lab's photon beam hall, Hall D, for the first time. To support these efforts, the JLab polarized target group is building a test laboratory to develop dynamic nuclear polarization techniques, as well as an apparatus to irradiate target material using electrons from JLab's injector test facility. In this talk, we will explore the development progress and plans for each of these efforts.

Maxwell, James [Thomas Jefferson National Accelera↗

The SPT-3G+ Experiment on the South Pole Telescope

Observations of the cosmic microwave background (CMB) offer an unparalleled opportunity to advance our understanding of fundamental physics. SPT-3G+ is an upgraded receiver for the arcminute-resolution South Pole Telescope (SPT) that plans to deploy in late 2028. SPT-3G+ will increase the CMB mapping speed of SPT by nearly an order of magnitude over the currently installed SPT-3G receiver. SPT-3G+ will have ~24,000 transition-edge sensor (TES) bolometers in two frequency bands with center frequencies at 95 GHz and 150 GHz that will be read out with microwave multiplexing. SPT-3G+ will measure the CMB lensing spectrum and galaxy clusters to constrain the growth of structure, dark matter, and dark energy. SPT-3G+ will also reach critical thresholds on inflationary constraints by combining data with BICEP/Keck, forming the South Pole Observatory (SPO). BICEP/Keck has deep degree-angular scale measurements but is currently delensing-limited, while SPT-3G+ will provide deep lensing measurements. Forecasts show that SPO will reach an uncertainty on the tensor-to-scalar ratio $r$ of $\sigma(r) \sim 1.2\times 10^{-3}$ by 2034. A detection at these levels would provide evidence of inflation and probe new physics at grand unified theory energy scales, while no detection would exclude large classes of models and shift the scientific paradigm describing the early universe. I will give an overview of SPT-3G+ including its design and current status.

Simon, Sara M. [Fermilab] (ORCID:0009000006683584)↗

Qubit Regularization of Quantum Field Theories

To study quantum field theories on a quantum computer, we must begin with Hamiltonians defined on a finite-dimensional Hilbert space and then take appropriate limits. This approach can be seen as a new type of regularization for quantum field theories, which we refer to as qubit regularization. A related finite-dimensional regularization, known as the D-theory approach, was proposed long ago as a general framework for all quantum field theories. In this framework, the dimensionality of the local Hilbert space at each spatial point can increase as needed through an additional flavor index. To reproduce asymptotically free QFTs, most studies assume that qubit-regularized theories require extending the local Hilbert space to infinity. However, contrary to this common belief, recent discoveries in (1+1) dimensions have revealed two examples where asymptotic freedom appears to emerge within a strictly finite-dimensional local Hilbert space through a novel renormalization group (RG) flow. These findings motivate further investigation into whether asymptotically free gauge theories could also emerge within a strictly finite-dimensional local Hilbert space. To support these explorations, we propose an orthonormal basis called the monomer-dimer-tensor-network (MDTN) basis and use it to construct new types of qubit-regularized lattice gauge theories.

Chandrasekharan, Shailesh [Duke Univ., Durham, NC ↗

Gravitational form factors of glueballs in Yang-Mills theory

This work presents preliminary results of the first determination of the energy-momentum tensor form factors of the scalar glueball, referred to as gravitational form factors (GFFs). The calculation has been carried out in lattice Yang-Mills theory at a single lattice spacing. Using variationally optimized operators, the matrix elements are extracted from ratios of three-point functions to two-point functions. The glueball GFFs and their kinematic dependence are compared to those of other hadrons from previous calculations.

Abbott, Ryan [Massachusetts Institute of Technolog↗

Confinement and Kink Entanglement Asymmetry on a Quantum Ising Chain

In this work, we explore the interplay of confinement, string breaking and entanglement asymmetry on a 1D quantum Ising chain. We consider the evolution of an initial domain wall and show that, surprisingly, while the introduction of confinement through a longitudinal field typically suppresses entanglement, it can also serve to increase it beyond a bound set for free particles. Our model can be tuned to conserve the number of domain walls, which gives an opportunity to explore entanglement asymmetry associated with link variables. We study two approaches to deal with the non-locality of the link variables, either directly or following a Kramers-Wannier transformation that maps bond variables (kinks) to site variables (spins). We develop a numerical procedure for computing the asymmetry using tensor network methods and use it to demonstrate the different types of entanglement and entanglement asymmetry.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

Ground state energy and magnetization curve of a frustrated magnetic system from real-time evolution on a digital quantum processor

Models of interacting many-body quantum systems that may realize new exotic phases of matter, notably quantum spin liquids, are challenging to study using even state-of-the-art classical methods such as tensor network simulations. Quantum computing provides a promising route for overcoming these difficulties to find ground states, dynamics, and more. In this paper, we argue that recently developed hybrid quantum-classical algorithms based on real-time evolution are promising methods for solving a particularly important model in the search for spin liquids, the antiferromagnetic Heisenberg model on the two-dimensional kagome lattice. We show how to construct efficient quantum circuits to implement time evolution for the model and to evaluate key observables on the quantum computer, and we argue that the method has favorable scaling with increasing system size. We then restrict to a 12-spin star plaquette from the kagome lattice and a related 8-spin system, and we give an empirical demonstration on these small systems that the hybrid algorithms can efficiently find the ground state energy and the magnetization curve. For these demonstrations, we use four levels of approximation: exact state vectors, exact state vectors with statistical noise from sampling, noisy classical emulators, and (for the 8-spin system only) real quantum hardware, specifically the Quantinuum H1-1 processor; for the noisy simulations and hardware demonstration, we also employ error mitigation strategies based on the symmetries of the Hamiltonian. Our results strongly suggest that these hybrid algorithms present a promising direction for studying quantum spin liquids and more generally for resolving important unsolved problems in condensed matter theory and beyond.

97 MATHEMATICS AND COMPUTING↗

Optimizing Distributed Training on Frontier for Large Language Models

Large language models (LLMs) have demonstrated remarkable success as foundational models, benefiting various downstream applications through fine-tuning. Loss scaling studies have demonstrated the superior performance of larger LLMs compared to their smaller counterparts. Nevertheless, training LLMs with billions of parameters poses significant challenges and requires considerable computational resources. For example, training a one trillion parameter GPT-style model on 20 trillion tokens requires a staggering 120 million exaflops. This research explores efficient distributed training strategies to extract this computation from Frontier, the world's first exascale supercomputer. We enable and investigate various model and data parallel training techniques, such as tensor parallelism, pipeline parallelism, and sharded data parallelism, to facilitate training a trillion-parameter model on Frontier. We empirically assess these techniques and their associated parameters to determine their impact on memory footprint, communication latency, and GPU's computational efficiency. We analyze the complex interplay among these techniques and find a strategy to combine them to achieve high throughput through hyperparameter tuning. We have identified efficient strategies for training large LLMs of varying sizes through empirical analysis and hyperparameter tuning. For 22 Billion, 175 Billion, and 1 Trillion parameters, we achieved GPU throughputs of 38.38%, 36.14%, and 31.96%, respectively. For the training of the 175 Billion parameter model and the 1 Trillion parameter model, we achieved 100% weak scaling efficiency on 1024 and 3072 Mi250X GPUs, respectively. We also achieved strong scaling efficiencies of 89% and 87% for these two models. We trained these models only tens of iterations instead of training till completion.

Yin, Junqi↗

The propagation of seismic waves, misinformation, and disinformation from the 2024-10-05 M 4.5 Iran earthquake

The 2024-10-05 Iran M 4.5 earthquake took place at a time of heightened tensions in the Middle East. We perform a discrimination and moment tensor analysis and identify a shallow-dipping, reverse fault source commensurate with the compressional setting of the Iranian interior. Nonetheless, the event's aftermath saw widespread dissemination of misinformation, and potentially active disinformation, concluding that it was in fact a test of an Iranian nuclear weapon. The 'evidence' for many of these claims was based on inaccurate interpretation of seismic data. In this paper, we analyze how geophysical 'fake news' propagated through social media (mainly Twitter/X) following this event, eventually gaining traction in mainstream, earned media. This event is an illustrative warning of how seismic data can be misinterpreted and/or manipulated in public discourse.

58 GEOSCIENCES↗

A User-Friendly GUI Tool for Automated Microstructural Analysis of Fiber-Reinforced Composites and Porous Structures

Understanding and quantifying microstructural features such as fiber orientation and porosity is critical for predicting the mechanical behavior and performance of fiber-reinforced polymer composites. Traditional manual analysis is time-consuming, subjective, and unsuitable for high-throughput datasets. We present a graphical user interface (GUI) application that automates the analysis of microscopy images to extract key microstructural metrics, including fiber orientation tensors, fiber orientation distribution, porosity and pore size distribution. The app integrates multiple image segmentation techniques including global and local thresholding, clustering, and region-based approaches, offering flexibility for different types of image qualities and features. Users can load microstructural images, select regions of interest and segmentation techniques tailored to their image dataset. It also addresses a critical challenge in fiber orientation analysis: the ambiguities caused by touching, overlapping, or partially cut fibers. It supports autorun examples for standardized workflows, enabling reproducible analysis and facilitating training and benchmarking. This tool significantly reduces manual intervention, enhances consistency, and accelerates data generation for structure–property modeling, process optimization, and digital materials research. The tool is intended for use by materials scientists, engineers, and researchers engaged in composite characterization, quality control, and machine learning-based microstructural studies.

Chawla, Komal [ORNL] (ORCID:0000000190327565)↗

A Unified Perspective on Poincaré and Galilei Relativity: II. General Relativity: A. Kinematics

Building on the first paper in this series (Paper I), a unified perspective on Poincaré and Galilei physics in a 5-dimensional spacetime setting is further pursued through a consideration of the kinematics of general relativity, with the gravitational dynamics to be addressed separately. The metric of the 5-dimensional affine spacetimes governed by the Bargmann groups considered in Paper I (central extensions of the Poincaré and Galilei groups) is generalized to curved spacetime by extending the usual 1 + 3 (traditionally ‘3 + 1’) formalism of general relativity on 4-dimensional spacetime to a 1 + 3 + 1 formalism, whose spacetime kinematics is shown to be consistent with that of the usual 1 + 3 formalism. Spacetime tensor laws governing the motion of an elementary classical material particle and the dynamics of a simple fluid are presented, along with their 1 + 3 + 1 decompositions; these reference the foliation of spacetime in a manner that partially reverts the Einstein perspective (accelerated fiducial observers, and geodesic material particles and fluid elements) to a Newton-like perspective (geodesic fiducial observers, and accelerated material particles and fluid elements subject to a gravitational force). These spacetime laws of motion for particles and fluids also suggest that a strong-field Galilei general relativity would involve a limit in which not only 𝑐 → ∞ but also 𝐺 → ∞ , such that 𝐺/𝑐 2 remains constant.

Bargmann group↗