Search NASA⌕ Search

SEARCH · Search NASA

Results for “Tensor”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 559 records · Page 31

Final Report for Center for Tokamak Transient Simulations at USU

Providing plasma fluid codes like NIMROD with continuum drift kinetic (CDK) physics that is quantitatively valid and computationally feasible throughout the spatial domain is difficult. Work at Utah State University (USU), in collaboration with the Center for Tokamak Transient Simulations (CTTS), focused on applying CDK closures in disruption-related calculations. Three examples where kinetic physics is paramount are (1) the electron stress tensor closure in Ohms law for accurately describing neoclassical tearing mode (NTM) evolution, (2) runaway electron (RE) density (nRE) and current (jRE) moments in NIMROD’s extended MHD model for self-consistent evolution of RE populations during disruptions and, (3) energetic ion effects on a myriad of MHD instabilities. While NTM simulations and continuum and PIC approaches to energetic ions in NIMROD have been a major goals of USU’s closure work for several years, the development of self-consistent CDK RE capability in NIMROD was started and extended considerably during the CTTS effort. Some goals of CDK RE in NIMROD are to explore the effects of the 2D relativistic phase space in 4D simulations and compare with NIMROD’s fluid RE model. Four publications and two PhD theses came out of the USU CTTS effort.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

Preserving Superconvergence of Spectral Elements for Curved Domains [Slides]

Finite Element Methods (FEM) and Spectral Element Methods (SEM) are crucial for solving partial differential equations (PDEs) on complex geometries. SEM offers superior accuracy due to potential superconvergence for simple domains. Challenges persist for domains with curved boundaries, restricting SEM’s advantages in real-world applications. A proposed solution is the introduction of a novel strategy to enhance accuracy and maintain superconvergence of SEM in curved domains. The strategy includes a mesh-generation procedure with geometrically refined elements near curved boundaries and a post-processing phase using the Adaptive Extended Stencil Finite Element Method (AES-FEM). The method, named AES-FEM post-processed Spectral Element Method (ApSEM), aligns the accuracy of non-tensor-product elements with superconvergent spectral elements.

97 MATHEMATICS AND COMPUTING↗

NEML2: A High Performance Library for Constitutive Modeling

NEML2, the New Engineering Material model Library, version 2, is an offshoot of NEML, an earlier material modeling code developed at Argonne National Laboratory. NEML2 extends the key philosophy of its predecessor, i.e., material models are flexible, modular, and can be built from smaller blocks. It also provides modern features that do not exist in the framework of its predecessor such as material model vectorization, automatic differentiation, device-portable just-in-time compilation, operator fusion, lazy tensor evaluation, etc. Moreover, NEML2 can seamlessly integrate with the popular machine learning package PyTorch to take advantage of modern and fast-growing machine learning techniques. In this fiscal year, the development of core library features and capabilities are complete. The purpose of this report is not to serve as a verbatim copy of the software API reference (which is available online at https://reverendbedford.github.io/neml2/). Instead, this report documents the motivation, implementation, design choices, and usage of each core capability as well as their applications in solving practical engineering problems. This report is compiled based on the NEML2 major release 2.0.0.

36 MATERIALS SCIENCE↗

Symbolic diagnostics to interpret and analyze neural network models

Embedded machine-learned models (EMLMs) have the promise to improve the predictive accuracy of engineering simulators in environments of national interest. EMLMs often comprise complex input-output maps (e.g., neural networks), which make them unamenable to rigorous analysis and generally difficult to interpret. In the face of decades of theory, this lack of interpretability is a significant barrier to building confidence in these models. This work outlines an approach to interpret EMLMs using sparse polynomial regression for comparison with theoretical understanding. To do so, we build on the concept of Locally Interpretable Model-agnostic Explanations (LIME) using physics-informed clustering, prototype selection, and library construction. While general, we demonstrate our method on tensor-basis neural networks used in Reynolds-Averaged Navier-Stokes simulations of hypersonic fluid flows. Results are presented for a simulated toy model and for direct numerical simulations (DNS) of turbulent flows over a flat plate.

97 MATHEMATICS AND COMPUTING↗

Adjoint waveform tomography of East Asia for improved waveform prediction

We present a preliminary version of the East Asia Tomography (EAT) model, an adjoint waveform tomography model of East and Southeast Asia. We used SPiRaL (Simmons et al., 2021) as our starting model and source parameters for 238 earthquakes from the Global Centroid Moment Tensor catalogue (Ekström et al., 2012). After 50 iterations on Lawrence Livermore National Laboratory’s Lassen supercomputer, we converge on a model with a minimum period of 50 seconds. The preliminary model shows improved slab structure compared to SPiRaL and shows significantly reduced misfit. In later versions of the model, we aim to harness techniques proposed in other studies to improve waveform predictions that travel through the ocean (e.g., Wehner et al., 2022) and use relative amplitude-based misfit functions (e.g., Tao et al., 2018) to better constrain Earth structure. We hope to iterate the current extent of the model to 25 seconds minimum period before iterating to shorter periods for a smaller subregion of the full model.

58 GEOSCIENCES↗

Developing Machine Learning Interatomic Potential for Fe-Cr-Ni Alloys

Accurate prediction of creep and fatigue behavior of stainless steel at elevated temperatures in hydrogen environment requires fundamental understanding of alloy-hydrogen interaction at cross-scale including bulk lattice and key defects such as vacancies, grain boundaries, surfaces, stacking faults, dislocations, and precipitates. This project aims to predict creep behavior of 347H stainless steel with H using machine learning interatomic potentials based on first-principles density functional theory simulations. The Moment Tensor Potentials platform is adopted for this work since it demonstrates a fine balance between model accuracy and computational efficiency. The potential is well trained based on large amount of high-fidelity density functional theory calculations. The validation is carried out by comparing various important properties including short range order, coefficient of thermal expansion, elastic properties, stacking fault energy, grain boundary energy, and surface energy. This work lays the foundation for reliable atomistic simulation of high temperature hydrogen attack of stainless steel.

density functional theory (DFT)↗

Multiplexing Focusing Analyzer for Efficient Stress-Strain Measurements

Statement of the problem or situation that is being addressed. Although thermal and cold neutron scattering is widely used and is critical for success in many areas of materials science and engineering, relatively low neutron fluxes severely limit applications of not only laboratory neutrons generators, but also large national neutron facilities. State-of-the-art thermal and cold neutron sources are large expensive national facilities, which serve diverse community of scientific and industrial users. The constant need to improve the instruments performance, stems from the fact that neutron methods are gaining in popularity, and becoming more and more powerful, while new neutron sources are not being constructed to keep pace with the developments and needs of the scientific community. Small research reactors at universities and National Labs, and laboratory-based neutron generators, are necessary not only for education and training, but also when samples cannot be transported to other facilities. However, the standard neutron techniques, which were developed for high-flux facilities, require much higher efficiencies to be used effectively with the low fluxes of small sources. Thus, the efficient use of neutron sources, such as with our proposed analyzer, is important for the progress and broader use of these neutron techniques. General statement of how this problem is being addressed. We propose to design and demonstrate novel diffractive optical device, which will enable very efficient residual stress neutron diffractometers. The proposed device will be a multi-foil analyzer, where each foil is constructed of focusing bent single crystals of Si. Such device will enable polychromatic residual stress neutron diffraction. At large national facilities, such as at Oak Ridge National Laboratory, these analyzers would enable very fast measurements for determining residual stress tensors, raster large samples or screen multiple samples. Commercial Applications and Other Benefits The outcome of this project would be the demonstration of commercial devices, novel neutron optical components, which could be utilized to improve the performance of existing instruments or build novel neutron scattering instruments at DOE neutron facilities and commercial laboratory neutron sources. These new devices will widen the scope of research conducted using neutrons and enable measurements not feasible at present. Summary for Members of Congress Thermal and cold neutron beams are a powerful materials science probe, which provide unique information about the structure of matter. The proposed innovations expand the reach of neutron-based investigations to new materials and industries by enabling new instrumentation capabilities, thereby greatly enhancing and expanding the role of small, laboratory-based neutron instrumentation, and improving education and training of neutron users.

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND ↗

Spectrally Stabilized Interface Capturing Formulation and Implementation in Nek5000/NekRS

This report documents the formulation of a novel level-set method for incompressible two-phase flows in the continuous Galerkin (CG) high order spectral element framework. The overall method hinges on a novel implementation of the spectral vanishing viscosity (SVV) operator for the stabilization of linear/non-linear hyperbolic problems. The multidimensional SVV convolution kernels, which in essence, have a similar effect as a high pass filter applied to the derivatives, are formulated by exploiting the tensor product form, analogous to the construction of the usual stiffness matrix system. The resulting kernels are directionally decoupled and ensure a linear, symmetric positive definite, elliptic matrix operator. The SVV formulation is demonstrated to provide a robust stabilizing mechanism through challenging linear and non-linear hyperbolic problems, including problems pertinent to the level-set formulation. The two-phase framework conceptualized herein is based on the conservative level-set (CLS) method which represents the interface between the fluids by the 0.5 iso-contour of the smoothed Heaviside function. The CLS method is augmented with a preconditioning procedure for interface normals using the signed distance function which precludes the manifestation of spurious oscillations in the vicinty of the interface. Further, the existing mixed explicit-implicit approach for the solution of Navier-Stokes equations in Nek5000, as described in Tomboulides et al, is augmented with a pressure coefficient splitting approach for the Poisson equation, which greatly accelerated the convergence of pressure solver for two-phase systems with large density ratio. The robustness and accuracy of the overall two-phase method is demonstrated through canonical challenging problems involving high density and viscosity ratios, with and without surface tension. The two-phase formulation is wholly implemented in Nek5000 and the SVV stabilization method is implemented in NekRS, which is the essential precursor to the two-phase framework, undergoing active development.

97 MATHEMATICS AND COMPUTING↗

Dispatch Manager for NEML2 Constitutive Model Calculations Embedded in MOOSE

This report describes the extended capabilities of the NEML2 constitutive modeling library, including a flexible and efficient work dispatching system designed to leverage both CPU and GPU resources. This enhancement addresses one of the primary computational challenges in large-scale simulations: the ability to distribute and execute batches of material model evaluations across heterogeneous computing devices. The new dispatch system introduces a modular set of dispatcher and scheduler classes that coordinate the flow of data and execution between devices. The dispatcher is responsible for efficiently packaging work, managing device-specific memory operations, and synchronizing results. This modularity allows for extensibility, making it straightforward to integrate additional computing backends in the future. From an implementation standpoint, the dispatcher system interfaces seamlessly with NEML2's existing models. They handle device-aware tensor operations, optimize memory transfers, and support asynchronous execution when applicable. This design ensures that batches of material points can be evaluated concurrently, substantially improving throughput compared to previous single-device or serial implementations. These improvements not only enhance the raw performance of NEML2 but also improve its usability in multiscale and high-fidelity simulations, where the simultaneous evaluation of large material point batches is critical. Benchmarks included in the report demonstrate the system’s scalability, highlighting its effectiveness when leveraging modern GPU architectures.

22 GENERAL STUDIES OF NUCLEAR REACTORS↗

Hybrid learning techniques for scientific data reduction with performance guarantees

The research initiatives supported by the U.S. Department of Energy (DOE) Grant DE-SC0022265 are fundamentally aimed at pioneering advanced machine learning (ML) techniques for scientific data compression within high-performance computing (HPC) environments. This comprehensive body of work addresses the critical challenge posed by the exponential growth of data generated by scientific simulations in domains such as fusion energy, climate modeling, and computational fluid dynamics (CFD). A core objective is to develop compression algorithms that achieve substantial data reduction—often by orders of magnitude—while rigorously ensuring the fidelity of both the primary data (PD) and scientifically crucial derived quantities of interest (QoI). The methodologies deployed under this grant integrate sophisticated deep learning architectures, prominently featuring autoencoders, advanced generative models like conditional diffusion, and hybrid learning techniques. Key innovations include the development of Guaranteed Autoencoders (GAE) and the Guaranteed Conditional Diffusion with Tensor Correction (GCDTC) framework, which provide explicit, instance-level error bounds on reconstructed data. Furthermore, specialized strategies such as nonlinear constraint satisfaction are employed to preserve the integrity of QoI, a vital requirement for the trustworthiness of downstream scientific analyses. This research also focuses on the design and implementation of scalable, GPU-accelerated software pipelines that seamlessly integrate into existing HPC workflows, ensuring both computational efficiency and practical applicability. The CAESAR framework, for example, unifies foundation and generative models to create an adaptive and efficient compression solution for spatio-temporal scientific data. Collectively, these efforts represent a significant advancement in mitigating the scientific data deluge, enabling more effective data management, accelerated scientific discovery, and optimized utilization of HPC resources.

97 MATHEMATICS AND COMPUTING↗

Final report- UFL - RAPIDS2: A SciDAC Institute for Computer Science, Data, and Artificial Intelligence

The research initiatives supported by the U.S. Department of Energy (DOE) Grant DE-SC0022265 are fundamentally aimed at pioneering advanced machine learning (ML) techniques for scientific data compression within high-performance computing (HPC) environments. This comprehensive body of work addresses the critical challenge posed by the exponential growth of data generated by scientific simulations in domains such as fusion energy, climate modeling, and computational fluid dynamics (CFD). A core objective is to develop compression algorithms that achieve substantial data reduction—often by orders of magnitude—while rigorously ensuring the fidelity of both the primary data (PD) and scientifically crucial derived quantities of interest (QoI). The methodologies deployed under this grant integrate sophisticated deep learning architectures, prominently featuring autoencoders, advanced generative models like conditional diffusion, and hybrid learning techniques. Key innovations include the development of Guaranteed Autoencoders (GAE) and the Guaranteed Conditional Diffusion with Tensor Correction (GCDTC) framework, which provide explicit, instance-level error bounds on reconstructed data. Furthermore, specialized strategies such as nonlinear constraint satisfaction are employed to preserve the integrity of QoI, a vital requirement for the trustworthiness of downstream scientific analyses. This research also focuses on the design and implementation of scalable, GPU-accelerated software pipelines that seamlessly integrate into existing HPC workflows, ensuring both computational efficiency and practical applicability. The CAESAR framework, for example, unifies foundation and generative models to create an adaptive and efficient compression solution for spatio-temporal scientific data. Collectively, these efforts represent a significant advancement in mitigating the scientific data deluge, enabling more effective data management, accelerated scientific discovery, and optimized utilization of HPC resources.

97 MATHEMATICS AND COMPUTING↗

Enabling Efficient Sparse Computations using Linear Algebra Aware Compilers

This project developed the LAPIS compiler framework, built on the Multilevel Intermediate Representation (MLIR), to optimize sparse linear algebra operations and support performance portability across diverse architectures. The main innovation of LAPIS is the Kokkos dialect, which allows for lowering codes from a high productivity language to different architectures in an elegant way. The dialect also allows the conversion of lower-level MLIR code to C++ Kokkos code, facilitating the integration of scientific machine learning (SciML) models into applications. To extend LAPIS for distributed memory architectures, a new partition dialect was created to manage the distribution of sparse tensors and express communication patterns for sparse linear algebra operations. This dialect also supports the distributed execution of operators and includes algorithmic optimizations to minimize communication to improve performance. The project also demonstrates that MLIR can enable effective linear algebra-level optimizations, improving performance on different GPUs for both sparse and dense linear algebra kernels. Key applications of LAPIS include sparse linear algebra and graph kernels, TenSQL, a relational database management solution built on GraphBLAS, and the development of subgraph isomorphism and monomorphism kernels, showcasing performance portability. In summary, the LAPIS framework supports productivity, performance, portability, and distributed memory execution, while also enabling linear algebra-level optimizations that are challenging in traditional programming languages, with successful applications ranging from simple sparse linear algebra to complex graph kernels.

97 MATHEMATICS AND COMPUTING↗

Quantum-Inspired Bayesian Sampling for Uncertainty Quantification and Machine Learning (Final Technical Report)

With increasing simulation and measurement data, machine learning and artificial intelligence have been widely used in computational decision-making of complex engineering systems. The resulting tools, such as uncertainty quantification solvers, reinforcement learning, and physics-informed machine learning, have achieved great success in critical DOE tasks such as material discovery and design, energy system modeling and control, and numerical weather and climate prediction. A core topic in scientific machine learning and artificial intelligence is Bayesian inference: given an observed data set, people want to estimate the posterior distribution of a (possibly large) number of hidden parameters. Due to the flexibility and weak assumptions, Bayesian sampling has been the mainstream Bayesian inference solvers despite the rapid progress of approximate Bayesian inference. Classical Bayesian sampling methods such as Markov-chain Monte Carlo suffer from a low-acceptance rate due to the random walk nature, therefore state-of-the-art techniques use Hamiltonian Monte Carlo and its variants to efficiently draw posterior samples in a high dimension. The key idea of Hamiltonian Monte Carlo and its variants is to simulate the Hamiltonian dynamics of a classical particle with a fixed mass, and their performance significantly degrades when the posterior distribution is highly spiky or has multiple modes. Leveraging the idea of quantum physics, this project has investigated new theory, algorithms and applications of Bayesian inference (especially Bayesian sampling). The main results include: (1) novel quantum-inspired Bayesian sampling methods that can lead to better accuracy for challenging multi-modal or spiky distributions, (2) more scalable machine learning framework leveraging tensor-compressed Bayesian inference, and (3) Bayesian and sampling approaches for verifying the robustness of continuous and binary neural networks.

97 MATHEMATICS AND COMPUTING↗

Frequency-Selectable Laser Source (FLS) for Cosmic Microwave Background Experiments

Cosmic Microwave Background (CMB) experiments measure remnant radiation from the early universe and use that data to determine fundamental properties of the universe. We can constrain key parameters, such as \textit{r}, the cosmic tensor-to-scalar ratio, and $N_\text{eff}$, the effective number of relativistic species, by analyzing the CMB power spectra. Improving our measurements of the CMB requires improving our instrument systematics, one of the most important of which is detector bandpass. Current experiments use a Fourier Transform Spectrometer (FTS) to measure bandpass. However, the FTS is systematics limited, and cannot achieve the accuracy needed to make improved CMB measurements. For this reason, we are developing a new instrument, the Frequency-Selectable Laser Source (FLS) to decrease the uncertainty in bandpass by an order of magnitude. In this paper, we describe work completed to support the version 2 upgrade to the FLS. Using ray-tracing software, we modeled the FLS optics to set physical tolerances for the new design. We also discuss the laser calibration, and future work to be completed in further development of the FLS upgrade.

Rosen-Turits, Gabriel M.↗

Technical Report on Adjoint Waveform Tomography of East Asia for Improved Waveform Prediction

We present a preliminary version of the East Asia Tomography (EAT) model, an adjoint waveform tomography model of East and Southeast Asia. We used SPiRaL (Simmons et al., 2021) as our starting model and source parameters for 250 earthquakes from the Global Centroid Moment Tensor catalogue (Ekström et al., 2012). Over 198 iterations, we have iterated the EAT model down to a minimum period of 35 seconds. We plan on continuing our iteration technique down to 30 seconds period before updating our misfit function to use a normalized cross correlation-based misfit functions (e.g., Tao et al., 2018) to better constrain Earth structure. We hope to iterate the current extent of the model to 25 seconds minimum period before iterating to shorter periods for a smaller subregion of the full model.

58 GEOSCIENCES↗

Improving the Performance of NEML2 with Modern Graph Compilation Backends

NEML2 vectorizes constitutive-model evaluation for large-scale multiphysics simulation, using PyTorch as its tensor backend so that a batch of material-point updates runs on CPU or GPU through a single implementation. In the two prior reports in this series it was a C++-native library, deployed through TorchScript tracing and just-in-time (JIT) compilation; it has since been rewritten from the ground up into a Python-native library deployed through Ahead-of-Time Inductor (AOTInductor), a modern PyTorch graph-compilation backend. The rewrite is driven by a persistent tension, not a language preference: NEML2 composes constitutive models at runtime from a registry of small, independently-authored pieces, and that flexibility is difficult to reconcile with the compile-time knowledge an efficient GPU kernel needs. This report documents the rewrite and the investment that accompanied it: the AOTInductor export pipeline that turns a Python-authored model into a portable, Python-free compiled artifact loadable from pure C++; the eager and compiled runtimes and the new implicit solver layer built on them; a head-to-head benchmark of legacy JIT against AOTInductor; the physics-model catalog and its worked examples; the developer tooling; and the corresponding overhaul of MOOSE’s NEML2 integration that lets MOOSE consume it. A central objective is to examine whether modern PyTorch graph-compilation backends are effective for MOOSE GPU integration. The benchmark answers directly: AOTInductor outperforms legacy JIT on every GPU scenario measured, by 1.0–4.5×. Modern graph-compilation backends are effective for MOOSE GPU integration, and AOTInductor specifically – not compilation in the abstract – is why.

Hu, Gary (Tianchen) [Argonne National Laboratory (↗

The Future Polarized Target Program at Jefferson Lab

Polarized targets have played a crucial role in Jefferson Lab's exploration of nuclear structure over the past four decades. The three original experimental halls have seen 19 separate installations of polarized solid or gas targets for use in the particle physics scattering experiments, and this trend will continue in the next decade. Five polarized target systems are in preparation for use at JLab in the coming years. Hall B will see the use of two solid polarized targets, one longitudinally polarized to the beam, the other transversely, as well as a novel 3He gas polarized target. In Hall C, new experiments will augment the tensor polarization in dynamically polarized solids. Plans are under development to bring a polarized solid target to Jefferson Lab's photon beam hall, Hall D, for the first time. To support these efforts, the JLab polarized target group is building a test laboratory to develop dynamic nuclear polarization techniques, as well as an apparatus to irradiate target material using electrons from JLab's injector test facility. In this talk, we will explore the development progress and plans for each of these efforts.

Maxwell, James [Thomas Jefferson National Accelera↗

The SPT-3G+ Experiment on the South Pole Telescope

Observations of the cosmic microwave background (CMB) offer an unparalleled opportunity to advance our understanding of fundamental physics. SPT-3G+ is an upgraded receiver for the arcminute-resolution South Pole Telescope (SPT) that plans to deploy in late 2028. SPT-3G+ will increase the CMB mapping speed of SPT by nearly an order of magnitude over the currently installed SPT-3G receiver. SPT-3G+ will have ~24,000 transition-edge sensor (TES) bolometers in two frequency bands with center frequencies at 95 GHz and 150 GHz that will be read out with microwave multiplexing. SPT-3G+ will measure the CMB lensing spectrum and galaxy clusters to constrain the growth of structure, dark matter, and dark energy. SPT-3G+ will also reach critical thresholds on inflationary constraints by combining data with BICEP/Keck, forming the South Pole Observatory (SPO). BICEP/Keck has deep degree-angular scale measurements but is currently delensing-limited, while SPT-3G+ will provide deep lensing measurements. Forecasts show that SPO will reach an uncertainty on the tensor-to-scalar ratio $r$ of $\sigma(r) \sim 1.2\times 10^{-3}$ by 2034. A detection at these levels would provide evidence of inflation and probe new physics at grand unified theory energy scales, while no detection would exclude large classes of models and shift the scientific paradigm describing the early universe. I will give an overview of SPT-3G+ including its design and current status.

Simon, Sara M. [Fermilab] (ORCID:0009000006683584)↗