Search NASA⌕ Search

SEARCH · Search NASA

Results for “Transformers”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 325 records · Page 18

Transformer-based operator learning framework for self-energy in strongly correlated systems

We introduce Σ-Attention, a transformer-based operator-learning framework for approximating the self-energy operator of strongly correlated electronic systems. By creating a batched dataset that combines results from three complementary approaches, i.e., many-body perturbation theory, strong-coupling expansion, and exact diagonalization, each effective in specific parameter regimes, Σ-Attention is applied to learn an accurate approximation for the self-energy operator that is valid across a wide range of parameter regimes. This hybrid strategy leverages the strengths of existing methods while relying on the transformer's ability to generalize beyond individual limitations. More importantly, the scalability of the transformer architecture allows the learned self-energy to be extended to systems with larger sizes, leading to much improved computational scaling. Using the one-dimensional Hubbard model, we demonstrate that Σ-Attention can accurately predict the Matsubara Green's function of large systems with a wide range of coupling strength. Our framework offers a promising and scalable pathway for studying strongly correlated systems with many possible generalizations.

Zhu, Yuanran↗

Unraveling the transformation pathway of the 𝛽 to 𝛾 phase transition in Ga 2 ⁢O 3 from atomistic simulations

Defect spinel 𝛾−Ga 2 ⁢O 3 is the least stable polymorph of Ga 2 ⁢O 3 , so its frequent appearance as a structural defect within or on the surface of monoclinic 𝛽−Ga 2 ⁢O 3 remains a mystery. Through first-principles calculations, we explore potential pathways for the phase transition from 𝛽−Ga 2⁢ O 3 to 𝛾−Ga 2 ⁢O 3 , and examine two key driving forces: tensile strain and Ga deficiency. When configurational entropy contributions to phase energies are included, the 𝛾 phase becomes energetically competitive with the 𝛽 phase, with the free energy difference between these phases diminishing even further under Ga-deficient conditions. Notably, a stability crossover occurs at room temperature at high vacancy concentrations ([V$^{3−}_{Ga}$]>3%) . A simple model 𝛽 → 𝛾 transformation pathway is identified, comprising two primary reactions, that enables the formation of the 𝛾 phase via simultaneous migration of Ga atoms from tetrahedral lattice sites to octahedral interstitial positions. The transformation barriers are prohibitively large in pristine Ga 2 ⁢O 3 , but can be substantially reduced by: (1) the presence of Ga vacancies, (2) elongational strains along the crystallographic 𝑎-axis, and (3) when volumetric relaxations are possible during transformation. These results elucidate prior experimental observations, where 𝛾−Ga 2⁢ O 3 is seen on damaged surfaces or in highly 𝑛-type 𝛽−Ga 2⁢ O 3 environments, which support Ga deficiency and mechanical strain. The insights into the driving forces and mechanisms of 𝛾−Ga 2⁢ O 3 formation enhance understanding of how localized strain and nonequilibrium defect concentrations may facilitate its formation from the 𝛽 phase.

Defects↗

Physics-informed transformation toward improving the machine-learned NLTE models of ICF simulations

The integration of machine-learning techniques into inertial confinement fusion (ICF) simulations has emerged as a powerful approach for enhancing computational efficiency. By replacing the costly nonlocal thermodynamic equilibrium (NLTE) model with machine-learning models, significant reductions in calculation time have been achieved. However, determining how to optimize machine-learning-based NLTE models in order to match ICF simulation dynamics remains challenging, underscoring the need for physically relevant error metrics and strategies to enhance model accuracy with respect to these metrics. Thus, we propose novel physics-informed transformations designed to emphasize energy transport, use these transformations to establish new error metrics, and demonstrate that they yield smaller errors within reduced principal-component spaces compared to conventional transformations. Published by the American Physical Society 2025

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

A low-temperature, one-step synthesis for monazite can transform monazite into a readily usable advanced nuclear waste form

It has been demonstrated that monazite-type materials are excellent candidates for nuclear waste forms, and hence, their facile synthesis is of great importance for the needed sequestration of existing nuclear waste. The synthesis of monazite, LaPO 4 , requires inconveniently high temperatures near 1000°C and generally involves the conversion of the presynthesized rhabdophane, LaPO 4 •nH 2 O, to the LaPO 4 monazite phase. During this structure transformation, the rhabdophane converts irreversibly to the thermodynamically stable monoclinic monazite structure. A low-temperature (185° to 260°C) mild hydrothermal acid-promoted synthesis of monazite is described that can both transform presynthesized rhabdophane or assemble reagents to the monoclinic monazite structure. The pH dependence of this reaction is detailed, and its applicability to the LnPO 4 (Ln = La, Ce, Pr, Nd, Sm-Gd), Ca 0.5 Th 0.5 PO 4 , and Sr 0.5 Th 0.5 PO 4 systems is discussed. The crystal growth of Ca 0.5 Th 0.5 PO 4 and Sr 0.5 Th 0.5 PO 4 is described, and their crystal structures were reported. In situ x-ray diffraction studies, performed as a function of temperature, provide insight into the structure transformation process.

Biogeochemistry↗

PyJMAK: An Open-Source Python Toolkit for Modeling Solid-State Metallurgical Phase Transformations

Accurate prediction of metallurgical phase transformations is an essential basis for autonomous optimization and rapid part qualification. Several methods can be used to estimate the evolution of phase fractions such as JMAK kinetics-based models, phase-field models, thermodynamic models, and data-driven machine learning models. Thermodynamic and phase-field-based methodologies solve multiphysics equations requiring numerous calibration parameters and significant computational resources. As a result, the computation domain is limited to a point or on order of micron-meters. The data-driven models rely on large datasets from experiments and simulations. While the JMAK model only provides information about phase fraction evolution, it can predict this evolution in near real-time using thermal history and thermodynamic data without restriction on the domain. JMAK models have been popularly used by researchers to model phase transformations occuring during additive manufacturing or over arbitrary temperature profiles. Commercial proprietary software such as Abaqus and Ansys or closed-source in-house implementations offer the ability to model JMAK based kinetics to predict phase transformation. However, these software packages are not open-source or freely available for use and development in conjunction with manufacturing machines, sensors, and machine learning algorithms. In addition, the use of the model is restricted by a license token. In contrast, given temperature profiles at multiple points in the domain, this Python-based PyJMAK model can compute phase evolution in parallel due to its stand-alone modular, voxel-based structure, and it can be executed on high-performance computing resources without any license restrictions.

Prabhune, Bhagya [Oak Ridge National Laboratory (O↗

Analyzing and Exploring Training Recipes for Large-Scale Transformer-Based Weather Prediction

Abstract The rapid rise of deep learning (DL) in numerical weather prediction (NWP) has led to a proliferation of models which forecast atmospheric variables with comparable or superior skill than traditional physics-based NWP. However, among these leading DL models, there is a wide variance in both the training settings and architecture used. Further, the lack of thorough ablation studies makes it hard to discern which components are most critical to success. In this work, we show that it is possible to attain high forecast skill even with relatively off-the-shelf architectures, simple training procedures, and moderate compute budgets. Specifically, we train a minimally modified Swin Transformer V2 (SwinV2) on ERA5 data and find that it attains superior skill in terms of mean-square errors of deterministic forecasts when compared against the European Centre for Medium-Range Weather Forecasts’ Integrated Forecasting System (IFS). Almost all DL–NWP systems share a core set of hyperparameters and design decisions. To aid and expedite future DL–NWP research, we present an in-depth, systematic exploration of different loss functions, model sizes and depths, patch sizes, and multistep training objectives. We also examine the model performance with metrics beyond the typical accuracy (ACC) and RMSE and investigate how the performance scales with model size. Through our open-source code, scoring pipelines, and models, we share our findings on key aspects of the training pipeline. These ablations reduce the necessity for expensive hyperparameter tuning and lower the barrier to entry for future DL–NWP research. Significance Statement This study investigates the potential of using large-scale transformer-based models for weather prediction, showing that it is possible to achieve high forecast accuracy with simpler, off-the-shelf architectures. By training a minimally modified SwinV2 transformer on ERA5 data, we show that the model achieves competitive forecast skill in terms of mean-square error for key variables, outperforming the European Centre for Medium-Range Weather Forecasts’ Integrated Forecasting System (IFS) at all lead times. Our findings suggest that effective training strategies, such as multistep fine-tuning and channel-weighted losses, significantly enhance the model’s performance. However, we also highlight that these improvements come with trade-offs in other areas, such as ensemble spread and high-frequency spatial detail. This work highlights the promise of deep learning in improving weather forecasts, which could lead to better preparedness and response to weather events, ultimately benefiting society by providing more reliable weather predictions.

Willard, Jared D. [Lawrence Berkeley National Labo↗

Sequence length scaling in vision transformers for scientific images on frontier

Vision Transformers (ViTs) are pivotal for foundational models in scientific imagery, including Earth science applications, due to their capability to process large sequence lengths. While transformers for text have inspired scaling sequence lengths in ViTs, adapting these for ViTs introduces unique challenges. We develop distributed sequence parallelism for ViTs, enabling them to handle up to 1M tokens. Our approach, leveraging DeepSpeed-Ulysses and Long-Sequence-Segmentation with model sharding, is the first to apply sequence parallelism in ViT training, achieving a 94% batch scaling efficiency on 2,048 AMD-MI250X GPUs. Evaluating sequence parallelism in ViTs, particularly in models up to 10B parameters, highlighted substantial bottlenecks. We countered these with hybrid sequence, pipeline, and flash attention strategies, to scale beyond single GPU memory limits. Our method significantly enhances climate modeling accuracy by 20% in temperature predictions, marking the first training of a vision transformer model to convergence with a sequence length of 188K tokens, using full self-attention.

Tsaris, Aristeidis (aris) [ORNL] (ORCID:0000000277↗

CurvilinearGrids.jl: A Julia package for curvilinear coordinate transformations

Finite-difference discretizations of partial differential equations are widespread throughout the scientific community. Oftentimes finite-differences are used to compute spatial gradients of fields on a discrete grid, which is typically a uniform or rectilinear Cartesian mesh. Arbitrary multidimensional geometry is difficult to discretize directly with finite differences, however, due to non-uniform grid spacing and non-orthogonality. Curvilinear coordinate transformations can be used as a strategy to enable arbitrary geometry. While these curvilinear transformations are straightforward, the governing PDEs require additional terms (metrics) and must adhere to strict conservation laws; these criteria complicate the application of the transformation and require careful implementation.

97 MATHEMATICS AND COMPUTING↗

Low-latency Jet Tagging for HL-LHC Using Transformer Architectures

Transformers are the state-of-the-art model architectures and widely used in application areas of machine learning. However the performance of such architectures is less well explored in the ultra-low latency domains where deployment on FPGAs or ASICs is required. Such domains include the trigger and data acquisition systems of the LHC experiments. We present a transformer-based algorithm for jet tagging built with the HGQ2 framework, which is able to produce a model with heterogeneous bitwidths for fast inference on FPGAs, as required in the trigger systems at the LHC experiments. The bitwidths are acquired during training by minimizing the total bit operations as an additional parameter. By allowing a bitwidth of zero, the model is pruned in-situ during training. Using this quantization-aware approach, our algorithm achieves state-of-the-art performance while also retaining permutation invariance which is a key property for particle physics applications. Due to the strength of transformers in representation learning, our work also serves as a stepping stone for the development of a larger foundation model for trigger applications.

Laatu, Lauri [Imperial Coll., London]↗

RTN-099: Photometric Transformation Relations for the LSST Data Preview 1

This technical note provides photometric transformation relations between the Vera C. Rubin Observatory's LSSTCam and LSSTComCam systems and other photometric systems. These transformations are derived using both synthetic and empirical data and are intended to support calibration and comparison across survey systems. We present both polynomial equations and lookup-table-based methods, depending on the available data and desired accuracy. The transformations are generally valid for stars with typical spectral energy distributions (SEDs), and caution should be used when applying them to objects with strong emission lines or atypical colors.

79 ASTRONOMY AND ASTROPHYSICS↗

RTN-125: Photometric Transformation Relations for the LSST Data Preview 2

This technical note provides photometric transformation relations between the NSF-DOE Vera C. Rubin Observatory's Data Preview 2 (DP2) and other photometric systems. These transformations are derived using both synthetic and empirical data and are intended to support calibration and comparison across survey systems. We present both polynomial equations and lookup-table-based methods, depending on the available data and desired accuracy. The transformations are generally valid for stars with typical spectral energy distributions (SEDs), and caution should be used when applying them to objects with strong emission lines or atypical colors.

79 ASTRONOMY AND ASTROPHYSICS↗

A transformer of closely spaced pulsed waveforms

Passive circuit, using diodes, transistors, and magnetic cores, transforms the voltage of repetitive positive or negative pulses. It combines a pulse transformer with switching devices to effect a resonant flux reset and can transform various pulsed waveforms that have a nonzero average value and are relatively cosely spaced in time.

Niedra, J.↗

Matrix transformations for spacecraft attitude determination

A common problem for experimental space physicists is the determination of the attitude matrix T which transforms vectors between representations in X and X' coordinate systems according to (vector V sub X) = (T sub XX')(vector V sub X'). A straightforward, simple, and efficient solution for the transformation matrix is a double-cross transformation. It is calculated from any two directions A and B, which are vectors normalized to unit length and are known in both X and X' coordinates. The B direction need be known only well enough to define the plane in which vectors A and B lie. The problem of the intersection of two cones as applicable to attitude solutions is also discussed.

Cauffman, D. P.↗

THREED: A computer program for three dimensional transformation of coordinates

Program THREED was developed for the purpose of a research study on the treatment of control data in lunar phototriangulation. THREED is the code name of a computer program for performing absolute orientation by the method of three-dimensional projective transformation. It has the capability of performing complete error analysis on the computed transformation parameters as well as the transformed coordinates.

Wong, K. W.↗

Reinvestigation of the olivine-spinel transformation in Ni2SiO4 and the incongruent melting of Ni2SiO4 olivine

The olivine-spinel transformation and the melting behavior of Ni2SiO4 were investigated over the PT ranges of 20-40 kbar, 650-1200 C, and 5-13 kbar, 1600-1700 C, respectively. It was confirmed that Ni2SiO4 olivine melts incongruently at high pressures and that it is a stable phase until melting occurs. The PT slope of the incongruent melting curve is approximately 105 bars/deg. The olivine-spinel transformation curve was shown to be a reversible univariant curve, and could be expressed by the linear equation P(bars) equals 23,300 + 11.8 x T(deg C). The transformation curve determined by Akimoto et al. (1965) is nearly parallel to that of the present work, but lies at pressures about 12% lower.

Ma, C.-B.↗

Application and sensitivity investigation of Fourier transforms for microwave radiometric inversions

Existing microwave radiometer technology now provides a suitable method for remote determination of the ocean surface's absolute brightness temperature. To extract the brightness temperature of the water from the antenna temperature equation, an unstable Fredholm integral equation of the first kind was solved. Fast Fourier Transform techniques were used to invert the integral after it is placed into a cross-correlation form. Application and verification of the methods to a two-dimensional modeling of a laboratory wave tank system were included. The instability of the Fredholm equation was then demonstrated and a restoration procedure was included which smooths the resulting oscillations. With the recent availability and advances of Fast Fourier Transform techniques, the method presented becomes very attractive in the evaluation of large quantities of data. Actual radiometric measurements of sea water are inverted using the restoration method, incorporating the advantages of the Fast Fourier Transform algorithm for computations.

Holmes, J. J.↗