Search NASA⌕ Search

SEARCH · Search NASA

Results for “Transformer model”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

Geothermal well testing pressure prediction by using a hybrid transformer model system: FORGE well use case

Geothermal has huge potential to become an indispensable component in achieving the goal of sustainable energy economy, given its capability to provide consistent baseload power to the electric grid. Injection tests are crucial in geothermal energy system as they naturally help to evaluate reservoir properties, understand fluid flow and even enhance reservoir performance. In this research, we developed a hybrid model system that integrates machine learning (ML) regression, a physics-based mathematical model, and transformer deep learning. Trained and validated using FORGE injection test dataset, this system can forecast the pressure variations both upward and downward over time. The pressure prediction achieved prediction accuracy within 3-6% variance of true pressure values. The system can significantly save time and reduce costs by testing only a few cycles and then using model predictions for further analysis, instead of conducting additional real injection cycle tests. The developed model system also holds promise for designing injection test processes and maintaining well production in geothermal energy. Presented at the IMAGE ‘25 Conference led by Shell.

FORGE↗

Driver Distraction Behavior Detection using a Vision Transformer Model based on Transfer Learning Strategy

Driver distraction behavior is one of the critical factors in traffic accidents. Thus, advanced driver state detection system has become the focus in the field of intelligent vehicle. However, in practical applications, insufficient samples of driving distraction behaviors bring great challenges to training a personalized behavior distraction detection model for a specific driver. To this end, a novel transformer model based on a transfer learning strategy is proposed in this paper to accurately recognize driver distraction behavior. Inspired by the effect of the transformer network in visual recognition, we firstly present a transformer behavior distraction detection system to identify the behavior categories that cause driver distraction. Then, for the specific driving dataset in practical application scenarios, the transfer learning strategy is introduced into the driver distraction detection model to further train the general transformer network. The effectiveness of the transformer based on the transfer learning strategy is validated compared with other traditional deep learning methods. The results show that the proposed detection method has better generalization ability and higher accuracy.

Fang, Zhenwu↗

MTL_TX: A Multi-Task Transformer Model for Improved Radiation Time-Series Estimation

Controlling radiation doses at potential radioactive facilities is critical to ensuring the safety of both personnel and the public. At the Thomas Jefferson National Accelerator Facility (JLab), multiple sensors are deployed around the three experimental halls to monitor key parameters, including single-beam current, energy levels, current leakage, and radiation values during accelerator operations. In this study, we developed a Multi-task Transformer model, MTL_TX, to accurately estimate radiation doses at sensor locations based on historical data, with the aim of enhancing safety in accelerator facilities and surrounding public areas. To improve estimation accuracy, we integrated two innovative components into the proposed model: hierarchical feature embedding (HFE) and multi-level decomposition attention (MDA). Additionally, the multi-task learning (MTL) framework effectively leverages correlations among multiple sensors, enabling individual estimations for each sensor. MTL_TX achieved outstanding results on data collected in 2018, with an MSE of 0.1464, an RMSE of 0.2353, and an R 2 score of 0.8584. Furthermore, when trained on 2018 data, MTL_TX exhibited excellent generalization capability to unseen datasets from 2016 to 2019, achieving an MSE of 0.1407, an RMSE of 0.2263, and an R 2 score of 0.8831. These results demonstrate a significant improvement over existing state-of-the-art models.

Transformer↗

Predicting Future Laboratory Fault Friction Through Deep Learning Transformer Models

Machine learning models using seismic emissions as input can predict instantaneous fault characteristics such as displacement and friction in laboratory experiments, and slow slip in Earth. Here, we address whether the seismic/acoustic emission (AE) from laboratory experiments contains information about future frictional behavior. The approach uses a convolutional encoder-decoder containing a transformer model in the latent space, similar to models used for natural language processing. We test the model limits using progressively larger AE input time windows and progressively larger output friction time windows. The results demonstrate that very near-term friction predictions are indeed contained in the AE signal, and predictions are progressively worse farther into the future. The future predictions by the model of impending failure in the near-term are remarkably robust. This first effort predicting future fault frictional behavior with machine learning will aid in guiding efforts for applications in Earth.

58 GEOSCIENCES↗

A scalable transformer model for real-time decision making in neutron scattering experiments

The U.S. Department of Energy's (DOE's) neutron research facilities at Oak Ridge National Laboratory (ORNL), including the High Flux Isotope Reactor (HFIR) and the Spallation Neutron Source (SNS), are a state-of-the-art neutron scattering facility that allows researchers to study the structure and dynamics of materials at the atomic scale. At the SNS, neutrons are measured using the time-of-flight (TOF) technique as they move through a neutron beamline to interact with a sample. Large volumes of neutron scattering data are collected and recorded in neutron event mode. Optimal productivity of the TOF instrument is limited due to the lack of real-time data analysis tools. The large amount of data generated by the experiments can be challenging to process and analyze in real time, particularly for experiments that require rapid feedback and adjustment of experimental parameters. The regular computer/workstation cannot keep up with the experiment speed to provide real-time feedback to adjust experimental parameters, so connecting the supercomputers available to the neutron facility is necessary to achieve real-time data analysis and experiment steering. To address this challenge, we exploit the Frontier supercomputer at Oak Ridge Leadership Computing Facility (OLCF) to train a scalable temporal fusion transformer model for real-time decision making of TOF neutron scattering experimentation. Here, in this paper, we present the results using Frontier to provide the processing power needed to rapidly process and analyze large volumes of single-crystal diffraction data collected at TOPAZ, a neutron time-of-flight Laue single-crystal diffractometer at the SNS.

97 MATHEMATICS AND COMPUTING↗

CovTransformer: A transformer model for SARS-CoV-2 lineage frequency forecasting

With hundreds of SARS-CoV-2 lineages circulating in the global population, there is an ongoing need for predicting and forecasting lineage frequencies and thus identifying rapidly expanding lineages. Accurate prediction would allow for more focused experimental efforts to understand pathogenicity of future dominating lineages and characterize the extent of their immune escape. Here, we first show that the inherent noise and biases in lineage frequency data make a commonly-used regression-based approach unreliable. To address this weakness, we constructed a machine learning model for SARS-CoV-2 lineage frequency forecasting, called CovTransformer, based on the transformer architecture. We designed our model to navigate challenges such as a limited amount of data with high levels of noise and bias. We first trained and tested the model using data from the UK and the USA, and then tested the generalization ability of the model to many other countries and US states. Remarkably, the trained model makes accurate predictions two months into the future with high levels of accuracy both globally (in 31 countries with high levels of sequencing effort) and at the US-state level. Our model performed substantially better than a widely used forecasting tool, the multinomial regression model implemented in Nextstrain, demonstrating its utility in SARS-CoV-2 monitoring. Assuming a newly emerged lineage is identified and assigned, our test using retrospective data shows that our model is able to identify the dominating lineages 7 weeks in advance on average before they became dominant. Overall, our work demonstrates that transformer models represent a promising approach for SARS-CoV-2 forecasting and pandemic monitoring.

60 APPLIED LIFE SCIENCES↗

Novel Deep Learning Transformer Model for Short to Sub‐Seasonal Streamflow Forecast

Accurate short-to-subseasonal streamflow forecasts are becoming crucial for effective water management in an increasingly variable climate. However, streamflow forecast remains challenging over extended lead times, uncertainty in meteorological inputs, and increased frequency and variability in extreme weather and climate events. We implemented a Future Time Series Transformer (FutureTST) model for streamflow forecasting that separately integrates past meteorological and streamflow data while incorporating future weather conditions. FutureTST achieves a mean Nash-Sutcliffe Efficiency (NSE) of 0.82 to 0.67 for 1- to 30-day streamflow forecasts. Incorporating upstream streamflow information improved forecast accuracy by up to 10%. During real-time forecast, FutureTST maintains higher forecast skills of 9.03 for 1-day and 5.74 for 14-day forecasts. In contrast, calibrated process-based hydrological model forecasts become unreliable beyond a 4-day lead time. Our findings demonstrate the potential of FutureTST as a reliable streamflow forecasting tool that offers a valuable addition to operational flood monitoring systems and climate-resilient decision-making.

Ambika, Anukesh Krishnankutty [Oak Ridge National ↗

Elasto-viscoplastic fast Fourier transform modeling framework for assessing microstructural effects on stress intensity factors characterizing fracture toughness

A large-strain elasto-viscoplastic fast Fourier transform (LS-EVPFFT) model with non-periodic (NP) velocity-based boundary conditions is adapted to simulate the sensitivity of stress intensity factors on microstructure for 304L stainless steel. The material was characterized via electron backscattered diffraction (EBSD) serial-sectioning to obtain a measured 3-D microstructural cell to perform simulations. The NP-LS-EVPFFT model, including the simulation setup and boundary conditions, was verified using a crystal plasticity finite element (CPFE) model. To this end, the generation of meshes of notched specimens was developed, which involved creating Python scripts for mesh “cutting” in Abaqus, and Sculpt scripts in Cubit for meshing of the measured microstructural cell processed with DREAM.3D. The complexity of the mesh preparation highlighted the advantages of the FFT-based model, which circumvents the mesh generation process. Given the efficiency of the FFT-based model, statistical distribution of stress intensity factors in function of crystal orientation at the crack tip, grain structure, and crystallographic texture surrounding the crack tip were predicted. Further, the distributions reveal about 10% variation of stress intensity factors with microstructure with the most significant sensitivity found to be the crystal orientation at the crack tip. The methodology developed in this work is discussed as a practical simulation tool for predicting the sensitivity of stress intensity factors on microstructural variability in metallic materials.

36 MATERIALS SCIENCE↗

MATEY: multiscale adaptive transformer models for spatiotemporal physical systems

Accurate representation of the multiscale features in spatiotemporal physical systems using vision transformer architectures requires extremely long, computationally prohibitive token sequences. To address this issue, we propose two novel adaptive tokenization schemes that dynamically adjust patch sizes based on local features: one ensures convergent behavior to uniform patch refinement, while the other offers better computational efficiency. Moreover, we present a set of spatiotemporal attention schemes, where the temporal or axial spatial dimensions are decoupled, to evaluate their baseline computational and data efficiencies and to determine whether adaptive tokenization can improve this performance. We assess the performance of the proposed multiscale adaptive model, MATEY, in a sequence of experiments. Compared to a full spatiotemporal attention scheme or a scheme that decouples only the temporal dimension, we find that fully decoupled axial attention is less efficient and expressive, requiring more training time and model parameters to achieve the same accuracy. The experiments on the adaptive tokenization schemes show that, compared to a uniformly refined model, the proposed schemes achieve comparable or improved accuracy at a much lower cost in the tested two-dimensional settings. While the asymptotic analysis suggests the potential for favorable scaling, empirical validation at substantially longer sequence lengths remains to be performed in future work. Finally, we demonstrate in two fine-tuning tasks featuring different physics that models pretrained on PDEBench data outperform the ones trained from scratch, especially in the low data regime with frozen attention.

adaptive tokenization↗

Synthesis of Correct Digital Controller Models from Specifications by Model Transformation (21-0320)

The design of high consequence controllers (in weapons systems, autonomy, etc.) that do what they are supposed to do is a significant challenge. Testing simply does not come close to meeting the requirements for assurance. Today circuit designers at Sandia (and elsewhere) typically capture the core behavior of their components using state models in tools such as STATEFLOW. They then check that their models meet certain requirements (e.g. “The system bus must not deadlock” or “both traffic lights at an intersection must not be green at the same time”) using tools called model checkers. If the model checker returns “yes” then the property is guaranteed to be satisfied by the model. However, there are several drawbacks to this industry practice: (1) there is a lot of detail to get right, this is particularly challenging when there are multiple components requiring complex coordination (2) any errors returned by the model checker have to be traced back through the design and fixed, necessitating rework, (3) there are severe scalability problems with this approach, particularly when dealing with concurrency. All this places high demands on the designers who now face not only an accelerated schedule but also controllers of increasing complexity. This report describes a new and fundamentally different approach to the construction of safety-critical digital controllers. Instead of directly constructing a complete model and then trying to verify it, the designer can start with an initial abstract (think “sketch”) model plus the requirements, from which a correct concrete model is automatically synthesized. There is no need for post-hoc verification of required functional properties. Having tool to carry this out will significantly impact the nation’s ability to ensure the safety of high-consequence digital systems. The approach has been implemented in a prototype tool, along with a suite of examples, including ones that reflect actual problems faced by designers. Our approach operates on a variant of Statecharts developed at Sandia called Qspecs. Statecharts are a widely used formalism for developing concurrent reactive systems, supporting scalability through allowing state models containing composite states, which are the serial or parallel composition of substates which can themselves contain statecharts. Statecharts enable an incremental style of development, in which states are progressively refined to incorporate greater detail in an incremental model of software development. Our approach formulates a set of constraints from the structure of the models and the requirements and propagates these constraints to a fixpoint. The solution to the constraints is an inductive invariant along with guards on the transitions. We also show how our approach extends to implementation refinement, decomposition, composition, and elaboration. We currently handle safety requirements written in LTL (Linear Temporal Logic)

42 ENGINEERING↗

CryoTransformer: a transformer model for picking protein particles from Cryo-EM micrographs

Cryo-electron microscopy (cryo-EM) is a powerful technique for determining the structures of large protein complexes. Picking single protein particles from cryo-EM micrographs (images) is a crucial step in reconstructing protein structures from them. However, the widely used template-based particle picking process requires some manual particle picking and is labor-intensive and time-consuming. Though machine learning and artificial intelligence (AI) can potentially automate particle picking, the current AI methods pick particles with low precision or low recall. The erroneously picked particles can severely reduce the quality of reconstructed protein structures, especially for the micrographs with low signal-to-noise-ratio (SNR).

59 BASIC BIOLOGICAL SCIENCES↗

Towards an astronomical foundation model for stars with a transformer-based model

ABSTRACT Rapid strides are currently being made in the field of artificial intelligence using transformer-based models like Large Language Models (LLMs). The potential of these methods for creating a single, large, versatile model in astronomy has not yet been explored. In this work, we propose a framework for data-driven astronomy that uses the same core techniques and architecture as used by LLMs. Using a variety of observations and labels of stars as an example, we build a transformer-based model and train it in a self-supervised manner with cross-survey data sets to perform a variety of inference tasks. In particular, we demonstrate that a single model can perform both discriminative and generative tasks even if the model was not trained or fine-tuned to do any specific task. For example, on the discriminative task of deriving stellar parameters from Gaia XP spectra, we achieve an accuracy of 47 K in Teff, 0.11 dex in log g, and 0.07 dex in [M/H], outperforming an expert XGBoost model in the same setting. But the same model can also generate XP spectra from stellar parameters, inpaint unobserved spectral regions, extract empirical stellar loci, and even determine the interstellar extinction curve. Our framework demonstrates that building and training a single foundation model without fine-tuning using data and parameters from multiple surveys to predict unmeasured observations and parameters is well within reach. Such ‘Large Astronomy Models’ trained on large quantities of observational data will play a large role in the analysis of current and future large surveys.

Leung, Henry W. (ORCID:0000000200362752)↗

Crystal plasticity modeling of strain-induced martensitic transformations to predict strain rate and temperature sensitive behavior of 304 L steels: Applications to tension, compression, torsion, and impact

This paper advances crystallographically-based Olson-Cohen (direct γ → α’) and deformation mechanism (indirect γ→ε→α’) phase transformation models for predicting strain-induced austenite to martensite transformation. Here, the advanced transformation models enable predictions of not only strain-path sensitive, but also of strain-rate and temperature sensitive deformation of polycrystalline stainless steels (SSs). The deformation of constituent grains in SSs is modeled as a combination of anisotropic elasticity, crystallographic slip, and phase transformation, while the hardening is based on the evolution of dislocation density and explicit shifts in phase fractions. Such grain-scale deformation is implemented within the meso-scale elasto-plastic self-consistent (EPSC) homogenization model, which is coupled with the implicit finite element (FE) method to provide a constitutive response at each FE integration point for solving boundary value problems at the macro-scale. Parameters pertaining to the hardening and transformation models within FEEPSC are calibrated and validated on a suite of data including flow curves and phase fractions for monotonic compression, tension, and torsion as a function of strain-rate and temperature for wrought and additively manufactured (AM) SS304L. To illustrate the potential and accuracy of the integrated multi-level FE-EPSC simulation framework, geometry, mechanical response, phase fractions, and texture evolution are simulated during gas-gun impact deformation of a cylinder and quasi-static tension of a notched specimen made of AM SS304L. Details of the simulation framework, comparison between experimental and simulation results, and insights from the results are presented and discussed.

304L steels↗

Phase transformation kinetics model for metals

We develop a new model for phase transformation kinetics in metals by generalizing the Levitas–Preston (LP) phase field model of martensite phase transformations (see Levitas and Preston (2002a,b) and Levitas et al. (2003)) to arbitrary pressure. Furthermore, we account for and track: the interface speed of the pressure-driven phase transformation, properties of critical nuclei, as well as nucleation at grain sites and on dislocations and homogeneous nucleation. The volume fraction evolution of each phase is described by employing KJMA (Kolmogorov, 1937; Johnson and Mehl, 1939; Avrami, 1939, 1940, 1941) kinetic theory. We then test our new model for iron under ramp loading conditions and compare our predictions for the α → ϵ iron phase transition to experimental data of Smith et al. (2013). In conclusion, more than one combination of material and model parameters (such as dislocation density and interface speed) led to good agreement of our simulations to the experimental data, thus highlighting the importance of having accurate microstructure data for the sample under consideration.

36 MATERIALS SCIENCE↗

labquake_future_prediction

The labquake_future_prediction code is a collection of python modules and scripts that serves as supporting information for the article “Predicting future laboratory fault friction through deep learning” for publication in the journal of “Geophysical Research Letters”. It is designed to predict laboratory fault slips in the immediate future by scanning continuous acoustic emission (AE) waveforms recorded in laboratory biaxial shear experiments. The predictions are made with a deep learning model based on convolutional encoder-decoder (CED) models and the Transformer model primarily developed for Natural Language Processing (NLP). The deep learning model is trained with the tensorflow package using publicly available laboratory data sets in standard binary file format in numpy. The utility functions for reading data files, configuring model hyperparameters, constructing the CED and Transformer models, training and testing of the models are defined in python module files. The workflow of training the models for labquake future predictions and the multiple GPU’s rapid model hyperparameter optimization as described in the journal article, are demonstrated in accompanying python script files and Jupyter notebooks.

Wang, Kun↗

Path-BigBird: An AI-Driven Transformer Approach to Classification of Cancer Pathology Reports

PURPOSE Surgical pathology reports are critical for cancer diagnosis and management. To accurately extract information about tumor characteristics from pathology reports in near real time, we explore the impact of using domain-specific transformer models that understand cancer pathology reports. METHODS We built a pathology transformer model, Path-BigBird, by using 2.7 million pathology reports from six SEER cancer registries. We then compare different variations of Path-BigBird with two less computationally intensive methods: Hierarchical Self-Attention Network (HiSAN) classification model and an offthe-shelf clinical transformer model (Clinical BigBird). We use five pathology information extraction tasks for evaluation: site, subsite, laterality, histology, and behavior. Model performance is evaluated by using macro and micro F 1 scores. RESULTS We found that Path-BigBird and Clinical BigBird outperformed the HiSAN in all tasks. Clinical BigBird performed better on the site and laterality tasks. Versions of the Path-BigBird model performed best on the two most difficult tasks: subsite (micro F 1 score of 72.53, macro F 1 score of 35.76) and histology (micro F 1 score of 80.96, macro F 1 score of 37.94). The largest performance gains over the HiSAN model were for histology, for which a Path-BigBird model increased the micro F 1 score by 1.44 points and the macro F 1 score by 3.55 points. Overall, the results suggest that a Path-BigBird model with a vocabulary derived from wellcurated and deidentified data is the best-performing model. CONCLUSION The Path-BigBird pathology transformer model improves automated information extraction from pathology reports. Although Path-BigBird outperforms Clinical BigBird and HiSAN, these less computationally expensive models still have utility when resources are constrained.

60 APPLIED LIFE SCIENCES↗