Search NASA⌕ Search

SEARCH · Search NASA

Results for “embedding model”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 181 records · Page 10

Insights from a coupled thermo-hydro-mechanical analysis of a layered high-temperature thermal energy storage reservoir

Coupled thermal-hydraulic-mechanical (THM) modeling is applied to investigate the performance of a seasonal high-temperature aquifer thermal energy storage operation based on data and conditions from current site investigations at the Geostorage Forsthaus pilot project in Bern (Switzerland). The model includes subhorizontal sand lenses of various lengths and dips that are embedded in a low permeability clay matrix. Thermal energy storage is simulated by seasonal injection and withdrawal of hot (up to 90 °C) water from a main well, with reservoir pressure regulated by two auxiliary wells at a distance of about 70 m from the main well. The results show how targeted injection into deeper permeable storage formations, along with active deep well pressure control, can effectively minimize geomechanical impact and the potential risk of damaging subsurface storage and sealing formations, or even surface facilities. With such pressure control, the subsurface mechanical responses are dominated by thermal strain and stress, which can be monitored with subsurface fiber optics. The study demonstrates how coupled THM modeling can be applied for the design of a safe and efficient thermal energy storage operation, and how subsurface fiber optic monitoring can be applied for performance confirmation, allowing for more confident operational forecasting.

Rutqvist, Jonny↗

PDF Entity Annotation Tool (PEAT)

While different text mining approaches – including the use of Artificial Intelligence (AI) and other machine based methods - continue to expand at a rapid pace, the tools used by researchers to create the labeled datasets required for training, modeling, and evaluation remain rudimentary. Labeled datasets contain the target attributes the machine is going to learn; for example, training an algorithm to delineate between images of a car or truck would generally require a set of images with a quantitative description of the underlying features of each vehicle type. Development of labeled textual data that can be used to build natural language machine learning models for scientific literature is not currently integrated into existing manual workflows used by domain experts. Published literature is rich with important information, such as different types of embedded text, plots, and tables that can all be used as inputs to train ML/natural language processing (NLP) models, when extracted and prepared in machine readable formats. Currently, both normalized data extraction of use to domain experts and extraction to support development of ML/NLP models are labor intensive and cumbersome manual processes. Automatic extraction of data and information from formats such as PDFs that are optimized for layout and human readability, not machine readability. The PDF (Portable Document Format) Entity Annotation Tool (PEAT) was developed with the goal of allowing users to annotate publications within their current print format, while also allowing those annotations to be captured in a machine-readable format. One of the main issues with traditional annotation tools is that they require transforming the PDF into plain text to facilitate the annotation process. While doing so lessens the technical challenges of annotating data, the user loses all structure and provenance that was inherent in the underlying PDF. Also, textual data extraction from PDFs can be an error prone process. Challenges include identifying sequential blocks of text and a multitude of document formats (multiple columns, font encodings, etc.). As a result of these challenges, using existing tools for development of NLP/ML models directly from PDFs is difficult because the generated outputs are not interoperable. We created a system that allows annotations to be completed on the original PDF document structure, with no plain text extraction. The result is an application that allows for easier and more accurate annotations. In addition, by including a feature that grants the user the ability to easily create a schema, we have developed a system that can be used to annotate text for different domain-centric schemas of relevance to subject matter experts. Different knowledge domains require distinct schemas and annotation tags to support machine learning.

97 MATHEMATICS AND COMPUTING↗

From minimum-viable-products to full models: a step-wise development of diagnostic forward models in support of design, analysis and modelling on the ST40 tokamak

Like most magnetic confined fusion experiments, the ST40 tokamak started off with a small subset of diagnostics and gradually increased the diagnostic set to include more complex and comprehensive systems. To make the most of each operational phase, forward models of various diagnostics are used and developed to aid design, provide consistency-checks during commissioning, test analysis methods, and build workflows to constrain high-level parameters to inform interpretation, theory and modelling. For new models and new analysis workflows, minimum-viable-products are released early, and their complexity is increased in a step-wise manner, facilitating the support of all programme phases on multiple parallel applications, while enabling learning opportunities and feedback loops. In this contribution we review the philosophy, scope and architecture of the framework under development. We discuss the details of some forward models, with examples on how they are used to aid diagnostic design, to investigate analysis methodologies through synthetic data, and how they are embedded in experimental analysis workflows. We compare previously published experimental results with new, more advanced analysis workflows employing more recent, detailed models and new diagnostic data, providing confirmation of the published material from the 2021–22 experimental campaign.

integrated data analysis↗

Transformer Neural Networks with Spatiotemporal Attention for Predictive Control and Optimization of Industrial Processes

In the context of real-time optimization and model predictive control of industrial systems, machine learning, and neural networks represent cutting-edge tools that hold promise for enhancing dynamic modeling. This work presents a novel transformer neural network architecture for real-time optimization and model predictive control. This network design includes a modified attention mechanism inspired by positional embedding attention from vision transformers and task-specific modifications to the input-output structure of the transformer’s decoder stack. Experiments were conducted using data from a 450 MW coal-fired power plant to evaluate this approach's effectiveness. The transformer neural network was compared with conventional recurrent models, including GRU and LSTM. The transformer exhibited a 6% increase in the R-squared (R2) value of predictions and an 83% reduction in mean squared error (MSE). Computation time was also reduced by 84% compared to conventional recurrent models.

Gallup, Ethan R.↗

Enhancing transfer learning in angle-resolved photoemission spectroscopy (ARPES) with spatially-aware representations via graph convolution

A recent application of machine learning has been to spatially-resolved angle-resolved photoemission spectroscopy (ARPES). Here we advance the state-of-the-art by applying representational learning to transform ARPES data into an embedding space of a pre-trained self-supervised learning model, thus enhancing the pipeline that improves the bandstructure classification and domain assignment/segmentation performance compared to a k-means clustering method. In the current iteration, the real-space information is entered into the domain assignment through the graph convolution method, which improves the transfer learning performance of the original self-supervised model. Lastly, an unsupervised automated tool is developed that incorporates these techniques to enable automatic domain assignment.

ARPES↗

Toward electron temperature profiles in hot-dense plasmas from x-ray spectral ensembles

High repetition rate laser systems enable new strategies for diagnosing plasma behavior with large datasets. Here, we define an ensemble technique that relies on randomized targeting of x-ray tracer micro-stripes. On each shot, a high-intensity laser pulse is focused on a solid target with Ti tracer stripes embedded in an Al foil, randomly targeting a micro-stripe, a portion of a stripe, or a gap between stripes. High-resolution, time-integrated x-ray spectrometers capture line emission from the portion of the micro-stripe that is heated to sufficiently high electron temperatures. Accumulation of many such cases is used to construct ensemble distributions of x-ray line intensities that encompass all relative offsets of the laser focus to the micro-stripe centers. Synthetic intensity distributions are likewise generated using collisional-radiative modeling. Bayesian fitting of modeled to measured intensity distributions establishes the most likely radial temperature profiles, enabling comparison to hydrodynamic models and calling into question the cylindrical symmetry of these micro-stripe-embedded systems. Ensemble techniques have significant potential for high-energy-density plasma diagnostics, especially with the advent of high repetition rate experiments.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

wa-hls4ml: A GNN Surrogate Model for hls4ml

Recent advancements in use of machine learning techniques on field-programmable gate arrays (FPGAs) have allowed for implementation of embedded neural networks with extremely low latency. This is invaluable for particle detectors at the Large Hadron Collider, where latency and used area must be strictly bounded. The hls4ml framework is a procedure for converting from trained machine learning model software, to a synthesis result that can be used on an FPGA. However, running the pipeline is a time-consuming procedure, and there is a strong risk of failure. In particular, it is possible that the model is unable to be converted into a synthesis result, or that the resource consumption of the model will exceed the resources of the target FPGA. To aid with this development, we introduce wa-hls4ml, a surrogate model which uses a graph neural network to emulate the structure of the source models. The goal is to estimate the chance of success and resource consumption of an arbitrary model when passed through the hls4ml procedure, without the time consumption of actually running the pipeline.

43 PARTICLE ACCELERATORS↗

Decode the Workload: Training Deep Learning Models for Efficient Compute Cluster Representation

Monitoring the status of a high throughput computing cluster running computationally intensive production jobs is a crucial yet challenging system administration task due to the complexity of such systems. To this end, we train autoencoders using the Linux kernel CPU metrics of the cluster. Additionally, we explore assisting these models with graph neural networks to share information across threads within a compute node. The models are compared in terms of their ability to: 1) Produce a compressed latent representation that captures the salient features of the input, 2) Detect anomalous activity, and 3) Make distinction between different kinds of jobs run at Jefferson Lab. The goal is to have a robust encoder whose compressed embeddings are used for several downstream tasks. We extend this study further by deploying these models in a human-in-the-loop production-based setting for the anomaly detection task and discuss the associated implementation aspects such as continual learning and the criterion to generate alarms. This study represents a first step in the endeavor towards building self-supervised large-scale foundation models for computing centers.

Mohammed, Ahmed↗

Pore Resolved Simulations of Joule Heating in Fibrous Media using an Embedded Boundary Method

Joule heating has been regarded as an energy-efficient and sustainable method for heating materials and gases at large scales. The modeling of local temperature effects at pore-resolved scales for such systems, however, has been difficult to achieve due to challenges in coupling thermo-chemical processes in complex porous media and in large representative volume elements (RVEs). To this end, we developed an electro-thermal model at the pore scale to study Joule heating effects in large heterogeneous systems with different microstructures. This was achieved using the level set method to implicitly delineate distinct regions within the domain, and an embedded boundary method to facilitate heat exchange across the fluid-solid interface. Moreover, we applied this method to investigate unsteady non-linear electro-thermal effects in non-woven fibrous graphite conductors for RVEs with characteristic lengths of 2 mm, with different fiber orientations, porosity (80% – 90%) and fiber diameters (10 – 20µm). The coupled equations were solved numerically and they produced peak temperatures greater than 2000 K resulting in heating rates as high as 80,000 K/s. Moreover, the results depended strongly on the microstructure of the fiber skeleton and current density. Geometries with large fibers (∼ 20µm) had the highest average and peak temperatures with the mean temperature increasing by 3.9 % while the peak temperature increased by 9.9 %. Anisotropic domains on the other hand had the lowest mean and peak temperatures with peak and mean temperatures of 2293 K and 1437.7K respectively representing a corresponding 12.1% and 5.1% drop in the temperatures. An increase in porosity from 80% to 90%, however, led to an increase in the peak temperature by 5.1%.

Joule heating↗

An Instrumented Capsule Design to Measure Thermal Conductivity in Miniature UO2 Specimens

Numerous separate effects irradiations of miniature nuclear fuel specimens have been conducted in the High Flux Isotope Reactor (HFIR) under the experimental platform designated as MiniFuel. MiniFuel is a static irradiation capability in which microstructural evolution and fuel performance phenomena are observed during postirradiation examination thereby offering a snapshot of the terminal fuel characteristics. This approach inherently requires fielding an irradiation where the experimental conditions are determined using predictive models and the pertinent outcomes are measured at the end of the test. Static irradiations can provide useful insights to the relationships between fuel performance and the pivotal irradiation conditions, namely temperature and burnup, but the ability to monitor fuel performance in situ would further support fuel development and qualification. To this end, an instrumented experiment design is being developed at Oak Ridge National Laboratory to capture thermal conductivity degradation and fission gas release during HFIR irradiation. These phenomena will be monitored using unique capsule designs that each target a different phenomenon. This paper details the thermal conductivity capsule (TCC) design and its expected performance envelope as determined using computer models. Each TCC will contain a miniature UO2 disc specimen (~0.5 mm thick × 5 mm diameter) sandwiched between metallic slugs with embedded thermocouples. The coupling of in situ temperature measurements, known thermal conductivity of the metallic components, and heat generation rates computed using high-fidelity neutronics models make the thermal conductivity measurement possible. This paper describes the reactor physics and heat transfer models used to predict the capsule’s performance and the methodology for calculating the fuel specimen’s thermal conductivity from the thermocouple measurements.

Gorton, Jacob [ORNL] (ORCID:0000000269806083)↗

STILGAR: Subsurface Models for Graymont Pleasant Gap Mine

The detection, location, and monitoring of underground structures are of great importance to national and global security. Tunnels and voids generate seismic signatures detectable at the surface, but using non-invasive seismic data to image near-surface presents several challenges in real-world applications. In this report, we describe the use of a dense surface seismic deployment to generate subsurface models of the Graymont Pleasant Gap mine - a single-layer mine with a complex structure embedded in a high-velocity P-wave limestone bedrock. Our approach consists of three key methods. We use P-wave arrival times from local blast events to perform a tomography inversion with the tomoTD method, constructing a P-wave velocity model of the subsurface. We model the layer above the mine using Rayleigh wave ellipticity and inversion techniques. We leverage ongoing anthropogenic activities to identify and locate noise sources both on the surface and within the subsurface. With this integrated approach we aim to overcome the challenges and enhance our ability to non-invasively characterize underground structures, contributing to improved seismic monitoring techniques.

58 GEOSCIENCES↗

A Graph Neural Network Surrogate Model for hls4ml

Recent advancements in use of machine learning (ML) techniques on field-programmable gate arrays (FPGAs) have allowed for the implementation of embedded neural networks with extremely low latency. This is invaluable for particle detectors at the Large Hadron Collider, where latency and used area are strictly bounded. The hls4ml framework is a procedure that converts trained ML model software to a synthesis result to can be used on an FPGA. However, running the pipeline is a time-consuming procedure, and there is a strong risk of failure. In particular, it may not be possible to successfully convert a model into a synthesis result, or the resource consumption of the model may exceed the resources of the target FPGA. To aid with this development, we introduce wa-hls4ml, a surrogate model using a graph neural network to emulate the structure of the source models. The goal is to estimate the chance of success and resource consumption of a given model when passed through the hls4ml pipeline, without needing to run the pipeline.

Plotnikov, Dennis↗

Predicting interface structure using the minima hopping method

Here, we adapt the minima hopping method (MHM) to the problem of interfacial structure prediction and apply it to study a canonical problem, the tilt grain boundaries in SrTiO 3 . Our method employs a hybrid approach by first exploring the potential energy surface (PES) of different grain boundary samplings with an empirical force field, among which the fifteen candidates with lower energies are then refined using ab initio density functional theory (DFT) calculations. During the exploratory stage, we bias the search using a local order parameter to primarily sample various reconstructions in the vicinity of the interface, while preserving the crystallinity of the bulk regions. We further enhance the search by incorporating initial structures with rigid body displacements to account for translational variations between bulk phases, enabling the MHM to effectively generate both stoichiometric and nonstoichiometric SrTiO 3 Σ⁢3(111)[110] and Σ⁢3(112)[110] grain boundaries. From an algorithmic standpoint, MHM outperforms earlier studies based on genetic algorithms (GA) by identifying more stable interfacial structures of several SrTiO 3 grain boundaries. The performance of the present implementation of the MHM approach is primarily limited by exploring an approximate description of the PES with a rather simple Buckingham potential. This limitation leads to variations in performance when compared to approaches utilizing more advanced surrogate PES models, such as direct DFT-PES sampling or GA with the embedded atom method (EAM). Despite the present limitations, the MHM approach is able to yield interfacial structures with comparable or lower interfacial energies in specific cases, such as Σ⁢3(111)[110] Γ=1, ±0.5 and Σ⁢3(112)[110] Γ= ±1, −2, underscoring the robustness of the MHM approach even with a simple approximation of the DFT PES. The MHM interfacial structure prediction method thus offers an efficient approach to understanding the grain boundaries and heterointerfaces at the atomic scale, providing an important prerequisite for effective materials design.

density functional theory↗

Bayesian Physics Informed Spatio-Temporal Network for Streamflow Data Imputation

Reliable reconstruction of incomplete streamflow records is critical for improving hydrological forecasting, flood preparedness, and water resource management. However, large observational gaps and uncertainties in governing physical parameters limit the accuracy of traditional statistical and machinelearning imputation frameworks. To address these challenges, we develop a Bayesian Physics-Informed Spatio-Temporal Network (BPI-STNet) that jointly captures spatial and temporal dependencies while enforcing hydrologic consistency through embedded physical constraints. The framework integrates a GraphSAGE-LSTM architecture to model spatial connectivity across gauges and temporal flow dynamics, coupled with a Bayesian update mechanism to estimate uncertain parameters in a simplified water-balance framework. Unlike conventional physics-informed networks that rely on sampling-based posterior estimation, BPI-STNet derives an analytic solution to the inverse problem, allowing closed-form Bayesian updates of uncertain parameters Λ={α,β,k} using Gaussian priors and likelihoods. Applied to daily observations from the Susquehanna River Basin (1980-2022), BPI-STNet achieves substantial improvements over a purely data-driven RGNN baseline, which reduced RMSE by 23 % and MAE by 9 %, and achieving an average NSE values up to 0.96. The results demonstrate that coupling Bayesian inference with physics-informed learning yields physically consistent, uncertainty-aware reconstructions that preserve the temporal persistence and statistical distribution of observed flows. The proposed framework establishes a generalizable paradigm for data-sparse hydrologic systems where both data fidelity and physical interpretability are essential.

Krishnan Kutty Ambika, Anukesh [ORNL] (ORCID:00000↗

Super Resolving Unrolled Neural Networks for Remote Sensing

In remote sensing systems, the capabilities of the system are constrained by the complex interactions between size, weight, and power (SWAP) of potential designs. In electro-optical (EO) systems, examples of these critical parameters include the system’s sensitivity and resolution. Those parameters can be increased by ever larger optical apertures and focal planes but at the cost of more SWAP. Multi-image super resolution (MISR) techniques allow resolution to be enhanced via computation rather than more sophisticated optical hardware. These algorithms combine multiple images together into a single, higher resolution image, trading temporal resolution and computation for spatial resolution. Fielded MISR techniques, such as Drizzle, can require several hundred images to create a single super resolved image, implying reduced temporal resolution, increased data acquisition load, and limiting mission applications. Iterative techniques, such as model-based image reconstruction and compressive sensing, have been shown to create super resolved images using fewer images than Drizzle. They do this by posing an optimization problem that balances accuracy between a highly accurate physical model and an image model. In the case of super resolution, the physical model is defined by the relation between low resolution input images and the desired high resolution output image. The image model encodes some assumptions about the super resolved image. These assumptions are meant to suppress reconstruction artifacts that arise due to deterministic physical model error, stochastic measurement noise, and potential undersampling. In practice, the performance of iterative methods are limited by imaging models compatible with optimization. Deep learning-based methods can effectively learn image models of arbitrary complexity, but lack the theoretical explainability and robustness of iterative techniques. Consensus equilibrium (CE) generalizes the iterative techniques beyond optimization, enabling blackbox algorithms such as traditional and neural image denoisers to be used as the image model. CE-based approaches retain much of the explainability and robustness of iterative techniques while allowing the expressiveness of machine learning image models to be used. Additionally, by unrolling iterations of CE with an embedded image denoiser, the image denoiser can be further trained and specialized to the specific application with potentially higher quality reconstructions. Under this project, we demonstrated the feasibility of training an unrolled neural network based upon CE. While we didn’t train one, we showed that the CE process is differentiable and its gradient can be tractably computed. We also explored the usage of a variants of CE akin to generative neural works. Most importantly, we applied the CE framework to a number of problems including non-blind deconvolution, upsampling, single-image super resolution, MISR, event-based sensing, and saturated deconvolution. Our MISR prototype creates high quality reconstructions with an order of magnitude fewer images than previous approaches and, critically, produces these reconstructions fast enough for practical usage.

47 OTHER INSTRUMENTATION↗

Evaluating Sea Breezes and Associated Convective Cloud Evolution in the Model Gray Zone

We characterize convective clouds associated with sea‐breeze circulations (SBC) using multi‐agency observations and multi‐case ensemble model simulations. The focus is on assessing convective cloud lifecycle properties and their merging behavior, as well as the environmental conditions they are embedded in, particularly SBC features. In total, 46 SBC days over the Houston‐Galveston region are selected and simulated using the Weather Research and Forecasting (WRF) model at a gray zone scale with a forecast‐like parameterization setup. Advanced techniques, including change‐point detection, a Lagrangian cloud tracking method, and a newly developed cell merging and splitting detection algorithm, are applied and/or developed for this study. Our findings indicate that the WRF model at 1 km grid spacing well represents the thermodynamic conditions over the region, as well as SBC timing and intensity. However, for the associated convective cells, WRF overestimates the 30‐dBZ echo top height, cell area, and maximum radar reflectivity compared to radar observations. This overestimation is potentially due to under‐resolved entrainment processes, an overestimated merging frequency, and the overestimation of updraft intensity. Furthermore, the model exhibits a deficiency in simulating congestus clouds, showing a more rapid transition from shallow to deep convection compared to observed behavior. Moreover, observations indicate stronger, deeper, and wider clouds when merging happens. Conversely, in simulations, the merging process does not necessarily lead to higher or longer‐lived cells, as many cases experience rapid and frequent merging and splitting which may result in more variance in convective updraft velocity during the convection lifetime.

54 ENVIRONMENTAL SCIENCES↗

Improved representation of black carbon mixing structures suggests stronger direct radiative heating

Black carbon significantly influences the Earth system because of its strong solar radiation absorption. However, its direct radiative effect remains poorly understood in current climate models, partly because current climate models oversimplify the diverse structures formed when black carbon mixes with other atmospheric components. Here we show that incorporating more realistic, multi-mixing-structure representations of black carbon increases the direct radiative effect. We find that aged black carbon particles, with thicker coatings and higher embedded fractions, enhance the direct radiative effect more efficiently. Using machine learning alongside the Community Earth System Model, we show that the direct radiative effect at the top of the atmosphere in regions with heavy black carbon pollution is 31.6% greater when multi-mixing structures are considered. These findings highlight the importance of modeling complex mixing structures of particle-resolved black carbon to accurately capture their warming impacts on global atmosphere, particularly in highly polluted regions.

54 ENVIRONMENTAL SCIENCES↗