Search NASA⌕ Search

SEARCH · Search NASA

Results for “training data”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 397 records · Page 22

Cloverleaf Data Artifacts for ArtIMis LDRD

This report summarizes the use of the open-source CloverLeaf/CloverLeaf3D mini-apps to generate synthetic data sets to train foundation models for the ArtIMis LDRD DI. These data artifacts are intended to be used by LANL collaborators and shared externally with our university and institutional partners. Note that CloverLeaf/CloverLeaf3D is not a LANL simulation code.

97 MATHEMATICS AND COMPUTING↗

Unsupervised domain adaptation for radioisotope identification in gamma spectroscopy

Training machine learning models for radioisotope identification using gamma spectroscopy remains an elusive challenge for many practical applications, largely stemming from the difficulty of acquiring and labeling large, diverse experimental datasets. Simulations can mitigate this challenge, but the accuracy of models trained on simulated data can deteriorate substantially when deployed to an out-of-distribution operational environment. In this study, we demonstrate that unsupervised domain adaptation (UDA) can improve the ability of a model trained on synthetic data to generalize to a new testing domain, provided unlabeled data from the target domain are available. Conventional supervised techniques are unable to utilize this data because the absence of isotope labels precludes defining a supervised classification loss. Instead, we first pretrain a spectral classifier using labeled synthetic data and subsequently leverage unlabeled target data to align the learned feature representations between the source and target domains. We compare a range of different UDA techniques, finding that minimizing the maximum mean discrepancy (MMD) between source and target feature vectors yields the most consistent improvement to testing scores. For instance, using a custom transformer-based neural network, we achieved a testing accuracy of $0.904 \pm 0.022$ on an experimental LaBr test set after performing unsupervised feature alignment via MMD minimization, compared to $0.754 \pm 0.014$ before alignment. Overall, our results highlight the potential of using UDA to adapt a radioisotope classifier trained on synthetic data for real-world deployment.

Lalor, Peter W.↗

SAXS Assistant: Automated SAXS analysis for structural discovery in biologics and polymeric nanoparticles

Small-angle x-ray scattering (SAXS) is a powerful technique for assessing macromolecular structure. High-throughput SAXS is limited by the time-consuming and, at times, subjective nature of SAXS data interpretation. Here, we present SAXS Assistant, a Python-based script that streamlines SAXS data analysis to extract features for machine learning (ML) and key structural parameters, including the Guinier radius of gyration (R g ), pair distance distribution function (PDDF)-derived R g , maximum particle dimension (D max ), and Kratky plots. The script builds upon BioXTAS RAW and validates reliability via Guinier/PDDF R g agreement, an important indicator of well-measured data sets. For assistance in D max estimation, a multilayer perceptron regressor was trained with 1940 data files from the Small Angle Scattering Biological Data Bank. The model achieved a test set performance R 2 = 0.90 and mean absolute error = 11.7 Å. Training exclusively with experimental data translates analyses from researchers, including experts in the field, to the ML model, which helps assess D max estimations from PDDF. Gaussian mixture model clustering was implemented to classify profiles into structural classes based on entries in the Small Angle Scattering Biological Data Bank. Users may therefore assess the similarity between experimental samples and known biomolecular shapes within the mapped repository entries. This probabilistic clustering aids in quantifying information from Kratky and generating shape-descriptive features. SAXS Assistant accelerates SAXS data analysis through enforced quality control, ML-ready outputs, and flags for low-confidence results. In addition to providing the ability to analyze large data sets at high throughput, this tool is versatile and may serve researchers in both biological and synthetic polymer research fields.

36 MATERIALS SCIENCE↗

Multi-task Parallelism for Robust Pre-training of Graph Foundation Models on Multi-source, Multi-fidelity Atomistic Modeling Data

Graph foundation models using graph neural networks promise sustainable, efficient atomistic modeling. To tackle challenges of processing multi-source, multi-fidelity data during pre-training, recent studies employ multi-task learning, in which shared message passing layers initially process input atomistic structures regardless of source, then route them to multiple decoding heads that predict data-specific outputs. This approach stabilizes pre-training and enhances a model’s transferability to unexplored chemical regions. Preliminary results on approximately four million structures are encouraging, yet questions remain about generalizability to larger, more diverse datasets and scalability on supercomputers. We propose a multi-task parallelism method that distributes each head across computing resources with GPU acceleration. Implemented in the open-source HydraGNN architecture, our method was trained on over 24 million structures from five datasets and tested on the Perlmutter, Aurora, and Frontier supercomputers, demonstrating efficient scaling on all three highly heterogeneous super-computing architectures.

Lupo Pasini, Massimiliano [ORNL] (ORCID:0000000249↗

Deep Learning for Full Waveform Inversion of Elastic Active-Source Seismic Data to Estimate P-Wave Velocity Models

Seismic imaging methods are critical for Global Security and Energy & Homeland Security missions and activities that rely on subsurface characterization, but traditional methods remain computationally expensive and require significant labor hours and expertise to execute. Within the past few years, machine learning (ML), namely deep learning (DL), has been used to develop data-driven end-to-end full waveform inversion (FWI) methods to estimate 2D P-wave velocity (Vp) models in a fraction of the time as conventional FWI. These methods, however, are trained on simplistic acoustic wave seismic data and Vp models that are not realistic nor representative of real-world observations, leaving a large gap between the state-of-the-art and deployable, feasible, and practical DL FWI methods. Here, we generate a synthetic active-source, 3D, elastic wave seismic data set and a variety of Vp models with realistic geologic structure for training DL FWI methods. We evaluate six different methods that have performed well for acoustic DL FWI or medical imaging tasks using our more realistic dataset. We find that these six trained models do not match the performance of published acoustic end-to-end DL FWI methods, indicating more training data may be needed, physics may need to be incorporated to achieve good accuracy at the sacrifice of the end-to-end advantage, and/or novel methods need to be developed to enable end-to-end DL FWI methods to perform well for real-world seismic data.

45 MILITARY TECHNOLOGY, WEAPONRY, AND NATIONAL DEF↗

Evaluation of Seismic Artificial Intelligence with Uncertainty

Artificial intelligence has transformed the seismic community with deep learning models (DLMs) that are trained to complete specific tasks within workflows. However, there is still a lack of robust evaluation frameworks for evaluating and comparing DLMs. Here, we address this gap by designing an evaluation framework that jointly incorporates two crucial aspects: performance uncertainty and learning efficiency. To target these aspects, we meticulously construct the training, validation, and test splits using a clustering method tailored to seismic data and enact an expansive training design to segregate performance uncertainty arising from stochastic training processes and random data sampling. The framework’s ability to guard against misleading declarations of model superiority is demonstrated through the evaluation of PhaseNet (Zhu and Beroza, 2018), a popular seismic phase picking DLM, under three training approaches. Our framework helps practitioners choose the best model for their problem and set performance expectations by explicitly analyzing model performance with uncertainty at varying budgets of training data.

58 GEOSCIENCES↗

Comparative Study of Large Language Model Architectures on Frontier

Large language models (LLMs) have garnered significant attention in both the AI community and beyond. Among these, the Generative Pre-trained Transformer (GPT) has emerged as the dominant architecture, spawning numerous variants. However, these variants have undergone pre-training under diverse conditions, including variations in input data, data preprocessing, and training methodologies, resulting in a lack of controlled comparative studies. Here we meticulously examine two prominent open-sourced GPT architectures, GPT-NeoX and LLaMA, leveraging the computational power of Frontier, the world’s first Exascale supercomputer. Employing the same materials science text corpus and a comprehensive end-to-end pipeline, we conduct a comparative analysis of their training and downstream performance. Our efforts culminate in achieving state-of-the-art performance on a challenging materials science benchmark. Furthermore, we investigate the computation and energy efficiency, and propose a computationally efficient method for architecture design. To our knowledge, these pre-trained models represent the largest available for materials science. Our findings provide practical guidance for building LLMs on HPC platforms.

Yin, Junqi↗

Driver Identification Midyear Report

First, we create a profile for each authorized driver based on their existing driving data. We then train a machine learning model on the driving data from this profile, yielding an individualized model for each driver. Finally during a drive, we pass the Controller Area Network (CAN) data to the model and authen ticate the driver’s identity in real-time. This verification or lack thereof could be used to alert supervisors of threats to their drivers or transported materials. Deviations from their normal driving behavior could indicate high-risk situations, medical events, or even insider threats.

98 NUCLEAR DISARMAMENT, SAFEGUARDS, AND PHYSICAL P↗

Sequential Inference of Hospitalization Electronic Health Records Using Probabilistic Models

In the dynamic hospital setting, decision support can be a valuable tool for improving patient outcomes. Data-driven inference of future outcomes is challenging in this dynamic setting, where long sequences such as laboratory tests and medications are updated frequently. This is due in part to heterogeneity of data types and mixed-sequence types contained in variable length sequences. In this work we design a probabilistic unsupervised model for multiple arbitrary-length sequences contained in hospitalization Electronic Health Record (EHR) data. The model uses a latent variable structure and captures complex relationships between medications, diagnoses, laboratory tests, neurological assessments, and medications. It can be trained on original data, without requiring any lossy transformations or time binning. Inference algorithms are derived that use partial data to infer properties of the complete sequences, including their length and presence of specific values. We train this model on data from subjects receiving medical care in the Kaiser Permanente Northern California integrated healthcare delivery system. The results are evaluated against held-out data for predicting the length of sequences and presence of Intensive Care Unit (ICU) in hospitalization bed sequences. Our method outperforms a baseline approach, showing that in these experiments the trained model captures information in the sequences that is informative of their future values.

97 MATHEMATICS AND COMPUTING↗

An efficient hybrid downscaling framework to estimate high-resolution river hydrodynamics

Flow depth and velocity are the most important hydrodynamic variables that govern various river functions, including water resources, navigation, sediment transport, and biogeochemical cycling. Existing high-resolution flow depth simulations rely on either computationally expensive river hydrodynamic models (RHMs) or data-driven models with formidable training costs, whereas data-driven modeling of flow velocity has rarely been explored. Here, using the hybrid Low-fidelity, Spatial analysis, and Gaussian process learning (LSG) model, we developed a downscaling approach to construct high-resolution flow depth and velocity from a two-dimensional (2-D) RHM simulation at coarse resolution. The LSG models were trained and tested in an urban watershed in Houston using two different hurricane-driven flood events. The high-resolution (as fine as 30 m resolution) and low-resolution (mostly 1000 m resolution) meshes include 664 724 and 14 536 grid cells, respectively. The results showed that through downscaling, the simulation errors were reduced to less than one-fourth and one-third of the errors of the low-resolution 2-D RHM for flow depth and velocity, respectively. Our analysis further revealed that the dominant uncertainty sources of the downscaled hydrodynamics are different, with flow velocity dominated by the dimensionality reduction error, which we reduced by using a regionalized training procedure. The downscaling approach achieves an 84-fold acceleration in computational time compared to the high-resolution 2-D RHM, making high-fidelity ensemble flood modeling feasible. More importantly, the developed method provides an opportunity to couple large-scale hydrodynamical processes with local physical, chemical, and biological processes in river models.

Tan, Zeli [Pacific Northwest National Laboratory (↗

TrojAI Alternate Analysis

In this portion of the TrojAI evaluation, we focus on the cyber-network-c2-mar2024 dataset. Recall that in this round ResNet18 and ResNet34 neural networks (NN) were trained on the USTC-TFC2016 dataset with the aim of distinguishing between benign versus botnet command and control (c2) packets. A range of bytes from each packet was reformatted into a 28x28 pixel image, and the collection of reformatted packets served as the training (and testing) data for the two ResNet models. For some of the data a trigger watermark was strategically placed to affect various inputs to the NNs. This watermarked, or poisoned, data in turn created a poisoned, or trojaned NN. The data were poisoned in different ways ultimately creating different trojaned NNs. This collection of trojaned NNs was combined with various versions of not trojaned NNs and served as the training and testing data for the performers. The performers’ task was to construct a classifier to distinguish between the trojaned and not trojaned models. It was previously noted that the performers struggled with the cyber-network-c2-mar2024 dataset, motivating this investigation of potential reasons the performers experienced challenges.

97 MATHEMATICS AND COMPUTING↗

Monitoring of Liquid Metal Reactor Heater Zones with Recurrent Neural Network Learning of Temperature Time Series

Advanced high-temperature fluid reactors (ARs), such as sodium fast reactors (SFRs) and molten salt cooled reactors (MSCRs) utilize high-temperature fluids at ambient pressure. To melt the fluid during reactor startup and prevent fluid freezing during cooldown, the thermal–hydraulic systems of such ARs include heater zones consisting of specific heaters with controllers, temperature sensors, and thermal insulation. The failure of heater zones due to insulation material degradation or improper installation, resulting in parasitic heat losses, can lead to fluid freezing. The detection of faults using a heat-transfer model is difficult because of a lack of knowledge of the experimental details. Data-driven machine learning of heater zone temperature time series offers a viable alternative. In this study, we benchmarked the performance of recurrent neural networks (RNNs) in an analysis of heat-up transient temperature time series of heater zones installed on a liquid sodium vessel. The RNN models include long short-term memory (LSTM) and gated recurrent unit (GRU) networks, as well as their bi-directional variants, BiLSTM and BiGRU. Anomalous temperature points were designated using a percentile-based threshold applied to residual fluctuations in the detrended temperature time series. Additionally, the impact of the exponentially weighted moving average (EWMA) method on detection accuracy was examined. The RNN models’ performance was assessed using precision, recall, and F 1 score metrics. Results demonstrated that RNN models effectively detect anomalies in temperature time series with the best models for each heater zone achieving F 1 scores of over 93%. To explain the variations in RNN model performance across different heater zones, we used Kullback–Leibler (KL) divergence to quantify the relative entropy between training and testing data, and the Detrended Fluctuation Analysis (DFA) to assess long-range temporal correlations. For datasets with strong long-range correlations and minimal relative entropy between training and testing data, GRU is the best-performing model. When the data exhibits weaker long-term correlations and a significant relative entropy between training and testing distributions, BiGRU shows the best performance. For the data sets with intermediate values of both KL divergence and DFA, the best performance is obtained with LSTM and BiLSTM, respectively.

gated recurrent unit↗

High temperature melting of dense molecular hydrogen from machine-learning interatomic potentials trained on quantum Monte Carlo

We present results and discuss methods for computing the melting temperature of dense molecular hydrogen using a machine learned model trained on quantum Monte Carlo data. In this newly trained model, we emphasize the importance of accurate total energies in the training. We integrate a two phase method for estimating the melting temperature with estimates from the Clausius–Clapeyron relation to provide a more accurate melting curve from the model. We make detailed predictions of the melting temperature, solid and liquid volumes, latent heat, and internal energy from 50 to 180 GPa for both classical hydrogen and quantum hydrogen. At pressures of roughly 173 GPa and 1635 K, we observe molecular dissociation in the liquid phase. Here, we compare with previous simulations and experimental measurements.

08 HYDROGEN↗

Simplified inelastic constitutive models for ASME Section III, Division 5 design by inelastic analysis

This report describes the development of simplified, universal constitutive model that captures the high temperature monotonic and cyclic behavior of a range of commonly-used high temperature materials. The goal of the work is to provide a simple, universal constitutive model to replace the current bespoke models for Grade 91, 316H, and Alloy 617 included in Nonmandatory Appendix HBB-Z of the ASME Boiler & Pressure Vessel Code, and to extend this model to cover Alloy 800H. We initiated this work in response to feedback from reactor vendors and other Code users requesting simplified models, compared to the current models, that are easier to implement and use in commercial finite element analysis software. This report describes the completion of this effort by developing a model to correct the defects in standard model forms presently used for high temperature material modeling, described in past work, developing and implementing new numerical methods to train this model against test data, and then actually training the model for the four materials. The report provides a complete mathematical description of the model along with the tabulated material coefficients for the four materials. The final step will be to formulate an ASME Code change to introduce the new models into the Code.

36 MATERIALS SCIENCE↗

Normalizing flows for domain adaptation when identifying Λ hyperon events

Here this study focuses on the application of a normalizing flow as a method of domain adaptation when classifying physics data. Normalizing flows offer a way to transform data points between two different distributions. The present study investigates a novel method of transforming latent representations of physics data to a normal distribution and then to a physics distribution again. The final distribution models a simulated distribution. After being transformed, the data can be classified by a neural network trained on labeled simulation data. The present study succeeds in training two normalizing flows that can transform between data (or simulation) and a Gaussian distribution.

47 OTHER INSTRUMENTATION↗

Multi-fidelity learning for interatomic potentials: low-level forces and high-level energies are all you need

The promise of machine learning interatomic potentials (MLIPs) has led to an abundance of public quantum mechanical (QM) training datasets. The quality of an MLIP is directly limited by the accuracy of the energies and atomic forces in the training dataset. Unfortunately, most of these datasets are computed with relatively low-accuracy QM methods, e.g. density functional theory with a moderate basis set. Due to the increased computational cost of more accurate QM methods, e.g. coupled-cluster theory with a complete basis set (CBS) extrapolation, most high-accuracy datasets are much smaller and often do not contain atomic forces. The lack of high-accuracy atomic forces is quite troubling, as training with force data greatly improves the stability and quality of the MLIP compared to training to energy alone. Because most datasets are computed with a unique level of theory, traditional single-fidelity (SF) learning is not capable of leveraging the vast amounts of published QM data. In this study, we apply multi-fidelity learning (MFL) to train an MLIP to multiple QM datasets of different levels of accuracy, i.e. levels of fidelity. Specifically, we perform three test cases to demonstrate that MFL with both low-level forces and high-level energies yields an extremely accurate MLIP—far more accurate than a SF MLIP trained solely to high-level energies and almost as accurate as a SF MLIP trained directly to high-level energies and forces. Therefore, MFL greatly alleviates the need for generating large and expensive datasets containing high-accuracy atomic forces and allows for more effective training to existing high-accuracy energy-only datasets. Indeed, low-accuracy atomic forces and high-accuracy energies are all that are needed to achieve a high-accuracy MLIP with MFL.

36 MATERIALS SCIENCE↗

Describing Point Defect Topology in 2D Energy Materials Through Computer Vision

Point defects such as vacancies and impurity atoms strongly impact the performance of 2D materials. Traditional efforts often rely on manual detection, a process that is time-intensive, prone to human error, and challenging to scale. Here we leverage machine learning (ML) methods to identify and quantify vacancies within 2D transition metal carbides (Ti3C2, MXenes), aiming to expedite detection while improving accuracy. MXenes exhibit valuable defect-defined electrochemical properties, but we currently lack statistical understanding of defect topology needed to fully harness these materials. Here we employ a convolutional neural network for semantic segmentation of experimental MXene images, opening an opportunity to conduct a rigorous statistical study on defect hierarchy while investigating local relaxation in the lattice. We show how the integration of ML can yield fundamental insight into point defects, providing a powerful tool that will play an increasingly crucial role in the future of materials science. ML is often not just a matter of straightforward application, and pretrained models proved ineffective in this case. Instead, we trained our own neural network (NN) and applied data augmentation techniques and fine-tuning to the training dataset. Since labeled microscopy data is often scarce, we developed training data from a previously published wide-frame MXene image, using customized Gaussian fitting to locate atomic positions. Our trained model was then applied to a large dataset of experimental images, enabling a statistical study of defect configurations across three samples prepared with different HF etchant concentrations (5%, 9.1%, and 12.5%), as shown in Fig. 1. This also allowed us to investigate local strain around vacancies, though we find that we are limited by the precision of measurements using high-angle annular dark field (HAADF) images, as shown in Fig. 2. This study demonstrates how ML enables large-scale, quantitative analysis of atomic defects - an otherwise infeasible task with traditional methods. While our NN was specialized for Ti3C2 MXenes, the pipeline we developed provides a foundation for future ML models tailored to other materials. Ultimately, we envision embedding the NN onto the microscope to give real-time feedback to the user. To make this a reality, continued work is necessary to fully understand the NN's capabilities and limitations. This study gets one step closer to our goals of automated experimentation moving away from traditional methods of manual labeling. As ML capabilities advance, we hope to continue adapting and applying these techniques in microscopy.

2D materials↗