Search NASA⌕ Search

SEARCH · Search NASA

Results for “Learning with errors”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 343 records · Page 19

Virtual Growth of SRF Materials

Niobium's native surface oxide affects SRF cavity and superconducting qubit performance, motivating interest in controlling its crystalline structure. We combine a literature-derived machine-learning analysis with temperature-dependent XRD to study crystalline ordering in Nb2O5. Random Forest models, trained on 74 processing conditions from 17 papers and validated by leave-one-group-out cross-validation, predicted broad crystallinity outcomes well (balanced accuracy 0.809), but struggled with specific polymorph identity (0.577). Annealing temperature was the dominant predictor across all targets; oxygen partial pressure showed negligible importance, reflecting narrow literature coverage rather than physical irrelevance. Temperature-dependent XRD on anodized and H2O2-treated Niobium showed structural evolution consistent with the machine learning predictions. Our model and overall approach provide a data-driven framework for identifying and optimizing conditions that promote crystallization in initially amorphous oxides. This framework can guide the selection of growth and post-annealing conditions for Nb surfaces by narrowing the experimental parameter space, thereby reducing trial-and-error efforts in developing oxide structures relevant to SRF applications.

Tilkin, Anthony [Fermilab]↗

Virtual Growth of SRF Materials: A Machine Learning Approach to Predict the Crystalline Structural Ordering in Nb Surface Oxides

Niobium's native surface oxide affects SRF cavity and superconducting qubit performance, motivating interest in controlling its crystalline structure. We combine a literature-derived machine-learning analysis with temperature-dependent XRD to study crystalline ordering in Nb2O5. Random Forest models, trained on 74 processing conditions from 17 papers and validated by leave-one-group-out cross-validation, predicted broad crystallinity outcomes well (balanced accuracy 0.809), but struggled with specific polymorph identity (0.577). Annealing temperature was the dominant predictor across all targets; oxygen partial pressure showed negligible importance, reflecting narrow literature coverage rather than physical irrelevance. Temperature-dependent XRD on anodized and H2O2-treated Niobium showed structural evolution consistent with the machine learning predictions. Our model and overall approach provide a data-driven framework for identifying and optimizing conditions that promote crystallization in initially amorphous oxides. This framework can guide the selection of growth and post-annealing conditions for Nb surfaces by narrowing the experimental parameter space, thereby reducing trial-and-error efforts in developing oxide structures relevant to SRF applications.

Tilkin, Anthony [Unlisted, US, IL; Fermilab]↗

Resolving the Coverage Dependence of Surface Reaction Kinetics with Machine Learning and Automated Quantum Chemistry Workflows

Microkinetic models for catalytic systems require estimation of many thermodynamic and kinetic parameters that can be calculated for isolated species and transition states using ab initio methods. However, the presence of nearby coadsorbates on the surface can dramatically alter these thermodynamic and kinetic parameters causing them to be dependent on species coverage fractions. As there are combinatorially many coadsorbed configurations on the surface, computing the coverage dependence of these parameters is far less straightforward. We present a framework for generating and applying machine learning models to predict coverage-dependent parameters for microkinetic models. Our toolkit enables automatic calculation and evaluation of coadsorbed configurations allowing us to sample 2,000 coadsorbed adsorbates and transition states (TSs) for a diverse set of 9 reactions on Cu(111), a challenging surface, with four possible coadsorbates. This dataset was then used to train subgraph isomorphic decision trees (SIDTs) to predict the stability and association energy of configurations. We were able to achieve mean absolute errors (MAEs) of 0.106 eV on adsorbates, 0.172 eV on TSs, and due to natural error cancellation in SIDTs for relative properties, 0.130 eV on reaction energies and 0.180 eV on activation barriers. In conclusion, we describe how to use these models to predict coverage-dependent corrections for adsorbates and TSs and demonstrate on H*, HO*, and O* comparing the generated SIDT model with an iteratively refined version.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Spectral Data Fusion From Handheld Laser-Induced Breakdown Spectroscopy (LIBS) and X-ray Fluorescence (XRF) Analyzers for Improved Detection of Cerium in a Simulated Dispersal Accident

Here, this work implements a mid-level data fusion methodology on spectral data from handheld X-ray fluorescence and laser-induced breakdown spectroscopy analyzers to quantify plutonium surrogate (CeO 2 ) contamination in soil samples for the first time. Spectral data from each analyzer were used independently to train supervised machine learning regressions to predict Ce concentration. Fused features from both data sets were then used to train the same models, comparing prediction performance by evaluating model precision and sensitivity. Fusing principal component scores from the two sensors yielded an order of magnitude improvement in precision and sensitivity of predictions made with an artificial neural network, compared to predictions made by models trained on independent sensor data. As a result, a boosted ensemble trained on the fused spectral features yielded an ideal predictor with root-mean-squared error on the order of 10 –6 and calculated limit of detection order 10 –5 wt %.

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND ↗

Comparing Classical and Machine Learning Force Fields for Modeling Deformation of Metal–Organic Frameworks Relevant for Direct Air Capture

Deformation of metal–organic frameworks (MOFs) induced by adsorbate molecules can affect adsorption properties such as capacity and selectivity, but most computational studies of MOFs assume framework rigidity to simplify calculations. Although flexible force fields (FFs) for MOFs have been parametrized for specific materials, the generality of FFs for reliably modeling adsorbate-induced deformation to accuracy nearing that of density functional theory (DFT) has not been established. This work confirms using DFT calculations that adsorbate-induced deformation can affect CO 2 and H 2 O adsorption energies in a considerable fraction of MOFs promising for direct air capture (DAC). We then benchmark the efficacy of several general-purpose FFs in describing adsorbate-induced deformation for DAC against DFT. Our results show that current classical FFs are insufficient for describing MOF deformation, especially in cases of interest for DAC where strong interactions exist between adsorbed molecules and MOF frameworks. Some emerging machine learning force fields (MLFFs) we tested, particularly CHGNet, MACE-MP-0, and Equiformer V2, appear to be more promising than the classical FF for emulating the deformation behavior described by DFT. The best performing FF (CHGNet), however, fails to achieve the accuracy required for practical predictions with a mean absolute adsorption energy error of 0.124 eV.

adsorption↗

Enhancing Fluid Flow Pressure and Saturation Prediction Accuracy and Reducing Uncertainty with Committee Machine – Illinois Basin Decatur Project (IBDP) as a Case Study

Presentation at the 17th International Conference on Greenhouse Gas Control Technologies GHGT-17 held in Calgary, Canada, October 20-24, 2024. Carbon capture and storage (CCS) is a way to play a critical role in the global transition to a low-emission economy. Current progress is hampered by a number of factors, among which the lack of risk-informed design tools and decision support frameworks is seen as a major roadblock. Significant interest exists in using artificial intelligence to accelerate CCS site feasibility studies, as well as to facilitate the permit application process. Existing works commonly train a single deep learning model. This work investigates the feasibility of using a conventional ensemble learning (committee machine) technique to further improve prediction accuracy. Ensemble-based algorithms generally improve over individual base learners in terms of robustness and accuracy. Deep ensembles, however, are time-consuming to create and train. A pragmatic question is whether small-sized ensembles may lead to prediction improvement. Here we evaluated the efficacy of an ensemble learning technique using the latent spectral model (LSM), an efficient deep neural operator algorithm, as base learners. Preliminary results, obtained using the Illinois Basin-Decatur Project (IBDP) carbon sequestration data/model, show that small-sized ensembles can improve prediction over the base learners, achieving prediction accuracy of ~1.6 psi root mean square error (RMSE) on pressure (relative the average reservoir pressure of 3150 psi), and less than 1.3% for saturation.

Sun, Alexander↗

Enhancing Fluid Flow Pressure and Saturation Prediction Accuracy and Reducing Uncertainty with Committee Machine – Illinois Basin Decatur Project (IBDP) as a Case Study

This is the conference paper accompanying an oral presentation at the 17th International Conference on Greenhouse Gas Control Technologies GHGT-17 held in Calgary, Canada, October 20-24, 2024. Carbon capture and storage (CCS) is a way to play a critical role in the global transition to a low-emission economy. Current progress is hampered by a number of factors, among which the lack of risk-informed design tools and decision support frameworks is seen as a major roadblock. Significant interest exists in using artificial intelligence to accelerate CCS site feasibility studies, as well as to facilitate the permit application process. Existing works commonly train a single deep learning model. This work investigates the feasibility of using a conventional ensemble learning (committee machine) technique to further improve prediction accuracy. Ensemble-based algorithms generally improve over individual base learners in terms of robustness and accuracy. Deep ensembles, however, are time-consuming to create and train. A pragmatic question is whether small-sized ensembles may lead to prediction improvement. Here we evaluated the efficacy of an ensemble learning technique using the latent spectral model (LSM), an efficient deep neural operator algorithm, as base learners. Preliminary results, obtained using the Illinois Basin-Decatur Project (IBDP) carbon sequestration data/model, show that small-sized ensembles can improve prediction over the base learners, achieving prediction accuracy of ~1.6 psi root mean square error (RMSE) on pressure (relative the average reservoir pressure of 3150 psi), and less than 1.3% for saturation.

Sun, Alexander↗

Latent space mapping: Revolutionizing predictive models for divertor plasma detachment control

The inherent complexity of boundary plasma, characterized by multi-scale and multi-physics challenges, has historically restricted high-fidelity simulations to scientific research due to their intensive computational demands. Consequently, routine applications such as discharge control and scenario development have relied on faster but less accurate empirical methods. This work introduces DivControlNN, a novel machine-learning-based surrogate model designed to address these limitations by enabling quasi-real-time predictions (i.e., ~ 0.2 ms) of boundary and divertor plasma behavior. Trained on over 70,000 2D UEDGE simulations from KSTAR tokamak equilibria, DivControlNN employs latent space mapping to efficiently represent complex divertor plasma states, achieving a computational speed-up of over 10 8 compared to traditional simulations while maintaining a relative error below 20% for key plasma property predictions. During the 2024 KSTAR experimental campaign, a prototype detachment control system powered by DivControlNN successfully demonstrated detachment control on its first attempt, even for a new tungsten divertor configuration and without any fine-tuning. These results highlight the transformative potential of DivControlNN in overcoming diagnostic challenges in future fusion reactors by providing fast, robust, and reliable predictions for advanced integrated control systems.

Artificial neural networks↗

Forecasting Battery Electrode Performance via Electrochemical Fluorescence Microscopy and Machine-Learning

Predicting lithium-ion battery performance is hindered by microscale electrode heterogeneities invisible to conventional diagnostics. Here, we combine electrochemical fluorescence microscopy (EFM), which maps electronic connectivity by visualizing an electrofluorophore reaction distribution, with a multitask ElasticNet regression to forecast discharge capacity from spatial heterogeneity. Analyzing 196 images from six pilot-scale LiNi 0.5 Mn 0.3 Co 0.2 O 2 cathodes with varying carbon loadings, we extract 62 descriptors that capture morphology and texture. A compact five-feature model predicts capacity across eight discharge rates, achieving a per-target R 2 of up to 0.63 and an overall R 2 of 0.92, with a mean absolute percentage error of less than 2%. This performance rivals impedance-based approaches while avoiding their reliance on postformation data and incomplete electronic network information. Our facile and rapid, image-driven method may enable electrode quality control upstream of costly cell assembly to offer a transformative tool for data-driven battery research and manufacturing.

battery electrodes↗

Comparing Machine Learning and Physics-Based Nanoparticle Geometry Determinations Using Far-Field Spectral Properties

Anisotropic metal nanostructures exhibit polarization-dependent light scattering, a property which has been widely studied and exploited to determine orientations of subwavelength structures using far-field microscopy. Here we explore the use of variational autoencoders (VAEs) to determine the geometries of gold nanorods (NRs) such as in-plane orientation and aspect ratio under linearly polarized dark-field illumination in an optical microscope. We enforce a shared latent space to connect two VAEs trained separately with polarized dark-field scattering spectra and electron microscopy images and achieve image prediction (shape, orientation, and size) of Au NRs using only polarized dark-field scattering spectra. We determine the geometrical parameters of orientational angle and aspect ratio quantitatively via both our dual-VAE and physics-based analysis on the input scattering spectra. We show that orientational angle prediction by dual-VAE performs well with only a small (~300 particle) training set, yielding a mean absolute error (MAE) of 14.4° and a concordance correlation coefficient (CCC) of 0.95. This performance is only marginally worse than the physics-based cos(2?) fitting approach between the scattering intensity and the polarizing angle, which achieves MAE of 8.78° and CCC of 0.99. Aspect ratio determination is also comparable for the dual-VAE and physics-based fitting comparison (MAE of 0.21 vs. 0.23 and CCC of 0.53 vs. 0.68). Here, this dual encoder-decoder architecture effectively exploits the structure-property relationships of plasmonic nanostructures to construct a cross-modal machine learning (ML) approach, providing a pathway to employ ML approaches to address other structure-property relationships in materials science.

Dark-field scattering↗

Latent space dynamics identification for interface tracking with application to shock-induced pore collapse

Capturing sharp, evolving interfaces remains a central challenge in reduced-order modeling, especially when data is limited and the system exhibits localized nonlinearities or discontinuities. Here, we propose LaSDI-IT (Latent Space Dynamics Identification for Interface Tracking), a data-driven framework that combines low-dimensional latent dynamics learning with explicit interface-aware encoding to enable accurate and efficient modeling of physical systems involving moving material boundaries. At the core of LaSDI-IT is a revised autoencoder architecture that jointly reconstructs the physical field and an indicator function representing material regions or phases, allowing the model to track complex interface evolution without requiring detailed physical models or mesh adaptation. The latent dynamics are learned through linear regression in the encoded space and generalized across parameter regimes using Gaussian process interpolation with greedy sampling. We demonstrate LaSDI-IT on the problem of shock-induced pore collapse in high explosives, a process characterized by sharp temperature gradients and dynamically deforming pore geometries. The method achieves relative prediction errors below 9% across the parameter space, accurately recovers key quantities of interest such as pore area and hot spot formation, and matches the performance of dense training with only half the data. This latent dynamics prediction was 10 6 times faster than the conventional high-fidelity simulation, proving its utility for multi-query applications. These results highlight LaSDI-IT as a general, data-efficient framework for modeling discontinuity-rich systems in computational physics, with potential applications in multiphase flows, fracture mechanics, and phase change problems.

Gaussian process↗

DNABERT-S: pioneering species differentiation with species-aware DNA embeddings

SUMMARY: We introduce DNABERT-S, a tailored genome model that develops species-aware embeddings to naturally cluster and segregate DNA sequences of different species in the embedding space. Differentiating species from genomic sequences (i.e. DNA and RNA) is vital yet challenging, since many real-world species remain uncharacterized, lacking known genomes for reference. Embedding-based methods are therefore used to differentiate species in an unsupervised manner. DNABERT-S builds upon a pre-trained genome foundation model named DNABERT-2. To encourage effective embeddings to error-prone long-read DNA sequences, we introduce Manifold Instance Mixup (MI-Mix), a contrastive objective that mixes the hidden representations of DNA sequences at randomly selected layers and trains the model to recognize and differentiate these mixed proportions at the output layer. We further enhance it with the proposed Curriculum Contrastive Learning (C2LR) strategy. Empirical results on 28 diverse datasets show DNABERT-S's effectiveness, especially in realistic label-scarce scenarios. For example, it identifies twice more species from a mixture of unlabeled genomic sequences, doubles the Adjusted Rand Index (ARI) in species clustering, and outperforms the top baseline's performance in 10-shot species classification with just a 2-shot training. AVAILABILITY AND IMPLEMENTATION: Model, codes, and data are publically available at https://github.com/MAGICS-LAB/DNABERT_S.

Zhou, Zhihan↗

A Machine Learning based Approach of Estimating Equivalent Circuit Model Parameters at Different SoCs of Li-ion Batteries from Voltage Relaxation

Abstract: In this study, an approach of estimating the equivalent circuit model (ECM) parameters for Li-ion batteries (LIBs) is proposed based on the voltage value at different intervals while relaxing the LIB after discharge. The typical approach for estimating ECM parameters of a LIB is to conduct electrochemical impedance spectroscopy (EIS) measurements at different frequencies and fit them to a predefined circuit model, which requires additional measuring arrangements and specialized devices. The proposed methodology utilizes four different voltages at 0s, 60s, 360s, and 1800s alongside the specific state of charge (SoC) value for a specific constant discharge current value of ~1C until the relaxation stage to train and evaluate three regression-based machine learning models— Support Vector Regression (SVR), Extreme Gradient Boosting (XGBoost), and Gaussian Process Regression (GPR)—for estimating the ECM parameters of the selected model. Bayesian optimization is employed for hyperparameter tuning to achieve optimal performance for all the regressor models, among which, the GPR provided the best performance with the root-mean-squared error (RMSE) of less than 4x10-4 on average for the resistive components and less than 0.27 for capacitive components with excellent R2 scores. The simplicity of the approach enables it to eliminate the need for sophisticated measuring equipment and computation power.

Sagar, Md. Samiul [The University of Alabama (UA)]↗

Dense Image Matching Uncertainty Estimation and Confidence Metrics

Dense stereo matching takes overlapping image pairs as input and outputs a disparity map which encodes pixel-by-pixel matches between the images. Recently, there has been an interest in ranking the quality, or even quantifying the accuracy, of disparity estimates. The proposed methods can be described as either uncertainty estimators or confidence metrics. Uncertainty estimators are a small minority of the research. However, they have the potential to be the most useful because they estimate disparity accuracy (in pixel units) that can be used to threshold matches or carried forward using error propagation. The majority of the research deals with confidence metrics which give an ordinal (or binary) ranking of a match’s quality relative to other matches. Confidence metrics do not have units and thus are useful primarily for thresholding matches from mismatches. The methods could also be described as handcrafted or deep-learning based. The majority of the research focused on outdoor driving scenes. Hence, our interest–application to a satellite semi-global matching pipeline–is a domain shift that may challenge deep-learning based methods. We conclude by recommending five handcrafted and two deep-learning based methods for evaluation in our pipeline.

97 MATHEMATICS AND COMPUTING↗

Transformation rate maps of dissolved organic carbon in the contiguous US

Riverine dissolved organic carbon (DOC) plays a vital role in regional and global carbon cycles. However, the processes of DOC conversion from soil organic carbon (SOC) and leaching into rivers are insufficiently understood, inconsistently represented, and poorly parameterized, particularly in land surface and Earth system models. As a first attempt to fill this gap, we propose a generic formula that directly connects SOC concentration with DOC concentration in headwater streams, where a single parameter, the transformation rate from SOC in the soil to DOC leaching flux (P r ), accounts for the overall processes governing SOC conversion to DOC and leaching from soils (along with runoff) into headwater streams. We then derive high-resolution P r maps over the contiguous US (CONUS) using SOC data from two different sources: the Harmonized World Soil Database v1.2 (HWSD) and SoilGrids 2.0. Both maps are developed following the same five major steps: (1) selecting independent catchments where observed riverine DOC data are available with reasonable quality; (2) estimating catchment-average SOC for the independent catchments; (3) estimating the P r values for these catchments based on the generic formula and catchment-average SOC; (4) developing a predictive model of P r with machine learning (ML) techniques and catchment-scale climate, hydrology, geology, and other attributes; and (5) deriving a national map of P r based on the ML model. For evaluation, we compare the DOC concentration derived using the P r map and the observed DOC concentration values at evaluation catchments. The resulting mean absolute scaled error and coefficient of determination are 0.73 and 0.47 for the HWSD-based model and 0.58 and 0.72 for the SoilGrids-based model, respectively, suggesting the effectiveness of the overall methodology. Efforts to constrain uncertainty and evaluate sensitivity of P r to different factors are discussed. To illustrate the use of such maps, we derive a riverine DOC concentration reanalysis dataset over CONUS. The two P r maps, robustly derived and empirically validated, lay a critical cornerstone for better simulating the terrestrial carbon cycle in land surface and Earth system models. Our findings not only set a foundation for improving our predictive understanding of the terrestrial carbon cycle at the regional and global scales, but also hold promises for informing policy decisions related to decarbonization and climate change mitigation. The data presented in this study are publicly available at https://doi.org/10.5281/zenodo.14563816 (Li et al., 2024).

54 ENVIRONMENTAL SCIENCES↗

CuXASNet: Rapid and accurate prediction of copper L-edge x-ray absorption spectra using machine learning

In this work, we have developed CuXASNet, a dense neural network that predicts simulated Cu -edge x-ray absorption spectra (XAS) from atomic structures. Featurization of the Cu local environment is performed using a component of M3GNet, a graph neural network developed for predicting the potential energy surface. CuXASNet is trained on simulated spectra from FEFF9 at the multiple scattering level of theory, and can predict the and edges for Cu sites to quantitative accuracy. To validate our approach, we compare 14 experimental spectra extracted from the literature with the predictions of CuXASNet. The agreement of CuXASNet with experiments is shown by an average mean absolute error of 0.125 and an average Spearman's correlation coefficient of 0.891, which is comparable to FEFF9's values of 0.131 and 0.898 for the same metrics. As such, CuXASNet can rapidly predict a large number of -edge XAS spectra at the same accuracy as FEFF9 simulations. This can be used as a drop-in replacement for multiple scattering codes for fast screening of candidate atomic structure models of a measured system. This model establishes a general framework for Cu XAS prediction, and can be extended to more computationally expensive levels of theory and to other transition metal edges.

36 MATERIALS SCIENCE↗

Toward a Generalizable Prediction Model of Molten Salt Mixture Density with Chemistry‐Informed Transfer Learning

Optimally designing applications of molten salts requires knowledge of their thermophysical properties over a wide range of temperatures and compositions. There exist significant gaps in existing databases and this data can be challenging to experimentally measure due to high temperatures, salt corrosivity, and salt hygroscopicity. Existing databases have been used to create Redlich–Kister (RK) models for mixture density showing improved accuracy with respect to ideal mixing assumptions, but these models require subcomponent data measurements for each new system, therefore lacking generality. In order to address generalizability and data sparsity, a transfer learning procedure is proposed to train deep neural networks (DNNs) using a combination of semi‐empirical relationships (RK), data from the thermophysical arm of the molten salt thermal properties database and universal ab initio properties of component mixtures taken from the joint automated repository for various integrated simulations (JARVIS) classical force‐field inspired descriptors database to predict density in molten salts. Herein, it is shown that DNNs predict molten salt density with an r 2 over 0.99 and a mean absolute percentage error under 1%, outperforming alternative methods.

inorganic materials↗

Explainable machine learning reveals that local structural motifs encode the thermodynamic state across the CuZr metallic glass-forming range

Metallic glasses derive their properties from the statistics of local atomic motifs rather than from long-range order, yet a quantitative, chemistry-specific link between motif populations and the underlying glassy state has remained elusive. In this work we combine large-scale molecular dynamics, Voronoi tessellation, deep neural networks, and SHapley Additive exPlanations (SHAP) to identify which local structural motifs define the glassy state of Cu—Zr metallic glasses. A dataset of 17,180 atomistic configurations spanning ten compositions (Cu 20 Zr 80 –Cu 80 Zr 20 ) and four quench rates (10 9 –10 12 K/s) is used to train a feed-forward neural network that regresses temperature across the 50–2000 K liquid–supercooled–glass range, achieving a mean absolute error of 19.89 K and R 2 = 0.9974, confirming that the local structural state is faithfully encoded in motif-level structure. SHAP analysis then reveals that a tightly coupled near-icosahedral family of motifs (coordination numbers (CN) 11–13, including the full icosahedron 001200 and its single-atom-perturbation sibling 10930) collectively encodes the thermodynamic state of the system across the full glass-forming range. The CN = 11–13 ordered members carry negative SHAP values at high populations, tracking the most deeply-quenched configurations, while 10930 shows the reversed signature consistent with its role as a soft-spot host whose population shrinks as the icosahedral network deepens. The analysis demonstrates that explainable machine learning can isolate the minimal motif vocabulary defining the glassy state and recovers the near-icosahedral building blocks previously identified by data-driven analyses of Cu—Zr. The approach provides a general, chemistry-specific route for characterizing the structural state of disordered materials.

36 MATERIALS SCIENCE↗