Search NASA⌕ Search

SEARCH · Search NASA

Results for “Prediction”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 127 records · Page 7

Gene-Metabolite Association Prediction with Interactive Knowledge Transfer Enhanced Graph for Metabolite Production

Identifying gene targets for enhancing metabolite production in metabolic engineering is challenging due to the vast research literature and the approximation in genome-scale metabolic model (GEM) simulations. Here, to address this, we propose the Gene-Metabolite Association Prediction task, which automates gene discovery for given metabolite-gene pairs, accompanied by a benchmark dataset of 2474 metabolites and 1947 genes for Saccharomyces cerevisiae (SC) and Issatchenkia orientalis (IO). This task is complicated by incomplete metabolic graphs and metabolic heterogeneity. We introduce an Interactive Knowledge Transfer mechanism based on Metabolism Graphs (IKT4Meta) to enhance prediction accuracy by integrating cross-metabolism knowledge. Using Pretrained Language Models (PLMs) to generate inter-graph links mitigates heterogeneity issues, while intra-graph links are propagated via these anchors. Gene-metabolite predictions are then performed on the enriched graphs integrating multiple microorganisms’ knowledge. Experiments show that IKT4Meta outperforms baselines by up to 12.3% in link prediction.

59 BASIC BIOLOGICAL SCIENCES↗

Part-scale microstructure prediction for laser powder bed fusion Ti-6Al-4V using a hybrid mechanistic and machine learning model

Laser powder bed fusion (LPBF) Ti-6Al-4V is widely studied for use in structural applications in aerospace and medical industries, but mechanical anisotropy and microstructural inhomogeneity prohibits its wider adoption. Although successful microstructure prediction models have been developed, a remaining challenge is their limited integration across length/time scales and validation by experimental studies. Here, this work proposes a physics-augmented machine learning surrogate model to unite predictions of LPBF temperature, β phase morphology and texture, and α/α’ formation into a single framework that is calibrated and validated with experiments. First, a phase field (PF) model of the martensitic β→α’ transformation is developed and calibrated using data from in-situ synchrotron cyclic heating/cooling studies quantifying the variation of α phase fraction with time. In parallel, an established finite difference-Monte Carlo (FDMC) model predicts the part-scale temperature profile and β grain formation during solidification. A dataset is developed using LPBF cyclic temperature descriptors from the FDMC model as inputs and corresponding α/α’ phase fraction and width from the PF model as outputs. Five machine learning (ML) regression models are tested and optimized, having mean absolute error in testing ≤ 4 %, and the k-nearest neighbors (KNN) model is selected as the best performing. The KNN model is called at the nodal level during post-processing of the FDMC model to replace and downscale the response of the PF model. The combined agility and accuracy of the hybrid FDMC-ML model enables part-scale microstructure predictions that can be further used for property predictions to accelerate AM process optimization.

36 MATERIALS SCIENCE↗

Predicting U.S. federal fleet electric vehicle charging patterns using internal combustion engine vehicle fueling transaction statistics

Utilizing fueling transactions from internal combustion engine vehicles (ICEVs), the authors estimated how frequently midday public charging would be required for U.S. federal fleet battery electric vehicles (BEVs). Fueling transaction summary statistics are more widely available than trip-level telematics data, making this methodology more accessible and transferable to other researchers and fleet managers considering BEV replacements. For example, readers can easily apply a linear model using only the count of back-to-back fueling events at gas stations over 57 straight-line miles apart to predict days exceeding range. This linear regression predicted binned days exceeding 250 miles at 80% accuracy on a hold-out test set from the same fleet as the training data and 66 % accuracy on a new fleet displaying different driving behaviors. The authors additionally provide linear equations for days exceeding 200 and 300 miles as alternative range estimates to account for differences in BEV range and temperature impacts. Beyond the single-feature linear models which readers can apply, the authors tuned and trained other machine learning models on a variety of fueling transaction statistics including consecutive transaction distances, transaction distance from garage, estimated miles traveled from fuel economy and fuel quantity, and transaction periodicity. Utilizing a subset of 1678 light-duty federal fleet vehicles which contained daily vehicle miles traveled (VMT) in addition to fueling statistics, the authors determined which fueling transaction statistics were most relevant in predicting driving days exceeding 250 miles (an approximation of BEV rated driving range). In support of the U.S. federal fleet transition to zero-emission vehicles (ZEVs), the authors used these statistics and machine learning models to predict the frequency of BEV midday charging. After training models on the subset with VMT, the authors predicted days exceeding rated range for 112,902 light-duty vehicles operating in similar circumstances in the federal fleet using a Support Vector Regressor (SVR). In conclusion, they then used the projections as part of the ZEV Planning and Charging (ZPAC) tool to identify optimal candidates for BEVs for the federal fleet. An anonymized version of ZPAC is included in the supplementary materials.

25 ENERGY STORAGE↗

Predicting RNA structure and dynamics with deep learning and solution scattering

Advanced deep learning and statistical methods can predict structural models for RNA molecules. However, RNAs are flexible, and it remains difficult to describe their macromolecular conformations in solutions where varying conditions can induce conformational changes. Small-angle x-ray scattering (SAXS) in solution is an efficient technique to validate structural predictions by comparing the experimental SAXS profile with those calculated from predicted structures. There are two main challenges in comparing SAXS profiles to RNA structures: the absence of cations essential for stability and charge neutralization in predicted structures and the inadequacy of a single structure to represent RNA’s conformational plasticity. We introduce a solution conformation predictor for RNA (SCOPER) to address these challenges. This pipeline integrates kinematics-based conformational sampling with the innovative deep learning model, IonNet, designed for predicting Mg 2+ ion binding sites. Validated through benchmarking against 14 experimental data sets, SCOPER significantly improved the quality of SAXS profile fits by including Mg 2+ ions and sampling of conformational plasticity. We observe that an increased content of monovalent and bivalent ions leads to decreased RNA plasticity. Therefore, carefully adjusting the plasticity and ion density is crucial to avoid overfitting experimental SAXS data. SCOPER is an efficient tool for accurately validating the solution state of RNAs given an initial, sufficiently accurate structure and provides the corrected atomistic model, including ions.

59 BASIC BIOLOGICAL SCIENCES↗

On the numerical sensitivity of cellular automata grain structure predictions to large thermal gradients and cooling rates

Cellular automata (CA) models of as-solidified grain structure, originally developed and applied to casting, have become a common means of predicting grain structure resulting from Additive Manufacturing (AM) processes. The majority of these models are based on the decentered octahedron approach, which attempts to correct for the effect of grid anisotropy on the prediction of competitive solidification of dendritic grains. However, AM solidification occurs under cooling rates ($\dot{T}$) and thermal gradients (G) that are orders of magnitude larger than those encountered in casting, and no systematic investigation on the effect of the CA model cell size (Δx) and time step (Δt) on AM microstructure predictions has been performed. Here, in this study, such an investigation is first performed via simulation of individual grains of various crystallographic orientations with a fixed, unidirectional G, showing that CA prediction of the steady-state undercooling matched the expected values based on the interfacial response function at small G and deviated from the expected values at large G. Simulation of competitive growth of multiple grains showed a weakening of the predicted texture as G and Δx became large. Simulation of solidification under AM conditions, where G and $\dot{T}$ vary spatially across the melt pools, showed that not only does grain selection weaken and deviate from expectations at large Δx, but grains with crystallographic $\langle$100$\rangle$ aligned with the grid directions are more adversely affected by the temperature field discontinuities than grains with other crystallographic orientations. Despite the fact that the exact grain competition results depended on Δt, the overall texture development was notably less sensitive to Δt than Δx, provided that a reasonable value of Δt is selected based on the ratio of Δx to the maximum local solidification velocity in the simulation domain. Finally, from the directional solidification and AM simulation results, an analysis of computational cost compared to simulation resolution is performed based on an equation derived to quantify the relatively inaccuracy in grain selection based on the model and temperature field inputs. From this analysis, it is concluded that there is a need for algorithmic improvements to improve CA grain competition accuracy for large G processing conditions as sufficiently small Δx to resolve the necessary competition is intractable for many AM processing conditions.

36 MATERIALS SCIENCE↗

Graph neural networks for CO 2 solubility predictions in Deep Eutectic Solvents

Deep Eutectic Solvents (DESs) are a promising class of solvents for CO 2 capture. DESs are complex mixtures that can be designed to optimize CO solubility and overall capture process efficiency. However, the vast design landscape of DES mixtures makes experimental investigation prohibitive; as such, there is a need for computational models that can quickly and efficiently navigate the design space and inform data collection efforts. In this work, we propose Graph Neural Network (GNN) models for predicting CO 2 solubility for DESs; the GNN leverages a mixture graph representation that captures the molecular structure of the DES components as well as their intermolecular interactions. Here, we compare the GNN framework against alternative architectures (neural networks, graph convolution networks, and random forests) and data representations (molecular fingerprints, sigma profiles, and graphs). We show that the proposed approach offers superior predictive performance; specifically, we show that solubility can be predicted reliably directly from molecular structure (without the need of using sigma profiles as proposed in previous studies). This result is important, as obtaining sigma profiles requires expensive density functional theory computations. We also explored the ability of GNNs to predict solubility for new DES mixtures and operating conditions. We found that the model extrapolates across temperature reliably. However, we also found deficiencies in the ability of the model to predict solubility for DES mixtures, pressures, and molar ratio not included in the training sets; we show that this is due to an inherent lack of chemical diversity in datasets available in the literature. The proposed computational capabilities can thus help navigate the design space of DES and inform data collection efforts. Our models, data, and benchmarks are shared as Python code implemented in Jupyter notebooks.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Predicting roughness effects in additively manufactured coolant channels with helical enhancements

Additive manufacturing (AM) is a promising technique for fabrication of complex geometries such as those expected to be utilized in the blanket, first wall, and divertor. In the case of cooling, metallic AM may be exploited to embed geometric enhancements (ribs, rifling, etc.) to improve cooling performance. However, due to the roughness of these unfinished internal AM surfaces, prediction of thermal hydraulic performance in such channels is difficult. In this work, we consider a methodology for predicting pressure drop and heat transfer in AM channels containing helical enhancements (e.g. helical ribs, twisted tapes) that allows the incorporation of roughness data through conventional pipe flow correlations. This methodology is tested using experimental friction factor and heat transfer coefficient data from high-pressure helium coolant flow measurements in AM stainless steel tubes fabricated by laser powder bed fusion. Both a featureless AM tube and one containing helical ribs were considered alongside a conventionally manufactured smooth tube. The AM surface roughness is obtained by profilometry and used to predict an equivalent sand-grain roughness, with this equivalent roughness confirmed through AM featureless tube measurements. Under the proposed methodology, this roughness information is incorporated into predictions of friction factor and Nusselt number for the rifled tube. Furthermore, these predictions agree well with experimental data across a large range of Reynolds numbers, encouraging the use of this methodology for thermal hydraulic analysis of similar systems and design of future coolant channel geometries.

Additive manufacturing↗

Regional-scale soil carbon predictions can be enhanced by transferring global-scale soil–environment relationships

Accurate modelling and mapping soil organic carbon are crucial for supporting soil health restoration and climate change mitigation at both regional and global scales. However, regional soil predictions often suffer from data scarcity and high prediction uncertainty. Utilizing a pre-trained global-to-regional soil carbon predictive model can be a potential solution to address this challenge. Despite its promise, how to construct and apply the global-scale model to enhance regional-scale soil carbon mapping remains largely unexplored. Here, we propose the Global Soil Carbon Pre-trained Model (GSoilCPM), a deep-learning-based domain adaptative model, to enhance regional-scale soil carbon predictions. Based on large amount of environmental covariate data and 106,167 soil samples across the globe, we verify our hypothesis of the effectiveness of this 'global-to-regional' modelling strategy. The pre-trained model can be then transferred and fine-tuned to bridge the regional- and global-scale soil–environment relationships. We applied and validated this modelling strategy in four regional-scale study areas, three in the Northern Hemisphere and one in the Southern Hemisphere, each with distinct environmental background. Compared to traditional modelling approaches as a baseline, four case studies all demonstrated significant improvement in prediction accuracy across diverse environments and varying data availabilities. The average percentage improvement across all regions is 10.93% (absolute values decreased by 1.20 g kg−1 averagely) in MAE and 29.04% (absolute values increased by 0.10 averagely) in CCC. The applicability and future horizons of using GSoilCPM were further discussed. We further reveal that regions with fewer soil samples or lower baseline accuracy benefit more from the pre-trained global model. Our findings highlight the advantages of leveraging the generalized knowledge from global models to enhance specifically localized soil modelling, positioning a potential paradigm shift in digital soil mapping, and far-reaching implications for soil monitoring and land management.

Deep learning↗

Self-supervised and multi-fidelity learning for extended predictive soil spectroscopy

Infrared spectroscopy is a cost-effective, non-destructive, and environmentally benign technology that is increasingly recognized as an important solution for meeting the global demand for soil data. While both near-infrared (NIR) and mid-infrared (MIR) diffuse reflectance spectroscopy enable rapid estimation of soil properties, they present a significant trade-off: NIR offers superior scalability and lower operational costs, whereas MIR provides higher analytical fidelity by capturing fundamental molecular vibrations. In this study, we propose a self-supervised, multi-fidelity learning framework designed to bridge this gap. Our approach leverages large-scale MIR spectral libraries to learn a compact, transferable latent representation, into which NIR spectra are subsequently aligned for downstream prediction. The workflow consists of pretraining a latent model on a large MIR library, adapting the representation using a smaller paired NIR–MIR dataset, and evaluating generalization on an independent external test set. Across a range of chemical and physical soil properties, we found that MIR-derived embeddings improved prediction accuracy relative to baseline models that used raw MIR inputs. Predictions derived from the spectrum conversion (NIR to MIR) task did not match the performance of the original MIR spectra but were similar or superior to predictive performance of NIR-only models, suggesting the unified spectral latent space can effectively leverage the larger and more diverse MIR dataset for prediction of soil properties not well represented in current NIR libraries.

54 ENVIRONMENTAL SCIENCES↗

A machine-learning-aided data recovery approach for predicting multi-material thermal behaviors in advanced test reactor capsules

Instrumented experiments conducted at test reactors are essential to the deployment of new advanced reactor systems. Designing new experiments and generating data on specific reactor conditions require significant investments in terms of both time and cost. Finite element analysis software can be used to create high-fidelity models of experiment environments in order to support the actual experiments, but computation time remains a concern in terms of applying outcomes to real-time usage of data (e.g., a digital twin [DT]). Here, the present research proposes a machine-learning (ML) aided approach to making temperature and displacement predictions based on the thickness of the outer gas gap on the experimental capsule used for in-pile demonstration of a novel new thermal conductivity probe in the Advanced Test Reactor (ATR). This capsule consisted of U10Zr fuel, a rodlet, sodium, and inner and outer capsules. Gas gaps existed between the fuel and the rodlet, and between the inner and the outer capsule. The learning data pertained to an experimental capsule's radial distributions of temperature and displacement, as obtained based on Abaqus and the physical features. For the first step of ML sequence, the temperature was predicted using three positional parameters. Next, the displacement was predicted using seven additional parameters. Each physical feature was normalized in order to be both nondimensional and standardized. The temperature and displacement predictions showed good agreement with the simulation results in all cases involving interpolation and extrapolation. Furthermore, data similarity enhancement increased the similarity between the training and the target data, thereby increasing the predictive accuracy of the ML models. In certain extrapolation cases involving limited original ML model accuracy, data similarity enhancement and data recovery was able to somewhat improve this accuracy.

11 NUCLEAR FUEL CYCLE AND FUEL MATERIALS↗

Comparison of URANS and LES predictions for the open phase of the OECD NEA CSNI fluid structure interaction CFD benchmark

The OECD NEA CSNI WGAMA CFD Task Group ran a benchmark in 2020 and 2021 to assess the predictive capabilities of coupled fluid structure interaction (FSI) CFD analysis methods. This paper presents the predictions made for the open phase of the benchmark using URANS and LES turbulence modelling approaches, and a comparison of the results to the experimental data. The benchmark comprised a channel containing two inline cylinders in cross-flow. The cylinders were fixed at one end, free at the other, and had measured resonant frequencies and damping properties. The URANS modelling used ANSYS Fluent 2-way coupled to ANSYS Mechanical. The LES modelling used Nek5000, 1-way coupled to Diablo. Comparisons with cross-channel velocity profiles are presented, both for the mean flow and its RMS. Comparisons are also made to the frequency spectra for point measurements of fluid velocity and pressure, and for the accelerations of the free end of each cylinder. URANS predicts the average velocity profiles relatively well, and is able to predict the velocity and acceleration spectra at the shedding frequency. However, the frequency content at the 4th harmonic of the shedding frequency is low in the URANS flow fields, and so does not excite accelerations at the resonant frequency of the cylinders. LES makes better predictions of the average profiles, and the velocity spectra agree well at both the shedding frequency and at higher frequencies. In conclusion, the 1-way coupled LES results show good agreement for acceleration spectra.

22 GENERAL STUDIES OF NUCLEAR REACTORS↗

Combined effects of horizontal and vertical resolution on reliable turbulence prediction at tidal energy sites: A systematic study in the Salish Sea, WA

Predicting turbulence characteristics with coastal ocean models is essential for tidal energy converter deployment. While large eddy simulation provides a detailed turbulence representation, computational limitations restrict its use to smaller domains. We systematically evaluate whether well-configured coastal models can provide reliable turbulence prediction through progressive refinement of 3D model representation. We implemented four model configurations (Levels 1–4) using terrain-following coordinates, isolating the impacts of horizontal resolution, vertical resolution, and layer distribution. We validated all configurations against field measurements from the Salish Sea, WA. Results show that tidal current velocity predictions remain unchanged regarding model configuration, but turbulence properties are sensitive to resolution refinement. Increasing vertical resolution alone proved insufficient; even with vertical sigma-levels rising from 11 to 41, significant underprediction persisted until finer horizontal resolution better captured bathymetric variations. The Level 4 configuration, incorporating geometric sigma-levels distribution, achieved turbulence prediction skill scores exceeding 0.90. Turbulence closure comparison revealed Mellor- Yamada 2.5 outperformed k-epsilon in TKE prediction (skill scores 0.84–0.94 versus 0.72–0.81) due to better boundary layer parameterization. This study shows that well-configured coastal models effectively bridge the gap between simplified tools and costly high-fidelity modeling, offering the tidal energy industry practical and cost-effective turbulence data at commercially relevant scales.

Marine Energy↗

Revisiting the validity of eddy viscosity models for predicting airflow over water waves

In this study, we revisit the validity of eddy viscosity models for predicting wave-induced airflow disturbances over ocean surface waves. We first derive a turbulence curvilinear model for the phase-averaged Navier–Stokes equations, extending the work of Cao, Deng & Shen (2020 J. Fluid Mech. 901, A27), by incorporating turbulence stress terms previously neglected in the linearised viscous curvilinear model. To verify our formulation, we perform a priori tests by numerically solving the model using mean wind and turbulence stress profiles from large-eddy simulations (LES) of airflow over waves across various wave ages. Results show that including turbulence stress terms improves wave-induced airflow predictions compared with the previous viscous curvilinear model. We further show that using a standard mixing-length eddy viscosity yields inaccurate predictions at certain wave ages, as it fails to capture wave-induced turbulence, which fundamentally differs from mean shear-driven turbulence. The LES data show that accurate representations of wave-induced stresses require a complex-valued eddy viscosity. The maximum magnitude of this eddy viscosity scales as ∼𝑢 𝜏 ⁢𝜁 𝑖𝑛𝑛𝑒𝑟 , where 𝑢 𝜏 is the friction velocity and 𝜁 𝑖𝑛𝑛𝑒𝑟 is the inner-layer thickness, the height at which the eddy-turnover time matches the wave advection time scale. This scaling aligns with the prediction by Belcher & Hunt (1993 J. Fluid Mech. 251, 109–148). Overall, the findings demonstrate that traditional eddy viscosity models are inadequate for capturing wave-induced turbulence. More sophisticated turbulence models are essential for the accurate prediction of airflow disturbances and form drag in wind–wave interaction models.

16 TIDAL AND WAVE POWER↗

Learning nonlinear operators in latent spaces for real-time predictions of complex dynamics in physical systems

Abstract Predicting complex dynamics in physical applications governed by partial differential equations in real-time is nearly impossible with traditional numerical simulations due to high computational cost. Neural operators offer a solution by approximating mappings between infinite-dimensional Banach spaces, yet their performance degrades with system size and complexity. We propose an approach for learning neural operators in latent spaces, facilitating real-time predictions for highly nonlinear and multiscale systems on high-dimensional domains. Our method utilizes the deep operator network architecture on a low-dimensional latent space to efficiently approximate underlying operators. Demonstrations on material fracture, fluid flow prediction, and climate modeling highlight superior prediction accuracy and computational efficiency compared to existing methods. Notably, our approach enables approximating large-scale atmospheric flows with millions of degrees, enhancing weather and climate forecasts. Here we show that the proposed approach enables real-time predictions that can facilitate decision-making for a wide range of applications in science and engineering.

97 MATHEMATICS AND COMPUTING↗

Impact of microkinetic modeling assumptions on predicted kinetics and mechanisms over undercoordinated sites

Accurate modeling of catalytic reactions on undercoordinated sites requires accounting for the structural and ensemble-specific nature of the active sites. This study examines how common microkinetic modeling (MKM) assumptions affect predicted kinetics and mechanisms on the stepped Pt(211) facet for the ethane dehydrogenation (EDH) and the ethane hydrogenolysis (EH). Six (211) MKMs were developed, differing in (i) the number of active sites represented, (ii) adsorbate site occupancy treatment, and (iii) inclusion of cross-facet interactions. These models are benchmarked against a particle-based microkinetic model (PB-MKM), which best represents step-edge behavior. MKM assumptions caused deviations in turnover frequencies exceeding ten orders of magnitude and led to contrasting mechanistic and selectivity predictions. Multi-site MKMs overestimate activity by inflating free site availability, single-site models underestimate activity, and uniform occupancy models overpredict coverage of multi-dentate intermediates, leading to reaction-specific artifacts. Overall, the Combined Site Edge Model (CSEM), a single-site MKM accounting for site occupancy and cross-facet interactions, most closely approximates PB-MKM predictions. All models predict similar kinetics when surfaces are clean or primarily occupied by monodentate species. This work provides practical guidance for selecting MKM frameworks for undercoordinated catalytic surfaces and highlights the critical role of modeling assumptions in catalytic predictions.

(211) facet↗

Machine learning prediction of enzyme optimum pH

The relationship between pH and enzyme catalytic activity, especially the optimal pH (pH opt ) at which enzymes function, is critical for biotechnological applications. Hence, computational methods to predict pH opt will enhance enzyme discovery and design by facilitating accurate identification of enzymes that function optimally at specific pH levels, and by elucidating sequence-function relationships. Here, in this study, we proposed and evaluated various machine learning methods for predicting pH opt , conducting extensive hyperparameter optimization and training over 11,000 model instances. Our results demonstrate that models utilizing language model embeddings markedly outperform other methods in predicting pHopt. We present EpHod, the best-performing model, to predict pHopt, making it publicly available to researchers. From sequence data, EpHod directly learns structural and biophysical features that relate to pH opt , including proximity of residues to the catalytic centre and the accessibility of solvent molecules. Overall, EpHod presents a promising advancement in pH opt prediction and will potentially speed up the development of enzyme technologies.

97 MATHEMATICS AND COMPUTING↗

Virtual node graph neural network for full phonon prediction

Understanding the structure-property relationship is crucial for designing materials with desired properties. The past few years have witnessed remarkable progress in machine-learning methods for this connection. However, substantial challenges remain, including the generalizability of models and prediction of properties with materials-dependent output dimensions. Here we present the virtual node graph neural network to address the challenges. By developing three virtual node approaches, we achieve Γ-phonon spectra and full phonon dispersion prediction from atomic coordinates. We show that, compared with the machine-learning interatomic potentials, our approach achieves orders-of-magnitude-higher efficiency with comparable to better accuracy. This allows us to generate databases for Γ-phonon containing over 146,000 materials and phonon band structures of zeolites. Additionally, our work provides an avenue for rapid and high-quality prediction of phonon band structures enabling materials design with desired phonon properties. The virtual node method also provides a generic method for machine-learning design with a high level of flexibility. In this study, the authors present a virtual node graph neural network to enable the prediction of material properties with variable output dimensions. This method offers fast and accurate predictions of phonon band structures in complex solids.

36 MATERIALS SCIENCE↗

Uncertainty quantification for molecular property predictions with graph neural architecture search

Graph Neural Networks (GNNs) have emerged as a prominent class of data-driven methods for molecular property prediction. However, a key limitation of typical GNN models is their inability to quantify uncertainties in the predictions. This capability is crucial for ensuring the trustworthy use and deployment of models in downstream tasks. To that end, we introduce AutoGNNUQ, an automated uncertainty quantification (UQ) approach for molecular property prediction. AutoGNNUQ leverages architecture search to generate an ensemble of high-performing GNNs, enabling the estimation of predictive uncertainties. Our approach employs variance decomposition to separate data (aleatoric) and model (epistemic) uncertainties, providing valuable insights for reducing them. In our computational experiments, we demonstrate that AutoGNNUQ outperforms existing UQ methods in terms of both prediction accuracy and UQ performance on multiple benchmark datasets, and generalizes well to out-of-distribution datasets. Additionally, we utilize t-SNE visualization to explore correlations between molecular features and uncertainty, offering insight for dataset improvement. AutoGNNUQ has broad applicability in domains such as drug discovery and materials science, where accurate uncertainty quantification is crucial for decision-making.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗