Search NASA⌕ Search

SEARCH · Search NASA

Results for “Predicting”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 505 records · Page 28

A Deep Multimodal Representation Learning Framework for Accurate Molecular Properties Prediction

Drug discovery is a complex and challenging process, requiring the optimization of candidate compounds to identify those with the potential to become safe and effective drugs. Predicting molecular properties is an indispensable step in the drug discovery pipeline. Traditionally, this process is costly and time-intensive, involving multiple rounds of experiments and clinical trials, rendering it impractical for every candidate compound. Deep learning techniques have emerged as a promising approach to drug discovery to reduce the cost and time required to identify novel drugs. However, prevalent research in deep learning models focused on predicting molecular properties has primarily fixated on single-modal models, which utilize a single modality of data, neglecting the potential benefits of combining different data modalities. To overcome this limitation, we introduce MRL-Mol: a deep \textbf{M}ultimodal \textbf{R}epresentation \textbf{L}earning framework for accurate \textbf{Mol}ecular properties prediction. MRL-Mol harnesses three data modalities: sequence, graph, and image, augmenting the depth of comprehension. Leveraging a large-scale unlabeled dataset~($\sim$1M unique molecules), we pretrain MRL-Mol to extract inter- and intra-modal information. Our study demonstrates the superior performance of MRL-Mol in predicting molecular properties across six benchmark datasets, including both classification and regression tasks. Notably, MRL-Mol outperforms other state-of-the-art molecular properties prediction models. These findings suggest that by combining information from multiple data modalities, MRL-Mol can comprehend molecules better than single-modal deep learning models and identify molecular properties with better accuracy.

Yang, Yuxin↗

Exascale Computing and Data Handling: Challenges and Opportunities for Weather and Climate Prediction

The emergence of exascale computing and artificial intelligence offer tremendous potential to significantly advance Earth system prediction capabilities. However, enormous challenges must be overcome to adapt models and prediction systems to use these new technologies effectively. A 2022 WMO report on exascale computing recommends “urgency in dedicating efforts and attention to disruptions associated with evolving computing technologies that will be increasingly difficult to overcome, threatening continued advancements in weather and climate prediction capabilities.” Further, the explosive growth in data from observations, model and ensemble output, and postprocessing threatens to overwhelm the ability to deliver timely, accurate, and precise information needed for decision-making. Artificial intelligence (AI) offers untapped opportunities to alter how models are developed, observations are processed, and predictions are analyzed and extracted for decision-making. Given the extraordinarily high cost of computing, growing complexity of prediction systems, and increasingly unmanageable amount of data being produced and consumed, these challenges are rapidly becoming too large for any single institution or country to handle. This paper describes key technical and budgetary challenges, identifies gaps and ways to address them, and makes a number of recommendations.

Atmosphere↗

The impact of kidney function on Alzheimer’s disease blood biomarkers: implications for predicting amyloid-β positivity

Impaired kidney function has a potential confounding effect on blood biomarker levels, including biomarkers for Alzheimer’s disease (AD). Given the imminent use of certain blood biomarkers in the routine diagnostic work-up of patients with suspected AD, knowledge on the potential impact of comorbidities on the utility of blood biomarkers is important. We aimed to evaluate the association between kidney function, assessed through estimated glomerular filtration rate (eGFR) calculated from plasma creatinine and AD blood biomarkers, as well as their influence over predicting Aβ-positivity. We included 242 participants from the Translational Biomarkers in Aging and Dementia (TRIAD) cohort, comprising cognitively unimpaired individuals (CU; n = 124), mild cognitive impairment (MCI; n = 58), AD dementia (n = 34), and non-AD dementia (n = 26) patients all characterized by [ 18 F] AZD-4694. Plasma samples were analyzed for Aβ42, Aβ40, glial fibrillary acidic protein (GFAP), neurofilament light chain (NfL), tau phosphorylated at threonine 181 (p-tau181), 217 (p-tau217), 231 (p-tau231) and N-terminal containing tau fragments (NTA-tau) using Simoa technology. Kidney function was assessed by eGFR in mL/min/1.73 m 2 , based on plasma creatinine levels, age, and sex. Participants were also stratified according to their eGFR-indexed stages of chronic kidney disease (CKD). We evaluated the association between eGFR and blood biomarker levels with linear models and assessed whether eGFR provided added predictive value to determine Aβ-positivity with logistic regression models. Biomarker concentrations were highest in individuals with CKD stage 3, followed by stages 2 and 1, but differences were only significant for NfL, Aβ42, and Aβ40 (not Aβ42/Aβ40). All investigated biomarkers showed significant associations with eGFR except plasma NTA-tau, with stronger relationships observed for Aβ40 and NfL. However, after adjusting for either age, sex or Aβ-PET SUVr, the association with eGFR was no longer significant for all biomarkers except Aβ40, Aβ42, NfL, and GFAP. When evaluating whether accounting for kidney function could lead to improved prediction of Aβ-positivity, we observed no improvements in model fit (Akaike Information Criterion, AIC) or in discriminative performance (AUC) by adding eGFR to a base model including each plasma biomarker, age, and sex. While covariates like age and sex improved model fit, eGFR contributed minimally, and there were no significant differences in clinical discrimination based on AUC values. We found that kidney function seems to be associated with AD blood biomarker concentrations. However, these associations did not remain significant after adjusting for age and sex, except for Aβ40, Aβ42, NfL, and GFAP. While covariates such as age and sex improved prediction of Aβ-positivity, including eGFR in the models did not lead to improved prediction for any biomarker. Our findings indicate that renal function, within the normal to mild impairment range, does not seem to have a clinically relevant impact when using highly accurate blood biomarkers, such as p-tau217, in a biomarker-supported diagnosis.

60 APPLIED LIFE SCIENCES↗

Application of artificial intelligence methods in the international roughness index prediction of rigid and composite pavements: a systematic review

The International Roughness Index (IRI) is a widely adopted metric for quantifying pavement roughness, directly influencing vehicle safety, ride comfort, and overall roadway performance. In recent years, the use of Machine Learning (ML) models for IRI prediction has gained momentum, with the goal of improving the allocation of maintenance and rehabilitation resources by enabling accurate assessments of pavement conditions. Most prior reviews, however, have concentrated on flexible pavements, leaving a notable gap regarding rigid and composite pavements. To address this gap, the present study conducts a systematic review of Artificial Intelligence (AI) methods applied to IRI prediction for rigid and composite pavements. Literature published between 2004 and 2025 is synthesized to highlight prevailing trends, methodological contributions, and directions for future research. Particular attention is given to the types of models employed, the datasets used for training and validation, and the role of input variables and data-processing strategies. Across the included studies, ensemble learning methods (especially gradient boosting variants such as XGBoost), artificial neural networks, and hybrid architectures frequently achieved high predictive skill, with several models reporting test-set coefficients of determination approaching 0.9–0.96, indicating strong potential for capturing the influence of traffic, pavement structure, and climatic factors. Since these results are obtained from heterogeneous datasets and evaluation protocols, they are interpreted qualitatively rather than as strict cross-study rankings. Analysis of input variables revealed that pavement age and initial IRI were included in 91% (21 of 23) and 78% (18 of 23) of studies, respectively. Climatic variables such as the freezing index appeared in 57% (13 of 23), while traffic-related factors were considered in 65% (15 of 23). The findings underscore the importance of standardized, high-quality datasets, such as those from the Long-Term Pavement Performance (LTPP) program, along with data consistency, model interpretability, computational efficiency, and replicability in enhancing IRI prediction. Future research should focus on incorporating input variable selection techniques to identify the most influential predictors, thereby improving accuracy and robustness. Integrating these approaches with advanced non-linear data-driven models, coupled with robust hyperparameter optimization, holds considerable promise for strengthening the reliability of IRI prediction and supporting resilient pavement management strategies.

42 ENGINEERING↗

Dataset for "Machine Learning Ensembles Can Enhance Hydrologic Predictions and Uncertainty Quantification" Willard et al. (2025).

This data release provides all data and code used in the paper " "Machine Learning Ensembles Can Enhance Hydrologic Predictions and Uncertainty Quantifications" Willard et al. (2025)" to model stream temperature, evaluate, and assess results. The associated manuscript explores the effect of different ensemble construction techniques across different common machine learning (ML) architectures for predictions in unmonitored basins. Modeling was done using long short-term memory (LSTM), gated recurrent unit (GRU), temporal convolution network (TCN), and extreme gradient boosting (XGBoost) models, and stream site coverage spans 1362 locations across the conterminous United States. The ensemble construction techniques investigated include ensemble by random weight initialization, differing hyperparameters, different random subsets of training data, different subselections of input features, different architectures, and Monte Carlo Dropout. The data is organized into these items items:Code repository and data for the paper " "Machine Learning Ensembles Can Enhance Hydrologic Predictions and Uncertainty Quantifications" Willard et al. (2025).Code: stream_temp_ml_regionalization.zip contains the code repositoryData to run the code:- data_dir.zip -- contains all files that should be moved to the "DATA_DIR" variable defined in the "set_env_vars.sh" script in the code repository- metadata_dir.zip -- contains all files that should be moved to the "METADATA_DIR" variable defined in the "set_env_vars.sh" script in the code repositoryData produced by the code and used in the paper:- outputs_dir.zip - contains model output and results (outputs_dir/results), model weights (outputs_dir/models), and all other outputs used for the paper including feature importances.To cite this code, please use the following BibTeX or MLA entries:bibtex:@misc{willard2025streamensembles,author = {Jared Willard and Charuleka Varadharajan},title = {Dataset for "Machine Learning Ensembles Can Enhance Hydrologic Predictions and Uncertainty Quantification"},year = {2024},doi = {10.15485/2527393},publisher = {ESS-DIVE Repository},url = {https://data.ess-dive.lbl.gov/datasets/doi:10.15485/2527393}}MLA: Willard, Jared, et al. Dataset for "Machine Learning Ensembles Can Enhance Hydrologic Predictions and Uncertainty Quantification". 2025. ESS-DIVE Repository, doi:10.15485/2448016.

54 ENVIRONMENTAL SCIENCES↗

Data and scripts from: “Denoising autoencoder for reconstructing sensor observation data and predicting evapotranspiration: noisy and missing values repair and uncertainty quantification”

This data package includes data and scripts from the manuscript “Denoising autoencoder for reconstructing sensor observation data and predicting evapotranspiration: noisy and missing values repair and uncertainty quantification”.The study addressed common challenges faced in environmental sensing and modeling, including uncertain input data, missing sensor observations, and high-dimensional datasets with interrelated but redundant variables. Point-scaled meteorological and soil sensor observations were perturbed with noises and missing values, and denoising autoencoder (DAE) neural networks were developed to reconstruct the perturbed data and further predict evapotranspiration. This study concluded that (1) the reconstruction quality of each variable depends on its cross-correlation and alignment to the underlying data structure, (2) uncertainties from the models were overall stronger than those from the data corruption, and (3) there was a tradeoff between reducing bias and reducing variance when evaluating the uncertainty of the machine learning models.This package includes:(1) Four ipython scripts (.ipynb): “DAE_train.ipynb” trains and evaluates DAE neural networks, “DAE_predict.ipynb” makes predictions from the trained DAE models, “ET_train.ipynb” trains and evaluates ET prediction neural networks, and “ET_predict.ipynb” makes predictions from trained ET models.(2) One python file (.py): “methods.py” includes all user-defined functions and python codes used in the ipython scripts.(3) A “sub_models” folder that includes five trained DAE neural networks (in pytorch format, .pt), which could be used to ingest input data before being fed to the downstream ET models in ‘ET_train.ipynb” or ‘ET_predict.ipynb’.(4) Two data files (.csv). Daily meteorological, vegetation, and soil data is in “df_data.csv”, where “df_meta.csv” contains the location and time information of “df_data.csv”. Each row (index) in “df_meta.csv” corresponds to each row in “df_data.csv”. These data files are formatted to follow the data structure requirements and be directly used in the ipython scripts, and they have been shuffled chronologically to train machine learning models. The meteorological and soil data was collected using point sensors between 2019-2023 at(4.a) Three shrub-dominated field sites in East River, Colorado (named “ph1”, “ph2” and “sg5” in “df_meta.csv”, where “ph1” and “ph2” were located at PumpHouse Hillslopes, and “sg5” was at Snodgrass Mountain meadow) and(4.b) One outdoor, mesoscale, and herbaceous-dominated experiment in Berkeley, California (named “tb” in “df_meta.csv”, short for Smartsoils Testbed at Lawrence Berkeley National Lab).- See "df_data_dd.csv" and "df_meta_dd.csv" for variable descriptions and the Methods section for additional data processing steps. See "flmd.csv" and "README.txt" for brief file descriptions.- All ipython scripts and python files are written in and require PYTHON language software.

54 ENVIRONMENTAL SCIENCES↗

Samoa Updater: An Application of the Levenberg-Marquardt Method to Update DELFIC Predictions Using Field Measurements

The US Department of Energy (DOE) Forensics Operations (DFO) is a member of the Ground Collections Task Force (GCTF), which is responsible for sample collection of radiological debris for attribution should a nuclear detonation ever occur in the United States. The DFO runs the Defense Land Fallout Interpretive Code (DELFIC) Fallout Planning Tool to predict the deposition of fallout from a nuclear detonation. This prediction is refined using the DELFIC Updater tool, which takes ground measurements and adjusts DELFIC inputs to minimize the difference between prediction and observation, yielding improved predictions of fallout in locations both measured and not yet measured. Samoa, a framework for uncertainty analysis and optimization, is used to improve DELFIC predictive fallout modeling. This new capability using Samoa, dubbed “Samoa Updater,” is compared with the current DELFIC Updater, a brute-force sampling approach. Samoa Updater uses the Levenberg– Marquardt (LM) method, a gradient-based nonlinear least squares approach that uses the functional shape of the input space to increase optimization speed. In simulated test cases Samoa Updater yields faster and more accurate solutions than the current Updater.

45 MILITARY TECHNOLOGY, WEAPONRY, AND NATIONAL DEF↗

Multiscale and Machine Learning Modeling for Process-informed Microstructure Prediction in Additively Manufactured Materials Using MALAMUTE

Advanced Materials and Manufacturing Technologies (AMMT) program under the Department of Energy Office of Nuclear Energy, aims to develop and qualify additively-manufactured materials for nuclear applications. The key challenges to these efforts are the microstructural variabilities observed on the AM products and their impact on the properties and performance of the material in extreme environments. AMMT is using a combination of high-through-put experimental and modeling techniques to accelerate the qualification efforts. Conventionally, in-situ and ex-situ characterizations and testing are performed to correlate different aspects of the AM process to the final product and its performance. However, adopting a trial-and-error approach to experimentally evaluate the vast range of process parameters required to capture the microstructural variabilities is cost-prohibitive. Modeling and simulation provide a comparatively inexpensive way to understand and correlate the microstructural evolution to the processing conditions. The modeling and simulation work-packages within the AMMT program aims to use physics-based and machine learning modeling capabilities to develop a digital twin for AM that can correlate the process conditions to the final product and establish a process-structure-property-performance (PSPP) correlation for AM materials. The melting and subsequent solidification that occurs during the AM process is a complex phenomenon that requires multiscale multiphysics analysis. Idaho National Laboratory’s (INL) Multiphysics Object-Oriented Simulation Environment (MOOSE), specifically the MOOSE Application Library for Advanced Manufacturing UTilitiEs (MALAMUTE) software, provides an ideal platform for developing the multiphysics multiscale model to explore the intricacies of the microstructural evolution during the AM processes within a single framework. Furthermore, given that such full-fidelity simulations can be computationally intensive, reduced order models are necessary to explore the PSPP space for AM materials in an efficient, reliable, and cost-effective way. This work package focuses on understanding the role of process variabilities on the various microstructural characteristics of the AM materials. Microstructures unique to AM materials, such as compositional micro-heterogeneity and dislocation cells, are of particular interest here since they can influence the creep properties and radiation performance. In fiscal year (FY) 24, we significantly advanced upon our work in the last fiscal year, both on physics-based and ML models. The alloy solidification model available in MOOSE has been extended to incorporate the thermodynamic properties and free energy relevant to 316SS. The model demonstrates the Cr segregation that occurs during solidifcation. It is demonstrated that rate of solidification and solute segregation is primarily influence by the cooling rate dictating the level of freezing. This work captures the microstructural variabilities at the subgrain level that are often missing in the part-scale models. With an aim to connect the microstructural evolution model to realistic process conditions, a reduced order model is developed for predicting the thermal conditions around meltpool from high-fidelity process simulations. Furthermore, machine learning approach is used to accelerate the temperature prediction during the AM process. In the following years, MALAMUTE will be used to connect different aspects of the models and quantitatively predict the microstructural evolution. The developed ML-based surrogate model will consider the process conditions as the input to predict the microstructural features in a cost-effective way. The generated microstructures can be used by other work packages under AMMT to evaluate the properties and environmental response of the material at the mesoscale. Thus, this work help identify the key microstructural features at the subgrain level that are significant in property/performance prediction of the AM products. This work will provide inputs to the large-scale process variability models to reevaluate and validate assumptions/simplifications made in the part-scale models. Furthermore, through active learning this work will help identify the data need from both modeling and experimental sides for development of a robust digital twin for AM.

36 MATERIALS SCIENCE↗

A Multiphysics Multiscale Simulation Platform for Damage, Environmental Degradation, and Life Prediction of CMCs in Extreme Environments

This project successfully developed a multiphysics, multiscale computational framework to enhance the design and development of CMCs, with a focus on modeling highly nonlinear, time-dependent damage mechanisms and material degradation under extreme conditions, such as those experienced in turbine service environments. The project made significant advances in improving our understanding of progressive damage, oxidative degradation, and time-dependent inelastic deformation in CMCs, with particular attention to the role of uncertainties in predictions. Key outcomes include the integration of advanced material characterization, uncertainty quantification, and multiphysics constitutive models to predict the behavior of CMCs over their service life. A novel multiscale methodology was employed, which integrated microscale constituent behaviors with structural-scale responses, enabling the manufacturing defects in the microstructure that are prone to damage nucleation. Through the development of DL algorithms, the project advanced the prediction of damage initiation and crack propagation, taking into account the defect morphology and statistical variations across multiple scales. The framework was rigorously validated using thermomechanical experiments, which tested CMCs under various mechanical loadings at elevated temperatures, further enhancing the model's predictive capability. Overall, the research outcomes have provided a more accurate, reliable method for predicting CMC component life, significantly advancing material design, and improving component reliability in extreme environments. This work has strong implications for the optimization of turbine components and other high-performance applications where CMCs are used.

03 NATURAL GAS↗

Materials Characterization, Prediction and Control Project: Summary Report on Data Analytics Framework

This report summarizes the activities performed under the data analytics Vertex in the Materials Characterization, Prediction and Control Project funded under laboratory directed research and development at Pacific Northwest National Laboratory. The data analytics Vertex developed models for associating global or local process parameters, microstructural features, and performance properties of friction-stir-processed 316L stainless steel plates. Statistical, machine learning, and deep learning models, as well as generative artificial intelligence approaches, were used to develop the associations between the process-structure-property data streams. These associations formed the basis for predicting global properties of parts manufactured under different process envelopes, providing a basis for predicting performance using data driven as well as physics-informed and physics-constrained approaches. Additionally, the associations were used to predict local process parameters and microstructural features of the product, predictive relationships that have the potential to form the basis of a control framework that could eventually modulate a friction-stir process to maintain product quality.

316L stainless steel↗

Leveraging Artificial Intelligence to Predict Novel Eutectic Alloys

The goal of this project was to train an artificial neural network (ANN) to predict the fractional composition and melting point of eutectic alloys using fundamental atomic properties as inputs. The fundamental properties considered include atomic number, atomic weight, atomic radius, valence electron concentration, electronegativity, and electron affinity. The project involved several phases, starting with data preparation, where phase diagram data was harvested from the ASM International database. Approximately 1300 binary eutectics were collected and cleaned to ensure relevance and accuracy. A regression model was selected for training, utilizing a rectified linear unit as the activation function. Various model configurations were evaluated for predictive accuracy, with validation techniques employed to ensure robustness. The model demonstrated predictive capabilities above random guessing and was able to achieve up to 11% accuracy under certain conditions. An ablative test identified atomic radius and valence electron concentration as critical inputs for model performance. Incorporating the melting point of atomic constituents improved accuracy significantly, although ultimately the model’s predictive capability still fell short of the 80% target. This report details the methodology, results, and implications of the research, contributing to the understanding of employing artificial intelligence to predict the phase transition behavior of eutectic alloys.

36 MATERIALS SCIENCE↗

Quantitative measurements of dislocations in metals for advancing predictive simulations

LLNL applications require scientists to predict how materials evolve under various thermomechanical conditions. While this is achieved through physics-based simulations, uncertainty in the predictions of mechanical properties remains a serious challenge that limits the predictive capabilities of models because we lack methods to compare predictions of atomic-scale defects (dislocations) with experimental measurements. High energy X-ray diffraction (HEXRD) is the most relevant technique that can provide the necessary statistical information on dislocations. However, this technique is not yet quantitative because we lack a precise understanding of the relationship between X-ray diffraction patterns and the underlying material dislocation content and arrangements. To address this need, we used our novel computational X-ray diffraction method to simulate the effect of dislocations on the diffraction patterns. We compared virtual and experimental diffraction patterns. Results allowed us to clearly establish the relationship between X-ray diffraction patterns and the underlying dislocation structures, proving that it is feasible to quantitatively measure dislocation statistics with HEXRD. This project delivered a method that can provide the missing piece to LLNL’s mechanical property simulations in advanced metals by obtaining experimentally long-needed quantitative dislocation data, which could fully enable predictive capabilities.

36 MATERIALS SCIENCE↗

Multiscale and Machine Learning Modeling for Process-informed Microstructure Prediction in Additively Manufactured Materials using MALAMUTE

The Advanced Materials and Manufacturing Technologies (AMMT) program under the Department of Energy Office of Nuclear Energy aims to develop and qualify additively manufactured materials for nuclear applications. One key challenge to this is the microstructural variability observed in the additively manufactured products and their impact on the properties and performance of the material in extreme environments. AMMT is using a combination of high-throughput experimental and modeling techniques to accelerate qualification. Conventionally, in-situ and ex-situ characterizations and testing are performed to correlate different aspects of the additive manufacturing process to the final product and its performance. However, adopting a trial-and-error approach to experimentally evaluate the vast range of process parameters required to capture microstructural variability is cost-prohibitive. Modeling and simulation provide a comparatively inexpensive way to understand and correlate the microstructural evolution to the processing conditions. The modeling and simulation work-packages within the AMMT program aims to use physics-based and machine learning models to develop a digital twin for additive manufacturing that can correlate the process conditions to the final product and establish a process-structure-property-performance (PSPP) correlation. The melting and subsequent solidification that occurs during the additive process is a complex phenomenon that requires multiscale multiphysics analysis. This work package focuses on understanding the role of process variabilities on the unique microstructural characteristics of additively manufactured materials. Microstructural features at the subgrain level, such as compositional micro-heterogeneity and dislocation cells, are of particular interest here since they can influence the creep properties and radiation performance. Idaho National Laboratory’s Multiphysics Object-Oriented Simulation Environment (MOOSE), specifically the MOOSE Application Library for Advanced Manufacturing UTilitiEs (MALAMUTE) software, provides an ideal platform for developing the multiphysics multiscale model to explore the intricacies of the microstructural evolution during the AM processes within a single framework. Furthermore, given that such full-fidelity simulations can be computationally intensive, reduced order models are necessary to explore the PSPP space for additively manufactured materials in an efficient, reliable, and cost-effective way. This work focuses on capturing the microstructural variabilities at the subgrain level that are often missing in the part-scale models. In fiscal year 2025, we significantly advanced upon our work in the last fiscal year, in terms of the predictive capabilities of the physics-based and ML models, by adding the capabilities to capture subgrain-level micro-segregation during solidification using phase-field model and to predict the time-dependent dynamics of the AM process through the MOGPAR model. The alloy solidification model in MOOSE incorporates the thermodynamic properties and free energy relevant to 316 stainless steel. The model demonstrates the Cr and Ni segregation that occurs during solidification, including that the rate of solidification. The microstructural evolution model is connected to the process conditions via the surrogate model developed in this work. This enables predictions of the final microstructure in conjunctions with the manufacturing process. This work supports AMMT's rapid qualification goals by laying the foundation for an efficient and cost-effective model establishing the PSPP correlation for AM. The generated microstructures and predicted micro-segregation can be used by other work packages under AMMT to evaluate the properties and environmental response of the material at the mesoscale. Thus, this work helps to identify the key microstructural features at the subgrain level that are significant in property and performance predictions of additively manufactured components. This work will also provide inputs to the large-scale process variability models to reevaluate and validate assumptions and simplifications made in the part-scale models. Furthermore, through active learning this work can help identify the data need from both modeling and experimental sides for development of a robust digital twin for additive manufacturing and accelerate the AMMT's qualification efforts.

36 - MATERIALS SCIENCE↗

Protein-ligand binding affinity prediction using multi-instance learning with docking structures

Recent advances in 3D structure-based deep learning approaches demonstrate improved accuracy in predicting protein-ligand binding affinity in drug discovery. These methods complement physics-based computational modeling such as molecular docking for virtual high-throughput screening. Despite recent advances and improved predictive performance, most methods in this category primarily rely on utilizing co-crystal complex structures and experimentally measured binding affinities as both input and output data for model training. Nevertheless, co-crystal complex structures are not readily available and the inaccurate predicted structures from molecular docking can degrade the accuracy of the machine learning methods. We introduce a novel structure-based inference method utilizing multiple molecular docking poses for each complex entity. Our proposed method employs multi-instance learning with an attention network to predict binding affinity from a collection of docking poses. We validate our method using multiple datasets, including PDBbind and compounds targeting the main protease of SARS-CoV-2. The results demonstrate that our method leveraging docking poses is competitive with other state-of-the-art inference models that depend on co-crystal structures. This method offers binding affinity prediction without requiring co-crystal structures, thereby increasing its applicability to protein targets lacking such data.

97 MATHEMATICS AND COMPUTING↗

Data from: Coupled machine learning-ecosystem ensemble models substantially improve predictions of nitrous oxide (N 2 O) fluxes from US croplands

Nitrous oxide (N₂O) is a potent and persistent greenhouse gas, with rising atmospheric concentrations driven in part by inefficient use of synthetic nitrogen (N) fertilizers in agriculture. Predicting soil N₂O emissions is challenging due to high spatial and temporal variability arising from complex soil biogeochemical processes. Process-based ecosystem models and standalone machine learning (ML) approaches without extensive site-specific calibration often miss high emission episodes. Here, we show how an Ensemble Modeling System (EMS) based on outputs from an ensemble of ecosystem models coupled to an ensemble of ML models can improve predictions and understanding of N2O fluxes from US cropland. Trained and validated on approximately 12,000 N2O chamber measurements at 17 U.S. Midwest sites (six crops, 35 management practices), the EMS accurately predicted daily fluxes of N2O at both training (R² = 0.84, RMSE = 16.4 g N ha⁻¹ d⁻¹) and held-out testing sites (R² = 0.84, RMSE = 6.2 g N ha⁻¹ d⁻¹). Analyses identified six dominant N₂O drivers: soil organic carbon (SOC), NH₄⁺, NO₃⁻, water-filled pore space (WFPS), soil temperature, and biomass production. Wet, warm soils produced large N₂O peaks only with sufficient SOC and mineral N; in low-SOC soils, fluxes remained low. Incorporating these drivers into process-based models might significantly improve their predictive capacity. The EMS demonstrates a strong potential to predict N₂O fluxes at unseen sites, enabling more reliable regional inventories, improved gap-filling where measurements are sparse, and enhanced understanding of mechanisms to advance targeted mitigation strategies in food, feed, and bioenergy crops.

agricultural sciences↗

Thermodynamics-guided machine learning model for predicting convective boundary layer height and its multi-site applicability

Accurate estimation of convective boundary layer height (CBLH) is vital for weather, climate, and air quality modeling. Machine learning (ML) shows promise in CBLH prediction, but input parameter selection often lacks physical grounding, limiting generalizability. This study introduces a novel ML framework for CBLH prediction, integrating thermodynamic constraints and the diurnal CBLH cycle as an implicit physical guide. Boundary layer growth is modeled as driven by surface heat fluxes and atmospheric heat absorption represented with the low tropospheric stability, using the diurnal cycle as input and output. TPOT and AutoKeras are employed to select optimal models, validated against Doppler lidar-derived CBLH data, achieving an R 2 of 0.84 across untrained years. Comparisons of eddy covariance (ECOR) and energy balance Bowen ratio (EBBR) flux measurements show the same prediction capability. Models trained on the ARM SGP C1 site with ECOR data and tested at E37 and E39 yield R 2 values of 0.79 and 0.81, respectively, demonstrating their adaptability. The ML model trained with all sites' data slightly enhances the performance compared with ML models trained over single-site data. The interquartile range for predicted CBLH is consistently narrower than that for DL-derived CBLH, reflecting lower variability in predicted CBLH compared to DL-derived CBLH, which is influenced by additional factors, which are not well represented with the model inputs. The model's generalizability across multiple sites at the ARM SGP site demonstrates its potential for transfer to greater distances, offering a scalable approach for enhancing boundary layer parameterization in atmospheric models.

Chu, Yufei [Stony Brook Univ., NY (United States)]↗

Crime prediction modeling

A study of techniques for the prediction of crime in the City of Los Angeles was conducted. Alternative approaches to crime prediction (causal, quasicausal, associative, extrapolative, and pattern-recognition models) are discussed, as is the environment within which predictions were desired for the immediate application. The decision was made to use time series (extrapolative) models to produce the desired predictions. The characteristics of the data and the procedure used to choose equations for the extrapolations are discussed. The usefulness of different functional forms (constant, quadratic, and exponential forms) and of different parameter estimation techniques (multiple regression and multiple exponential smoothing) are compared, and the quality of the resultant predictions is assessed.

Source record↗

Demonstration of the range over which the Langley Research Center digital computer charring ablation program (CHAP) can be used with confidence: Comparisons of CHAP predictions and test data for three ablation materials

Comparisons of ablation calculations with the charring ablation computer code and ablation test data are presented over a wide range of environmental conditions in air for three materials: low-density nylon phenolic, Avcoat 5026-39HC/G, and a filled silicon elastomer. Heat fluxes considered range from over 500 Btu/sq ft-sec to less than 50 Btu/sq ft-sec. Pressures range from 0.5 atm to .004 atm. Enthalpies range from about 2000 Btu/lb to 18000 Btu/lb. Predictions of recession, pyrolysis penetration, and thermocouple responses are considered. Recession predictions for nylon phenolic are good as steady state is approached, but strongly transient cases are underpredicted. Pyrolysis penetrations and thermocouple responses are very well predicted. Recession amounts for Avcoat and silicone elastomer are less well predicted, although high heat flux cases near steady state are fairly satisfactory. Pyrolysis penetrations and thermocouple responses are very well predicted.

Moyer, C. B.↗