Search NASA⌕ Search

SEARCH · Search NASA

Results for “molecular properties”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

PySIDT: Subgraph Isomorphic Decision Trees for Molecular Property Prediction

Accurate molecular property prediction is important across all fields of chemistry. Deep neural networks (DNNs) have become increasingly popular due to their ability to train automatically, avoiding the incredibly tedious process of constructing and extending traditional property estimation schemes. However, DNNs require large amounts of training data, are challenging to interpret, require large amounts of memory to load even during inference, and have severe difficulties incorporating qualitative chemical knowledge, which are often desired for molecular property prediction tasks. Here, in this study, we present PySIDT (https://github.com/zadorlab/PySIDT), a software for training and running inference on Subgraph Isomorphic Decision Trees (SIDTs). SIDTs are graph-based decision trees made of nodes associated with molecular substructures. Inference is done by descending target molecular structures down the decision tree to nodes with matching subgraph isomorphic substructures and making predictions based on the final (most specific) nodes matched. SIDTs scale down well to dataset sizes much smaller than is feasible for DNNs. As trees of molecular substructures, SIDTs are inherently readable and easy to visualize, making them easy to analyze. They are also straightforward to extend and retrain, facilitate uncertainty estimation, and enable easy integration of expert knowledge. We demonstrate the SIDT approach discussing its application to a diverse range of molecular prediction tasks: rate coefficient estimation, diffusion coefficient estimation, thermochemistry estimation, transition state bond stretch prediction, p K a prediction, stability of molecular structures, stability of surface structures, and prediction of surface lateral interaction energetics. Additionally, we demonstrate the power of the SIDT algorithms in two direct learning curve vanilla comparisons with the popular DNN-based software Chemprop and the popular gradient boosted trees-based software XGBoost on enthalpy of formation and rate coefficient prediction tasks. In particular, in the enthalpy of formation case, vanilla PySIDT is able to outperform vanilla Chemprop and XGBoost across the full range of training/validation set sizes out to 11,560 data points.

Johnson, Matthew Sean [Sandia National Laboratorie↗

Active Causal Machine Learning for Molecular Property Prediction

Predicting properties from molecular structures is paramount to design tasks in medicine, materials science, and environmental management. However, design rules derived from the structure-property relationships using correlative data-driven methods fail to elucidate underlying causal mechanisms controlling chemical phenomena. This preliminary work proposes a workflow to actively learn robust cause-effect relations between structural features and molecular property for a broad chemical space utilizing smaller subsets, entailing partial information.

Fox, Zach↗

A Deep Multimodal Representation Learning Framework for Accurate Molecular Properties Prediction

Drug discovery is a complex and challenging process, requiring the optimization of candidate compounds to identify those with the potential to become safe and effective drugs. Predicting molecular properties is an indispensable step in the drug discovery pipeline. Traditionally, this process is costly and time-intensive, involving multiple rounds of experiments and clinical trials, rendering it impractical for every candidate compound. Deep learning techniques have emerged as a promising approach to drug discovery to reduce the cost and time required to identify novel drugs. However, prevalent research in deep learning models focused on predicting molecular properties has primarily fixated on single-modal models, which utilize a single modality of data, neglecting the potential benefits of combining different data modalities. To overcome this limitation, we introduce MRL-Mol: a deep \textbf{M}ultimodal \textbf{R}epresentation \textbf{L}earning framework for accurate \textbf{Mol}ecular properties prediction. MRL-Mol harnesses three data modalities: sequence, graph, and image, augmenting the depth of comprehension. Leveraging a large-scale unlabeled dataset~($\sim$1M unique molecules), we pretrain MRL-Mol to extract inter- and intra-modal information. Our study demonstrates the superior performance of MRL-Mol in predicting molecular properties across six benchmark datasets, including both classification and regression tasks. Notably, MRL-Mol outperforms other state-of-the-art molecular properties prediction models. These findings suggest that by combining information from multiple data modalities, MRL-Mol can comprehend molecules better than single-modal deep learning models and identify molecular properties with better accuracy.

Yang, Yuxin↗

Dashboard for Visualizing Molecular Property Prediction Machine Learning Results

This is a dashboard for exploring the results of machine learning models for predicting molecular properties from molecular structure. It includes tools for: 1. Modifying molecules to observe the change in predicted properties 2. Exploring the relationship between molecular structure and predicted properties 3. Recommending structurally similar molecules with improved properties 4. Exploring the impact of data subsampling on model performance metrics

Xu, Audrey↗

Evaluating uncertainty-based active learning for accelerating the generalization of molecular property prediction

Deep learning models have proven to be a powerful tool for the prediction of molecular properties for applications including drug design and the development of energy storage materials. However, in order to learn accurate and robust structure–property mappings, these models require large amounts of data which can be a challenge to collect given the time and resource-intensive nature of experimental material characterization efforts. Additionally, such models fail to generalize to new types of molecular structures that were not included in the model training data. The acceleration of material development through uncertainty-guided experimental design has the promise to significantly reduce the data requirements and enable faster generalization to new types of materials. To evaluate the potential of such approaches for electrolyte design applications, we perform comprehensive evaluation of existing uncertainty quantification methods on the prediction of two relevant molecular properties - aqueous solubility and redox potential. We develop novel evaluation methods to probe the utility of the uncertainty estimates for both in-domain and out-of-domain data sets. Finally, we leverage selected uncertainty estimation methods for active learning to evaluate their capacity to support experimental design.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Active deep kernel learning of molecular properties from structural embeddings

As vast databases of chemical identities become increasingly available, the challenge shifts to how we effectively explore and leverage these resources to study molecular properties. This paper presents an active learning approach for molecular discovery using deep kernel learning (DKL), demonstrated on the QM9 dataset. DKL links structural embeddings directly to properties, creating organized latent spaces that prioritize relevant property information. By iteratively recalculating embedding vectors in alignment with target properties, DKL uncovers concentrated maxima representing key molecular properties and reveals unexplored regions with potential for innovation. This approach underscores DKL’s potential in advancing molecular research and discovery.

Artificial neural networks↗

Uncertainty quantification for molecular property predictions with graph neural architecture search

Graph Neural Networks (GNNs) have emerged as a prominent class of data-driven methods for molecular property prediction. However, a key limitation of typical GNN models is their inability to quantify uncertainties in the predictions. This capability is crucial for ensuring the trustworthy use and deployment of models in downstream tasks. To that end, we introduce AutoGNNUQ, an automated uncertainty quantification (UQ) approach for molecular property prediction. AutoGNNUQ leverages architecture search to generate an ensemble of high-performing GNNs, enabling the estimation of predictive uncertainties. Our approach employs variance decomposition to separate data (aleatoric) and model (epistemic) uncertainties, providing valuable insights for reducing them. In our computational experiments, we demonstrate that AutoGNNUQ outperforms existing UQ methods in terms of both prediction accuracy and UQ performance on multiple benchmark datasets, and generalizes well to out-of-distribution datasets. Additionally, we utilize t-SNE visualization to explore correlations between molecular features and uncertainty, offering insight for dataset improvement. AutoGNNUQ has broad applicability in domains such as drug discovery and materials science, where accurate uncertainty quantification is crucial for decision-making.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Multitask methods for predicting molecular properties from heterogeneous data

Data generation remains a bottleneck in training surrogate models to predict molecular properties. We demonstrate that multitask Gaussian process regression overcomes this limitation by leveraging both expensive and cheap data sources. In particular, we consider training sets constructed from coupled-cluster (CC) and density functional theory (DFT) data. We report that multitask surrogates can predict at CC-level accuracy with a reduction in data generation cost by over an order of magnitude. Of note, our approach allows the training set to include DFT data generated by a heterogeneous mix of exchange–correlation functionals without imposing any artificial hierarchy on functional accuracy. More generally, the multitask framework can accommodate a wider range of training set structures—including the full disparity between the different levels of fidelity—than existing kernel approaches based on Δ-learning although we show that the accuracy of the two approaches can be similar. Consequently, multitask regression can be a tool for reducing data generation costs even further by opportunistically exploiting existing data sources.

Chemistry↗

Molecular Properties and Chemical Transformations Near Interfaces

The properties of bulk water and aqueous solutions are known to change in the vicinity of an interface and/or in a confined environment, including the thermodynamics of ion selectivity at interfaces, transition states and pathways of chemical reactions, and nucleation events and phase growth. Here we describe joint progress in identifying unifying concepts about how air, liquid, and solid interfaces can alter molecular properties and chemical reactivity compared to bulk water and multicomponent solutions. Furthermore, we also discuss progress made in interfacial chemistry through advancements in new theory, molecular simulation, and experiments.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Ecological theory applied to environmental metabolomes reveals compositional divergence despite conserved molecular properties

Stream and river systems transport and process substantial amounts of dissolved organic matter (DOM) from terrestrial and aquatic sources to the ocean, with global biogeochemical implications. However, the underlying mechanisms affecting the spatiotemporal organization of DOM composition are under-investigated. To understand the principles governing DOM composition, we leverage the recently proposed synthesis of metacommunity ecology and metabolomics, termed ‘meta-metabolome ecology.’ Applying this novel approach to a freshwater ecosystem, we demonstrated that despite similar molecular properties across metabolomes, metabolite identity significantly diverged due to environmental filtering and variations in putative biochemical transformations. We refer to this phenomenon as ‘thermodynamic redundancy,’ which is analogous to the ecological concept of functional redundancy. We suggest that under thermodynamic redundancy, divergent metabolomes can support equivalent biogeochemical function just as divergent ecological communities can support equivalent ecosystem function. As these analyses are performed in additional ecosystems, potentially generalizable principles, like thermodynamic redundancy, can be revealed and provide insight into DOM dynamics.

Danczak, Robert E.↗

Q4Q: Quantum Computation for Quantum Prediction of Materials and Molecular Properties

We aim at exploring the potential of quantum computation to solve practical problems that are currently targeted by “traditional” high performance computing (HPC). Along the way, we will develop theoretical frameworks and algorithmic strategies. We propose computational research activities that will advance the application of quantum computers to selected materials/molecular properties: (1) Band structure of solids; (2) Molecular electronic/vibronic properties; (3) Free energies and phase transitions.

36 MATERIALS SCIENCE↗

Theoretical studies of chemical reactions related to the formation and growth of polycyclic aromatic hydrocarbons (PAH) and molecular properties of their key intermediates (Final Progress Report)

The formation mechanisms of polycyclic aromatic hydrocarbons, (PAHs) – organic molecules carrying fused benzene rings – are of great interest to scientists and engineers due to their importance in combustion chemistry and astrochemistry. On Earth, PAHs are largely produced in incomplete combustion of fossil fuel and are considered as critical precursors to unwanted soot particles leading to combustion inefficiency and causing air pollution along with detrimental health effects. Simple PAH molecules initially formed in the gas phase, are further involved in a build-up process in combustion flames leading to larger PAH, bowl-shaped nanostructures, fullerenes, and solid-phase species including carbonaceous dust, graphene particles, and soot. In deep space, PAH and their derivatives are potential key intermediates and nucleation sites leading eventually to carbonaceous nanoparticles (“interstellar grains”). Therefore, the understanding of the key processes in the synthesis of PAHs along with their precursors and their degradation mechanisms in combustion systems and in interstellar, circumstellar, and planetary atmospheric environments will provide critical insights into how complex aromatic structures, carbonaceous nanoparticles, and fullerenes are formed and destroyed. Achieving this understanding is an important step in the development of the efficient combustion processes and of the ecofriendly devices with reduced environmental pollution as well as technological strategies for the production of hydrogen and solid carbon through thermal or plasma-assisted pyrolysis of natural gas and biomass. Also, the understanding of the key processes of PAH and soot growth will help in our comprehension of chemical evolution in the universe. Detailed information on the mechanisms and reliable rate constants of the key elementary chemical reactions involved in PAH formation and destruction processes and in inception of soot particles is often missing, with the main deficiencies being the absence of temperature- and pressure-dependent rate constants for the broad range of conditions occurring in various terrestrial and interstellar processes and the lack of data on the reaction products and their branching ratios. Complementary to experimental studies, these gaps in knowledge can be filled by using quantum chemical calculations of reaction potential energy surfaces providing us with accurate energies of reaction products, intermediates, and transition states, revealing the reaction mechanism, and giving the molecular properties required to compute rate constants for relevant reaction steps and product branching ratios using the RRKM-Master Equation (ME) method. Molecular dynamics (MD) simulations can be used in cases when a reaction rate cannot be properly described by statistical theories. During the terminal renewal project period we employed these ab initio/RRKM-ME and MD approaches to complete our studies on several key reactions relevant to the formation/growth of PAH and inception of soot particles including (1) the reaction mechanism and kinetics of the resonance stabilized fulvenallenyl radical with propargyl and C 3 H 4 isomers; (2) the reaction mechanism and kinetics for the C + indene and C 2 + styrene reactions producing naphthyl or azulenyl radicals in low-temperature environments; (3) the MD study of non-equilibrium dimerization of acepyrene and coronene and its radical. The information derived from our theoretical calculations contributed to a better fundamental understanding of the reaction mechanisms and provide missing critical kinetic data to improve combustion models of hydrocarbon fuels and astrochemical models of the growth of carbonaceous molecules and particles in cold molecular clouds, circumstellar envelopes, and planetary atmospheres.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Relating Molecular Properties to the Persistence of Marine Dissolved Organic Matter with Liquid Chromatography–Ultrahigh-Resolution Mass Spectrometry

Marine dissolved organic matter (DOM) contains a complex mixture of small molecules that eludes rapid biological degradation. Spatial and temporal variations in the abundance of DOM reflect the existence of fractions that are removed from the ocean over different time scales, ranging from seconds to millennia. However, it remains unknown whether the intrinsic chemical properties of these organic components relate to their persistence. Here, we elucidate and compare the molecular compositions of distinct DOM fractions with different lability along a water column in the North Atlantic Gyre. Our analysis utilized ultrahigh-resolution Fourier transform ion cyclotron resonance mass spectrometry at 21 T coupled to liquid chromatography and a novel data pipeline developed in CoreMS that generates molecular formula assignments and metrics of isomeric complexity. Clustering analysis binned 14 857 distinct molecular components into groups that correspond to the depth distribution of semilabile, semirefractory, and refractory fractions of DOM. The more labile fractions were concentrated near the ocean surface and contained more aliphatic, hydrophobic, and reduced molecules than the refractory fraction, which occurred uniformly throughout the water column. These findings suggest that processes that selectively remove hydrophobic compounds, such as aggregation and particle sorption, contribute to variable removal rates of marine DOM.

54 ENVIRONMENTAL SCIENCES↗

Assessing the effect of regularization on the molecular properties predicted by SCAN and self-interaction corrected SCAN meta-GGA

Recent regularization of the SCAN meta-GGA functional (rSCAN) has simplified the numerical complexities of the SCAN functional, alleviating SCAN's stringent demand on the numerical integration grids to some extent. The regularization of rSCAN, however, results in the breaking of some constraints such as the uniform electron gas limit, the slowly varying density limit, and coordinate scaling of the iso-orbital indicator. Here, we assess the effects of regularization on the electronic, structural, vibrational, and magnetic properties of molecules by comparing the SCAN and rSCAN predictions. The properties studied include atomic energies, atomization energies, ionization potentials, electron affinities, barrier heights, infrared intensities, dissociation and reaction energies, spin moments of molecular magnets, and isomer ordering of water clusters. Our results show that rSCAN requires less dense numerical grids and gives very similar results to those of SCAN for all properties examined with the exception of atomization energies, which are worsened in rSCAN. We also examine the performance of self-interaction-corrected (SIC) rSCAN with respect to SIC-SCAN using the Perdew–Zunger (PZ) SIC method. The PZSIC method uses orbital densities to compute one-electron self-interaction errors and places an even more stringent demand on numerical grids. Our results show that SIC-rSCAN gives marginally better performance than SIC-SCAN for almost all properties studied in this work with numerical grids that are on average half or less as dense as that needed for SIC-SCAN.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Molecular property prediction for very large databases with natural language processing: a case study in ionic liquid design

The prospect of using artificial intelligence (AI) to accurately screen very large databases of compounds for multiple properties has yet to be realized. Here, we explore this possibility using ionic liquids (ILs) which offer unique physicochemical properties and excellent tunability, making them highly versatile solvents for various research applications. Screening millions of potential ILs for the best perfomance for use in specific tasks with experimental methods alone however, is impractical. Further, traditional’ physics-based computational chemistry is hindered by high computational cost. To address this challenge, we leverage a natural language processing (NLP)-based molecular embedding technique with advanced machine learning (ML) models to predict seven key IL properties: viscosity, density, ionic conductivity, surface tension, melting temperature, toxicity, and water solubility. Comprehensive datasets for these properties are obtained, then NLP featurization with Mol2vec is compared with other featurization techniques such as 2D Morgan fingerprints, and 3D quantum chemistry-derived sigma profiles. NLP-based featurization exhibited the best predictive performance, achieving the highest R 2 and lowest RMSE values for all the studied IL properties. Further, we present case studies of how ILs might be screened using combined property criteria for practical cases – lignocellulosic biomass processing, CO 2 capture, and optimal electrolytes for batteries – screening a novel database of ∼10.6 million generated feasible ILs. The results introduce NLP as a powerful tool for engineering many designer solvents with desirable properties for task specific applications.

Mohan, Mood [Oak Ridge National Laboratory (ORNL),↗

YAP1 Dysfunction Promotes Molecular Properties Linked to Breast Cancer Susceptibility

YAP1 is a cotranscription factor that promotes malignant and stem cell properties in cancer. We previously found that YAP1 dysregulation is associated with aging in human mammary epithelia. With increased age, YAP1 expression changes in luminal epithelial cells, the prospective breast cancer cell of origin. Because age is a significant risk factor for breast cancer, we tested whether YAP1 dysregulation acted early in cancer progression by conferring cellular states associated with increased cancer susceptibility. In this study, we find that with increased age and genetic risk for developing cancer, human breast tissues showed significantly increased YAP1 expression, and cultured primary human mammary epithelial cells (HMEC) showed significantly increased expression of both YAP1 and its transcriptional targets. Increased YAP1 expression in cultured HMEC induced gene expression changes associated with increased cancer susceptibility, such as genes associated with stem cell states, increased telomerase activity, breast cancer progression, and increased age and genetic breast cancer risk. Furthermore, overexpression of YAP1 in post-stasis HMEC—finite lifespan cells that have bypassed a retinoblastoma-mediated senescence barrier—promoted properties related to increased growth potential. We found that YAP1 dysregulation in finite epithelial cells allows for access to gene programs and functions that are typically thought to be restricted to stem cells. We hypothesize that YAP1 acts early in breast cancer progression, long before the development of a tumor, to impose cancer-susceptible molecular states.

YAP1↗