Search NASA⌕ Search

SEARCH · Search NASA

Results for “SMILES”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

Ensemble Spread Behavior in Coupled Climate Models: Insights From the Energy Exascale Earth System Model Version 1 Large Ensemble

AbstractAssessing uncertainty in future climate projections requires understanding both internal climate variability and external forcing. For this reason, single‐model initial condition large ensembles (SMILEs) run with Earth System Models (ESMs) have recently become popular. Here we present a new 20‐member SMILE with the Energy Exascale Earth System Model version 1 (E3SMv1‐LE), which uses a “macro” initialization strategy choosing coupled atmosphere/ocean states based on inter‐basin contrasts in ocean heat content (OHC). The E3SMv1‐LE simulates tropical climate variability well, albeit with a muted warming trend over the twentieth century due to overly strong aerosol forcing. The E3SMv1‐LE's initial climate spread is comparable to other (larger) SMILEs, suggesting that maximizing inter‐basin ocean heat contrasts may be an efficient method of generating ensemble spread. We also compare different ensemble spread across multiple SMILEs, using surface air temperature and OHC. The Community Earth system Model version 1, the only ensemble which utilizes a “micro” initialization approach perturbing only atmospheric initial conditions, yields lower spread in the first ∼30 years. The E3SMv1‐LE exhibits a relatively large spread, with some evidence for anthropogenic forcing influencing spread in the late twentieth century. However, systematic effects of differing “macro” initialization strategies are difficult to detect, possibly resulting from differing model physics or responses to external forcing. Notably, the method of standardizing results affects ensemble spread: control simulations for most models have either large background trends or multi‐centennial variability in OHC. This spurious disequlibrium behavior is a substantial roadblock to understanding both internal climate variability and its response to forcing.

Stevenson, Samantha↗

Thermochemical Data for Furan-based Monomer Candidates for Frontal Ring-Opening Metathesis Polymerization (FROMP)

This dataset includes 471 furan-based monomer candidates for frontal ring-opening metathesis polymerization (FROMP) and relevant thermochemistry as calculated with density functional theory (DFT). The monomer candidates were combinatorically enumerated using Diels-Alder reactions of furan derivatives as dienes and four types of dienophiles (alkenes, alkynes, allenes, and benzynes). Common substituents were enumerated for the dienophile classes, and methyl substitution on the diene was explored. We used the SMILES arbitrary target specification (SMARTS) language to produce monomers and ring-opened structures from diene and dienophile precursor SMILES, and we studied the ring-opening reaction using a homodesmotic equation with ethene. RDKit conformers were initially generated from SMILES, then optimized with GFN2-xTB. The two conformers lowest in energy were then optimized with DFT using the wb97x-D3 functional, def2-TZVP basis set, and def2/J auxiliary basis set. Gibbs free energy corrections were obtained through frequency calculations. Structures with imaginary frequencies below -50 cm^{-1} were excluded from this work, and smaller imaginary modes were flipped to be positive for free energy calculations. Modes below 50 cm^{-1} were treated with the modified rigid rotor approximation, and all thermochemical values were calculated at T=200C. The CSV file contains the monomer SMILES, the free energy of reaction for Diels-Alder addition (G_DA_200), and the enthalpy of the ring-opening reaction (H_RO_200). All energies are given in kcal/mol. An interactive HTML is also included to visualize the monomers in this dataset.

Chua, Lauren↗

A Comprehensive Machine Learning Model for Metal–Ligand Binding Prediction: Applications in Chemistry and Biology

A machine-learning (ML) model that predicts metal–ligand binding constants was developed using the open-source Chemprop software. The model was trained on over 30,000 experimental log K 1 values, which include both protonation and metal–ligand stability constants, comprising over 3500 ligands and 10 2 metal ions from 73 total elements, thus generalizing beyond existing limited approaches, which focus only on specific metals or ligand families. The best-performing model included a combination of SMILES-based molecular representations along with descriptors for the metal ion and experimental conditions. It had an external test R 2 value of 0.942, and MAE value of 0.834. A “SMILES-only” simpler version also produced accurate predictions and preserved the binding trends, serving as a quick and easily accessible alternative for users without computational expertise. The SMILES-only model performed comparably to density functional theory (DFT) calculations but utilized a fraction of the computational resources. The model was successfully applied across diverse domains, including bioinorganic chemistry, heavy metal remediation, and sensor development and demonstrated its effectiveness as a rapid and reliable screening tool for both academic and industrial uses.

Ligands↗

Evaluating the Use of Foundational Chemical Language Models in Multimodal Graph Fusion

Rapid and accurate prediction of the physicochemical properties of molecules given their structures remains a key challenge in cheminformatics. Machine learning approaches offer high-throughput options, but the optimality of inductive biases and data representations are up for debate. For example, BERT-based masked language models (MLMs) can be trained in a self-supervised way on hundreds of millions to billions of readily available SMILES strings. Another option is graph neural networks (GNNs), which can operate directly on molecular structures. Yet, generating accurate molecular geometry is computationally expensive, leading to a relative scarcity in data compared to SMILES strings. It is attractive to combine these two paradigms by pre-training an LM on a large corpus of SMILES strings and embedding these representation into a geometric graph neural network. Despite the promise of such an approach, and contrary to previous studies, we find mixed results with the combination of the LMs and GNNs on several molecule datasets. In particular, we found evidence for improvement on the FreeSolv and QM7 benchmarks, but degraded performance on the ESOL, LIPO and QM9 datasets compared to a GNN baseline.

Francel, Collin [University of Alabama]↗

Comparison of Machine Learning Approaches for Prediction of the Equivalent Alkane Carbon Number for Microemulsions Based on Molecular Properties

The chemical properties of oils are vital in the design of microemulsion systems. The hydrophilic–lipophilic difference equation used to predict microemulsions’ phase behavior expresses the oils’ physiochemical properties as the equivalent alkane carbon number (EACN). The experimental determination of EACN requires knowledge of the temperature dependence of the microemulsion system and the effects of different surfactant concentrations. Thus, the experimental determination is time-intensive and tedious, requiring days to months for proper separations. Furthermore, the experiments require high purity of chemicals because microemulsions are sensitive to impurities. Our work focuses on the quick and reliable predictions of the EACN with machine learning (ML) models. Due to the immaturity of ML chemical predictions, we compare three graph neural networks (GNNs) and a gradient-boosted tree algorithm, known as XGBoost. The GNNs use the molecular structures represented as simplified molecular-input line-entry system (SMILES) codes for the initial input, which allows us to assess whether geometry optimization is necessary for reliable results. The XGBoost model also begins with the SMILES representations of the molecules but uses molecular descriptors instead of geometry optimizations. As a result, the best model tested (crystal graph convolutional neural network with Merck molecular force field-94) has an error of 1.15 EACN units of the true EACN for unknown data with the errors skewed toward zero and an R² score of 0.9

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Changing windstorm characteristics over the US Northeast in a single model large ensemble

Abstract Extreme windstorms pose a significant hazard to infrastructure and public safety, particularly in the highly populated US Northeast (NE). However, the influence climate change and changing land use will have on these events remains unclear. A large ensemble generated using the Max-Planck Institute (MPI) Earth system model is used to generate projections of NE windstorms under different shared socioeconomic pathways (SSPs) and to attribute changes to projected land use land cover (LULC) change, externally forced changes and internal climate variability. To reduce the influence of coarse grid cell resolution and uncertainties in surface roughness lengths, windstorms are identified using simultaneous widespread exceedance of local 99th percentile 10 m wind speeds (U 99 ). Projected declines in forest cover in the NE and the resulting reductions in surface roughness length under SSP3-7.0 lead to projections of large increases in U 99 and derived windstorm intensity and scale. However, these projected changes in regional LULC under SSP3-7.0 are unprecedented in a historical context and may not be realistic. After corrections are applied to remove the influence of LULC on wind speeds, regionally averaged U 99 exhibit declines for most of the single model initial-condition large ensemble (SMILE) members which are broadly proportional to the radiative forcing and global air temperature increase in the SSPs, with a median value of −0.15 ms −1 °C −1 . While weak cyclones are projected to decline in frequency in the NE, intense cyclones and the resulting windstorms and indices of socioeconomic loss do not. Where present, significant trends in these loss indices are positive, and some MPI SMILE members generate future windstorms that are unprecedented in the historical period.

Coburn, Jacob (ORCID:0000000309538117)↗

Evaluation of historical precipitation interannual variability in CMIP6 over the United States

Interannual precipitation variability profoundly influences society via its effects on agriculture, water resources, infrastructure, and disaster risks. In this study, we use daily in situ precipitation observations from the global historical climatology network-daily (GHCN-D) to assess the ability of 21 Coupled Model Intercomparison Project Phase 6 (CMIP6) models, including the 50-member fifth-generation Canadian Earth System Model single model initial-condition large ensemble (CanESM5_SMILE), to realistically simulate historical interannual precipitation variability trends within 17 regions of the contiguous United States (CONUS). We assess how accurately the CMIP6 simulations align with observational data across annual, summer, and winter periods, focusing on four key hydrometeorological metrics, including interannual precipitation variability, relative interannual precipitation variability (coefficient of variation), annual mean precipitation, and annual wet day frequency. Our findings reveal that CMIP6 ensemble members generally reproduce the spatial patterns of observed trends in annual mean precipitation. In most regions, models agree well with the signs of observed changes in annual mean precipitation, though discrepancies in trend magnitude are evident. Further, observed trends in winter mean precipitation broadly exhibit a spatial pattern similar to that of the observed annual mean. However, analysis of the CanESM5_SMILE shows that trends in precipitation variability may primarily be the result of model-simulated internal variability, suggesting caution in interpreting multi-model single-realization ensemble results. Challenges in accurately simulating interannual precipitation variability underscore the need for ongoing model refinement and validation to enhance climate projections, especially in regions vulnerable to extreme precipitation events.

54 ENVIRONMENTAL SCIENCES↗

A Variational Autoencoder Model Toward Molecular Structure Representation Learning of Fuels

Here, in this work, a Variational Autoencoder (VAE)-based data-driven modeling framework is developed with the overarching goal of enabling fuel design. The VAE model is trained on a large dataset with several chemical species to learn a compressed latent space molecular representation. Chemical structure in the form of Simplified Molecular Input Line Entry System (SMILES) string is fed as input, encoded into the VAE latent space, and decoded back to the SMILES string using Long Short-Term Memory (LSTM) networks. Complexities of the VAE training loss function are thoroughly examined by varying the weightage (beta (𝜷) parameter) of the latent space regularization term, thereby assessing the balance between reconstruction accuracy and validity, and focusing on both accurate molecular structure reconstruction and latent space consistency. Two different strategies for 𝜷 variation are evaluated: linear annealing and cyclic annealing. In addition, the impact of total correlation adjustment and hierarchical priors is also studied with regard to the balance between reconstruction fidelity and latent space regularization, and potential issues such as posterior collapse, over-regularization, and poor disentanglement of latent variables. Overall, the best performance of the model is achieved with hierarchical priors and incrementally increasing 𝜷 from 0 to a threshold value of 0.25 over 75 epochs. The generative VAE model can be readily coupled with Quantitative Structure–Property Relationship (QSPR) analysis to develop an integrated end-to-end framework for fuel-property prediction and molecular design of novel promising fuels.

fuel design↗

Equivariant Graph Attention Network - 3D Conformers & Feature Fusion

EGAN-3F (Equivariant Graph Attention Network - 3D Conformers & Feature Fusion) presents an innovative approach for predicting binding affinity between small molecules and protein targets, a fundamental task in drug discovery. Traditional structure-based methods often depend on protein-ligand complex structures obtained from crystallography or molecular docking. In contrast, ligand-only machine learning models using 1D or 2D representations such as SMILES have been developed to predict binding affinity without structural information about the target; however, their accuracy is often limited due to the lack of 3D ligand information. EGAN-3F addresses this limitation by integrating spatially aware graph learning with traditional descriptor-based features. We systematically investigate how combining 2D and 3D molecular representations enhances binding affinity prediction from SMILES strings. This approach underscores the importance of modeling conformational diversity and incorporating chemically meaningful descriptors to improve predictive accuracy. The key innovation of EGAN-3F lies in its ability to achieve robust ligand-based binding affinity predictions without requiring protein-ligand complex structures, effectively bridging the gap between purely structural and ligand-only modeling paradigms.

Shim, Heesung [Lawrence Livermore National Laborat↗

HT Model Dataset

This website contains the dataset that was used for writing the manuscript "HT Model: Using the Molecular Transformer for predicting hydrotreating reactions" (PNNL-SA-186589) The dataset includes a collection of hydrotreating reactions compiled from 41 peer-reviewed literature sources. These sources contain experimental data related to hydrotreating reactions. These reactions involve the reaction of chemical compounds with hydrogen gas in the presence of a catalyst to remove heteroatoms or to convert specific functional groups. The dataset contains reactions both with and without reaction conditions. Reaction conditions refer to the specific parameters under which the reaction takes place, such as temperature and pressure. For each reaction, the dataset includes both SMILES and SELFIES representations. SMILES (Simplified Molecular Input Line Entry System) and SELFIES (SELF-referencIng Embedded Strings) are two popular notations used to represent chemical structures in a compact and standardized format. The dataset was created with the aim of training a predictive model, specifically using the Molecular Transformer architecture.

bioprocessing, hydrotreating, deep learning algori↗

Bloom filters for molecules

Abstract Ultra-large chemical libraries are reaching 10s to 100s of billions of molecules. A challenge for these libraries is to efficiently check if a proposed molecule is present. Here we propose and study Bloom filters for testing if a molecule is present in a set using either string or fingerprint representations. Bloom filters are small enough to hold billions of molecules in just a few GB of memory and check membership in sub milliseconds. We found string representations can have a false positive rate below 1% and require significantly less storage than using fingerprints. Canonical SMILES with Bloom filters with the simple FNV (Fowler-Noll-Voll) hashing function provide fast and accurate membership tests with small memory requirements. We provide a general implementation and specific filters for detecting if a molecule is purchasable, patented, or a natural product according to existing databases at https://github.com/whitead/molbloom .

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Creation of Polymer Datasets with Targeted Backbones for Screening of High-Performance Membranes for Gas Separation

A simple approach was developed to computationally construct a polymer dataset by combining simplified molecular-input line-entry system (SMILES) strings of a targeted polymer backbone and a variety of molecular fragments. This method was used to create 14 polymer datasets by combining seven polymer backbones and molecules from two large molecular datasets (MOSES and QM9). Polymer backbones that were studied include four polydimethylsiloxane (PDMS) based backbones, poly(ethylene oxide) (PEO), poly(allyl glycidyl ether) (PAGE), and polyphosphazene (PPZ). The generated polymer datasets can be used for various cheminformatics tasks, including high-throughput screening for gas permeability and selectivity. This study utilized machine learning (ML) models to screen the polymers for CO2/CH4 and CO2/N2 gas separation using membranes. Several polymers of interest were identified. Here the results highlight that employing an ML model fitted to polymer selectivities leads to higher accuracy in predicting polymer selectivity compared to using the ratio of predicted permeabilities.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Deep Learning Approaches for Predicting the Surface Tension of Ionic Liquids

Ionic liquids (ILs) are a novel class of solvents that have attracted significant attention due to their unique and tunable properties. Among their physiochemical characteristics, surface tension plays a critical role in various industrial applications including electrolytes, heat transfer fluids, and separation processes. However, because of the exploratory nature of IL design and the vast combinatorial space of possible anion–cation pairs, the experimental determination of these properties is often impractical, being both time-consuming and costly. To overcome these challenges, computational approaches are increasingly employed to develop accurate predictive models that can accelerate IL discovery and design. In this study, we present two deep learning (DL) models for predicting the surface tension of ILs across a broad temperature range at a constant pressure. The models use simplified molecular input line entry system, SMILES, representations of ILs to extract molecular features as inputs. Both DL models demonstrate excellent agreement with experimental data, achieving an R 2 value of 0.990 and a root-mean-square error of 0.792 mN/m. In conclusion, these results offer valuable insights for the rapid screening and rational design of ILs with tailored surface tension values.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Database of Nonaqueous Proton-Conducting Materials

This work presents the assembly of 48 papers, representing 74 different compounds and blends, into a machine-readable database of nonaqueous proton-conducting materials. SMILES was used to encode the chemical structures of the molecules, and we tabulated the reported proton conductivity, proton diffusion coefficient, and material composition for a total of 3152 data points. The data spans a broad range of temperatures ranging from -70 to 260 °C. To explore this landscape of nonaqueous proton conductors, DFT was used to calculate the proton affinity of 18 unique proton carriers. The results were then compared to the activation energy derived from fitting experimental data to the Arrhenius equation. It was found that while the widely recognized positive correlation between the activation energy and proton affinity may hold among closely related molecules, this correlation does not necessarily apply across a broader range of molecules. This work serves as an example of the potential analyses that can be conducted using literature data combined with emerging research tools in computation and data science to address specific materials design problems.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Attention-based functional-group coarse-graining: a deep learning framework for molecular prediction and design

Machine learning (ML) offers considerable promise for the design of new molecules and materials. In real-world applications, the design problem is often domain-specific, and suffers from insufficient data, particularly labeled data, for ML training. In this study, we report a data-efficient, deep-learning framework for molecular discovery that integrates a coarse-grained functional-group representation with a self-attention mechanism to capture intricate chemical interactions. Our approach exploits group-contribution concepts to create a graph-based intermediate representation of molecules, serving as a low-dimensional embedding that substantially reduces the data demands typically required for training. Using a self-attention mechanism to learn the subtle but highly relevant chemical context of functional groups, the method proposed here consistently outperforms existing approaches for predictions of multiple thermophysical properties. In a case study focused on adhesive polymer monomers, we train on a limited dataset comprising only 6,000 unlabeled and 600 labeled monomers. The resulting chemistry prediction model achieves over 92% accuracy in forecasting properties directly from SMILES strings, exceeding the performance of current state-of-the-art techniques. Furthermore, the latent molecular embedding is invertible, enabling the design pipeline to automatically generate new monomers from the learned chemical subspace. We illustrate this functionality by targeting several properties, including high and low glass transition temperatures (Tg), and demonstrate that our model can identify new candidates with values that surpass those in the training set. The ease with which the proposed framework navigates both chemical diversity and data scarcity offers a promising route to accelerate and broaden the search for functional materials.

Han, Ming [Univ. of Chicago, IL (United States)]↗

A database of thermally activated delayed fluorescent molecules auto-generated from scientific literature with ChemDataExtractor

A database of thermally activated delayed fluorescent (TADF) molecules was automatically generated from the scientific literature. It consists of 25,482 data records with an overall precision of 82%. Among these, 5,349 records have chemical names in the form of SMILES strings which are represented with 91% accuracy; these are grouped in a subsidiary database. Each data record contains one of the following four properties: maximum emission wavelength (λ EM ), photoluminescence quantum yield (PLQY), singlet-triplet energy splitting (ΔE ST ), and delayed lifetime (τ D ). The databases were created through text mining using ChemDataExtractor, a chemistry-aware natural-language-processing toolkit, which has been adapted for TADF research. The text-mined corpus consisted of 2,733 papers from the Royal Society of Chemistry and Elsevier. To the best of our knowledge, these databases are the first databases that have been auto-generated for TADF molecules from existing publications. The databases have been publicly released for experimental and computational applications in the TADF research field.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

A curated benchmark for cofolding models on kinase conformational states

Abstract Protein kinases are critical drug targets, requiring therapeutics that can modulate their active and inactive conformational states. While cofolding models can generate global folds directly from kinase sequences and ligand SMILES strings, these models have not yet been tested on their ability to recover ligand-induced-fit conformational states of the kinase proteins. Here, we introduce KinConfBench, a curated benchmark of 2225 high-quality human kinase chains to evaluate the ability of four state-of-the-art cofolding models—Boltz-2, Chai-1, Protenix, and RoseTTAFold-All-Atom—to recover both canonical and rare conformational states. We show that geometric success metrics of a ligand pose in the active site do not correlate strongly with the correct kinase conformational state, motivating a new set of dynamical benchmarks for assessing cofolding models. While all four cofolding models achieve ~60–80% prediction accuracy for kinase conformational classification, they exhibit severe mode collapse when performing multiple inferences, show negligible structural diversity in sampling induced-fit motions, and display a prevalent “apo-drift” in which most cofolding models predominantly predict the kinase to be in its ligand-free state. Our results highlight that capturing ligand-induced protein conformational diversity, not just geometric fit, is critical for next-generation structure-based drug discovery.

Sun, Kunyang↗

Uncovering novel liquid organic hydrogen carriers: a systematic exploration of chemical compound space using cheminformatics and quantum chemical methods

We present a comprehensive, in silico-based discovery approach to identifying novel liquid organic hydrogen carrier (LOHC) candidates using cheminformatics methods and quantum chemical calculations. We screened over 160 billion molecules from ZINC15 and GDB-17 chemical databases for structural similarity to known LOHCs and employed a data-driven selection criterion connecting molecular features with dehydrogenation enthalpy. This scoring criterion effectively predicts dehydrogenation enthalpies from SMILES strings, streamlining the LOHC screening process. After rigorous screening and down-selection, we compiled a database of 3000 dehydrogenation reactions for the most promising LOHC candidates, setting the stage for future selection based on kinetics and catalysis. This work demonstrates the significant impact of integrating quantum chemistry and cheminformatics in materials discovery, accelerating the selection process while reducing experimental efforts and time. By proposing new molecules as prospective LOHC candidates, our study provides a valuable resource for researchers and engineers in the development of advanced LOHC systems and showcases a successful approach for high-throughput discovery, contributing to more efficient and sustainable energy storage solutions.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗