Search NASA⌕ Search

SEARCH · Search NASA

Results for “molecular descriptors”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

Predicting permeation of compounds across the outer membrane of P. aeruginosa using molecular descriptors

The ability Gram-negative pathogens have at adapting and protecting themselves against antibiotics has increasingly become a public health threat. Data-driven models identifying molecular properties that correlate with outer membrane (OM) permeation and growth inhibition while avoiding efflux could guide the discovery of novel classes of antibiotics. Here we evaluate 174 molecular descriptors in 1260 antimicrobial compounds and study their correlations with antibacterial activity in Gram-negative Pseudomonas aeruginosa. The descriptors are derived from traditional approaches quantifying the compounds’ intrinsic physicochemical properties, together with, bacterium-specific from ensemble docking of compounds targeting specific MexB binding pockets, and all-atom molecular dynamics simulations in different subregions of the OM model. Using these descriptors and the measured inhibitory concentrations, we design a statistical protocol to identify predictors of OM permeation/inhibition. We find consistent rules across most of our data highlighting the role of the interaction between the compounds and the OM. An implementation of the rules uncovered in our study is shown, and it demonstrates the accuracy of our approach in a set of previously unseen compounds. Our analysis sheds new light on the key properties drug candidates need to effectively permeate/inhibit P. aeruginosa, and opens the gate to similar data-driven studies in other Gram-negative pathogens.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Machine learning-based discovery of molecular descriptors that control polymer gas permeation

While machine learning has found increasing use in predicting the properties of polymeric materials with only a knowledge of chain architecture, determining the molecular factors underpinning properties (“interpretable AI”) has remained less well explored. We show that encoding chain chemistry in commonly employed formats, e.g., binary-valued fingerprints, leads to uniqueness issues during the hashing process to save storage space. This is because the hashing algorithm can map several chemical moieties into the same bit. These issues carry over into the ML algorithms, especially for “inverse” design and interpretable AI, and cannot be avoided by changing the length of the fingerprint. Using MACCS key featurizations of monomer repeats resolves some of these issues, and we show that a few substructures consistently appear in top features for maximizing permeability across several gases and ML models. These are carbon–carbon double bonds (as in polyacetylenes) especially when they are associated with methyl groups (found in branching architectures). Here these results, derived from the limited data set of ~ 500 polymers with experimental gas permeation data, are in agreement with physical insight and thus provide a robust foundation which could further enable study of these material classes through detailed experiments and simulations.

36 MATERIALS SCIENCE↗

Open-source generation of sigma profiles: impact of quantum chemistry and solvation treatment on machine learning performance

The combination of machine learning (ML) models with chemistry-related tasks requires the description of molecular structures in a machine-readable way. The nature of these so-called molecular descriptors has a direct and major impact on the performance of ML models and remains an open problem in the field. Structural descriptors like SMILES strings or molecular graphs lack size-independence and can be memory intensive. Machine-learned descriptors can be of low dimensionality and constant size but lack physical significance and human interpretability. Sigma profiles, which are unnormalized histograms of the surface charge distributions of solvated molecules, combine physical significance with low dimensionality and size-independence, making them a suitable candidate for a universal molecular descriptor. However, their widespread adoption in ML applications requires open access to sigma profile generation, which is currently not available. This work details the development of OpenSPGen – an open-source tool for generating sigma profiles. Also presented are studies on the effect of different settings on the efficacy of the generated sigma profiles at predicting thermophysical material properties when used as inputs to a Gaussian process as a simple surrogate ML model. We find that a higher level of theory does not translate to more accurate results. We also provide further recommendations for sigma profile calculation and use in ML models.

Salih, Fathya Y. M. [University of Notre Dame, IN ↗

EvoDiffMol: evolutionary diffusion framework for 3D molecular design with optimized properties

Designing molecules with specific target properties remains a fundamental challenge in computational chemistry. While existing approaches show promise, most rely on simplified representations like SMILES strings or 2D graphs that lack essential three-dimensional geometric information. We present EvoDiffMol, a computational framework that integrates evolutionary algorithms with three-dimensional diffusion models for property-driven molecular generation. The method operates through adaptive evolutionary optimization, where population-based selection guides the generation process toward desired property landscapes. EvoDiffMol supports both unconstrained molecular design and scaffold-constrained generation that preserves fixed substructures while optimizing complementary regions. Comprehensive evaluation demonstrates exceptional performance, achieving the highest drug-likeness score (0.94) among all compared state-of-the-art methods while maintaining excellent validity, uniqueness, and novelty. Beyond single property optimization, the framework demonstrates flexible multi-property optimization capabilities, simultaneously controlling multiple molecular descriptors including synthetic accessibility, lipophilicity, topological polar surface area, and clinically relevant ADMET properties such as cardiotoxicity (hERG) and intestinal permeability (Caco-2). This adaptability spans from simple descriptors to practical pharmaceutical endpoints without requiring complete model retraining. The framework achieves precise control over target property values, generating molecules with properties closely matching specified targets for both single and multiple descriptors. Scaffold-constrained experiments preserve fixed molecular cores while maintaining effective property optimization. The three-dimensional representation offers advantages in maintaining structural validity during iterative optimization, with potential for geometry-aware applications in materials science and drug discovery.

3D molecular generation↗

Comparison of Machine Learning Approaches for Prediction of the Equivalent Alkane Carbon Number for Microemulsions Based on Molecular Properties

The chemical properties of oils are vital in the design of microemulsion systems. The hydrophilic–lipophilic difference equation used to predict microemulsions’ phase behavior expresses the oils’ physiochemical properties as the equivalent alkane carbon number (EACN). The experimental determination of EACN requires knowledge of the temperature dependence of the microemulsion system and the effects of different surfactant concentrations. Thus, the experimental determination is time-intensive and tedious, requiring days to months for proper separations. Furthermore, the experiments require high purity of chemicals because microemulsions are sensitive to impurities. Our work focuses on the quick and reliable predictions of the EACN with machine learning (ML) models. Due to the immaturity of ML chemical predictions, we compare three graph neural networks (GNNs) and a gradient-boosted tree algorithm, known as XGBoost. The GNNs use the molecular structures represented as simplified molecular-input line-entry system (SMILES) codes for the initial input, which allows us to assess whether geometry optimization is necessary for reliable results. The XGBoost model also begins with the SMILES representations of the molecules but uses molecular descriptors instead of geometry optimizations. As a result, the best model tested (crystal graph convolutional neural network with Merck molecular force field-94) has an error of 1.15 EACN units of the true EACN for unknown data with the errors skewed toward zero and an R² score of 0.9

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Similarity Downselection: Finding the n Most Dissimilar Molecular Conformers for Reference-Free Metabolomics

Computational methods for creating in silico libraries of molecular descriptors (e.g., collision cross sections) are becoming increasingly prevalent due to the limited number of authentic reference materials available for traditional library building. These so-called “reference-free metabolomics” methods require sampling sets of molecular conformers in order to produce high accuracy property predictions. Due to the computational cost of the subsequent calculations for each conformer, there is a need to sample the most relevant subset and avoid repeating calculations on conformers that are nearly identical. The goal of this study is to introduce a heuristic method of finding the most dissimilar conformers from a larger population in order to help speed up reference-free calculation methods and maintain a high property prediction accuracy. Finding the set of the n items most dissimilar from each other out of a larger population becomes increasingly difficult and computationally expensive as either n or the population size grows large. Because there exists a pairwise relationship between each item and all other items in the population, finding the set of the n most dissimilar items is different than simply sorting an array of numbers. For instance, if you have a set of the most dissimilar n = 4 items, one or more of the items from n = 4 might not be in the set n = 5. An exact solution would have to search all possible combinations of size n in the population exhaustively. We present an open-source software called similarity downselection (SDS), written in Python and freely available on GitHub. SDS implements a heuristic algorithm for quickly finding the approximate set(s) of the n most dissimilar items. We benchmark SDS against a Monte Carlo method, which attempts to find the exact solution through repeated random sampling. We show that for SDS to find the set of n most dissimilar conformers, our method is not only orders of magnitude faster, but it is also more accurate than running Monte Carlo for 1,000,000 iterations, each searching for set sizes n = 3–7 out of a population of 50,000. We also benchmark SDS against the exact solution for example small populations, showing that SDS produces a solution close to the exact solution in these instances. Using theoretical approaches, we also demonstrate the constraints of the greedy algorithm and its efficacy as a ratio to the exact solution.

97 MATHEMATICS AND COMPUTING↗

The Influence of Chemical Descriptors on the Reactivity of Potential Hypergolic Fuels With Hydrogen Peroxide

The state-of-the-art for storable hypergolic bipropellant systems is monomethylhydrazine (MMH) and mixed oxides of nitrogen (MON). While MMH and MON provide fast and highly reliable hypergolic ignition with excellent propulsive performance, they suffer from safety concerns impacting their production, cost, and supply to various test and launch facilities. In contrast, identifying and demonstrating low-toxicity fuels hypergolic with rocket grade hydrogen peroxide (H2O2) has been challenging with long ignition delay (IDT) and low propulsive performance affecting their deployment in flight systems. This work describes a preliminary observational study to assess the effects of molecular descriptors suspected to predict the reactivity and hypergolicity of compounds within the same chemical class (i.e., thioamide). The approach centered on evaluating functional groups, electrostatic potential, number of carbons, dipole moment, and adiabatic ionization energy. Thiourea is hypergolic with 89.2 wt.% H2O2 (minimum IDT of 21.9 ms) and was selected as a core fuel to study the reactivity of derived fuels with different functionalities. Additionally, the chemical properties were calculated with the Gaussian16 suite. The results showed that group substitution and generally increased number of carbons on thiourea increases IDT. It was observed that a more positively charged sulfur atom in the thioamide does not increase the reactivity towards H2O2. The dipole moment analysis revealed an overall trend of decreasing IDT due to increasing dipole moment. Finally, a lower adiabatic ionization energy related to increased IDTs for most cases. Additional chemical classes comprising a wider range of functionalities and electrophilic attributes are being investigated in on-going work.

Propellant↗

Mechanistic studies of small molecule ligands selective to RNA single G bulges

Abstract Small-molecule RNA binders have emerged as an important pharmacological modality. A profound understanding of the ligand selectivity, binding mode, and influential factors governing ligand engagement with RNA targets is the foundation for rational ligand design. Here, we report a novel class of coumarin derivatives exhibiting selective binding affinity towards single G RNA bulges. Harnessing the computational power of all-atom Gaussian accelerated molecular dynamics simulations, we unveiled a rare minor groove binding mode of the ligand with a key interaction between the coumarin moiety and the G bulge. This predicted binding mode is consistent with results obtained from structure-activity relationship studies and transverse relaxation measurements by nuclear magnetic resonance spectroscopy. We further generated 444 molecular descriptors from 69 coumarin derivatives and identified key contributors to the binding events, such as charge state and planarity, by lasso (least absolute shrinkage and selection operator) regression. Our work deepened the understanding of RNA-small molecule interactions and integrated a new framework for the rational design of selective small-molecule RNA binders.

Biochemistry & Molecular Biology↗

Prediction of hydration energies of adsorbates at Pt(111) and liquid water interfaces using machine learning

Aqueous phase heterogeneous catalysis is important to various industrial processes, including biomass conversion, Fischer–Tropsch synthesis, and electrocatalysis. Accurate calculation of solvation thermodynamic properties is essential for modeling the performance of catalysts for these processes. Explicit solvation methods employing multiscale modeling, e.g., involving density functional theory and molecular dynamics have emerged for this purpose. Although accurate, these methods are computationally intensive. This study introduces machine learning (ML) models to predict solvation thermodynamics for adsorbates on a Pt(111) surface, aiming to enhance computational efficiency without compromising accuracy. In particular, ML models are developed using a combination of molecular descriptors and fingerprints and trained on previously published water–adsorbate interaction energies, energies of solvation, and free energies of solvation of adsorbates bound to Pt(111). These models achieve root mean square error values of 0.09 eV for interaction energies, 0.04 eV for energies of solvation, and 0.06 eV for free energies of solvation, demonstrating accuracy within the standard error of multiscale modeling. Feature importance analysis reveals that hydrogen bonding, van der Waals interactions, and solvent density, together with the properties of the adsorbate, are critical factors influencing solvation thermodynamics. Furthermore, these findings suggest that ML models can provide rapid and reliable predictions of solvation properties. This approach not only reduces computational costs but also offers insights into the solvation characteristics of adsorbates at Pt(111)–water interfaces.

Adsorption↗

Spatiotemporal 4D Whole-cell Modeling of a Minimal Autotroph Reveals Central Carbon Metabolism Regulated Locally by Protein Megacomplexes via Post-translational Modifications under Light Disturbance

Photosynthetic microorganisms rely on multiple pathways in central carbon metabolism to adapt to fluctuating light and energy availability across diel cycles. Mechanistic insight into the regulatory dynamics of this adaptation requires integrating processes spanning disparate timescales, from rapid redox-dependent post-translational modifications (PTMs) to slower changes in protein expression and metabolic pathway usage. To address this complexity beyond genome-based inference and traditional modeling, we develop a whole-cell four-dimensional (3D + time) model of the marine cyanobacterium Prochlorococcus marinus MED4 that explicitly represents the spatial organization of enzymatic and molecular processes in central carbon metabolism under light perturbation. We employ a perturbation-based research design to experimentally generate time-series, multi-omics measurements that provide molecular descriptors and cryo-ET derived 3D segmented volumes as constraints for this dynamic 4D framework. The integration of experiments and modeling across defined light regimes enables quantitative validation of system-level responses and forecasting under distinct light disturbances. We test the hypothesis that light-dependent redox PTMs regulating the structural assembly of a protein megacomplex, the “dark complex,” modulate metabolic flux at a conserved regulatory node of the Calvin–Benson cycle (CBC) in cyanobacteria. Our model shows that subcellular spatial organization buffers rapid light-induced changes in thylakoid reaction rates, which are followed by redox-PTM-mediated sequestration or release of CBC enzymes in the dark complex, ultimately impacting carbon fixation dynamics within carboxysomes. Comparison with an equivalently parameterized well-mixed stochastic model demonstrates that post-translational regulation not only buffers transcriptional noise and diffusion-driven fluctuations but also stabilizes phenotypic outcomes, underscoring the importance of spatial heterogeneity in phenotypic robustness. This ability to probe adaptive, spatiotemporally resolved mechanisms in photosynthetic machinery and central carbon metabolism addresses a critical gap in genotype-to-phenotype inference and expands modeling and design capabilities for understudied or genetically intractable autotrophs such as P. marinus MED4.

Johnson, Connah G.↗

Quantum Chemistry-Driven Machine Learning Approach for the Prediction of the Surface Tension and Speed of Sound in Ionic Liquids

Ionic liquids (ILs) have unique solvent properties and have thus garnered significant interest. However, exhaustive experimental determination of the physicochemical properties of ILs is unrealistic due to the large structural diversity of anions and cations, their high cost, the requirements of elevated temperature and pressure, and the time required. To circumvent these experimental costs, computational approaches to accurately calculate these properties have emerged. Here in the present study, we present a demonstration of two machine learning (ML) models for the prediction of two critical IL physical properties, the surface tension and the speed of sound, across a wide range of temperatures and pressures. The models make use of molecular descriptors derived from the COSMO-RS, a quantum chemical-based model. The ML models show excellent agreement with experimental observations, with an R2 value of 0.96–0.99 and RMSE of 1.71 mN/m and 16.12 m/s for the surface tension and speed of sound, respectively. This work paves the way for the development of COSMO-RS-informed ML models for the prediction of IL properties which can help to further optimize and accelerate technology development for ILs.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

The Use of Machine Learning Models for Predicting the Dielectric Strength of Gases

Technological advancements in high voltage systems have pushed sulfur hexafluoride (SF6) to its operational limits. Furthermore, this gas has other drawbacks including a high liquefaction temperature and a high global warming potential. Therefore, there has been an urgent need to find alternative gases with high dielectric strength (DS). In this work, density functional theory (DFT) is used to calculate molecular descriptors that are fed into an artificial neural network (ANN) and a random forest (RF). These machine learning (ML) models are then used to predict the DS for hundreds of molecules. A finite element model (FEM) is also used to calculate the electric field profile of multiple simple electrode geometries as the applied voltage to the system is increased. Results indicate that the random forest model has better generalization to unseen data than the neural network. The highest DS value predicted by the RF was 2.16 relative to the experimental DS of SF6. The results also demonstrate how choosing a gas with a higher DS and a geometry with minimal edges and corners can significantly increase the operating voltage of an electrical system. Due to its superior generalization, the RF represents the most promising path toward an accurate DS predictor once sufficient experimental data are available.

Mileski, Matthew [AFIT]↗

Artificial Neural Network Models for Octane Number and Octane Sensitivity: A Quantitative Structure Property Relationship Approach to Fuel Design

Octane sensitivity (OS), defined as the research octane number (RON) minus the motor octane number (MON) of a fuel, has gained interest among researchers due to its effect on knocking conditions in internal combustion engines. Compounds with a high OS enable higher efficiencies, especially within advanced compression ignition engines. RON/MON must be experimentally tested to determine OS, requiring time, funding, and specialized equipment. Thus, predictive models trained with existing experimental data and molecular descriptors (via quantitative structure-property relationships (QSPRs)) would allow for the preemptive screening of compounds prior to performing these experiments. Here, the present work proposes two methods for predicting the OS of a given compound: using artificial neural networks (ANNs) trained with QSPR descriptors to predict RON and MON individually to compute OS (derived octane sensitivity (dOS)), and using ANNs trained with QSPR descriptors to directly predict OS. Twenty-five ANNs were trained for both RON and MON and their test sets achieved an overall 6.4% and 5.2% error, respectively. Twenty-five additional ANNs were trained for both dOS and OS; dOS calculations were found to have 15.3% error while predicting OS directly resulted in 9.9% error. A chemical analysis of the top QSPR descriptors for RON/MON and OS is conducted, highlighting desirable structural features for high-performing molecules and offering insight into the inner mathematical workings of ANNs; such chemical interpretations study the interconnections between structural features, descriptors, and fuel performance showing that connectivity, structural diversity, and atomic hybridization consistently drive fuel performance.

09 BIOMASS FUELS↗

Developing a SARS-CoV-2 main protease binding prediction random forest model for drug repurposing for COVID-19 treatment

The coronavirus disease 2019 (COVID-19) global pandemic resulted in millions of people becoming infected with the severe acute respiratory syndrome coronavirus 2 (SARS-CoV-2) virus and close to seven million deaths worldwide. It is essential to further explore and design effective COVID-19 treatment drugs that target the main protease of SARS-CoV-2, a major target for COVID-19 drugs. In this study, machine learning was applied for predicting the SARS-CoV-2 main protease binding of Food and Drug Administration (FDA)-approved drugs to assist in the identification of potential repurposing candidates for COVID-19 treatment. Ligands bound to the SARS-CoV-2 main protease in the Protein Data Bank and compounds experimentally tested in SARS-CoV-2 main protease binding assays in the literature were curated. These chemicals were divided into training (516 chemicals) and testing (360 chemicals) data sets. To identify SARS-CoV-2 main protease binders as potential candidates for repurposing to treat COVID-19, 1188 FDA-approved drugs from the Liver Toxicity Knowledge Base were obtained. A random forest algorithm was used for constructing predictive models based on molecular descriptors calculated using Mold2 software. Model performance was evaluated using 100 iterations of fivefold cross-validations which resulted in 78.8% balanced accuracy. The random forest model that was constructed from the whole training dataset was used to predict SARS-CoV-2 main protease binding on the testing set and the FDA-approved drugs. Model applicability domain and prediction confidence on drugs predicted as the main protease binders discovered 10 FDA-approved drugs as potential candidates for repurposing to treat COVID-19. Our results demonstrate that machine learning is an efficient method for drug repurposing and, thus, may accelerate drug development targeting SARS-CoV-2.

Research & Experimental Medicine↗

Advanced Modeling and Process-Materials Co-Optimization Strategies for Swing Adsorption Based Gas Separations

This project devised a computational framework for simultaneously co-optimizing pressure swing adsorption process designs along with the sorbent materials (specifically, metal-organic frameworks) to be employed in the associated packed bed columns. The materials optimization aspect involved search over a design space that can describe the material’s molecular structure, while the process optimization aspect considered various process degrees of freedom for steps arising in various cycle configurations. This framework was demonstrated on the separation of nitrogen and carbon dioxide, which arises ubiquitously in a multitude of post-combustion carbon capture and “blue” hydrogen production applications. Our results led to metal-organic framework molecular descriptor choices that are predicted to outperform standard structures used in practice, providing guidance for future metal-organic framework synthesis efforts.

20 FOSSIL-FUELED POWER PLANTS↗

From Structured Solvents to Hybrid Materials (SS2HM) for Chemically Selective Capture and Electromagnetic Release of CO 2 : Mechanisms, Stability and Interfaces (Final Report)

The goal of this research program was to develop high capacity sorbents amenable for alternative regeneration approaches for direct air capture (DAC) of CO 2 . In particular, the research aimed to develop an understanding of CO 2 binding mechanism, thermal and oxidative stability, and regeneration energetics of functionalized ionic liquids (ILs), deep eutectic solvents (DESs), and porous materials. ILs and DESs are high-dielectric solvents with structural tunability that permits the rational-design for energy-efficient regeneration approaches based on electromagnetic (EM) field and moisture-swing. By further incorporating these solvents into polymeric capsules and other structural supports, multi-scale interfaces for targeted CO 2 and energy transfers were achieved. Aspects related to CO 2 capacity, selectivity, stability, dielectric properties, and binding energies were examined through experimental and computational design to identify molecular descriptors to inform future design of structured solvents and hybrid materials for DAC. Enclosed final report details the key findings, science advancements, and workforce development efforts from this project.

36 MATERIALS SCIENCE↗