Search NASA⌕ Search

SEARCH · Search NASA

Results for “Molecular discovery”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

Determining best practices for using genetic algorithms in molecular discovery

Genetic algorithms (GAs) are a powerful tool to search large chemical spaces for inverse molecular design. However, GAs have multiple hyperparameters that have not been thoroughly investigated for chemical space searches. In this tutorial, we examine the general effects of a number of hyperparameters, such as population size, elitism rate, selection method, mutation rate, and convergence criteria, on key GA performance metrics. Here, we show that using a self-termination method with a minimum Spearman’s rank correlation coefficient of 0.8 between generations maintained for 50 consecutive generations along with a population size of 32, a 50% elitism rate, three-way tournament selection, and a 40% mutation rate provides the best balance of finding the overall champion, maintaining good coverage of elite targets, and improving relative speedup for general use in molecular design GAs.

36 MATERIALS SCIENCE↗

Fingerprinting Interactions between Proteins and Ligands for Facilitating Machine Learning in Drug Discovery

Molecular recognition is fundamental in biology, underpinning intricate processes through specific protein–ligand interactions. This understanding is pivotal in drug discovery, yet traditional experimental methods face limitations in exploring the vast chemical space. Computational approaches, notably quantitative structure–activity/property relationship analysis, have gained prominence. Molecular fingerprints encode molecular structures and serve as property profiles, which are essential in drug discovery. While two-dimensional (2D) fingerprints are commonly used, three-dimensional (3D) structural interaction fingerprints offer enhanced structural features specific to target proteins. Machine learning models trained on interaction fingerprints enable precise binding prediction. Recent focus has shifted to structure-based predictive modeling, with machine-learning scoring functions excelling due to feature engineering guided by key interactions. Notably, 3D interaction fingerprints are gaining ground due to their robustness. Various structural interaction fingerprints have been developed and used in drug discovery, each with unique capabilities. This review recapitulates the developed structural interaction fingerprints and provides two case studies to illustrate the power of interaction fingerprint-driven machine learning. The first elucidates structure–activity relationships in β2 adrenoceptor ligands, demonstrating the ability to differentiate agonists and antagonists. The second employs a retrosynthesis-based pre-trained molecular representation to predict protein–ligand dissociation rates, offering insights into binding kinetics. Despite remarkable progress, challenges persist in interpreting complex machine learning models built on 3D fingerprints, emphasizing the need for strategies to make predictions interpretable. Binding site plasticity and induced fit effects pose additional complexities. Interaction fingerprints are promising but require continued research to harness their full potential.

3D structural interaction fingerprints↗

Generative Electrolyte Solvent and Formulation Discovery

Molecular mixtures and/or formulations are of great importance in fields ranging from materials science to pharmaceuticals to chemistry. In batteries, electrolytes are complex molecular mixtures consisting of multiple salts and solvents and additives at different concentrations that dictate battery capacity, safety, and cycle life, among others. Unfortunately, due to the complex composition and infinite design space as well as the conflicting property requirements, electrolyte design is the rate-determining step in the design of next generation battery chemistries. In this work, we develop a transformer-based generative AI model − ElectrolyteGPT − capable of generating solvents and electrolyte formulations to satisfy a wide range of desired property requirements. First, we curate an electrolyte-relevant database and develop a new line notation for formulations. Then, we show that ElectrolyteGPT can generate solvents and formulations conditioned on a wide range of important electrolyte properties such as ionic conductivity, oxidative stability, Coulombic efficiency, viscosity, and more. Finally, we experimentally synthesize the generated solvents and fabricate the electrolyte formulations and show that they can meet the desired property requirements and enable longterm cycling in energy-dense anode-free lithium metal batteries. Our work showcases the ability of generative models to address challenges in molecular mixture design for next generation batteries.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Molecular Glue Discovery: Current and Future Approaches

The intracellular interactions of biomolecules can be maneuvered to redirect signaling, reprogram the cell cycle, or decrease infectivity using only a few dozen atoms. Such "molecular glues," which can drive both novel and known interactions between protein partners, represent an enticing therapeutic strategy. Here, we review the methods and approaches that have led to the identification of small-molecule molecular glues. We first classify current FDA-approved molecular glues to facilitate the selection of discovery methods. We then survey two broad discovery method strategies, where we highlight the importance of factors such as experimental conditions, software packages, and genetic tools for success. In conclusion, we hope that this curation of methodologies for directed discovery will inspire diverse research efforts targeting a multitude of human diseases.

59 BASIC BIOLOGICAL SCIENCES↗

Autonomous Molecular Structure Imaging with High-Resolution Atomic Force Microscopy for Molecular Mixture Discovery

Due to its single-molecule sensitivity, high-resolution atomic force microscopy (HR-AFM) has proved to be a valuable and uniquely advantageous tool to study complex molecular mixtures, which hold promise for developing clean energy and achieving environmental sustainability. However, significant challenges remain to achieve the full potential of the sophisticated and time-consuming experiments. Automation combined with machine learning (ML) and artificial intelligence (AI) is key to overcoming these challenges. Here we present Auto-HR-AFM, an AI tool to automatically collect HR-AFM images of petroleum-based mixtures. In this study, we trained an instance segmentation model to teach Auto-HR-AFM how to recognize features in HR-AFM images. Auto-HR-AFM then uses that information to optimize the imaging by adjusting the probe-molecule distance for each molecule in the run. Auto-HR-AFM is the initial tool that will lead to fully automated scanning probe microscopy (SPM) experiments, from start to finish. This automation will allow SPM to become a mainstream characterization technique for complex mixtures, an otherwise unattainable target.

36 MATERIALS SCIENCE↗

Machine learning-based discovery of molecular descriptors that control polymer gas permeation

While machine learning has found increasing use in predicting the properties of polymeric materials with only a knowledge of chain architecture, determining the molecular factors underpinning properties (“interpretable AI”) has remained less well explored. We show that encoding chain chemistry in commonly employed formats, e.g., binary-valued fingerprints, leads to uniqueness issues during the hashing process to save storage space. This is because the hashing algorithm can map several chemical moieties into the same bit. These issues carry over into the ML algorithms, especially for “inverse” design and interpretable AI, and cannot be avoided by changing the length of the fingerprint. Using MACCS key featurizations of monomer repeats resolves some of these issues, and we show that a few substructures consistently appear in top features for maximizing permeability across several gases and ML models. These are carbon–carbon double bonds (as in polyacetylenes) especially when they are associated with methyl groups (found in branching architectures). Here these results, derived from the limited data set of ~ 500 polymers with experimental gas permeation data, are in agreement with physical insight and thus provide a robust foundation which could further enable study of these material classes through detailed experiments and simulations.

36 MATERIALS SCIENCE↗

Discovery of molecular hydrogen fluorescence in the diffuse interstellar medium

The first detection of molecular hydrogen fluorescence in the diffuse interstellar medium is reported. Using the Berkeley UVX Shuttle Spectrometer, H2 Lyman band fluorescence has been observed in four directions, each with high significance. Molecular hydrogen fluorescence is detected in all directions that have previously been found to contain significant CO emission. A simple equilibrium model has been developed that includes attenuation of the incident UV radiation field by H2 line and dust continuum absorption. Evidence is found that the gas in the CO emission portions of the clouds may be clumpy, with a filling factor less than 0.2 and an average density greater than 30/cu cm in most cases.

Martin, Christopher↗

Discovery of hydrogen storage molecules using large language models and machine learning

Accelerating the discovery of new molecules with targeted properties is a central challenge in molecular design. In this contribution, we present an AI-driven molecular discovery framework that integrates Large Language Models (LLMs) for generative molecular design with Machine Learning (ML)-based screening to identify novel Liquid Organic Hydrogen Carrier (LOHC) candidates. Using the developed framework, LOHC molecules were systematically generated, evaluated, and refined iteratively, combining LLM-guided molecular generation and ML-predicted hydrogenation enthalpies (Δ H ), under physicochemical property constraints such as optimal melting points (MP), desired hydrogen storage capacity (wt% H 2 ), and synthetic accessibility (SA) scores. This approach enabled the discovery of 42 new LOHC candidates in two distinct campaigns, one seeded with experimentally known and another with previously computationally identified LOHCs, respectively. Although we began with different numbers of starting molecules (31 vs . 7 seed molecules), both runs yielded a comparable number of viable candidates, suggesting an influence of chemically intuitive seed molecule selection for success. Selected LOHC molecules, such as 3-methyl pyridine, 1-ethylnapthalene, 1,1-diphenylethane, and benzofuran, were experimentally tested and compared with benchmark LOHCs (toluene and 9-ethylcarbazole) for hydrogenation using a series of commercial supported metal catalysts. The order of conversion into fully hydrogenated products at 200 °C was 3-methyl pyridine (100%) > 9-ethyl carbazole (86.4%) > 2,3-benzofuran (74%) > 1,1-diphenylethane (66.9%) > 1-ethylnapthalene (66.7%) > toluene (57%), further validating the AI-guided molecular design. This study demonstrates promise of LLM-driven molecular design in conjunction with ML-based screening for accelerated discovery and design of molecules.

Harb, Hassan [Argonne National Laboratory (ANL), A↗

Active deep kernel learning of molecular properties from structural embeddings

As vast databases of chemical identities become increasingly available, the challenge shifts to how we effectively explore and leverage these resources to study molecular properties. This paper presents an active learning approach for molecular discovery using deep kernel learning (DKL), demonstrated on the QM9 dataset. DKL links structural embeddings directly to properties, creating organized latent spaces that prioritize relevant property information. By iteratively recalculating embedding vectors in alignment with target properties, DKL uncovers concentrated maxima representing key molecular properties and reveals unexplored regions with potential for innovation. This approach underscores DKL’s potential in advancing molecular research and discovery.

Artificial neural networks↗

CACTUS: Chemistry Agent Connecting Tool Usage to Science

Large language models (LLMs) have shown remarkable potential in various domains but often lack the ability to access and reason over domain-specific knowledge and tools. In this article, we introduce Chemistry Agent Connecting Tool-Usage to Science (CACTUS), an LLM-based agent that integrates existing cheminformatics tools to enable accurate and advanced reasoning and problem-solving in chemistry and molecular discovery. We evaluate the performance of CACTUS using a diverse set of open-source LLMs, including Gemma-7b, Falcon-7b, MPT-7b, Llama3-8b, and Mistral-7b, on a benchmark of thousands of chemistry questions. Our results demonstrate that CACTUS significantly outperforms baseline LLMs, with the Gemma-7b, Mistral-7b, and Llama3-8b models achieving the highest accuracy regardless of the prompting strategy used. Moreover, we explore the impact of domain-specific prompting and hardware configurations on model performance, highlighting the importance of prompt engineering and the potential for deploying smaller models on consumer-grade hardware without a significant loss in accuracy. By combining the cognitive capabilities of open-source LLMs with widely used domain-specific tools provided by RDKit, CACTUS can assist researchers in tasks such as molecular property prediction, similarity searching, and drug-likeness assessment.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Laboratory Astrophysics Needs of the Herschel Space Observatory

The science teams of the Herschel Space Observatory have identified a number of areas where laboratory study is required for proper interpretation of Herschel observational data. The most critical is the collection and compilation of laboratory data on spectral line frequencies, transition probabilities and energy levels for the known astrophysical atomic and molecular species in 670 to 57 micron wavelength range of Herschel. The second most critical need is the compilation of collisional excitation cross sections for the species known to dominate the energy balance in the ISM and the temperature dependent chemical reaction rates. On the theoretical front, chemical and radiative transfer models need to be prepared in advance to assess calibration and identify instrument anomalies. In the next few years there will be a need to incorporate spectroscopists and theoretical chemists into teams of astronomers so that the spectroscopic surveys planned can he properly calibrated and rapidly interpreted once the data becomes available. The science teams have also noted that the enormous prospects for molecular discovery will be greatly handicapped by the nearly complete lack of spectroscopic data for anything not already well known in the ISM. As a minimum, molecular species predicted to exist by chemical models should be subjected to detailed laboratory study to ensure conclusive detections. This has the greatest impact on any astrobiology program that might be proposed for Herschel. Without a significant amount of laboratory work in the very near future Herschel will not be prepared for many planned observations, much less addressing the open questions in molecular astrophysics.

Pearson, J. C.↗

Biology of bone and how it orchestrates the form and function of the skeleton

The principal role of the skeleton is to provide structural support for the body. While the skeleton also serves as the body's mineral reservoir, the mineralized structure is the very basis of posture, opposes muscular contraction resulting in motion, withstands functional load bearing, and protects internal organs. Although the mass and morphology of the skeleton is defined, to some extent, by genetic determinants, it is the tissue's ability to remodel--the local resorption and formation of bone--which is responsible for achieving this intricate balance between competing responsibilities. The aim of this review is to address bone's form-function relationship, beginning with extensive research in the musculoskeletal disciplines, and focusing on several recent cellular and molecular discoveries which help understand the complex interdependence of bone cells, growth factors, physical stimuli, metabolic demands, and structural responsibilities. With a clinical and spine-oriented audience in mind, the principles of bone cell and molecular biology and physiology are presented, and an attempt has been made to incorporate epidemiologic data and therapeutic implications. Bone research remains interdisciplinary by nature, and a deeper understanding of bone biology will ultimately lead to advances in the treatment of diseases and injuries to bone itself.

Review↗

Attention-based functional-group coarse-graining: a deep learning framework for molecular prediction and design

Machine learning (ML) offers considerable promise for the design of new molecules and materials. In real-world applications, the design problem is often domain-specific, and suffers from insufficient data, particularly labeled data, for ML training. In this study, we report a data-efficient, deep-learning framework for molecular discovery that integrates a coarse-grained functional-group representation with a self-attention mechanism to capture intricate chemical interactions. Our approach exploits group-contribution concepts to create a graph-based intermediate representation of molecules, serving as a low-dimensional embedding that substantially reduces the data demands typically required for training. Using a self-attention mechanism to learn the subtle but highly relevant chemical context of functional groups, the method proposed here consistently outperforms existing approaches for predictions of multiple thermophysical properties. In a case study focused on adhesive polymer monomers, we train on a limited dataset comprising only 6,000 unlabeled and 600 labeled monomers. The resulting chemistry prediction model achieves over 92% accuracy in forecasting properties directly from SMILES strings, exceeding the performance of current state-of-the-art techniques. Furthermore, the latent molecular embedding is invertible, enabling the design pipeline to automatically generate new monomers from the learned chemical subspace. We illustrate this functionality by targeting several properties, including high and low glass transition temperatures (Tg), and demonstrate that our model can identify new candidates with values that surpass those in the training set. The ease with which the proposed framework navigates both chemical diversity and data scarcity offers a promising route to accelerate and broaden the search for functional materials.

Han, Ming [Univ. of Chicago, IL (United States)]↗

Accurate Dehydrogenation Enthalpies Dataset for Liquid Organic Hydrogen Carriers

This contribution presents a comprehensive extension of the QM9 dataset (originally at 133 K molecules) with the calculation of G4MP2 enthalpies for 9,841 molecules, featuring up to nine heavy atoms. We present QM9-LOHC, a (de)hydrogenation dataset of 10,373 reactions, including a minimum of 5.5% weight hydrogen storage capacity in line with the Department of Energy standards for Liquid Organic Hydrogen Carriers (LOHC). By utilizing the accurate quantum chemical method G4MP2 we expand the QM9 database and explore new avenues for the exploration of hydrogen storage technologies (electrochemical LOHCs, alkali metal-LOHCs, and mixtures of LOHCs). The QM9-LOHC dataset, with its focus on reactions that vary only by hydrogen saturation levels, provides a needed data resource for advancing the design and optimization of both conventional and innovative LOHC systems, and high-fidelity data for molecular discovery.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Deep kernel methods learn better: from cards to process optimization

Abstract The ability of deep learning methods to perform classification and regression tasks relies heavily on their capacity to uncover manifolds in high-dimensional data spaces and project them into low-dimensional representation spaces. In this study, we investigate the structure and character of the manifolds generated by classical variational autoencoder (VAE) approaches and deep kernel learning (DKL). In the former case, the structure of the latent space is determined by the properties of the input data alone, while in the latter, the latent manifold forms as a result of an active learning process that balances the data distribution and target functionalities. We show that DKL with active learning can produce a more compact and smooth latent space which is more conducive to optimization compared to previously reported methods, such as the VAE. We demonstrate this behavior using a simple cards dataset and extend it to the optimization of domain-generated trajectories in physical systems. Our findings suggest that latent manifolds constructed through active learning have a more beneficial structure for optimization problems, especially in feature-rich target-poor scenarios that are common in domain sciences, such as materials synthesis, energy storage, and molecular discovery. The Jupyter Notebooks that encapsulate the complete analysis accompany the article.

97 MATHEMATICS AND COMPUTING↗

Introducing Molecular Hypernetworks for Discovery in Multidimensional Metabolomics Data

Orthogonal separations of data from high-resolution mass spectrometry can provide insight into sample composition and address challenges of complete annotation of molecules in untargeted metabolomics. “Molecular networks” (MNs), as used in the Global Natural Products Social Molecular Networking platform, are a prominent strategy for exploring and visualizing molecular relationships and improving annotation. MNs are mathematical graphs showing the relationships between measured multidimensional data features. MNs also show promise for using network science algorithms to automatically identify targets for annotation candidates and to dereplicate features associated with a single molecular identity. Here, this paper introduces “molecular hypernetworks” (MHNs) as more complex MN models able to natively represent multiway relationships among observations. Compared to MNs, MHNs can more parsimoniously represent the inherent complexity present among groups of observations, initially supporting improved exploratory data analysis and visualization. MHNs also promise to increase confidence in annotation propagation, for both human and analytical processing. We first illustrate MHNs with simple examples, and build them from liquid chromatography- and ion mobility spectrometry-separated MS data. We then describe a method to construct MHNs directly from existing MNs as their “clique reconstructions”, demonstrating their utility by comparing examples of previously published graph-based MNs to their respective MHNs.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Data-Driven Discovery of Linear Molecular Probes with Optimal Selective Affinity for PFAS in Water

Approaches to tackle the wide and growing variety of highly persistent per- and polyfluoroalkyl substances (PFAS) are of pressing global need because of their detrimental human health effects, such as cancer, birth defects, and hormone imbalance. Sensitive, selective, and easy-to-use real-time sensors to monitor and detect PFAS and sorbents to extract them are critical to meeting government-mandated environmental concentrations. In this work, we combine all-atom molecular dynamics simulations, enhanced sampling, deep representational learning, and Bayesian optimization to perform high-throughput virtual screening for highly sensitive and selective molecular probes. Our molecular design space consists of 3850 linear hydrocarbon chains with varying degrees of halogenation with and without amine- and phosphine-based headgroups. By employing a data-driven search process, we efficiently explore the molecular design space to optimize the sensitivity to perfluorooctanesulfonic acid (PFOS) as a prototypical PFAS analyte and selectivity relative to a sodium dodecyl sulfate (SDS) interferent. We calculate 504 Gibbs free energies of probe-analyte and probe-interferent interactions and identify probes with PFOS association free energies of up to (-ΔG PFOS ) = 9.8 ± 0.2 kJ/mol and selectivities relative to SDS of (-ΔΔG PFOS–SDS ) = 3.1 ± 1.5 kJ/mol. A C 11 Br 23 P(CH 3 ) 2 probe containing 11 backbone brominated carbons and a tertiary phosphine headgroup possesses the most sensitive binding constant to PFOS within the defined search space of K b PFOS = 177.4 ± 12.7, and a semibrominated probe C 5 H 11 C 7 Br 14 N(CH 3 ) 2 containing 12 backbone carbons and a tertiary amine headgroup possesses the highest selectivity relative to SDS of K b PFOS /K b SDS = 4.6 ± 1.7. A retrospective analysis of our data to extract interpretable design rules reveals that the sensitivity of linear hydrogenated probes increases by approximately 1 kJ/mol per C–C bond. The addition or removal of halogen atoms and amine or phosphine headgroups produces nonmonotonic changes in both sensitivity and selectivity with changes to the sensitivity of up to 2.5 kJ/mol. Finally, this work places empirical limitations on the performance of a wide range of linear probes for PFOS detection and offers a generic strategy for high-throughput computational screening to promote selective and sensitive binding.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Discovery of optical molecular emission from the bipolar nebula surrounding HD 44179

Spectrophotometry and spectropolarimetry with HD 44179 are presented. These measurements reveal that the very broad bump evident in previous low-resolution spectra possesses a large amount of structure, including groups of narrow emission lines and several diffuse features. A reduction in polarization, but constant position angle, through the bump indicates that this emission originates within the nebula itself and merely dilutes the polarized scattered starlight. A few very weak atomic emission lines are detected, but the overall feature, which strongly resembles the emission spectra of some molecules, remains unidentified. Constraints on the excitation mechanism are discussed.

Schmidt, G. D.↗