Search NASA⌕ Search

SEARCH · Search NASA

Results for “Chemical space”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

Navigating Large Chemical Spaces Using Graph Theory and Integer Programming

Navigating and analyzing large chemical spaces are necessary to accelerate the design and discovery of new molecules and chemical processes. In this work, we introduce a computational framework that integrates graph theory and integer programming to enable the efficient navigation of large chemical spaces. Our framework represents the chemical space as a graph, wherein nodes represent molecules and edges represent the degree of similarity or connectivity based on domain-specific information. Using the graph representation, we identify representative molecules by computing the so-called minimum dominating set (MDS), which in our context is the minimum set of molecules that is connected to all other molecules. We present a suite of solution strategies for the MDS problem including heuristic and rigorous integer programming (IP) approaches. We show that these approaches allow us to capture physicochemical properties and domain-specific logic and constraints, facilitating the identification of molecules with the target properties. We demonstrate the effectiveness of the proposed approach by navigating the chemical space of per- and polyfluoroalkyl substances (PFAS); this comprises approximately 15,000 molecular structures. We compare our framework against traditional dimensionality reduction and clustering methods such as t-SNE and K-means clustering.

Chemical structure↗

Charting the chemical space of Zintl phases with graph neural networks and bonding insights

A large number of Zintl phases have been discovered by solid-state chemists driven by empirical knowledge, chemical intuition and in some cases, through serendipitous accidents. These discoveries have only scratched the surface, given the vast compositional and structural diversity that Zintl phases can accommodate. The large chemical space of Zintl phases, as well as intermetallic compounds in general, remain under-explored. Here, we use graph neural networks and the upper bound energy minimization approach to efficiently scan a large chemical space of >90 000 hypothetical Zintl phases and accurately discover 1810 new thermodynamically stable phases with 90% precision, as validated with first-principles calculations. We show that our approach is more than 2× more accurate in predicting DFT stability than M3GNet (40% precision) on the same dataset. Using a random forest model and SHAP analysis, we demonstrate the critical role of ionic bonding in the thermodynamic stability of Zintl phases. Our results not only expand the known chemical landscape of Zintl phases but also highlight the efficacy of machine learning frameworks combined with domain knowledge in uncovering chemically meaningful insights across complex intermetallics.

36 MATERIALS SCIENCE↗

SmileyLlama: modifying large language models for directed chemical space exploration

Here we show that large language models (LLMs) can be transformed via supervised fine-tuning of engineered prompts into SmileyLlama for exploring the chemical space of drug molecules. We benchmark SmileyLlama against pretrained LLMs and chemical language models trained from scratch for generating valid and novel drug-like molecules, and use direct preference optimization to both improve SmileyLlama’s adherence to a prompt and as part of the iMiner reinforcement learning framework to predict molecules with optimized three-dimensional conformations and high binding affinity to drug targets. By training an LLM to speak directly as a chemical language model, while retaining most of its natural language capabilities, we show that SmileyLlama can reliably generate molecules with user-specified properties rather than acting only as a chatbot with knowledge of chemistry or as a virtual assistant. While SmileyLlama is geared toward drug discovery, the supervised fine-tuning/direct preference optimization/LLM framework can be extended to other chemical, biological and materials applications.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Nuclear quantum effects in molecular liquids across chemical space

Abstract Nuclear quantum effects (NQEs) influence many physical and chemical phenomena, particularly those involving light atoms or occurring at low temperatures. However, their impact has been carefully quantified in few systems-like water-and is rarely considered more broadly. Here we use path-integral molecular dynamics to systematically investigate NQEs on thermophysical properties of 92 organic liquids at ambient conditions. Depending on chemical constitution, we find substantial impact across thermal expansivity, compressibility, dielectric constant, enthalpy of vaporization, and notably molar volume, which shows consistent, positive quantum-classical differences up to 5%; similar, less pronounced trends manifest as isotope effects from deuteration. Using data-driven analysis, we identify three features-molar mass, classical hydrogen density, and classical thermal expansivity-that accurately predict NQEs and facilitate understanding of how characteristics like branching and heteroatom content influence behavior. This work highlights the broad relevance of NQEs in molecular liquids, while also providing a conceptual and practical framework to anticipate their impact.

Science & Technology - Other Topics↗

Stochastic machine learning via sigma profiles to build a digital chemical space

This work establishes a different paradigm on digital molecular spaces and their efficient navigation by exploiting sigma profiles. To do so, the remarkable capability of Gaussian processes (GPs), a type of stochastic machine learning model, to correlate and predict physicochemical properties from sigma profiles is demonstrated, outperforming state-of-the-art neural networks previously published. The amount of chemical information encoded in sigma profiles eases the learning burden of machine learning models, permitting the training of GPs on small datasets which, due to their negligible computational cost and ease of implementation, are ideal models to be combined with optimization tools such as gradient search or Bayesian optimization (BO). Gradient search is used to efficiently navigate the sigma profile digital space, quickly converging to local extrema of target physicochemical properties. While this requires the availability of pretrained GP models on existing datasets, such limitations are eliminated with the implementation of BO, which can find global extrema with a limited number of iterations. A remarkable example of this is that of BO toward boiling temperature optimization. Holding no knowledge of chemistry except for the sigma profile and boiling temperature of carbon monoxide (the worst possible initial guess), BO finds the global maximum of the available boiling temperature dataset (over 1,000 molecules encompassing more than 40 families of organic and inorganic compounds) in just 15 iterations (i.e., 15 property measurements), cementing sigma profiles as a powerful digital chemical space for molecular optimization and discovery, particularly when little to no experimental data is initially available.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Active causal learning for decoding chemical complexities with targeted interventions

Abstract Predicting and enhancing inherent properties based on molecular structures is paramount to design tasks in medicine, materials science, and environmental management. Most of the current machine learning and deep learning approaches have become standard for predictions, but they face challenges when applied across different datasets due to reliance on correlations between molecular representation and target properties. These approaches typically depend on large datasets to capture the diversity within the chemical space, facilitating a more accurate approximation, interpolation, or extrapolation of the chemical behavior of molecules. In our research, we introduce an active learning approach that discerns underlying cause-effect relationships through strategic sampling with the use of a graph loss function. This method identifies the smallest subset of the dataset capable of encoding the most information representative of a much larger chemical space. The identified causal relations are then leveraged to conduct systematic interventions, optimizing the design task within a chemical space that the models have not encountered previously. While our implementation focused on the QM9 quantum-chemical dataset for a specific design task—finding molecules with a large dipole moment—our active causal learning approach, driven by intelligent sampling and interventions, holds potential for broader applications in molecular, materials design and discovery.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Labels as a feature: Network homophily for systematically annotating human GPCR drug-target interactions

Machine learning has revolutionized drug discovery by enabling the exploration of vast, uncharted chemical spaces essential for discovering novel patentable drugs. Despite the critical role of human G protein-coupled receptors in FDA-approved drugs, exhaustive in-distribution drug-target interaction testing across all pairs of human G protein-coupled receptors and known drugs is rare due to significant economic and technical challenges. This often leaves off-target effects unexplored, which poses a considerable risk to drug safety. In contrast to the traditional focus on out-of-distribution exploration (drug discovery), we introduce a neighborhood-to-prediction model termed Chemical Space Neural Networks that leverages network homophily and training-free graph neural networks with labels as features. We show that Chemical Space Neural Networks’ ability to make accurate predictions strongly correlates with network homophily. Thus, labels as features strongly increase a machine learning model’s capacity to enhance in-distribution prediction accuracy, which we show by integrating labeled data during inference. We validate these advancements in a high-throughput yeast biosensing system (3773 drug-target interactions, 539 compounds, 7 human G protein-coupled receptors) to discover novel drug-target interactions for FDA-approved drugs and to expand the general understanding of how to build reliable predictors to guide experimental verification.

Hansson, Frederik G↗

High entropy oxides prediction and discovery by the Mixed Enthalpy-Entropy Descriptor

The vast, high-dimensional composition space of high-entropy oxides (HEOs) offers exceptional opportunities for functional materials discovery, yet it also poses a fundamental challenge: the rational and efficient prediction of stable, synthesizable compositions and the corresponding structure–property relationships. Despite growing interest, the field still lacks broadly applicable, physically grounded descriptors capable of navigating various large chemical spaces. Here, we introduce a Mixed Enthalpy–Entropy Descriptor (MEED) that enables rapid, first-principles–based prediction of HEOs synthesizability across diverse chemistries. Using MEED, we perform high-throughput screening of two distinct HEO families: rocksalt oxides and perovskite oxides. The predicted top candidates in each family were experimentally validated. MEED reveals unifying thermodynamic and structural principles governing stability across both chemical compositions and polymorphs, providing mechanistic insight into the formation of high-entropy phases. This work significantly broadens the accessible chemical design space for HEOs and establishes a data-efficient framework for accelerating the discovery of next-generation functional materials.

Yu, Liping [University of Central Florida]↗

On the Prospect of Chemically Transferable Coarse-Grained Electronic Models for Soft Materials

Electronic coarse-graining (ECG) methods predict quantum-mechanical electronic properties directly from coarse-grained (CG) molecular configurations, enabling electronic predictions at mesoscale length scales. Here, we present a diagnostic assessment of the feasibility of chemically transferable ECG models across a broad polymer-relevant chemical space using all-atom, united-atom, and Martini-scale representations. While high-resolution ECG models achieve near-quantitative accuracy, we show that chemically transferable ECG at the Martini resolution fails because the CG force field does not sample the same configurational distribution of local molecular structure as that underlying the DFT-parameterized ECG model. We demonstrate that our proposed Element-Count-Label (ECL) representation, which augments Martini beads with explicit stoichiometric data, significantly improves chemical generalization across diverse polymer chemistries. However, we find that even with improved chemical resolution, the model cannot recover electronic property distributions that are absent from the configurational space sampled by the CG force field. These results demonstrate that chemically transferable ECG requires future Martini-like force fields to explicitly preserve quantum chemistry–compatible local molecular structure in addition to thermodynamic and structural fidelity.

Kidder, Katherine M [Department of Chemistry; Univ↗

PubChemLite Plus Collision Cross Section (CCS) Values for Enhanced Interpretation of Nontarget Environmental Data

Finding relevant chemicals in the vast (known) chemical space is a major challenge for environmental and exposomics studies leveraging nontarget high resolution mass spectrometry (NT-HRMS) methods. Chemical databases now contain hundreds of millions of chemicals, yet many are not relevant. This article details an extensive collaborative, open science effort to provide a dynamic collection of chemicals for environmental, metabolomics, and exposomics research, along with supporting information about their relevance to assist researchers in the interpretation of candidate hits. The PubChemLite for Exposomics collection is compiled from ten annotation categories within PubChem, enhanced with patent, literature and annotation counts, predicted partition coefficient (logP) values, as well as predicted collision cross section (CCS) values using CCSbase. Monthly versions are archived on Zenodo under a CC-BY license, supporting reproducible research, and a new interface has been developed, including historical trends of patent and literature data, for researchers to browse the collection. This article details how PubChemLite can support researchers in environmental and exposomics studies, describes efforts to increase the availability of experimental CCS values, and explores known limitations and potential for future developments. The data and code behind these efforts are openly available.

PubChem↗

Scaling Ensembles of Data-Intensive Quantum Chemical Calculations for Millions of Molecules

Deep learning models are efficient computational tools that can accelerate the inverse design of molecules with desired functional properties by generating predictions at a fraction of the time required by traditional quantum chemical approaches. To ensure that a model maintains accuracy and transferability across broad regions of the chemical space explored during the inverse design, it must be trained on massively large volumes of simulation data. This requires running large-scale ensemble quantum chemical calculations on high-performance computing (HPC) systems for data collection. However, the efficient execution of such large ensemble calculations and the management of large volumes of output data require tools that can judiciously utilize computational resources and manage metadata overhead on the file system. Therefore, we present a high-performance, scalable, ensemble management framework for performing data-intensive quantum chemical electronic structure calculations for organic molecules. This framework provides abstractions to plug different ab initio, first principles, and first principles-based semi-empirical methods and executes them efficiently at large scale on HPC systems. It dynamically distributes tasks to resources and uses tiered storage for managing large collections of files. We employed this framework to process over ten million organic molecules and generate open-source datasets that provide UV-vis absorption spectra by running time-dependent density-functional tight-binding calculations. It is the largest database containing molecular optical spectra that were simulated with quantum chemical methods in a consistent manner.

Mehta, Kshitij↗

Improved 140 Nd Production for the 140 Nd/ 140 Pr In Vivo Generator through Target Recycling and Radiochemical Optimization

Theranostic strategies that utilize f-block therapeutic radionuclides, including 161 Tb, 177 Lu, 225 Ac, and 227 Th, suffer from a shortage of positron emission tomography (PET) imaging counterparts in the same chemical space and often rely on 68 Ga as a surrogate. The 140 Nd/ 140 Pr in vivo PET generator, which belongs to the f-block, may address this issue and can be produced via the 141 Pr(p,2n) 140 Nd production route by using medium-energy cyclotrons. However, impurities in the target material, including stable Nd, and the inherent difficulty of adjacent lanthanide separations limit the achievable radionuclidic and chemical purity of 140 Nd. In this work, we address these challenges through the purification and recycling of praseodymium target material and optimization of Nd/Pr separation. The resulting purified 140Nd was evaluated using DOTA and Macropa chelators via radiolabeling and in vitro stability studies. A target material purification and recycling method was developed for the monoisotopic 141 Pr starting material to remove stable Nd impurities, yielding 90.3 ± 4.7% (n = 3) recovery. The purified 141 Pr was isolated as Pr 6 O 11 and irradiated with 24 MeV protons (20.07 MeV at the target surface) at 20 μA for 4 h, which produced 1417.0 ± 83.4 MBq (38.3 ± 2.2 mCi) of 140 Nd at the end of bombardment (EOB). The produced 140 Nd was purified through an optimized DGA normal method to recover 71.6 ± 6.3% pure 140 Nd. The amount of stable Nd reduced progressively in each target purification cycle from >340 ppm without purification to <250 ppb after three cycles, while other measured metallic impurities were below 30 ppb. This improvement in target purity was reflected in the direct increase of apparent molar activity (AMA), when purified 140 Nd was evaluated with DOTA and Macropa chelators. AMA of [ 140 Nd]Nd-DOTA and [ 140 Nd]Nd-Macropa increased from 70.3 MBq/μmol (1.9 mCi/μmol) and 74 MBq/μmol (2.0 mCi/μmol) to 8025.3 MBq/μmol (216.9 mCi/μmol) and 8473.0 MBq/μmol (229.0 mCi/μmol), respectively, after the third target purification cycle. Further evaluation of chelator-labeled 140 Nd showed that [ 140 Nd]Nd-DOTA was stable in phosphate-buffered saline (PBS), saline, human serum, and mouse serum, whereas [140Nd]Nd-Macropa was stable in all except human serum. This work established a practical methodological advance for the production of 140 Nd/ 140 Pr in vivo PET generators, combining optimized target recycling and radiochemical separation to enable scaled-up and high-molar activity 140 Nd suitable for preclinical imaging. These advances support broader development of 140 Nd/ 140 Pr as a robust PET analogue, especially for f-block therapeutics.

Irradiation↗

NCAP: Noncanonical Amino Acid Parameterization Software for CHARMM Potentials

Noncanonical Amino Acids (NCAAs) provide numerous avenues for introduction of novel functionality to peptides and proteins. NCAAs can be incorporated through solid phase synthesis or genetic code expansion in conjugation with heterologous expression of the encoded protein modification. Due to the difficulty of synthesis, wide chemical space and lack of empirically resolved structures modeling the effects of NCAA mutation is critical for rational protein design. To evaluate the structural and functional perturbations NCAAs introduce we utilize molecular potentials that describe the forces in protein structure. Most potentials such as CHARMM are designed to model canonical residues but can be parameterized in include novel NCAAs. Here, in this work, we introduce NCAP a software package to generate CHARMM compatible parameters from quantum chemical calculation. Unlike currently available tools NCAP is designed to recognize NCAA structure and automatically bridge the gap between DFT calculations and potential parameters. For our software we discuss workflow, validation against canonical parameter sets and comparison to published NCAA-protein structures.

59 BASIC BIOLOGICAL SCIENCES↗

Transferable predictions of energetic and structural properties for refractory solid solution alloys across chemical compositions

We present a data-efficient approach to train graph neural networks (GNNs) on density functional theory (DFT) data for accurate and transferable predictions of energetic and structural properties of refractory solid solution alloys in the niobium-tantalum-vanadium (Nb-Ta-V) chemical space. We start by training the GNN model only on DFT data that describes refractory binary alloys niobium-tantalum (Nb-Ta), niobium-vanadium (Nb-V), and tantalum-vanadium (Ta-V) to predict formation enthalpy and root mean squared displacement. Once trained, the GNN predictions are tested on DFT data describing refractory ternary alloys Nb-Ta-V. While, unsurprisingly, direct transferability from binary to ternary is not sufficiently accurate, augmenting the training with only 1% of the available ternary data (uniformly distributed across the entire range of chemical compositions) improves significantly the quality of the GNN predictions. For comparison, we assess the transferability in the opposite direction by training GNN models on ternary Nb-Ta-V data and making predictions on binaries Nb-Ta, Nb-V, and Ta-V, which exhibits notably higher predictive errors. The proposed methodology, which favors transferability from lower-component to higher-component alloys, offers an efficient path towards avoiding the curse of dimensionality incurred when collecting DFT data for discovery and design of multi-component disordered alloys.

Density functional theory calculations↗

Learning Molecular Mixture Property Using Chemistry-Aware Graph Neural Network

Recent advances in machine learning (ML) are expediting materials discovery and design. One significant challenge facing ML for materials is the expansive combinatorial space of potential materials formed by diverse constituents and their flexible configurations. This complexity is particularly evident in molecular mixtures, a frequently explored space for materials, such as battery electrolytes. Owing to the complex structures of molecules and the sequence-independent nature of mixtures, conventional ML methods have difficulties in modeling such systems. Here, we present MolSets, a specialized ML model for molecular mixtures, to overcome the difficulties. Representing individual molecules as graphs and their mixture as a set, MolSets leverages a graph neural network and the deep sets architecture to extract information at the molecular level and aggregate it at the mixture level, thus addressing local complexity while retaining global flexibility. We demonstrate the efficacy of MolSets in predicting the conductivity of lithium battery electrolytes and highlight its benefits in the virtual screening of the combinatorial chemical space. Published by the American Physical Society 2024

Zhang, Hengrui (ORCID:0000000231831654)↗