SEARCH · Search NASA
Results for “chemical space”
Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.
Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.
Chespa: Streamlining Expansive Chemical Space Evaluation of Molecular Sets
Thousands of chemical properties can be calculated for small molecules, which can be used to place the molecules within the context of a broader “chemical space.” These definitions vary based on compounds of interest and the goals for the given chemical space definition. Here, we introduce a customizable (i.e., modular) Python module, chespa, built to easily assess different chemical space definitions through cluster-ing of compounds in these spaces and visualize trends of these clusters. To demonstrate this, chespa currently streamlines prediction of vari-ous molecule descriptors (predicted chemical properties, molecular substructures, AI-based chemical space, and chemical class ontology) in order to test 6 different chemical space definitions. Furthermore, we investigated how these varying definitions trend with mass spectrometry (MS)-based observability, i.e., the ability of a molecule to be observed with MS (e.g., as a function of the molecule ionizability), using an example data set from the U.S. EPA's Non-Targeted Analysis Collaborative Trial (ENTACT), where blinded samples had been analyzed previously, providing 1,398 data points. Improved understanding of observability would offer many advantages in small molecule identifica-tion, such as (i) a priori selection of experimental conditions based on suspected sample composition, (ii) the ability to reduce the number of candidate structures during compound identification by removing those less likely to ionize, and, in turn, (iii) a reduced false discovery rate and increased confidence in identifications. Factors controlling observability are not fully understood, making prediction of this property non-trivial and a prime candidate for chemical space analysis. Chespa is available at github.com/pnnl/chespa.
ChemoGraph: Interactive Visual Exploration of the Chemical Space
Exploratory analysis of the chemical space is an important task in the field of cheminformatics. For example, in drug discovery research, chemists investigate sets of thousands of chemical compounds in order to identify novel yet structurally similar synthetic compounds to replace natural products. Manually exploring the chemical space inhabited by all possible molecules and chemical compounds is impractical, and therefore presents a challenge. To fill this gap, we present ChemoGraph, a novel visual analytics technique for interactively exploring related chemicals. In ChemoGraph, we formalize a chemical space as a hypergraph and apply novel machine learning models to compute related chemical compounds. It uses a database to find related compounds from a known space and a machine learning model to generate new ones, which helps enlarge the known space. Moreover, ChemoGraph highlights interactive features that support users in viewing, comparing, and organizing computationally identified related chemicals. With a drug discovery usage scenario and initial expert feedback from a case study, we demonstrate the usefulness of ChemoGraph.
Navigating Transition-Metal Chemical Space: Artificial Intelligence for First-Principles Design
Conspectus The variability of chemical bonding in open-shell transition-metal complexes not only motivates their study as functional materials and catalysts but also challenges conventional computational modeling tools. Here, tailoring ligand chemistry can alter preferred spin or oxidation states as well as electronic structure properties and reactivity, creating vast regions of chemical space to explore when designing new materials atom by atom. Although first-principles density functional theory (DFT) remains the workhorse of computational chemistry in mechanism deduction and property prediction, it is of limited use here. DFT is both far too computationally costly for widespread exploration of transition-metal chemical space and also prone to inaccuracies that limit its predictive performance for localized d electrons in transition-metal complexes. These challenges starkly contrast with the well-trodden regions of small-organic-molecule chemical space, where the analytical forms of molecular mechanics force fields and semiempirical theories have for decades accelerated the discovery of new molecules, accurate DFT functional performance has been demonstrated, and gold-standard methods from correlated wavefunction theory can predict experimental results to chemical accuracy. The combined promise of transition-metal chemical space exploration and lack of established tools has mandated a distinct approach. In this Account, we outline the path we charted in exploration of transition-metal chemical space starting from the first machine learning (ML) models (i.e., artificial neural network and kernel ridge regression) and representations for the prediction of open-shell transition-metal complex properties. The distinct importance of the immediate coordination environment of the metal center as well as the lack of low-level methods to accurately predict structural properties in this coordination environment first motivated and then benefited from these ML models and representations. Once developed, the recipe for prediction of geometric, spin state, and redox potential properties was straightforwardly extended to a diverse range of other properties, including in catalysis, computational “feasibility”, and the gas separation properties of periodic metal–organic frameworks. Interpretation of selected features most important for model prediction revealed new ways to encapsulate design rules and confirmed that models were robustly mapping essential structure–property relationships. Encountering the special challenge of ensuring that good model performance could generalize to new discovery targets motivated investigation of how to best carry out model uncertainty quantification. Distance-based approaches, whether in model latent space or in carefully engineered feature space, provided intuitive measures of the domain of applicability. With all of these pieces together, ML can be harnessed as an engine to tackle the large-scale exploration of transition-metal chemical space needed to satisfy multiple objectives using efficient global optimization methods. In practical terms, bringing these artificial intelligence tools to bear on the problems of transition-metal chemical space exploration has resulted in ML-model assessments of large, multimillion compound spaces in minutes and validated new design leads in weeks instead of decades.
Navigating Large Chemical Spaces Using Graph Theory and Integer Programming
Navigating and analyzing large chemical spaces are necessary to accelerate the design and discovery of new molecules and chemical processes. In this work, we introduce a computational framework that integrates graph theory and integer programming to enable the efficient navigation of large chemical spaces. Our framework represents the chemical space as a graph, wherein nodes represent molecules and edges represent the degree of similarity or connectivity based on domain-specific information. Using the graph representation, we identify representative molecules by computing the so-called minimum dominating set (MDS), which in our context is the minimum set of molecules that is connected to all other molecules. We present a suite of solution strategies for the MDS problem including heuristic and rigorous integer programming (IP) approaches. We show that these approaches allow us to capture physicochemical properties and domain-specific logic and constraints, facilitating the identification of molecules with the target properties. We demonstrate the effectiveness of the proposed approach by navigating the chemical space of per- and polyfluoroalkyl substances (PFAS); this comprises approximately 15,000 molecular structures. We compare our framework against traditional dimensionality reduction and clustering methods such as t-SNE and K-means clustering.
Discovery of structure–property relations for molecules via hypothesis-driven active learning over the chemical space
The discovery of the molecular candidates for application in drug targets, biomolecular systems, catalysts, photovoltaics, organic electronics, and batteries necessitates the development of machine learning algorithms capable of rapid exploration of chemical spaces targeting the desired functionalities. Here, we introduce a novel approach for active learning over the chemical spaces based on hypothesis learning. We construct the hypotheses on the possible relationships between structures and functionalities of interest based on a small subset of data followed by introducing them as (probabilistic) mean functions for the Gaussian process. This approach combines the elements from the symbolic regression methods, such as SISSO and active learning, into a single framework. The primary focus of constructing this framework is to approximate physical laws in an active learning regime toward a more robust predictive performance, as traditional evaluation on hold-out sets in machine learning does not account for out-of-distribution effects which may lead to a complete failure on unseen chemical space. Here, we demonstrate it for the QM9 dataset, but it can be applied more broadly to datasets from both domains of molecular and solid-state materials sciences.
Charting the chemical space of Zintl phases with graph neural networks and bonding insights
A large number of Zintl phases have been discovered by solid-state chemists driven by empirical knowledge, chemical intuition and in some cases, through serendipitous accidents. These discoveries have only scratched the surface, given the vast compositional and structural diversity that Zintl phases can accommodate. The large chemical space of Zintl phases, as well as intermetallic compounds in general, remain under-explored. Here, we use graph neural networks and the upper bound energy minimization approach to efficiently scan a large chemical space of >90 000 hypothetical Zintl phases and accurately discover 1810 new thermodynamically stable phases with 90% precision, as validated with first-principles calculations. We show that our approach is more than 2× more accurate in predicting DFT stability than M3GNet (40% precision) on the same dataset. Using a random forest model and SHAP analysis, we demonstrate the critical role of ionic bonding in the thermodynamic stability of Zintl phases. Our results not only expand the known chemical landscape of Zintl phases but also highlight the efficacy of machine learning frameworks combined with domain knowledge in uncovering chemically meaningful insights across complex intermetallics.
Toward DMC Accuracy Across Chemical Space with Scalable Δ-QML
In the past decade, quantum diffusion Monte Carlo (DMC) has been demonstrated to successfully predict the energetics and properties of a wide range of molecules and solids by numerically solving the electronic many-body Schrödinger equation. With O(N 3 ) scaling with the number of electrons N, DMC has the potential to be a reference method for larger systems that are not accessible to more traditional methods such as CCSD(T). Assessing the accuracy of DMC for smaller molecules becomes the stepping stone in making the method a reference for larger systems. We show that when coupled with quantum machine learning (QML)-based surrogate methods, the computational burden can be alleviated such that quantum Monte Carlo (QMC) shows clear potential to undergird the formation of high-quality descriptions across chemical space. We discuss three crucial approximations necessary to accomplish this: the fixed-node approximation, universal and accurate references for chemical bond dissociation energies, and scalable minimal amons-set-based QML (AQML) models. Numerical evidence presented includes converged DMC results for over 1000 small organic molecules with up to five heavy atoms used as amons and 50 medium-sized organic molecules with nine heavy atoms to validate the AQML predictions. Finally, numerical evidence collected for Δ-AQML models suggests that already modestly sized QMC training data sets of amons suffice to predict total energies with near chemical accuracy throughout chemical space.
SmileyLlama: modifying large language models for directed chemical space exploration
Here we show that large language models (LLMs) can be transformed via supervised fine-tuning of engineered prompts into SmileyLlama for exploring the chemical space of drug molecules. We benchmark SmileyLlama against pretrained LLMs and chemical language models trained from scratch for generating valid and novel drug-like molecules, and use direct preference optimization to both improve SmileyLlama’s adherence to a prompt and as part of the iMiner reinforcement learning framework to predict molecules with optimized three-dimensional conformations and high binding affinity to drug targets. By training an LLM to speak directly as a chemical language model, while retaining most of its natural language capabilities, we show that SmileyLlama can reliably generate molecules with user-specified properties rather than acting only as a chatbot with knowledge of chemistry or as a virtual assistant. While SmileyLlama is geared toward drug discovery, the supervised fine-tuning/direct preference optimization/LLM framework can be extended to other chemical, biological and materials applications.
Mixed Enthalpy–Entropy Descriptor for the Rational Design of Synthesizable High-Entropy Materials Over Vast Chemical Spaces
The practically unlimited high-dimensional composition space of high-entropy materials (HEMs) has emerged as an exciting platform for functional material design and discovery. However, the identification of stable and synthesizable HEMs and robust design rules remains a daunting challenge. Here, we propose a mixed enthalpy–entropy descriptor (MEED) that enables highly efficient, robust, high-throughput prediction of synthesizable HEMs across vast chemical spaces from first-principles. The MEED is based on two parameters: the relative formation enthalpy with respect to the most stable competing compound and the spread of the point-defect formation energy spectrum. The former measures the relative synthesizability of an HEM to its most stable competing phase, going beyond the conventional thermodynamic understanding. Further, the latter gauges the relative entropy forming ability of an HEM, entailing no sampling over numerous alloy configurations. By applying the MEED to two structurally distinct representative material systems (i.e., 3D rocksalt carbides and 2D layered sulfides), we not only successfully identify all experimentally reported HEMs within these systems but also reveal a cutoff criterion for assessing their relative synthesizability within each system. By the MEED, tens of new high-entropy carbides and 2D high-entropy sulfides are also predicted, which have the potential for a wide variety of applications such as coating in aerospace devices, energy conversion and storage, and flexible electronics.
Expansion of bond dissociation prediction with machine learning to medicinally and environmentally relevant chemical space
Bond dissociation energetics underpin the thermodynamics of chemical transformations where bonds are broken or formed and can also be used to predict reaction rates and selectivities.
QM7-X, a comprehensive dataset of quantum-mechanical properties spanning the chemical space of small organic molecules
We introduce QM7-X, a comprehensive dataset of 42 physicochemical properties for ≈4.2 million equilibrium and non-equilibrium structures of small organic molecules with up to seven non-hydrogen (C, N, O, S, Cl) atoms. To span this fundamentally important region of chemical compound space (CCS), QM7-X includes an exhaustive sampling of (meta-)stable equilibrium structures—comprised of constitutional/structural isomers and stereoisomers, e.g., enantiomers and diastereomers (including cis-/trans- and conformational isomers)—as well as 100 non-equilibrium structural variations thereof to reach a total of ≈4.2 million molecular structures. Computed at the tightly converged quantum-mechanical PBE0+MBD level of theory, QM7-X contains global (molecular) and local (atom-in-a-molecule) properties ranging from ground state quantities (such as atomization energies and dipole moments) to response quantities (such as polarizability tensors and dispersion coefficients). By providing a systematic, extensive, and tightly-converged dataset of quantum-mechanically computed physicochemical properties, we expect that QM7-X will play a critical role in the development of next-generation machine-learning based models for exploring greater swaths of CCS and performing in silico design of molecules with targeted properties.
Nuclear quantum effects in molecular liquids across chemical space
Abstract Nuclear quantum effects (NQEs) influence many physical and chemical phenomena, particularly those involving light atoms or occurring at low temperatures. However, their impact has been carefully quantified in few systems-like water-and is rarely considered more broadly. Here we use path-integral molecular dynamics to systematically investigate NQEs on thermophysical properties of 92 organic liquids at ambient conditions. Depending on chemical constitution, we find substantial impact across thermal expansivity, compressibility, dielectric constant, enthalpy of vaporization, and notably molar volume, which shows consistent, positive quantum-classical differences up to 5%; similar, less pronounced trends manifest as isotope effects from deuteration. Using data-driven analysis, we identify three features-molar mass, classical hydrogen density, and classical thermal expansivity-that accurately predict NQEs and facilitate understanding of how characteristics like branching and heteroatom content influence behavior. This work highlights the broad relevance of NQEs in molecular liquids, while also providing a conceptual and practical framework to anticipate their impact.
Stochastic machine learning via sigma profiles to build a digital chemical space
This work establishes a different paradigm on digital molecular spaces and their efficient navigation by exploiting sigma profiles. To do so, the remarkable capability of Gaussian processes (GPs), a type of stochastic machine learning model, to correlate and predict physicochemical properties from sigma profiles is demonstrated, outperforming state-of-the-art neural networks previously published. The amount of chemical information encoded in sigma profiles eases the learning burden of machine learning models, permitting the training of GPs on small datasets which, due to their negligible computational cost and ease of implementation, are ideal models to be combined with optimization tools such as gradient search or Bayesian optimization (BO). Gradient search is used to efficiently navigate the sigma profile digital space, quickly converging to local extrema of target physicochemical properties. While this requires the availability of pretrained GP models on existing datasets, such limitations are eliminated with the implementation of BO, which can find global extrema with a limited number of iterations. A remarkable example of this is that of BO toward boiling temperature optimization. Holding no knowledge of chemistry except for the sigma profile and boiling temperature of carbon monoxide (the worst possible initial guess), BO finds the global maximum of the available boiling temperature dataset (over 1,000 molecules encompassing more than 40 families of organic and inorganic compounds) in just 15 iterations (i.e., 15 property measurements), cementing sigma profiles as a powerful digital chemical space for molecular optimization and discovery, particularly when little to no experimental data is initially available.
High-Resolution Mass Spectrometry for Human Exposomics: Expanding Chemical Space Coverage
Not Available
Exploring the Chemical Space of Linear Alkane Pyrolysis via Deep Potential GENerator
Not provided.