Search NASA⌕ Search

SEARCH · Search NASA

Results for “generative molecular design”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

Enhancing generative molecular design via uncertainty-guided fine-tuning of variational autoencoders

In recent years, deep generative models have been successfully applied to various molecular design tasks, particularly in the life and materials sciences. One critical challenge for pre-trained generative molecular design (GMD) models is to fine-tune them to be better suited for downstream design tasks that aim at optimizing specific molecular properties. However, redesigning and training an existing effective generative model from scratch for each new design task are impractical. Furthermore, the black-box nature of typical downstream tasks that involve property prediction makes it nontrivial to optimize the generative model in a task-specific manner. In this work, we propose an uncertainty-guided fine-tuning strategy that can effectively enhance a pre-trained variational autoencoder (VAE) for GMD through performance feedback in an active learning setting. The strategy begins by quantifying the model uncertainty of the generative model using an efficient active subspace-based UQ (uncertainty quantification) scheme. Next, the decoder diversity within the characterized model uncertainty class is explored to expand the viable space of molecular generation. The low-dimensionality of the active subspace makes this exploration tractable using a black-box optimization scheme, which in turn enables us to identify and leverage a diverse set of high-performing models to generate enhanced molecules. Empirical results across six target molecular properties using multiple VAE-based generative models demonstrate that our uncertainty-guided fine-tuning strategy consistently leads to improved models that outperform the original pre-trained models.

97 MATHEMATICS AND COMPUTING↗

Discovery of hydrogen storage molecules using large language models and machine learning

Accelerating the discovery of new molecules with targeted properties is a central challenge in molecular design. In this contribution, we present an AI-driven molecular discovery framework that integrates Large Language Models (LLMs) for generative molecular design with Machine Learning (ML)-based screening to identify novel Liquid Organic Hydrogen Carrier (LOHC) candidates. Using the developed framework, LOHC molecules were systematically generated, evaluated, and refined iteratively, combining LLM-guided molecular generation and ML-predicted hydrogenation enthalpies (Δ H ), under physicochemical property constraints such as optimal melting points (MP), desired hydrogen storage capacity (wt% H 2 ), and synthetic accessibility (SA) scores. This approach enabled the discovery of 42 new LOHC candidates in two distinct campaigns, one seeded with experimentally known and another with previously computationally identified LOHCs, respectively. Although we began with different numbers of starting molecules (31 vs . 7 seed molecules), both runs yielded a comparable number of viable candidates, suggesting an influence of chemically intuitive seed molecule selection for success. Selected LOHC molecules, such as 3-methyl pyridine, 1-ethylnapthalene, 1,1-diphenylethane, and benzofuran, were experimentally tested and compared with benchmark LOHCs (toluene and 9-ethylcarbazole) for hydrogenation using a series of commercial supported metal catalysts. The order of conversion into fully hydrogenated products at 200 °C was 3-methyl pyridine (100%) > 9-ethyl carbazole (86.4%) > 2,3-benzofuran (74%) > 1,1-diphenylethane (66.9%) > 1-ethylnapthalene (66.7%) > toluene (57%), further validating the AI-guided molecular design. This study demonstrates promise of LLM-driven molecular design in conjunction with ML-based screening for accelerated discovery and design of molecules.

Harb, Hassan [Argonne National Laboratory (ANL), A↗

Expanding the Domain of Applicability of Machine Learning Models with Limited Data for Drug Property Prediction

Accurate machine learning models for predicting small molecule interactions with biological targets are essential for therapeutic discovery, biothreat response, and computational drug design, but their performance is often limited for understudied targets with sparse experimental data. To address this challenge, we developed and evaluated methods to improve molecular property prediction under low-data conditions, using the NimA-related kinase (NEK) family as a proof-of-concept. This work focused on two complementary goals within the ATOM Modeling PipeLine (AMPL) and the Generative Molecular Design (GMD) loop: expanding model applicability through transfer learning, representation learning, feature scaling, sampling strategies, and active-learning-inspired compound selection; and enabling efficient virtual screening to prioritize compounds that balance predicted activity, design objectives, and synthetic accessibility.

organic↗

Generative Electrolyte Solvent and Formulation Discovery

Molecular mixtures and/or formulations are of great importance in fields ranging from materials science to pharmaceuticals to chemistry. In batteries, electrolytes are complex molecular mixtures consisting of multiple salts and solvents and additives at different concentrations that dictate battery capacity, safety, and cycle life, among others. Unfortunately, due to the complex composition and infinite design space as well as the conflicting property requirements, electrolyte design is the rate-determining step in the design of next generation battery chemistries. In this work, we develop a transformer-based generative AI model − ElectrolyteGPT − capable of generating solvents and electrolyte formulations to satisfy a wide range of desired property requirements. First, we curate an electrolyte-relevant database and develop a new line notation for formulations. Then, we show that ElectrolyteGPT can generate solvents and formulations conditioned on a wide range of important electrolyte properties such as ionic conductivity, oxidative stability, Coulombic efficiency, viscosity, and more. Finally, we experimentally synthesize the generated solvents and fabricate the electrolyte formulations and show that they can meet the desired property requirements and enable longterm cycling in energy-dense anode-free lithium metal batteries. Our work showcases the ability of generative models to address challenges in molecular mixture design for next generation batteries.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Recent progress in atomic-scale controlled plasma processing

Atomic-scale control in plasma processing is becoming increasingly critical for fabricating of advanced semiconductor devices, particularly as the industry shifts toward three-dimensional (3D) architectures and high-aspect-ratio (HAR) structures. This review presents a comprehensive overview of recent developments in atomic-scale controlled plasma processes, organized along two key directions: the hierarchical structure of plasma–surface interactions and the generational evolution of atomic layer processing (ALP) technologies. We examined the gas phase, where molecular design enables selective generation of ions and radicals; the boundary layer, where transport phenomena govern species delivery into nanoscale features, and the surface, where temperature-dependent reactions and cyclic processing determine etching selectivity and precision. Building on this foundation, we outline five generations of ALP—from thermal atomic layer deposition to transport-aware, temporally and structurally decoupled processes—highlighting the increasing sophistication of process control. The review further explores the transition from empirical recipe development to science-based, data-driven methodologies. By integrating quantum-chemical modeling, advanced diagnostics, and machine learning, we demonstrated how predictive models can link plasma species composition to process outcomes, enabling autonomous and adaptive control strategies. Finally, this review discusses the broader societal implications of plasma process innovation through the E4 quartet: energy and resource efficiency, environmental sustainability, evolutionary advancement, and educational promotion. These principles guide the development of sustainable and intelligent atomic-scale manufacturing technologies that are not only technically advanced but also socially responsible.

Ishikawa, Kenji [Nagoya Univ. (Japan)] (ORCID:0000↗

EvoDiffMol: evolutionary diffusion framework for 3D molecular design with optimized properties

Designing molecules with specific target properties remains a fundamental challenge in computational chemistry. While existing approaches show promise, most rely on simplified representations like SMILES strings or 2D graphs that lack essential three-dimensional geometric information. We present EvoDiffMol, a computational framework that integrates evolutionary algorithms with three-dimensional diffusion models for property-driven molecular generation. The method operates through adaptive evolutionary optimization, where population-based selection guides the generation process toward desired property landscapes. EvoDiffMol supports both unconstrained molecular design and scaffold-constrained generation that preserves fixed substructures while optimizing complementary regions. Comprehensive evaluation demonstrates exceptional performance, achieving the highest drug-likeness score (0.94) among all compared state-of-the-art methods while maintaining excellent validity, uniqueness, and novelty. Beyond single property optimization, the framework demonstrates flexible multi-property optimization capabilities, simultaneously controlling multiple molecular descriptors including synthetic accessibility, lipophilicity, topological polar surface area, and clinically relevant ADMET properties such as cardiotoxicity (hERG) and intestinal permeability (Caco-2). This adaptability spans from simple descriptors to practical pharmaceutical endpoints without requiring complete model retraining. The framework achieves precise control over target property values, generating molecules with properties closely matching specified targets for both single and multiple descriptors. Scaffold-constrained experiments preserve fixed molecular cores while maintaining effective property optimization. The three-dimensional representation offers advantages in maintaining structural validity during iterative optimization, with potential for geometry-aware applications in materials science and drug discovery.

3D molecular generation↗

Applying Mechanistic Co-Design Towards Room-Temperature Organic, Optical, and Multiple-Center Molecular Qubits

The next generation of qubits, the building blocks of quantum technologies, are designed using molecules. Molecular qubits have natural advantages over solid-state equivalents in biological and chemical sensing. Chemistry allows the properties of molecular qubits to be fine-tuned, such as the ability of multiple qubits to work together using qubit-qubit coupling. Molecules and structures containing multiple qubits, made from both inorganic and organic elements, are developed in this work to bring quantum technologies closer to room-temperature operation. The goal is achieved by understanding how the core of the qubit, known as spin, interacts with the vibrating parts of the molecule, other spins, and the surrounding environment. A combined approach involving theory, spectroscopy, and synthesis is used to achieve the design goal.

36 MATERIALS SCIENCE↗

Polymer Deconstruction and Redesign Strategies for Plastics Recycling

Advancing plastics recycling requires both the selective deconstruction of existing polymers and the design of new materials that enable efficient reuse without loss of performance. This perspective highlights an integrated approach that is rooted in polymer chemistry, catalysis, and process engineering which can enable a circular plastics economy. Here, we outline recent advances in catalytic, solvolytic, and enzymatic pathways for plastic deconstruction, and examine the molecular design principles driving next-generation recyclable-by-design and bio-based polymers. Despite these advances, major knowledge gaps remain in understanding the evolution of polymer morphology and catalyst structure during deconstruction, assessing deconstruction processes with realistic polymers, and offering redesigned polymers with competitive cost and environmental advantage over conventional plastics. United States Department of Energy (U.S. DOE) national laboratories offer unique capabilities to address these challenges through in situ and operando characterization, high-throughput experimentation, environmental studies, technoeconomic and life cycle assessment, scale-up support, and collaboration networks. Advances made in understanding plastic deconstruction mechanisms and structure-property correlations of redesigned polymers inform emerging research directions including autonomous experimentation, real-time feedback-enabled process optimization, and protein engineering for enzymatic depolymerization.

36 MATERIALS SCIENCE↗

Evaluation of δ-Phase ZrH1.4 to ZrH1.7 Thermal Neutron Scattering Laws Using Ab Initio Molecular Dynamics Simulations

Zirconium hydride is commonly used for next-generation reactor designs due to its excellent hydrogen retention capacity at temperatures below 1000 K. These types of reactors operate at thermal neutron energies and require accurate representation of thermal scattering laws (TSLs) to optimize moderator performance and evaluate the safety indicators for reactor design. In this work, we present an atomic-scale representation of sub-stoichiometric ZrH2−x(0.3≤x≤0.6), which relies on ab initio molecular dynamics (AIMD) in tandem with velocity auto-correlation (VAC) analysis to generate phonon density of states (DOS) for TSL development. The novel NJOY+NCrystal tool, developed by the European Spallation Source community, was utilized to generate the TSL formulations in the A Compact ENDF (ACE) format for its utility in neutron transport software. First, stoichiometric zirconium hydride cross sections were benchmarked with experiments. Then sub-stoichiometric zirconium hydride TSLs were developed. Significant deviations were observed between the new δ-phase ZrH2−x TSLs and the TSLs in the current ENDF release. It was also observed that varying the hydrogen vacancy defect concentration and sites did not cause as significant a change in the TSLs (e.g., ZrH1.4 vs. ZrH1.7) as was caused by the lattice transformation from ϵ- to δ-phase.

42 ENGINEERING↗

Spatiotemporal 4D Whole-cell Modeling of a Minimal Autotroph Reveals Central Carbon Metabolism Regulated Locally by Protein Megacomplexes via Post-translational Modifications under Light Disturbance

Photosynthetic microorganisms rely on multiple pathways in central carbon metabolism to adapt to fluctuating light and energy availability across diel cycles. Mechanistic insight into the regulatory dynamics of this adaptation requires integrating processes spanning disparate timescales, from rapid redox-dependent post-translational modifications (PTMs) to slower changes in protein expression and metabolic pathway usage. To address this complexity beyond genome-based inference and traditional modeling, we develop a whole-cell four-dimensional (3D + time) model of the marine cyanobacterium Prochlorococcus marinus MED4 that explicitly represents the spatial organization of enzymatic and molecular processes in central carbon metabolism under light perturbation. We employ a perturbation-based research design to experimentally generate time-series, multi-omics measurements that provide molecular descriptors and cryo-ET derived 3D segmented volumes as constraints for this dynamic 4D framework. The integration of experiments and modeling across defined light regimes enables quantitative validation of system-level responses and forecasting under distinct light disturbances. We test the hypothesis that light-dependent redox PTMs regulating the structural assembly of a protein megacomplex, the “dark complex,” modulate metabolic flux at a conserved regulatory node of the Calvin–Benson cycle (CBC) in cyanobacteria. Our model shows that subcellular spatial organization buffers rapid light-induced changes in thylakoid reaction rates, which are followed by redox-PTM-mediated sequestration or release of CBC enzymes in the dark complex, ultimately impacting carbon fixation dynamics within carboxysomes. Comparison with an equivalently parameterized well-mixed stochastic model demonstrates that post-translational regulation not only buffers transcriptional noise and diffusion-driven fluctuations but also stabilizes phenotypic outcomes, underscoring the importance of spatial heterogeneity in phenotypic robustness. This ability to probe adaptive, spatiotemporally resolved mechanisms in photosynthetic machinery and central carbon metabolism addresses a critical gap in genotype-to-phenotype inference and expands modeling and design capabilities for understudied or genetically intractable autotrophs such as P. marinus MED4.

Johnson, Connah G.↗

Leveraging Natural Language Processing and Generative Models in Molecular Chemistry: Property Prediction and Novel Compound Generation

The accurate prediction of molecular properties is important for the rational design and the advancement of green chemistry and sustainable materials research. However, the predictive power of traditional computational chemistry methods is limited due to computational restrictions. Here, in this study, we examine an alternative approach to the accurate prediction of properties of organic compounds: natural language processing (NLP)-based molecular embedding. Using viscosity, partition coefficient (log P), and enthalpy of vaporization as test properties through a survey of comprehensive datasets comprising 5695 data points for viscosity, 25 870 data points for log P, and 2296 data points for enthalpy of vaporization. These are important properties for the design of greener, safer, and sustainable chemical processes. Models were trained using NLP methods such as Mol2vec and fine-tuned ChemBERTa, and results were compared with traditional input featurization techniques such as Morgan fingerprints and quantum chemistry derived sigma profiles and DFT features. Among the various machine learning models, Mol2vec demonstrated superior predictive capabilities, achieving the highest correlation coefficient (R 2 = 0.945) and lowest RMSE (0.106 mPa s) for viscosity, as well as high accuracy for log P and enthalpy of vaporization predictions. These findings establish the Mol2vec featurization technique, graph-convolutional neural networks (GCNN), and fine-tuned ChemBERTa model as powerful tools for predictive modeling of organic compounds properties, offering a significant improvement over previously used featurization techniques and opening up strategies for very-high-throughput computational screening. Finally, we integrated ML models with hybrid language-model-based generative adversarial networks (LM-GAN) to generate novel molecular sequences with desirable properties for different research applications. The ability to computationally design solvents with lower viscosity, lower log P, and lower enthalpy of vaporization offers a data-driven route to accelerating the discovery of sustainable alternatives to traditionally toxic solvents.

ChemBERTa↗

Bond-centric modular design of protein assemblies

Directional interactions that generate regular coordination geometries are a powerful means of guiding molecular and colloidal self-assembly, but implementing such high-level interactions with proteins remains challenging due to their complex shapes and intricate interface properties. Here we describe a modular approach to protein nanomaterial design inspired by the rich chemical diversity that can be generated from the small number of atomic valencies. We design protein building blocks using deep learning-based generative tools, incorporating regular coordination geometries and tailorable bonding interactions that enable the assembly of diverse closed and open architectures guided by simple geometric principles. Experimental characterization confirms the successful formation of more than 20 multicomponent polyhedral protein cages, two-dimensional arrays and three-dimensional protein lattices, with a high (10%–50%) success rate and electron microscopy data closely matching the corresponding design models. Due to modularity, individual building blocks can assemble with different partners to generate distinct regular assemblies, resulting in an economy of parts and enabling the construction of reconfigurable networks for designer nanomaterials.

Biomaterials – proteins↗

Structural basis for varying drug resistance of SARS-CoV-2 M pro E166 variants

ABSTRACT Severe acute respiratory syndrome coronavirus 2 (SARS-CoV-2) main protease (M pro ) has an essential role in the virus lifecycle and, accordingly, it is a target for antiviral drugs. Multiple studies have identified an M pro mutation (E166V) that confers strong resistance to clinically relevant inhibitors, including nirmatrelvir, but the underlying mechanism is not fully understood. Here, we report on crystal structures of SARS-CoV-2 M pro E166V in complex with nirmatrelvir, ensitrelvir, and bofutrelvir. The structures suggest that resistance is caused in part by the loss of a direct hydrogen bond and also, especially for nirmatrelvir, by a steric clash with the substituted valine residue. In comparison, the binding of bofutrelvir shows greater flexibility, which may help alleviate this steric effect and allow bofutrelvir to fit the mutant active site despite the loss of a direct polar contact. Thermal stability analyses also corroborate E166V most severely affecting the binding of nirmatrelvir and, to lesser and different extents, ensitrelvir and bofutrelvir. We further show that E166I causes even more severe nirmatrelvir resistance, whereas E166A and E166L have much milder effects. These studies shed light on the molecular mechanisms of a key M pro drug resistance mutation and may help inform the design of next-generation inhibitors. IMPORTANCE Using a combination of high-resolution X-ray crystallographic and biochemical analyses, we reveal the molecular mechanisms by which a mutation in the severe acute respiratory syndrome coronavirus 2 main protease (M pro ) confers strong resistance against clinically relevant antiviral drugs that inhibit M pro activity. The results presented here may help inform the design of next-generation inhibitors to combat the problem of therapy resistance.

Microbiology↗

Attention-based functional-group coarse-graining: a deep learning framework for molecular prediction and design

Machine learning (ML) offers considerable promise for the design of new molecules and materials. In real-world applications, the design problem is often domain-specific, and suffers from insufficient data, particularly labeled data, for ML training. In this study, we report a data-efficient, deep-learning framework for molecular discovery that integrates a coarse-grained functional-group representation with a self-attention mechanism to capture intricate chemical interactions. Our approach exploits group-contribution concepts to create a graph-based intermediate representation of molecules, serving as a low-dimensional embedding that substantially reduces the data demands typically required for training. Using a self-attention mechanism to learn the subtle but highly relevant chemical context of functional groups, the method proposed here consistently outperforms existing approaches for predictions of multiple thermophysical properties. In a case study focused on adhesive polymer monomers, we train on a limited dataset comprising only 6,000 unlabeled and 600 labeled monomers. The resulting chemistry prediction model achieves over 92% accuracy in forecasting properties directly from SMILES strings, exceeding the performance of current state-of-the-art techniques. Furthermore, the latent molecular embedding is invertible, enabling the design pipeline to automatically generate new monomers from the learned chemical subspace. We illustrate this functionality by targeting several properties, including high and low glass transition temperatures (Tg), and demonstrate that our model can identify new candidates with values that surpass those in the training set. The ease with which the proposed framework navigates both chemical diversity and data scarcity offers a promising route to accelerate and broaden the search for functional materials.

Han, Ming [Univ. of Chicago, IL (United States)]↗

Bridging Atomic Solvation Environment with Electrochemical Properties for the Bis(trifluoromethylsulfonyl)imide-Based Divalent Cation Electrolytes for the Next-Generation Energy Storage Systems

A deep molecular-level understanding of the multivalent electrolyte and its correlation with the electrochemical properties is crucial for designing optimized electrolytes for next-generation rechargeable batteries. Comprehensive knowledge of the atomic level of the solvation structure and its connection with electrochemical stability and ion transport properties is especially critical. However, the interaction of these three components coupled with clear atomistic insights is lacking in the literature. Here, our current contribution evaluates representative electrolytes with the bis(trifluoromethanesulfonyl)imide (TFSI) anions for multivalent cations of Mg, Ca, and Zn, at different ionic conditions with and without a cosolvated environment in ether-based solvent. Two critical problems are investigated: first, resolving the solvation structures in the electrolyte solutions as a function of concentrations through pair distribution function analysis and the corresponding electrochemical transport properties; second, unmasking the quantitative correlation of the atomistic environment with both electrochemical kinetics and cation dependence. We discovered that the magnesium- and calcium-based electrolytes display versatile coordination lengths but poor average anodic stability due to ion pairing with TFSI - . On the contrary, the zinc-based electrolytes show the shortest solvent coordination lengths, shielding the Zn cation from rigid solvent interactions and resulting in the highest anodic stabilities. Calcium-based electrolytes exhibit the longest and most concentration-independent coordination lengths. This work provides valuable insights into the molecular structural and electrochemical features of diverse multivalent electrolyte systems with cations in various solvation environments, emphasizing the importance of the solvation structure and construction in designing high-performance electrolytes.

cation coordination↗