Search NASA⌕ Search

SEARCH · Search NASA

Results for “generative molecular design”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

Enhancing generative molecular design via uncertainty-guided fine-tuning of variational autoencoders

In recent years, deep generative models have been successfully applied to various molecular design tasks, particularly in the life and materials sciences. One critical challenge for pre-trained generative molecular design (GMD) models is to fine-tune them to be better suited for downstream design tasks that aim at optimizing specific molecular properties. However, redesigning and training an existing effective generative model from scratch for each new design task are impractical. Furthermore, the black-box nature of typical downstream tasks that involve property prediction makes it nontrivial to optimize the generative model in a task-specific manner. In this work, we propose an uncertainty-guided fine-tuning strategy that can effectively enhance a pre-trained variational autoencoder (VAE) for GMD through performance feedback in an active learning setting. The strategy begins by quantifying the model uncertainty of the generative model using an efficient active subspace-based UQ (uncertainty quantification) scheme. Next, the decoder diversity within the characterized model uncertainty class is explored to expand the viable space of molecular generation. The low-dimensionality of the active subspace makes this exploration tractable using a black-box optimization scheme, which in turn enables us to identify and leverage a diverse set of high-performing models to generate enhanced molecules. Empirical results across six target molecular properties using multiple VAE-based generative models demonstrate that our uncertainty-guided fine-tuning strategy consistently leads to improved models that outperform the original pre-trained models.

97 MATHEMATICS AND COMPUTING↗

Discovery of hydrogen storage molecules using large language models and machine learning

Accelerating the discovery of new molecules with targeted properties is a central challenge in molecular design. In this contribution, we present an AI-driven molecular discovery framework that integrates Large Language Models (LLMs) for generative molecular design with Machine Learning (ML)-based screening to identify novel Liquid Organic Hydrogen Carrier (LOHC) candidates. Using the developed framework, LOHC molecules were systematically generated, evaluated, and refined iteratively, combining LLM-guided molecular generation and ML-predicted hydrogenation enthalpies (Δ H ), under physicochemical property constraints such as optimal melting points (MP), desired hydrogen storage capacity (wt% H 2 ), and synthetic accessibility (SA) scores. This approach enabled the discovery of 42 new LOHC candidates in two distinct campaigns, one seeded with experimentally known and another with previously computationally identified LOHCs, respectively. Although we began with different numbers of starting molecules (31 vs . 7 seed molecules), both runs yielded a comparable number of viable candidates, suggesting an influence of chemically intuitive seed molecule selection for success. Selected LOHC molecules, such as 3-methyl pyridine, 1-ethylnapthalene, 1,1-diphenylethane, and benzofuran, were experimentally tested and compared with benchmark LOHCs (toluene and 9-ethylcarbazole) for hydrogenation using a series of commercial supported metal catalysts. The order of conversion into fully hydrogenated products at 200 °C was 3-methyl pyridine (100%) > 9-ethyl carbazole (86.4%) > 2,3-benzofuran (74%) > 1,1-diphenylethane (66.9%) > 1-ethylnapthalene (66.7%) > toluene (57%), further validating the AI-guided molecular design. This study demonstrates promise of LLM-driven molecular design in conjunction with ML-based screening for accelerated discovery and design of molecules.

Harb, Hassan [Argonne National Laboratory (ANL), A↗

Expanding the Domain of Applicability of Machine Learning Models with Limited Data for Drug Property Prediction

Accurate machine learning models for predicting small molecule interactions with biological targets are essential for therapeutic discovery, biothreat response, and computational drug design, but their performance is often limited for understudied targets with sparse experimental data. To address this challenge, we developed and evaluated methods to improve molecular property prediction under low-data conditions, using the NimA-related kinase (NEK) family as a proof-of-concept. This work focused on two complementary goals within the ATOM Modeling PipeLine (AMPL) and the Generative Molecular Design (GMD) loop: expanding model applicability through transfer learning, representation learning, feature scaling, sampling strategies, and active-learning-inspired compound selection; and enabling efficient virtual screening to prioritize compounds that balance predicted activity, design objectives, and synthetic accessibility.

organic↗

Generative Electrolyte Solvent and Formulation Discovery

Molecular mixtures and/or formulations are of great importance in fields ranging from materials science to pharmaceuticals to chemistry. In batteries, electrolytes are complex molecular mixtures consisting of multiple salts and solvents and additives at different concentrations that dictate battery capacity, safety, and cycle life, among others. Unfortunately, due to the complex composition and infinite design space as well as the conflicting property requirements, electrolyte design is the rate-determining step in the design of next generation battery chemistries. In this work, we develop a transformer-based generative AI model − ElectrolyteGPT − capable of generating solvents and electrolyte formulations to satisfy a wide range of desired property requirements. First, we curate an electrolyte-relevant database and develop a new line notation for formulations. Then, we show that ElectrolyteGPT can generate solvents and formulations conditioned on a wide range of important electrolyte properties such as ionic conductivity, oxidative stability, Coulombic efficiency, viscosity, and more. Finally, we experimentally synthesize the generated solvents and fabricate the electrolyte formulations and show that they can meet the desired property requirements and enable longterm cycling in energy-dense anode-free lithium metal batteries. Our work showcases the ability of generative models to address challenges in molecular mixture design for next generation batteries.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Recent progress in atomic-scale controlled plasma processing

Atomic-scale control in plasma processing is becoming increasingly critical for fabricating of advanced semiconductor devices, particularly as the industry shifts toward three-dimensional (3D) architectures and high-aspect-ratio (HAR) structures. This review presents a comprehensive overview of recent developments in atomic-scale controlled plasma processes, organized along two key directions: the hierarchical structure of plasma–surface interactions and the generational evolution of atomic layer processing (ALP) technologies. We examined the gas phase, where molecular design enables selective generation of ions and radicals; the boundary layer, where transport phenomena govern species delivery into nanoscale features, and the surface, where temperature-dependent reactions and cyclic processing determine etching selectivity and precision. Building on this foundation, we outline five generations of ALP—from thermal atomic layer deposition to transport-aware, temporally and structurally decoupled processes—highlighting the increasing sophistication of process control. The review further explores the transition from empirical recipe development to science-based, data-driven methodologies. By integrating quantum-chemical modeling, advanced diagnostics, and machine learning, we demonstrated how predictive models can link plasma species composition to process outcomes, enabling autonomous and adaptive control strategies. Finally, this review discusses the broader societal implications of plasma process innovation through the E4 quartet: energy and resource efficiency, environmental sustainability, evolutionary advancement, and educational promotion. These principles guide the development of sustainable and intelligent atomic-scale manufacturing technologies that are not only technically advanced but also socially responsible.

Ishikawa, Kenji [Nagoya Univ. (Japan)] (ORCID:0000↗

Automatic Evolution of Molecular Nanotechnology Designs

This paper describes strategies for automatically generating designs for analog circuits at the molecular level. Software maps out the edges and vertices of potential nanotechnology systems on graphs, then selects appropriate ones through evolutionary or genetic paradigms.

Globus, Al↗

EvoDiffMol: evolutionary diffusion framework for 3D molecular design with optimized properties

Designing molecules with specific target properties remains a fundamental challenge in computational chemistry. While existing approaches show promise, most rely on simplified representations like SMILES strings or 2D graphs that lack essential three-dimensional geometric information. We present EvoDiffMol, a computational framework that integrates evolutionary algorithms with three-dimensional diffusion models for property-driven molecular generation. The method operates through adaptive evolutionary optimization, where population-based selection guides the generation process toward desired property landscapes. EvoDiffMol supports both unconstrained molecular design and scaffold-constrained generation that preserves fixed substructures while optimizing complementary regions. Comprehensive evaluation demonstrates exceptional performance, achieving the highest drug-likeness score (0.94) among all compared state-of-the-art methods while maintaining excellent validity, uniqueness, and novelty. Beyond single property optimization, the framework demonstrates flexible multi-property optimization capabilities, simultaneously controlling multiple molecular descriptors including synthetic accessibility, lipophilicity, topological polar surface area, and clinically relevant ADMET properties such as cardiotoxicity (hERG) and intestinal permeability (Caco-2). This adaptability spans from simple descriptors to practical pharmaceutical endpoints without requiring complete model retraining. The framework achieves precise control over target property values, generating molecules with properties closely matching specified targets for both single and multiple descriptors. Scaffold-constrained experiments preserve fixed molecular cores while maintaining effective property optimization. The three-dimensional representation offers advantages in maintaining structural validity during iterative optimization, with potential for geometry-aware applications in materials science and drug discovery.

3D molecular generation↗

Applying Mechanistic Co-Design Towards Room-Temperature Organic, Optical, and Multiple-Center Molecular Qubits

The next generation of qubits, the building blocks of quantum technologies, are designed using molecules. Molecular qubits have natural advantages over solid-state equivalents in biological and chemical sensing. Chemistry allows the properties of molecular qubits to be fine-tuned, such as the ability of multiple qubits to work together using qubit-qubit coupling. Molecules and structures containing multiple qubits, made from both inorganic and organic elements, are developed in this work to bring quantum technologies closer to room-temperature operation. The goal is achieved by understanding how the core of the qubit, known as spin, interacts with the vibrating parts of the molecule, other spins, and the surrounding environment. A combined approach involving theory, spectroscopy, and synthesis is used to achieve the design goal.

36 MATERIALS SCIENCE↗

Applying Contamination Modelling to Spacecraft Propulsion Systems Designs and Operations

Molecular and particulate contaminants generated from the operations of a propulsion system may impinge on spacecraft critical surfaces. Plume depositions or clouds may hinder the spacecraft and instruments from performing normal operations. Firing thrusters will generate both molecular and particulate contaminants. How to minimize the contamination impact from the plume becomes very critical for a successful mission. The resulting effect from either molecular or particulate contamination of the thruster firing is very distinct. This paper will discuss the interconnection between the functions of spacecraft contamination modeling and propulsion system implementation. The paper will address an innovative contamination engineering approach implemented from the spacecraft concept design, manufacturing, integration and test (I&T), launch, to on- orbit operations. This paper will also summarize the implementation on several successful missions. Despite other contamination sources, only molecular contamination will be considered here.

Chen, Philip T.↗

Polymer Deconstruction and Redesign Strategies for Plastics Recycling

Advancing plastics recycling requires both the selective deconstruction of existing polymers and the design of new materials that enable efficient reuse without loss of performance. This perspective highlights an integrated approach that is rooted in polymer chemistry, catalysis, and process engineering which can enable a circular plastics economy. Here, we outline recent advances in catalytic, solvolytic, and enzymatic pathways for plastic deconstruction, and examine the molecular design principles driving next-generation recyclable-by-design and bio-based polymers. Despite these advances, major knowledge gaps remain in understanding the evolution of polymer morphology and catalyst structure during deconstruction, assessing deconstruction processes with realistic polymers, and offering redesigned polymers with competitive cost and environmental advantage over conventional plastics. United States Department of Energy (U.S. DOE) national laboratories offer unique capabilities to address these challenges through in situ and operando characterization, high-throughput experimentation, environmental studies, technoeconomic and life cycle assessment, scale-up support, and collaboration networks. Advances made in understanding plastic deconstruction mechanisms and structure-property correlations of redesigned polymers inform emerging research directions including autonomous experimentation, real-time feedback-enabled process optimization, and protein engineering for enzymatic depolymerization.

36 MATERIALS SCIENCE↗

Molecular beam generator

Vacuum deposition generator has nozzle and aperature designed especially for beams of heavy organic molecules. Deposition rates are from 6 to 15 angstroms per minute.

Richmond, R. G.↗

Evaluation of δ-Phase ZrH1.4 to ZrH1.7 Thermal Neutron Scattering Laws Using Ab Initio Molecular Dynamics Simulations

Zirconium hydride is commonly used for next-generation reactor designs due to its excellent hydrogen retention capacity at temperatures below 1000 K. These types of reactors operate at thermal neutron energies and require accurate representation of thermal scattering laws (TSLs) to optimize moderator performance and evaluate the safety indicators for reactor design. In this work, we present an atomic-scale representation of sub-stoichiometric ZrH2−x(0.3≤x≤0.6), which relies on ab initio molecular dynamics (AIMD) in tandem with velocity auto-correlation (VAC) analysis to generate phonon density of states (DOS) for TSL development. The novel NJOY+NCrystal tool, developed by the European Spallation Source community, was utilized to generate the TSL formulations in the A Compact ENDF (ACE) format for its utility in neutron transport software. First, stoichiometric zirconium hydride cross sections were benchmarked with experiments. Then sub-stoichiometric zirconium hydride TSLs were developed. Significant deviations were observed between the new δ-phase ZrH2−x TSLs and the TSLs in the current ENDF release. It was also observed that varying the hydrogen vacancy defect concentration and sites did not cause as significant a change in the TSLs (e.g., ZrH1.4 vs. ZrH1.7) as was caused by the lattice transformation from ϵ- to δ-phase.

42 ENGINEERING↗

Spatiotemporal 4D Whole-cell Modeling of a Minimal Autotroph Reveals Central Carbon Metabolism Regulated Locally by Protein Megacomplexes via Post-translational Modifications under Light Disturbance

Photosynthetic microorganisms rely on multiple pathways in central carbon metabolism to adapt to fluctuating light and energy availability across diel cycles. Mechanistic insight into the regulatory dynamics of this adaptation requires integrating processes spanning disparate timescales, from rapid redox-dependent post-translational modifications (PTMs) to slower changes in protein expression and metabolic pathway usage. To address this complexity beyond genome-based inference and traditional modeling, we develop a whole-cell four-dimensional (3D + time) model of the marine cyanobacterium Prochlorococcus marinus MED4 that explicitly represents the spatial organization of enzymatic and molecular processes in central carbon metabolism under light perturbation. We employ a perturbation-based research design to experimentally generate time-series, multi-omics measurements that provide molecular descriptors and cryo-ET derived 3D segmented volumes as constraints for this dynamic 4D framework. The integration of experiments and modeling across defined light regimes enables quantitative validation of system-level responses and forecasting under distinct light disturbances. We test the hypothesis that light-dependent redox PTMs regulating the structural assembly of a protein megacomplex, the “dark complex,” modulate metabolic flux at a conserved regulatory node of the Calvin–Benson cycle (CBC) in cyanobacteria. Our model shows that subcellular spatial organization buffers rapid light-induced changes in thylakoid reaction rates, which are followed by redox-PTM-mediated sequestration or release of CBC enzymes in the dark complex, ultimately impacting carbon fixation dynamics within carboxysomes. Comparison with an equivalently parameterized well-mixed stochastic model demonstrates that post-translational regulation not only buffers transcriptional noise and diffusion-driven fluctuations but also stabilizes phenotypic outcomes, underscoring the importance of spatial heterogeneity in phenotypic robustness. This ability to probe adaptive, spatiotemporally resolved mechanisms in photosynthetic machinery and central carbon metabolism addresses a critical gap in genotype-to-phenotype inference and expands modeling and design capabilities for understudied or genetically intractable autotrophs such as P. marinus MED4.

Johnson, Connah G.↗

Leveraging Natural Language Processing and Generative Models in Molecular Chemistry: Property Prediction and Novel Compound Generation

The accurate prediction of molecular properties is important for the rational design and the advancement of green chemistry and sustainable materials research. However, the predictive power of traditional computational chemistry methods is limited due to computational restrictions. Here, in this study, we examine an alternative approach to the accurate prediction of properties of organic compounds: natural language processing (NLP)-based molecular embedding. Using viscosity, partition coefficient (log P), and enthalpy of vaporization as test properties through a survey of comprehensive datasets comprising 5695 data points for viscosity, 25 870 data points for log P, and 2296 data points for enthalpy of vaporization. These are important properties for the design of greener, safer, and sustainable chemical processes. Models were trained using NLP methods such as Mol2vec and fine-tuned ChemBERTa, and results were compared with traditional input featurization techniques such as Morgan fingerprints and quantum chemistry derived sigma profiles and DFT features. Among the various machine learning models, Mol2vec demonstrated superior predictive capabilities, achieving the highest correlation coefficient (R 2 = 0.945) and lowest RMSE (0.106 mPa s) for viscosity, as well as high accuracy for log P and enthalpy of vaporization predictions. These findings establish the Mol2vec featurization technique, graph-convolutional neural networks (GCNN), and fine-tuned ChemBERTa model as powerful tools for predictive modeling of organic compounds properties, offering a significant improvement over previously used featurization techniques and opening up strategies for very-high-throughput computational screening. Finally, we integrated ML models with hybrid language-model-based generative adversarial networks (LM-GAN) to generate novel molecular sequences with desirable properties for different research applications. The ability to computationally design solvents with lower viscosity, lower log P, and lower enthalpy of vaporization offers a data-driven route to accelerating the discovery of sustainable alternatives to traditionally toxic solvents.

ChemBERTa↗

Bond-centric modular design of protein assemblies

Directional interactions that generate regular coordination geometries are a powerful means of guiding molecular and colloidal self-assembly, but implementing such high-level interactions with proteins remains challenging due to their complex shapes and intricate interface properties. Here we describe a modular approach to protein nanomaterial design inspired by the rich chemical diversity that can be generated from the small number of atomic valencies. We design protein building blocks using deep learning-based generative tools, incorporating regular coordination geometries and tailorable bonding interactions that enable the assembly of diverse closed and open architectures guided by simple geometric principles. Experimental characterization confirms the successful formation of more than 20 multicomponent polyhedral protein cages, two-dimensional arrays and three-dimensional protein lattices, with a high (10%–50%) success rate and electron microscopy data closely matching the corresponding design models. Due to modularity, individual building blocks can assemble with different partners to generate distinct regular assemblies, resulting in an economy of parts and enabling the construction of reconfigurable networks for designer nanomaterials.

Biomaterials – proteins↗