Search NASA⌕ Search

SEARCH · Search NASA

Results for “Generalizable models”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 145 records · Page 8

Combining ToF‐SIMS and Multivariate Analysis to Resolve Active Sites on Ni‐Based HER Catalysts

Unambiguous identification of active sites in heterogeneous catalysis remains a major challenge, particularly for materials with ultrathin, chemically mixed surface layers. Here, we demonstrate a generalizable approach that combines time-of-flight secondary ion mass spectrometry (ToF-SIMS) with multivariate statistical analysis (principal component analysis [PCA] and multivariate curve resolution [MCR]) to resolve catalytically relevant motifs at the nanoscale. Using Ni electrodes as a model system, PCA distinguished hydroxide-enriched domains from oxide- and metal-rich regions, while MCR decomposed depth profiles and 3D images into hydroxide, oxide, and metallic layers with nanometer resolution. A unique secondary-ion fragment, NiO 3 H 3 − (m/z 108.94), emerged as a marker of hydroxide-rich environments and correlated with hydrogen evolution reaction (HER) activity across a series of Ni electrodes. Complementary density functional theory (DFT) calculations revealed that Ni(OH) 2 clusters adjacent to metallic Ni offer the most favorable water dissociation energetics, establishing the structural origin of the marker. While illustrated here for Ni-based HER, this workflow provides a broadly applicable framework to isolate and rank near-surface patterns that govern catalytic activity, thereby extending ToF-SIMS from a qualitative probe to a predictive tool for active site identification.

HER active sites↗

Strangers in a foreign land: ‘Yeastizing’ plant enzymes

Abstract Expressing plant metabolic pathways in microbial platforms is an efficient, cost‐effective solution for producing many desired plant compounds. As eukaryotic organisms, yeasts are often the preferred platform. However, expression of plant enzymes in a yeast frequently leads to failure because the enzymes are poorly adapted to the foreign yeast cellular environment. Here, we first summarize the current engineering approaches for optimizing performance of plant enzymes in yeast. A critical limitation of these approaches is that they are labour‐intensive and must be customized for each individual enzyme, which significantly hinders the establishment of plant pathways in cellular factories. In response to this challenge, we propose the development of a cost‐effective computational pipeline to redesign plant enzymes for better adaptation to the yeast cellular milieu. This proposition is underpinned by compelling evidence that plant and yeast enzymes exhibit distinct sequence features that are generalizable across enzyme families. Consequently, we introduce a data‐driven machine learning framework designed to extract ‘yeastizing’ rules from natural protein sequence variations, which can be broadly applied to all enzymes. Additionally, we discuss the potential to integrate the machine learning model into a full design‐build‐test cycle.

59 BASIC BIOLOGICAL SCIENCES↗

Distinguishing Orbiting and Infalling Dark Matter Particles with Machine Learning

Dark matter halos are typically defined as spheres that enclose some overdensity, but these sharp, somewhat arbitrary boundaries introduce nonphysical artifacts such as backsplash halos, pseudo-volution, and an incomplete accounting of halo mass. A more physically motivated alternative is to define halos as the collection of particles that are physically orbiting within their potential well. However, existing methods to classify particles as orbiting or infalling suffer from trade-offs between accuracy, computational cost, and generalizability across cosmologies. We present an efficient, yet accurate, supervised machine learning approach using decision trees. The classification is based on only the particle radii and velocities at two epochs. Compared to detailed analysis of particle trajectories, we find that our model matches the classification of 97% of particles. Consequently, we are able to quickly and accurately reproduce the density profiles of the orbiting and infalling components out to many virial radii. We demonstrate that our model generalizes to a significantly different cosmology that lies outside the training data set. We make publicly available both our final model and the code to train similar models.

79 ASTRONOMY AND ASTROPHYSICS↗

Roadmap for transforming heterogeneous catalysis with artificial intelligence

Artificial intelligence (AI) is poised to transform heterogeneous catalysis, opening avenues for catalytic materials discovery. By uncovering intricate patterns in high-dimensional data, AI has been reshaping our pursuit of sustainable catalytic processes across the energy, environmental and chemical sectors. This promise, however, hinges on overcoming fundamental barriers, including limitations in data availability and quality, challenges in the generalizability and interpretability of data-augmented decisions, and the persistent gap between in silico predictions and experiments. Furthermore, we outline a forward-looking roadmap for deeply integrating AI into heterogeneous catalysis with an AI-ready data ecosystem, multimodal foundation models, and ultimately autonomous laboratories to accelerate the development of next-generation catalytic technologies via AI-empowered human–machine collaboration.

Computational methods↗

The unique architecture of umbrella toxins permits a two-tiered molecular bet-hedging strategy for interbacterial antagonism

Bacteria exist in competitive and rapidly changing environments in which the nature of future threats cannot be easily predicted. Streptomyces coelicolor produces three antibacterial umbrella particles that harbor distinct polymorphic toxin domains and an overlapping set of six diversified lectins. Here, we show that the exquisite specificity of umbrella particles derives from lectin-mediated species-specific binding to previously undescribed hypervariable surface glycoconjugates. A cryo-electron microscopy (cryo-EM) structure of one such lectin in complex with its oligosaccharide substrate defines the molecular basis for targeting through the coordinated recognition of multiple glycan features. Biochemical and genetic studies of several target species, in conjunction with lectin-swapping experiments, support a model whereby S. coelicolor umbrella toxin diversification at the levels of lectin composition and toxin polymorphism represents a unique, two-tiered bet-hedging strategy. Bioinformatic analyses support this as a means by which the unusual architecture of umbrella toxins offers Streptomyces a generalizable strategy to antagonize an unpredictable array of competitors.

59 BASIC BIOLOGICAL SCIENCES↗

The Second Skin: A Wearable Sensor Suite that Enables Real-Time Human Biomechanics Tracking Through Deep Learning

Objective: Real-time determination of human kinematics and kinetics could advance biomechanics research and enable valuable applications of biofeedback and generalizable exoskeleton control. Here, this work aims to investigate a taskindependent, user-independent method for obtaining precise realtime joint state estimation across lower-body joints during a wide variety of tasks. Methods: We developed a generalizable sensing approach using a suit comprised of inertial measurement units (IMUs) and pressure insoles. With the suit, we collected a dataset of 33 tasks commonly performed during construction and hazardous waste cleanup (N = 10). We then trained deep learning user-independent, task-agnostic models to estimate joint lowerbody kinematics and dynamics using only worn sensor data. We likewise computed joint kinematics and dynamics analytically from sensor data to serve as a comparison tool for model results. Results: Our models achieved overall angle estimation root-meansquared-errors (RMSE) of 6.56±.92°, 8.60±1.01°, 7.58±.89°, and 6.00±.73° compared to 13.9±.1.3°, 15.31±1.0°, 10.76±.70°, and 7.56±.48° via analytical methods at the lower back, hip, knee, and ankle, respectively. Likewise, our models achieved overall normalized moment estimation RMSEs of .207±.069 Nm/kg, .242±.044 Nm/kg, .202±.038 Nm/kg, and .193±.034 Nm/kg compared to .306±.036 Nm/kg, .407±.021 Nm/kg, 1.18 ±.022 Nm/kg, and 1.73±.071 Nm/kg via analytical methods at the lower back, hip, knee, and ankle, respectively. Conclusion: These results are comparable to other state-of-the-art wearable sensing systems, establishing deep learning as a viable sensing approach that generalizes to new users and tasks. Significance: This work shows promise for enabling accurate real-world biomechanical data collection and enhancement of biofeedback systems and wearable robot control.

Casey, Ryan T. F. [Georgia Institute of Technology↗

Generalizable, fast, and accurate DeepQSPR with fastprop

Abstract Quantitative Structure–Property Relationship studies (QSPR), often referred to interchangeably as QSAR, seek to establish a mapping between molecular structure and an arbitrary target property. Historically this was done on a target-by-target basis with new descriptors being devised to specifically map to a given target. Today software packages exist that calculate thousands of these descriptors, enabling general modeling typically with classical and machine learning methods. Also present today are learned representation methods in which deep learning models generate a target-specific representation during training. The former requires less training data and offers improved speed and interpretability while the latter offers excellent generality, while the intersection of the two remains under-explored. This paper introduces , a software package and general Deep-QSPR framework that combines a cogent set of molecular descriptors with deep learning to achieve state-of-the-art performance on datasets ranging from tens to tens of thousands of molecules. provides both a user-friendly Command Line Interface and highly interoperable set of Python modules for the training and deployment of feedforward neural networks for property prediction. This approach yields improvements in speed and interpretability over existing methods while statistically equaling or exceeding their performance across most of the tested benchmarks. is designed with Research Software Engineering best practices and is free and open source, hosted at github.com/jacksonburns/fastprop.

Burns, Jackson W. (ORCID:0000000206579426)↗

AI-assisted rapid crystal structure generation towards a target local environment

In material design, traditional crystal structure prediction approaches are expensive as they require extensive structural sampling through expensive energy minimization methods. Emerging artificial intelligence (AI) generative models have shown great promise in rapidly generating realistic crystals, but they typically handle only a few tens of atoms per unit cell. To overcome this limitation, we introduce a symmetry-informed approach, the Local Environment Geometry-Oriented Crystal Generator (LEGO-xtal). Our method generates initial structures using AI models trained on an augmented dataset, and then optimizes them using structure descriptors rather than energy-based optimization. We demonstrate its effectiveness by expanding from 25 known low-energy sp2 carbon allotropes to over 1700, all within 0.5 eV/atom of the ground-state energy of graphite. This framework offers a generalizable strategy for the targeted design of materials with modular building blocks, such as metal-organic frameworks and battery materials.

Ridwan, Osman Goni [University of North Carolina a↗

Commutative Algebra Modeling in Materials Science – A Case Study on Metal–Organic Frameworks (MOFs)

Metal-organic frameworks (MOFs) are a class of important crystalline and highly porous materials whose hierarchical geometry and chemistry hinder interpretable predictions in materials properties. Commutative algebra is a branch of abstract algebra that has been rarely applied in data and material sciences. We introduce the first ever commutative algebra modeling and prediction in materials science. Specifically, category-specific commutative algebra (CSCA) is proposed as a new framework for MOF representation and learning. It integrates element-based categorization with multiscale algebraic invariants to encode both local coordination motifs and global network organization of MOFs. These algebraically consistent, chemically aware representations enable compact, interpretable, and data efficient modeling of MOF properties such as Henry’s constants and uptake capacities for common gases. Compared to traditional geometric and graph-based approaches, CSCA achieves comparable or superior predictive accuracy while substantially improving interpretability and stability across data sets. By aligning commutative algebra with the chemical hierarchy, the CSCA establishes a rigorous and generalizable paradigm for understanding structure and property relationships in porous materials and provides a nonlinear algebra-based framework for data-driven material discovery.

Khaemba, Caleb S.↗

Scenario storyline discovery for complex multi-actor human-natural systems

This poster was presented at the AGU Fall Meeting 2024. Abstract:Scenario analysis is a useful tool for assessing the impacts of future conditions or alternative strategies. However, the common practice of focusing on a small number of predetermined scenarios can limit our understanding of key uncertainties, and fail to represent diverse stakeholder impacts. Exploratory modeling approaches have been developed to address these issues by simulating a wide range of possible futures and system perspectives. A challenge with these approaches is that they often involve large ensemble experiments which limit interpretability and usability. We recently introduced the FRamework for Narrative Storylines and Impact Classification (FRNSIC; pronounced ``forensic''), a scenario discovery framework that helps users identify scenario storylines that capture key system dynamics and as well as important outcomes. In this poster presentation, we present training materials to support the generalizable application of the framework to other multi-actor systems with complex dynamics. Specifically, we will present a step-by-step methodological typology of tools and methods that can be used to generate and classify plausible states of the world on key metrics and consequential dynamics. The typology will also discuss potential implications of these choices and their applicability to different systems.

Colorado River↗

Autonomous alloy composition optimization using molecular dynamics guided by a large language model

Here, we present an autonomous materials discovery framework that couples a large language model (LLM) with molecular dynamics (MD) simulations to optimize Fe–Cr–Mn alloy compositions for tensile strength. Starting from six distinct compositions, the LLM operated as an intelligent agent, iteratively proposing changes based on prior simulation results and constraints. Over 50 iterations per case, the LLM adaptively explored the composition space, identifying high-strength regions, not easily accessible by conventional methods. The highest strength, 18.7 GPa, was achieved with Fe 71 Cr 25 Mn 4 composition, identified from a Fe 75 Cr 20 Mn 5 starting point. The LLM autonomously adjusted its strategy in real time, demonstrating closed-loop decision-making using commodity hardware. This approach showcases the potential of LLMs as scientific co-pilots, capable of accelerating materials discovery and generalizable to other domains like biology and drug design.

Autonomy↗

Catalytically Active Site Mapping Realized through Energy Transfer Modeling

Abstract The demands of a sustainable chemical industry are a driving force for the development of heterogeneous catalytic platforms exhibiting facile catalyst recovery, recycling, and resilience to diverse reaction conditions. Homogeneous‐to‐heterogeneous catalyst transitions can be realized through the integration of efficient homogeneous catalysts within porous matrices. Herein, we offer a versatile approach to understanding how guest distribution and evolution impact the catalytic performance of heterogeneous host–guest catalytic platforms by implementing the resonance energy transfer (RET) concept using fluorescent model systems mimicking the steric constraints of targeted catalysts. Using the RET‐based methodology, we mapped condition‐dependent guest (re)distribution within a porous support on the example of modular matrices such as metal–organic frameworks (MOFs). Furthermore, we correlate RET results performed on the model systems with the catalytic performance of two MOF‐encapsulated catalysts used to promote CO 2 hydrogenation and ring‐closing metathesis. Guests are incorporated using aperture‐opening encapsulation, and catalyst redistribution is not observed under practical reaction conditions, showcasing a pathway to advance catalyst recyclability in the case of host–guest platforms. These studies represent the first generalizable approach for mapping the guest distribution in heterogeneous host–guest catalytic systems, providing a foundation for predicting and tailoring the performance of catalysts integrated into various porous supports.

Thompson, William J.↗

Catalytically Active Site Mapping Realized through Energy Transfer Modeling

Abstract The demands of a sustainable chemical industry are a driving force for the development of heterogeneous catalytic platforms exhibiting facile catalyst recovery, recycling, and resilience to diverse reaction conditions. Homogeneous‐to‐heterogeneous catalyst transitions can be realized through the integration of efficient homogeneous catalysts within porous matrices. Herein, we offer a versatile approach to understanding how guest distribution and evolution impact the catalytic performance of heterogeneous host–guest catalytic platforms by implementing the resonance energy transfer (RET) concept using fluorescent model systems mimicking the steric constraints of targeted catalysts. Using the RET‐based methodology, we mapped condition‐dependent guest (re)distribution within a porous support on the example of modular matrices such as metal–organic frameworks (MOFs). Furthermore, we correlate RET results performed on the model systems with the catalytic performance of two MOF‐encapsulated catalysts used to promote CO 2 hydrogenation and ring‐closing metathesis. Guests are incorporated using aperture‐opening encapsulation, and catalyst redistribution is not observed under practical reaction conditions, showcasing a pathway to advance catalyst recyclability in the case of host–guest platforms. These studies represent the first generalizable approach for mapping the guest distribution in heterogeneous host–guest catalytic systems, providing a foundation for predicting and tailoring the performance of catalysts integrated into various porous supports.

Thompson, William J.↗

Machine learning for reparameterization of multi-scale closures

Scientific machine learning (ML) is becoming increasingly useful in learning closure models for multi-scale physics problems; however, many ML approaches require a vast array of training data and can struggle with generalization and interpretability. Here, rather than learning an entire closure operator, we adopt an existing reduced-dimension model of the microphysics and learn an optimal re-parameterization of the solver. We demonstrate two approaches for training the reduced dimension closure model (1) an a priori method that optimizes the closure parameterization and the neural network parameters separately and (2) an a posteriori method that simultaneously optimizes both. Using the simulation of biomass pyrolysis as a motivating example, we show that the a posteriori method achieves better target losses and is less dependent on training dataset size for generalizability. We then demonstrate the impact that implementing this reparameterization has at the macroscale, showing improved predictive performance with no modification to the underlying macroscale solvers.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

HydraGNN_Predictive_GFM_2026 - Ensemble of predictive graph foundation models for atomistic materials modeling

This release contains data and parameters of HydraGNN-based graph foundation models trained as a result of the work published in the pre-print "Exascale Multi-Task Graph Foundation Models for Imbalanced, Multi-Fidelity Atomistic Data" by M. Lupo Pasini et al. (https://arxiv.org/abs/2604.15380). We jointly train on 16 open first-principles datasets (544+ million structures covering 85+ elements) using a multi-task architecture with per-dataset heads and a scalable ADIOS2/DDStore data pipeline. On Frontier, we execute six large-scale DeepHyper hyperparameter optimization campaigns in FP64 and promote the top-performing message-passing models to sustained 2,048-node training, yielding a PaiNN-based lead model. The version of HydraGNN used to generate the outputs provided in this release is HydraGNN v5.0 (https://github.com/ORNL/HydraGNN/releases/tag/v5.0) The list of datasets used for the training of the graph foundation model is the following: 1) Alexandria [1] 2) ANI1x [2] 3) MPTrj [3] 4) Open Catalyst 2020 (OC20) [4] 5) Open Catalyst 2022 (OC22) [5] 6) Open Catalyst 2025 (OC25) [6] 7) Open Direct ir Capture 2023 (ODAC23) [7] 8) Open Materials 2024 (OMat24) [8] 9) Open Molecules 2025 (OMol25) [9] 10) OMol25-neutral (subset of OMol25 that contains only molecules with zero total charge) 11) OMol25-non-neutral (subset of OMol25 that contains only molecules with non-zero total charge) 12) Open Polymers 2026 (OPoly2026) [10] 13) Nabla2DFT [11] 14) QCML [12] 15) QM7X [reference 13] 16) transition1x [14] Dataset references: [1] J. Schmidt et al., “A dataset of 175k stable and metastable materials calculated with the PBEsol and SCAN functionals,” Scientific Data, vol. 9, p. 64, 2022. [2] J. S. Smith et al., “The ANI-1ccx and ANI-1x data sets, coupled-cluster and density functional theory properties for molecules,” Scientific Data, vol. 7, p. 134, 2020. [Online]. Available: https: //www.nature.com/articles/s41597-020-0473-z [3] A. Jain et al., “Commentary: The Materials Project: A materials genome approach to accelerating materials innovation,” APL Materials, vol. 1, no. 1, p. 011002, 07 2013. [Online]. Available: https://doi.org/10.1063/1.4812323 [4] L. Chanussot et al., “Open catalyst 2020 (oc20) dataset and community challenges,” ACS Catalysis, vol. 11, no. 10, pp. 6059–6072, 2021. [Online]. Available: https://doi.org/10.1021/acscatal.0c04525 [5] K. Tran et al., “Open catalyst 2022 (oc22) dataset and challenges for oxidation electrocatalysts,” ACS Catalysis, vol. 13, no. 5, pp. 3066–3084, 2023. [Online]. Available: https://doi.org/10.1021/acscatal.2c05426 [6] S. J. Sahoo et al., “The open catalyst 2025 (oc25) dataset and models for solid-liquid interfaces,” arXiv preprint arXiv:2509.17862, 2025. [Online]. Available: https://arxiv.org/abs/2509.17862 [7] A. Sriram et al., “The open DAC 2023 dataset and challenges for sorbent discovery in direct air capture,” ACS Central Science, vol. 10, no. 5, pp. 923–941, 2024. [8] L. Barroso-Luque et al., “Open materials 2024 (omat24) inorganic materials dataset and models,” 2024. [Online]. Available: https://arxiv.org/abs/2410.12771 [9] D. S. Levine et al., “The open molecules 2025 (OMol25) dataset, evaluations, and models,” 2025. [Online]. Available: https://arxiv.org/abs/2505.08762 [10] D. S. Levine et al., The open polymers 2026 (OPoly26) dataset and evaluations,” arXiv preprint arXiv:2512.23117, 2025. [Online]. Available: https://arxiv.org/abs/2512.23117 [11] K. Khrabrov et al., “Nabla2dft: A universal quantum chemistry dataset of drug-like molecules and a benchmark for neural network potentials,” in NeurIPS 2024 Datasets and Benchmarks Track, 2024. [Online]. Available: https://openreview.net/forum?id=ElUrNM9U8c [12] S. Ganscha et al., “The QCML dataset, quantum chemistry reference data from 33.5M DFT and 14.7B semi-empirical calculations,” Scientific Data, vol. 12, p. 406, 2025. [13] J. Hoja et al., “QM7-X, a comprehensive dataset of quantum-mechanical properties spanning the chemical space of small organic molecules,” Scientific Data, vol. 8, p. 43, 2021. [Online]. Available: https://www.nature.com/articles/s41597-021-00812-2 [14] M. Schreiner et al., “Transition1x - a dataset for building generalizable reactive machine learning potentials,” Scientific Data, vol. 9, p. 779, 2022. The folder "datasets_ADIOS2_format" contains the set of pre-processed datasets in Adaptable I/O System (ADIOS) format (https://www.exascaleproject.org/research-project/adios/) that have been used for the development and training of GFMs in this work. The "datasets_ADIOS2_format" directory contains 2 sub-directories, one for the version "v1" of the datasets and one for the version "v2" of the datasets. The version "v1" of the datasets provides values of the total energy as they are extracted from the original data as it was released by the respective institutions. The version "v2" of the datasets provides values of the energy that have been realigned. The realignment was performed by training a linear regression model that predicts the total energy as a function of the chemical composition of the atomistic structure, and then subtract such prediction from the original value of the total energy. Both folders "v1" and "v2" contain 16 sub-directories, each corresponding to an ADIOS2-formatted dataset The folder "DeepHyper-results" contains the configurational files and model's parameters for all the 186 HPO trials that were successfully completed by the scalable hyperparameter optimization (HPO) runs on Frontier. The content of the folder "DeepHyper-results" I structured as follows: 1) task-list.txt: list of mpnn name, jobid, and deephyper task id 2) gfm_${MPNN}_${JOBID}_0.${TASKID}: run directory with checkpoint files 3) gfm_${MPNN}: deephyper summary directory (*.csv) for each specific MPNN type 4) deephyper-experiment-${JOBID}: output and error logs for each job The file "deephyper-sorted.csv" contains the details of each HydraGNN model built and tested by HPO, obtained by merging the (*.csv) filed from each HPO run executed. Out of all the HPO trials, we selected 10 to continue the training of the respective HydraGNN models. Due to limited computational budget available in the LRN070 allocation we could not complete the training till convergence for all these 10 selected models. The folder "models" contains multiple sub-folders, one per each HydraGNN model trained. Each model sub-folder contains the parameters of each HydraGNN model, with multiple checkpoint-restarts. The list of sub-folders are as follows: 1) multidataset_hpo-BEST1-fp64 2) multidataset_hpo-BEST2-fp64 3) multidataset_hpo-BEST3-fp64 4) multidataset_hpo-BEST4-fp64 5) multidataset_hpo-BEST5-fp64 6) multidataset_hpo-BEST6-fp64 7) multidataset_hpo-BEST7-fp64 8) multidataset_hpo-BEST8-fp64 9) multidataset_hpo-BEST9-fp64 10) multidataset_hpo-BEST10-fp64 Within each one of these folders, additional auxiliary log files are provided with descriptions about how the training proceeded. The lead PaiNN-model is contained inside "multidataset_hpo-BEST6-fp64". The file "mlp_branch_weights" contains the parameters of the multi-layer perceptron (MLP) used to reconcile the predictions of the 16 output decoding heads of the HydragNN architectures. The MLP takes in input the chemical composition of the atomistic structure and predicts averaging weights to linearly mix the predictions of each output decoding head toward consolidating them into a single one. The folder "1.1billion-structure-inference" contains 1.1 billion atomistic structures randomly generated. Each structures is associated with energy and forces predicted with the lead-PaiNN model combined with the MLP model for reconciliation of the multi-branch predictions generated by the 16 output decoding heads. The folder "1.1billion-structure-inference" contains 9,300 (*.tar.gz) subdirectories, one per Frontier compute node used to execute the inference at exascale. Once uncompressed, each (*.tar.gz) subdirectory contains an ADIOS2 (*.bp) file container, where each atomistic structure is stored as a PyTorch-Geometric Data object. The file "export_dataset_environment_variables.sh" contains the environment variables that need to be set before running the HydraGNN code to reproduce the results provided in this dataset release. The code that can be used to load the ADIOS2 files, load HydraGNN models, and run inference is available at: https://github.com/ORNL/HydraGNN/releases/tag/v5.0

36 MATERIALS SCIENCE↗

Atomate2: modular workflows for materials science

High-throughput density functional theory (DFT) calculations have become a vital element of computational materials science, enabling materials screening, property database generation, and training of “universal” machine learning models. While several software frameworks have emerged to support these computational efforts, new developments such as machine learned force fields have increased demands for more flexible and programmable workflow solutions. This manuscript introduces atomate2, a comprehensive evolution of our original atomate framework, designed to address existing limitations in computational materials research infrastructure. Key features include the support for multiple electronic structure packages and interoperability between them, along with generalizable workflows that can be written in an abstract form irrespective of the DFT package or machine learning force field used within them. Our hope is that atomate2's improved usability and extensibility can reduce technical barriers for high-throughput research workflows and facilitate the rapid adoption of emerging methods in computational material science.

97 MATHEMATICS AND COMPUTING↗

Controlled Acidity Gradients Enable CO 2 Reduction to Formic Acid (Not Formate) by Molecular Electrocatalysts

Neutral or basic conditions are commonly required for the selective electrochemical reduction of CO 2 , leading to the accumulation of carbonate salts and the generation of formate rather than formic acid. A generalizable strategy for obtaining formic acid (not formate) in the electroreduction of CO 2 with molecular catalysts is introduced, based on controlling acidity gradients using a dual-electrolyte cell with a proton-exchange membrane. This approach uses anodic water oxidation as the source of protons and electrons for CO 2 reduction to formic acid, while mitigating H 2 evolution near the cathode and avoiding carbonate formation. Mechanistic studies, including systems modeling, provide insight into the origin of the formic acid selectivity and guide the broader implementation of this strategy in molecular electrocatalysis for CO 2 utilization.

Alcohols↗

Preemptive optimization of a clinical antibody for broad neutralization of SARS-CoV-2 variants and robustness against viral escape

Most previously authorized clinical antibodies against severe acute respiratory syndrome coronavirus 2 (SARS-CoV-2) have lost neutralizing activity to recent variants due to rapid viral evolution. To mitigate such escape, we preemptively enhance AZD3152, an antibody authorized for prophylaxis in immunocompromised individuals. Using deep mutational scanning (DMS) on the SARS-CoV-2 antigen, we identify AZD3152 vulnerabilities at antigen positions F456 and D420. Through two iterations of computational antibody design that integrates structure-based modeling, machine-learning, and experimental validation, we co-optimize AZD3152 against 24 contemporary and previous SARS-CoV-2 variants, as well as 20 potential future escape variants. Our top candidate, 3152-1142, restores full potency (100-fold improvement) against the more recently emerged XBB.1.5+F456L variant that escaped AZD3152, maintains potency against previous variants of concern, and shows no additional vulnerability as assessed by DMS. This preemptive mitigation demonstrates a generalizable approach for optimizing existing antibodies against potential future viral escape.

59 BASIC BIOLOGICAL SCIENCES↗