Search NASA⌕ Search

SEARCH · Search NASA

Results for “machine learning potential”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 541 records · Page 30

A Machine Learning Ready Dataset of Acoustic Power Maps for Detection of Active Region Emergence

The development of an accurate forecast for solar eruptive activity has become increasingly important in order to prevent any potential impact on activities in space and the Earth's environment. It is therefore crucial to detect active regions before they appear on the solar surface and create early warning capabilities for upcoming Space Weather disturbances. In this work, 9TB of solar data (SDO/HMI dopplergrams, magnetograms and continuum intensity maps) involving the emergence of 61 NOAA solar active regions since 2010 were processed using the NASA HECC capabilities. An acoustic power maps time-series dataset was created (for four different frequency ranges and processed to take into account the solar sphere geometric effect ) which can be used for understanding the dynamics of the solar surface and train a variety of ML models. The calculated acoustic power maps carry precursor information associated with the decrease in continuum intensity on the solar surface, verifying older helioseismology research. Our results show that a Long Short-Term Memory (LSTMs) model, with a modest layer depth and the right hyperparameters tuned, when trained on this solar acoustic power maps dataset can predict without false negatives a drop in intensity (associated with the emergence of the active region), up to 18 hours in advance.

SMD↗

Developing a Prototype Methodology to Rank CO2-EOR Wells and Assess Their Reuse Potential for Geologic Carbon Storage

This paper presents a prototype methodology to assess the possible transition of Class II carbon dioxide-enhanced oil recovery (CO2-EOR) wells to Class VI wells. The focus is on wellbore construction materials—casing, cement, tubing, and the packer—and includes comprehensive workflows to evaluate these materials, with primary emphasis on compliance with Environmental Protection Agency (EPA) Class VI well construction and conversion guidelines. These workflows systematically assess material properties and performance criteria to ensure regulatory compliance and optimize long-term wellbore integrity and functionality. Utilizing Python scripts and JavaScript Object Notation (JSON) representations, the study automates checks on digitized Texas Railroad Commission (TRRC) data to rank wells based on workflow criteria. By emphasizing critical factors such as casing integrity, cementing techniques, tubing compatibility, and packer selection, the methodology helps well owners and operators prioritize wells for potential reuse as CO2 injection wells. Given limitations in digitized data, manual user verification is required in some sections. Future improvements include integrating non-digitized data through web scraping and machine learning techniques. This research serves as a practical guide for stakeholders, supporting environmental compliance and sustainable well operations.

geologic carbon sequestration↗

Potential applications of microbial genomics in nuclear non-proliferation

As nuclear technology evolves in response to increased demand for diversification and decarbonization of the energy sector, new and innovative approaches are needed to effectively identify and deter the proliferation of nuclear arms, while ensuring safe development of global nuclear energy resources. Preventing the use of nuclear material and technology for unsanctioned development of nuclear weapons has been a long-standing challenge for the International Atomic Energy Agency and signatories of the Treaty on the Non-Proliferation of Nuclear Weapons. Environmental swipe sampling has proven to be an effective technique for characterizing clandestine proliferation activities within and around known locations of nuclear facilities and sites. However, limited tools and techniques exist for detecting nuclear proliferation in unknown locations beyond the boundaries of declared nuclear fuel cycle facilities, representing a critical gap in non-proliferation safeguards. Microbiomes, defined as “characteristic communities of microorganisms” found in specific habitats with distinct physical and chemical properties, can provide valuable information about the conditions and activities occurring in the surrounding environment. Microorganisms are known to inhabit radionuclide-contaminated sites, spent nuclear fuel storage pools, and cooling systems of water-cooled nuclear reactors, where they can cause radionuclide migration and corrosion of critical structures. Microbial transformation of radionuclides is a well-established process that has been documented in numerous field and laboratory studies. These studies helped to identify key bacterial taxa and microbially-mediated processes that directly and indirectly control the transformation, mobility, and fate of radionuclides in the environment. Expanding on this work, other studies have used microbial genomics integrated with machine learning models to successfully monitor and predict the occurrence of heavy metals, radionuclides, and other process wastes in the environment, indicating the potential role of nuclear activities in shaping microbial community structure and function. Results of this previous body of work suggest fundamental geochemical-microbial interactions occurring at nuclear fuel cycle facilities could give rise to microbiomes that are characteristic of nuclear activities. These microbiomes could provide valuable information for monitoring nuclear fuel cycle facilities, planning environmental sampling campaigns, and developing biosensor technology for the detection of undisclosed fuel cycle activities and proliferation concerns.

59 BASIC BIOLOGICAL SCIENCES↗

Open Science for Life in Space: Data Sharing and Tools for Knowledge Discovery

The next era in human space exploration is rapidly approaching and will require the use of countermeasures to deep space health hazards. The development of countermeasures (or, the re-purposing of existing agents) will be highly dependent on our understanding of basic biological responses to space stressors (e.g. ionizing radiation, altered gravitational fields, altered day-night cycles, confinement, isolation, hostile-closed environments, distance-duration from Earth, exposure to celestial regolith, etc.). The fast-growing array of space biological data, which in the past was simply archived after minimal analysis, holds great potential if it can be reorganized and formatted for Open Science. Organizing the data for such analysis is a challenge because of its diverse nature (molecular, cellular, tissue, imaging, whole organism and behavior). We will discuss here several strategies that NASA’s Biological and Physical Science Division has put in place to maximize the return on investment for spaceflight bioscience data. Open Science, as a scientific philosophy, is the concept that the more people who have access to the data, the more knowledge will be gained from it. This guiding principle led NASA to develop GeneLab in 2015. GeneLab houses spaceflight and relevant ground-based multi-omics data, and has grown to ~400 transcriptomic, proteomic, metabolomic and epigenomic datasets from plant, rodent, small animal, and microbial space experiments. GeneLab provides users with various tools for data analysis and a visualization portal that allows users to interact with gene expression data from space-related ‘omics experiments. Open Science is also about building scientific communities, and with this spirit in mind, GeneLab has spawned several Analysis Working Groups (AWGs), comprised of more than 200 volunteer scientists. The AWGs initially provided feedback on the processing pipeline and metadata ‘omics standards for GeneLab. Over the last few years, they have become a community-driven science enterprise, engaging in large meta-analysis of GeneLab datasets, resulting in 10 publications (beyond the originally submitted research). Overall, the Open Science nature of GeneLab has resulted in a high degree of data re-use, resulting in 38 additional publications derived from the original 67 publication over the past four years. The enormous success and knowledge gained from GeneLab has led to a collection of sister NASA “Open Science Data Repositories (OSDR)” and research support groups. These include the NASA Ames Life Sciences Data Archive (ALSDA), the NASA Biological Institutional Scientific Collection (NBISC), and the Biospecimen Sharing Program (BSP). All are adopting the GeneLab data architecture system to maximize open-access, find-ability, accessibility, interoperability, and reusability (FAIR). ALSDA collects and curates phenotypic-physiological bioimaging-behavioral data from space and space-relevant non-human experiments, oftentimes coming from the same omics-associated experimental datasets found in GeneLab. Since 2021, a community of ~100 researchers have rallied around ALSDA, to provide feedback in a new ALSDA AWG focused on phenotypic-physiological investigation-sample-assay metadata standards (e.g., Micro-Computed Tomography, Light/Fluorescence Microscopy, Western Blot, Flow Cytometry, Novel Object Recognition, Elevated Plus Maze, etc. of ~50 assays collected). These standards are part of a new single point-of-entry data submission portal for all non-human Space Biology and Human Research Program principal investigators, to submit, curate, and share their research data. With open-access space biological data now collected and curated together with rich metadata, and with the potential for linkage to “big data” from the international biological and medical communities (NIH, EBI, etc.), the artificial intelligence and machine learning (AI/ML) era has started for Space Biology. Several other talks will cover these topics in this conference.

life sciences↗

Open Science for Life in Space: Data Sharing and Tools for Knowledge Discovery

The next era in human space exploration is rapidly approaching and will require the use of countermeasures to deep space health hazards. The development of countermeasures (or, there-purposing of existing agents) will be highly dependent on our understanding of basic biological responses to space stressors (e.g. ionizing radiation, altered gravitational fields, altered day-night cycles, confinement, isolation, hostile-closed environments, distance-duration from Earth, exposure to celestial regolith, etc.). The fast-growing array of space biological data, which in the past was simply archived after minimal analysis, holds great potential if it can be reorganized and formatted for Open Science. Organizing the data for such analysis is a challenge because of its diverse nature (molecular, cellular, tissue, imaging, whole organism and behavior). We will discuss here several strategies that NASA's Biological and Physical Science Division has put in place to maximize the return on investment for spaceflight bioscience data. Open Science, as a scientific philosophy, is the concept that the more people who have access to the data, the more knowledge will be gained from it. This guiding principle led NASA to develop GeneLab in 2015. GeneLab houses spaceflight and relevant ground-based multi-omics data, and has grown to ~400 transcriptomatic, proteomic, metabolomic and epigenomic datasets from plant, rodent, small animal, and microbial space experiments. GeneLab provides users with various tools for data analysis and a visualization portal that allows users to interact with gene expression data from space-related 'omics experiments. Open Science is also about building scientific communities, and with this spirit in mind, GeneLab has spawned several Analysis Working Groups (AWGs), comprised of more than 200 volunteer scientists. The AWGs initially provided feedback on the processing pipeline and metadata 'omics standards for GeneLab. Over the last few years, they have become a community-driven science enterprise, engaging in large meta-analysis of GeneLab datasets, resulting in 10 publications (beyond the originally submitted research). Overall, the Open Science nature of GeneLab has resulted in a high degree of data-use, resulting in 40 enabled publications by open data. The enormous success and knowledge gained from GeneLab has led to a collection of sister NASA "Open Science Data Repositories (OSDR)" and research support groups. These include the NASA Ames Life Sciences Data Archive (ALSDA), the NASA Biological Institutional Scientific Collection (NBISC), and the Biospecimen Sharing Program (BSP). All are adopting the GeneLab data architecture system to maximize open-access, find-ability, accessibility, interoperability, and reusability (FAIR). ALSDA collects and curates phenotypic-physiological bioimaging-behavioral data from space and space-relevant non-human experiments, oftentimes coming from the same omics-associated experimental datasets found in GeneLab. Since 2021, a community of ~100 researchers have rallied around ALSDA, to provide feedback in a new ALSDA AWG focused on phenotypic-physiological investigation-sample-assay metadata standards (e.g., Micro-Computed Tomography, Light/Flourescence Microscopy, Western Blot, Flow Cytometry, Novel Object Recognition, Elevated Plus Maze, etc. of ~50 assays collected). These standards are part of a new single point-of-entry data submission portal for all non-human Space Biology and Human Research Program principal investigators, to submit, curate, and share their research data. With open-access space biological data now collected and curated together with rich metadata, and with the potential for linkage to "big data" from the international biological and medical communities (NIH, EBI, etc.), the artificial intelligence and machine learning (AI/ML) era has started for Space Biology.

omics↗

Parameterization of Vertical Cloud Distribution from C3M and MERRA Data Using ML Method

Clouds play a key role in regulating the hydrological cycle and the Earth's radiative energy budget. However, global climate models (GCMs) with a horizontal grid spacing on the order of 100 km have limitations in representing sub-grid cloud dynamics with spatial scales on the order of 1 km, leading to potential uncertainties in cloud radiative feedback on the global scale. In our research, we will leverage the capabilities of Deep Machine Learning (DML) methods to construct parameterizations of sub-grid volumetric cloud fraction (VCF), which is the frequency of occurrence on a grid volume accumulated in the horizontal and vertical directions. Our investigation delves into the intricate relationship between VCF obtained from the NASA CALIPSO-CloudSat-CERES-MODIS (CCCM) satellite observation data and 3-D MERRA-2 reanalysis meteorological profiling data (e.g., wind, relative humidity, temperature). Through a comprehensive one-year data training utilizing the Sequence to Sequence DML method, we have successfully disentangled the complicated cloud formation dynamics across diverse meteorological conditions through a day-to-day analysis framework. Preliminary findings reveal promising statistical agreements in geographical and vertical distributions and seasonal variations of volumetric cloud fraction between ML prediction and satellite measurements. These results underscore the aptitude of our DML model to discern underlying cloud physical processes and accurately represent sub-grid cloud formation dynamics. Additionally, we have also employed trained neural network to analyze uncertainties arising from errors in meteorological data, further enhancing the robustness of our VCF parameterization.

Shan Zeng↗

Machine Learned Force Field Modeling of Metal Organic Frameworks for CO2 Direct Air Capture

Metal organic frameworks (MOFs) are a large class of porous materials and have garnered significant interest due to their large surface areas and their tunable physical and chemical properties. Numerous prior studies have been performed to screen large databases of this material class for promising DAC sorbent materials. These studies have often relied on classical model potentials. While density functional theory (DFT) calculations have been shown to be very accurate for modeling the interaction of CO2 with MOFs, such calculations are too computationally demanding for statistically significant adsorption predictions. To overcome this barrier, we developed methods for training models to achieve DFT-level accuracy for the forces and energies associated with MOF flexibility and CO2 adsorption using machine learned force fields (MLFFs). These methods were parametrized based on DFT calculations of CO2 in a flexible MOF and used to predict MOF structural properties as well as CO2 adsorption in several MOFs.

Findley, John↗

Day-Ahead Probabilistic Forecasting of Net-Load and Demand Response Potentials with High Penetration of Behind-the-Meter Solar-plus-Storage

The goal of this project is to develop advanced methods for day-ahead net-load forecasting, by leveraging the state-of-the-art machine learning techniques. The developed models produce both point and probabilistic forecasts for a variety of use cases, and are versatile to work with different types of data sets. The innovation lies in the novel design of the architectures, leveraging the most recent advances in machine learning that have not been explored in power systems, accompanied by techniques in the broader artificial intelligence fields such as fuzzy systems. This project has achieved the following accomplishments: (1) preprocessing of over 10 data sets covering varying geographical regions, time horizons, and system levels, which form a robust foundation for training and evaluating forecasting models across a wide range of realistic grid scenarios; (2) development of an interactive web app that enables exploratory analysis of load and generation data, and supports better understanding of data trends, anomalies, and correlations, facilitating model development and stakeholder engagement; (3) implementation of over 10 benchmark models for point and probabilistic forecasting, which include a mix of conventional machine learning methods and state-of-the-art deep learning approaches, providing a comprehensive baseline for performance comparison and validation of the proposed models; (4) development of a fuzzy system based gradient boosting model, tailored for small (less than 3 years) data sets, which achieves a mean absolute percentage error (MAPE) of 4% for point forecasting and a 20% improvement in average pinball loss for probabilistic forecasting; (5) development of a Transformer (a state-of-the-art deep learning architecture) based neural network model, tailored for large (3 years or more) data sets, which achieves a MAPE of 2% for point forecasting and a 20% improvement in average pinball loss for probabilistic forecasting; (6) development of a methodology for quantifying DR potential, and extensions of the previous models for multi-target forecasting of net load and DR potential, which achieve a MAPE of 10% for DR potential.

24 POWER TRANSMISSION AND DISTRIBUTION↗

Using Coordinated, Multi-Agent Platforms for Dynamic Ocean Worlds Science

Planetary science missions have the opportunity to enhance science return through deployment of autonomous capabilities designed to dynamically respond to new information. Future outer solar system missions to ocean worlds in particular would benefit from this technology - intelligent science payloads (ISP) - because it would allow for a coordinated, near real-time response to ephemeral ‘events’ such as plumes, tectonism, surface implantation, volatile releases, thermal and magnetic anomalies, or radiation, as well as increasing the cadence and coverage of data collection. Prioritization and decision-making frameworks from ISP could be deployed at various scales - from analysis onboard a spacecraft with multiple instruments – to coordinated analyses among separate spacecraft in an e.g., distributed systems mission (DSM) composed of multiple SmallSats. Goddard’s Intelligent Science Payload team is developing an agile autonomous architecture for an icy ocean worlds DSM concept. Our goals are to coordinate data collection and onboard data analysis, and to make autonomous decisions for new data collection and analysis based on science priorities between multiple spacecraft with variable instrumentation and orbits. We use a range of data analysis tools to coordinate the DSM response, spanning from observations of data over a specified threshold to more computationally intensive machine learning algorithms (ML). ML algorithms here currently focus on determining the composition of an ocean world using mass spectrometry, and specifically methods for understanding ‘novelties’ and potential biosignatures. These algorithms could be used to quickly process and analyze onboard data that would be significantly delayed in downlink due to long communication delays for outer solar system missions in order to make dynamic science observations. Our ocean worlds case study ISP architecture is intended as an ‘agile’ and modular framework that could be used as a whole or as particular modules based on mission needs.

Distributed Systems↗

Social Bias in AI and its Implications

Previous studies have documented many different types of biases that exist in artificial intelligence (AI) and machine learning (ML) systems. We reviewed the literature on AI and ML bias with a focus on social implications and found that bias in AI and ML can potentially have harmful social impacts on individuals and/or groups of people. By affecting people differently according to characteristics such as race, gender, or sexual orientation, AI and ML systems may lead to harm by exacerbating social inequities. We recount examples of issues that have occurred in systems that use technology that might be used at NASA and elsewhere so that similar issues might be identified and mitigated in future systems. We also provide interested parties with a gateway into existing work on social bias in AI and ML systems.

artificial intelligence (AI)↗

A Simple Panel System to Overcome Interface Challenges for Retrofits: Preprint

Retrofitting buildings is usually an expensive and labor-intensive process. Weatherization measures can improve comfort and energy affordability to some extent, but deep energy retrofits are needed to optimize performance and comfort, and to achieve significant energy cost savings. Barriers to deep energy retrofits include a limited supply of skilled labor, different building types, planning complexity, split incentives, and a long or non-existent ROI horizon. The "Simple Panel System" (SPS) workflow developed and demonstrated in this effort streamlines deep energy retrofits by applying advanced site capture, machine learning, and mixed reality to panelized construction. The result is a one-stop, product-independent solution for rapidly scalable retrofits with the potential to reduce construction time and project costs by 50%. Soft costs are reduced by more than 66%, total costs by more than 50%, and field construction time by more than 50% - not to mention the reduction in construction waste, improvement in working conditions, and the ability to scale without an influx of skilled labor. This paper presents the SPS and the preliminary results and findings from the pilot project.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗

Optimizing Air Traffic - Integrating Artificial Intelligence and Machine Learning in Flight Path Planning and 3D Airspace Visualization for Air Traffic Control

Air Traffic Control (ATC) systems are vital components of the National Airspace System (NAS). ATC, Airport Traffic Control Towers (ATCT), and Terminal Radar Approach Control (TRACON) are responsible for directing all flights departing from and arriving at airports, managing our nation’s airspace, preventing potential accidents, and ensuring that every flight is accounted for. However, these systems often face challenges in effectively monitoring the skies. Issues such as poor communication between operators, difficulty in performing operations, and the constant need for vigilance frequently burden ATC operators. Additionally, the projected increase in air traffic in the coming years will only exacerbate the stress associated with this role. To address these issues, we propose a system that assists ATC operators in situations such as handovers, emergencies, and routing aircraft to avoid weather hazards. Our solution includes an Artificial Intelligence (AI) and Machine Learning (ML)-based Flight Pathways Planning System (FPPS) designed to find the fastest and most optimal routes for aircraft, taking into account weather conditions, restricted terrain, and Extended-Range Twin-Engine Operational Performance Standards (ETOPS) ratings. The proposed Predictive Weather Planning Model, included in FPPS, adjusts routes based on real-time and forecasted weather conditions. Additionally, our NVIDIA Omniverse 3D Visualization System offers a highly interactive environment for better visualization and a clear view of the airspace. By incorporating these systems, the roles of ATC, ATCT, and TRACON operators will become more manageable and less stressful, equipping them to efficiently handle the growing density of airspace.

Regina Ayoubi↗

Anomalous electroweak physics unraveled via evidential deep learning

The ever-growing ecosystem of beyond standard model (BSM) calculations and parametrizations has motivated the development of systematic methods for making quantitative cross-comparisons over the wide range of possible models, especially with controllable uncertainties. In this setting, the language of uncertainty quantification (UQ) furnishes useful metrics for assessing statistical overlaps and discrepancies among BSM and related models. In this study, we leverage recent machine learning (ML) developments in evidential deep learning (EDL) for UQ to separate data (aleatoric) and knowledge (epistemic) uncertainties in a model-discrimination setting. We construct several potentially BSM-motivated scenarios for the anomalous electroweak interaction (AEWI) of neutrinos with nucleons in deep inelastic scattering ( v DIS). These scenarios are then quantitatively mapped, as a demonstration, alongside Monte Carlo replicas of the CT18 PDFs used to calculate the $\varDelta \chi ^{2}$ statistic for a typical multi-GeV v DIS experiment, CDHSW. Our framework effectively highlights areas of model agreement and provides a classification of out-of-distribution (OOD) samples. By offering the opportunity to quantitatively understand model overlaps, the approach presented in this work can help facilitate efficient BSM model exploration and exclusion for future New Physics searches.

AI↗

Machine-Learned Committor Functions for Reactive Molecular Dynamics

Reactive molecular dynamics (MD) is a powerful tool for atomistic-scale modeling of a diverse range of chemical processes. However, scaling these simulations to large systems and long times scales remains a challenge because of the complexity of the potential energy function required. The authors previously developed a heuristic approach, called REACTER, that incorporates reactivity in MD simulations in a less general but much more computationally efficient manner. REACTER uses standard, fixed valence force fields as the underlying potentialenergy surface for describing all interatomic interactions but adds a procedure for enforcing user-defined reactions that occur when certain geometric constraints on relative atomic positions are satisfied. Further, these bonding changes can be accepted or rejected with a probability related tothe local thermal energy. This work seeks to generalize this approach by replacing the set of user defined geometric constraints and energetic criteria with a committor function that specifies the probability of a reaction occurring on the basis of the local atomic configuration. The committor function is a useful mathematical tool for modeling rare events but, unfortunately, is very difficult to compute for realistic systems in a general way. This work describes a method for approximating the committor function using a machine learning approach, specifically a deep neural network trained with data from reactive MD and DFT-based dynamics simulations. This network is coupled to the existing REACTER protocol, as implemented in the LAMMPS MD package, and used to make on-the-fly predictions of reaction probabilities without the more extensive user input previously required. The new method is demonstrated using the polymerization of polystyrene as a case study. Although very dependent on the quality and quantity of training data, machine-learned committor functions show promise as a method for incorporating reaction probability from higher level calculations into highly scalable MD simulations.

polymer simulations↗

LandScan mosaic enables high-resolution gridded population estimates with explicit uncertainty

Gridded population datasets represent high-resolution distributions of human occupancy, enabling informed decision-making across a broad range of fields. These data products are valuable for assessing environmental risk, urban development, disaster preparedness and resource allocation—areas where accurate population estimates directly enhance policy effectiveness and optimize resource distribution. Despite the importance of gridded population datasets, traditional population modeling approaches often overlook inherent uncertainties in the estimation process. This limitation can create a false sense of certainty in population estimates, potentially leading to flawed decisions by those who rely on the data. To address this methodological gap, we introduce a probabilistic machine learning modeling framework, LandScan Mosaic, that explicitly incorporates uncertainty into the population modeling process. Our approach systematically quantifies uncertainty in three key modeling parameters of the LandScan HD gridded population dataset: building use types, floor counts, and occupancy rates. By employing Monte Carlo simulations, we propagate these uncertainties through the modeling process, yielding probability distributions of population counts in place of deterministic point estimates. We demonstrate the practical application of this framework in Iloilo City, Philippines, using structured decision-making techniques and our probabilistic estimates to identify and prioritize areas most affected by projected flooding, supporting targeted interventions that address both economic and social risks. In doing so, we propose a population-specific approach for incorporating confidence into structured decision making processes. Through a comparative analysis with conventional deterministic approaches and point estimate approaches, including LandScan HD and WorldPop, we evaluate how the incorporation of machine learning and uncertainty influences decision rankings. This research advances population distribution modeling by offering a robust, quantitative approach that explicitly accounts for uncertainty in the underlying data, along with guidance for how users can apply uncertainty in their decision-making.

Environmental sciences↗

Operando probing dynamic migration of copper carbonyl during electrocatalytic CO2 reduction

Single crystals and shape-controlled nanocrystals are well known to exhibit facet-dependent catalytic properties. However, few studies have investigated how those nanocrystals evolve and (de)activate during reactions, calling for the development of nanoscale time-resolved operando methods. In this context, we have designed Cu nanocubes as a model system to elucidate the underlying driving force of dynamic nanocatalyst reconstruction during the CO2 reduction reaction (CO2RR). Operando electrochemical liquid-cell scanning transmission electron microscopy (EC-STEM) and synchrotron-based X-ray spectroscopy reveal the size- and potential-dependent complete transformation from (100)-oriented Cu@Cu2O nanocubes to polycrystalline metallic Cu nanograins under CO2RR conditions. In addition, machine learning-assisted operando four-dimensional STEM reveals that large Cu nanograins derived from nanocubes form mainly crystalline domains, while their smaller counterparts are more amorphous due to faster evolution kinetics. In situ Raman spectroscopy and density functional theory calculations suggest that CO drives the ejection of single Cu atoms, resulting in few-nanometre Cu clusters and the surface migration of highly mobile copper carbonyl (Cu–CO) species. Combined, these multimodal operando methods and theoretical approaches pave the way for understanding the complex structural evolution of energy-related nanocatalysts under electrochemical conditions.

Yang, Yao↗

A foundation model for atomistic materials chemistry

Atomistic simulations of matter, especially those that leverage first-principles (ab initio) electronic structure theory, provide a microscopic view of the world, underpinning much of our understanding of chemistry and materials science. Over the last decade or so, machine-learned force fields have transformed atomistic modeling by enabling simulations of ab initio quality over unprecedented time and length scales. However, early machine-learning (ML) force fields have largely been limited by (i) the substantial computational and human effort required to develop and validate potentials for each particular system of interest and (ii) a general lack of transferability from one chemical system to the next. Here, we show that it is possible to create a general-purpose atomistic ML model, trained on a public dataset of moderate size, that is capable of running stable molecular dynamics for a wide range of molecules and materials. We demonstrate the power of the MACE-MP-0 model-and its qualitative and at times quantitative accuracy-on a diverse set of problems in the physical sciences, including properties of solids, liquids, gases, chemical reactions, interfaces, and even the dynamics of a small protein. The model can be applied out of the box as a starting or "foundation" model for any atomistic system of interest and, when desired, can be fine-tuned on just a handful of application-specific data points to reach ab initio accuracy. Establishing that a stable force-field model can cover almost all materials changes atomistic modeling in a fundamental way: experienced users obtain reliable results much faster, and beginners face a lower barrier to entry. Foundation models thus represent a step toward democratizing the revolution in atomic-scale modeling that has been brought about by ML force fields.

Batatia, Ilyes↗

Learning Molecular Mixture Property Using Chemistry-Aware Graph Neural Network

Recent advances in machine learning (ML) are expediting materials discovery and design. One significant challenge facing ML for materials is the expansive combinatorial space of potential materials formed by diverse constituents and their flexible configurations. This complexity is particularly evident in molecular mixtures, a frequently explored space for materials, such as battery electrolytes. Owing to the complex structures of molecules and the sequence-independent nature of mixtures, conventional ML methods have difficulties in modeling such systems. Here, we present MolSets, a specialized ML model for molecular mixtures, to overcome the difficulties. Representing individual molecules as graphs and their mixture as a set, MolSets leverages a graph neural network and the deep sets architecture to extract information at the molecular level and aggregate it at the mixture level, thus addressing local complexity while retaining global flexibility. We demonstrate the efficacy of MolSets in predicting the conductivity of lithium battery electrolytes and highlight its benefits in the virtual screening of the combinatorial chemical space. Published by the American Physical Society 2024

Zhang, Hengrui (ORCID:0000000231831654)↗