Search NASA⌕ Search

SEARCH · Search NASA

Results for “Machine Learning Model”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 379 records · Page 21

Discovery of Ternary Antimonides A–Al–Sb (A = Rb or Cs) with Desired Structural Motifs Guided by Machine Learning

Specific structural motifs in inorganic solids are often related to their targeted physical properties. For many classes of solids, such as Zintl phases and polar intermetallics, the crystal structures are diverse and not easy to predict. Various antimonides that are potential thermoelectric materials were proposed to be synthesizable on the basis of their estimated formation energies. Their structures were broadly classified as clathrate, channel, layered, or network through a machine learning model trained on existing ternary phases and features based on elemental properties using the sure independence screening and sparsifying operator algorithm. Through experimental validation, three new ternary antimonides were synthesized and confirmed to form layered structures: tetragonal RbAlSb 2 and CsAlSb 2 , which are isopointal but not isotypic to LiBSi 2 ; and monoclinic Rb 2 Al 2 Sb 3 , which adopts the Na 2 Al 2 Sb 3 -type structure. Finally, reinvestigation of the related compound Cs 2 In 2 Sb 3 revealed a low thermal conductivity and p-type semiconducting behavior.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Trust Not Verify? The Critical Need for Data Curation Standards in Materials Informatics

The importance of data curation has been recognized in multiple areas of research; however, the discussion of this important issue is only beginning to emerge in materials science. In this Perspective, we highlight the benefits of using the standardized data curation protocols in materials science and discuss current gaps in accurate and reproducible data reporting using case studies drawn from high-impact materials science papers and well-known databases such as the Crystallography Open Database (COD) and the Cambridge Structural Database (CSD). We argue that both experimental and computational materials scientists need to embrace a culture of rigorous data curation as part of modern research data management. We propose a sample data curation pipeline for materials chemistry and illustrate its use by creating two new materials chemistry databases. Here, we hope that this perspective will serve to catalyze further discussion and promote the continuous development of rigorous data curation practices within the materials science research community. We posit that adherence to best practices of data curation will promote and enhance the reliability, reproducibility, and integrity of materials research and enable the development of reliable AI and machine learning models that critically depend on the use of quality data.

Chemical structure↗

Data Generation for Machine Learning Interatomic Potentials and Beyond

The field of data-driven chemistry is undergoing an evolution, driven by innovations in machine learning models for predicting molecular properties and behavior. Recent strides in ML-based interatomic potentials have paved the way for accurate modeling of diverse chemical and structural properties at the atomic level. The key determinant defining MLIP reliability remains the quality of the training data. A paramount challenge lies in constructing training sets that capture specific domains in the vast chemical and structural space. This Review navigates the intricate landscape of essential components and integrity of training data that ensure the extensibility and transferability of the resulting models. We delve into the details of active learning, discussing its various facets and implementations. We outline different types of uncertainty quantification applied to atomistic data acquisition and the correlations between estimated uncertainty and true error. The role of atomistic data samplers in generating diverse and informative structures is highlighted. Furthermore, we discuss data acquisition via modified and surrogate potential energy surfaces as an innovative approach to diversify training data. The Review also provides a list of publicly available data sets that cover essential domains of chemical space.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Computational Discovery of Intermolecular Singlet Fission Materials Using Many-Body Perturbation Theory

Intermolecular singlet fission (SF) is the conversion of a photogenerated singlet exciton into two triplet excitons residing on different molecules. SF has the potential to enhance the conversion efficiency of solar cells by harvesting two charge carriers from one high-energy photon, whose surplus energy would otherwise be lost to heat. The development of commercial SF-augmented modules is hindered by the limited selection of molecular crystals that exhibit intermolecular SF in the solid state. Computational exploration may accelerate the discovery of new SF materials. The GW approximation and Bethe–Salpeter equation (GW+BSE) within the framework of many-body perturbation theory is the current state-of-the-art method for calculating the excited-state properties of molecular crystals with periodic boundary conditions. In this Review, we discuss the usage of GW+BSE to assess candidate SF materials as well as its combination with low-cost physical or machine learned models in materials discovery workflows. We demonstrate three successful strategies for the discovery of new SF materials: (i) functionalization of known materials to tune their properties, (ii) finding potential polymorphs with improved crystal packing, and (iii) exploring new classes of materials. In addition, three new candidate SF materials are proposed here, which have not been published previously.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Using Active Learning to Rapidly Develop Machine Learned Diffusion Coefficients of CO 2 Conversion Reagents in Metal–Organic Frameworks

Here, we used a combined molecular dynamics/active learning (AL) approach to create machine learning models that can predict the diffusion coefficient of epichlorohydrin and chloropropene carbonate, the reactant and product of a common CO 2 cycloaddition reaction, in metal–organic frameworks (MOFs). Nanoporous MOFs are effective catalysts for the cycloaddition of CO 2 to epoxides. The diffusion rates within nanoporous catalysts can control the rate of reaction as the reactants and products must diffuse to the active sites within the MOF and then out of the nanoporous material for reusability. However, the diffusion process is routinely ignored when searching for new materials in catalytic applications. Here we verified improvement during the AL process by consistently tracking metrics on the same groups of MOFs to ensure consistency. Metal identity was found to have little impact on diffusion rates, while structural features like pore limiting diameter act as a threshold where a minimum value is needed for high diffusion rates. We identified the MOFs with the highest epichlorohydrin and chloropropene carbonate diffusion coefficients which can be used for further studies of reaction energetics.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

One Pot Synthesis of Cyan Emitting CdZnSSe Quantum Dots for Human Centric Lighting

A one pot synthesis of blue and green emissive CdZnSSe quantum dots (QDs) from thio- and selenoureas and Cd and Zn carboxylates is optimized using high throughput robotic optimization. A large set of spectral data (N = 192) is used to train machine learning models that accurately predict the photoluminescence emission wavelength (λmax) and full-width half-maximum, and the relative photoluminescence quantum yield (PLQY) from the S:Se and Zn:Cd stoichiometries and reaction time. ZnS shells are deposited on the crude QD heterostructures using 4-tert-butylbenzyl mercaptan, a more reactive source of sulfide that enables shell growth below the temperature where ion diffusion in the QD can broaden its optical spectrum (≤275 °C). These optimized procedures provide gram quantities of blue-green emitting QDs (PLQY = 85–99%) in a single reaction vessel. A solid state lighting device (4260 K) that incorporates cyan emissive QDs achieved a higher luminous efficacy of 179 lm/W and melanopic daylight efficiency ratio (0.71) than existing commercial human centric lighting devices.

Jordan, Abraham J↗

Computational Investigation of a CO 2 Conversion Strategy via Diels–Alder Reaction in a Carbon Capture Solvent

Molecular-level insights into reactive separations are crucial for the design of new conversion pathways of carbon dioxide (CO 2 ). This work explores a postulated pathway that directs CO 2 to undergo inverse-electron-demand Diels–Alder reactions to produce heterocycles using the CO 2 chemically fixed on water-lean solvent molecules. Density functional theory calculations are applied to evaluate the lowest unoccupied molecular orbital (LUMO) energies of three types of reactants (1,3-butadiene, 1,3-cyclohexadiene, and 1,2,4,5-tetrazine) with various functional substituents. These calculations also provide a data set (5.8k data) for developing a machine learning model to efficiently predict LUMO energies. A computational screening of LUMO energies for an additional 47k diene and tetrazine candidates is performed, and a list of candidates with lowered LUMO energies by electron-withdrawing substituents is provided. These candidates are further examined by their reaction energy barriers computed from the interatomic potential or density functional theory. Two major energy barriers are identified, one for the proton transfer within the water-lean solvent and the other for the CO 2 transfer from the solvent molecule to the reactant candidate (diene or tetrazine). The functional substituents have a more significant impact on the second barrier but a very slight one on the first barrier. This exploratory work demonstrates a new possibility for guiding experimental efforts toward the chemical conversion of fixated CO 2 to value-added compounds.

Chemical reactions↗

Distributed Acoustic Sensing to Estimate the Permeability

Optical fiber in a borehole can be interrogated with distributed acoustic sensors (DAS) to capture fracture displacements with the potential to map surrounding fracture networks. We designed a laboratory experiment to test the capability of DAS to determine borehole flow characteristics, and we show that for the first time DAS can be used to remotely estimate permeability. Optical fiber was wrapped around a bead filled pipe and the pressure drop and flow velocity were measured to directly calculate permeability. A machine learning model using statistical features from continuous DAS estimated the bulk permeability. Fluid interactions with the permeable material demonstrate insufficient resolution using DAS amplitude-based measurements for estimating pressure drop to infer permeability. Variations in the spectral domain relate DAS measurements to the pressure drop and provide consistent permeability estimates. Resolution with DAS is sufficient to estimate permeability and provides a reliable method to monitor at depth in borehole conditions.

58 GEOSCIENCES↗

Anomalous lattice thermal conductivity increase with temperature in cubic GeTe correlated with strengthening of second-nearest neighbor bonds

Understanding thermal transport mechanisms in phase change materials is critical to elucidating the microscopic picture of phase transitions and advancing thermal energy conversion and storage. Experiments consistently show that cubic phase germanium telluride (GeTe) has an unexpected increase in lattice thermal conductivity with rising temperature. Despite its ubiquity, resolving its origin has remained elusive. In this work, we carry out temperature-dependent lattice thermal conductivity calculations for cubic GeTe through efficient, high-order machine-learned models and additional corrections for coherence effects. We corroborate the calculated phonon properties with our inelastic X-ray scattering measurements. Our calculated lattice thermal conductivity values agree well with experiments and show a similar increasing trend. Through additional bonding strength calculations, we propose that a major contributor to the increasing lattice thermal conductivity is the strengthening of second-nearest neighbor interactions. The findings herein serve to deepen our understanding of thermal transport in phase-change materials.

36 MATERIALS SCIENCE↗

Active learning of ternary alloy structures and energies

Abstract Machine learning models with uncertainty quantification have recently emerged as attractive tools to accelerate the navigation of catalyst design spaces in a data-efficient manner. Here, we combine active learning with a dropout graph convolutional network (dGCN) as a surrogate model to explore the complex materials space of high-entropy alloys (HEAs). We train the dGCN on the formation energies of disordered binary alloy structures in the Pd-Pt-Sn ternary alloy system and improve predictions on ternary structures by performing reduced optimization of the formation free energy, the target property that determines HEA stability, over ensembles of ternary structures constructed based on two coordinate systems: (a) a physics-informed ternary composition space, and (b) data-driven coordinates discovered by the Diffusion Maps manifold learning scheme. Both reduced optimization techniques improve predictions of the formation free energy in the ternary alloy space with a significantly reduced number of DFT calculations compared to a high-fidelity model. The physics-based scheme converges to the target property in a manner akin to a depth-first strategy, whereas the data-driven scheme appears more akin to a breadth-first approach. Both sampling schemes, coupled with our acquisition function, successfully exploit a database of DFT-calculated binary alloy structures and energies, augmented with a relatively small number of ternary alloy calculations, to identify stable ternary HEA compositions and structures. This generalized framework can be extended to incorporate more complex bulk and surface structural motifs, and the results demonstrate that significant dimensionality reduction is possible in thermodynamic sampling problems when suitable active learning schemes are employed.

Chemistry↗

Systematic softening in universal machine learning interatomic potentials

Machine learning interatomic potentials (MLIPs) have introduced a new paradigm for atomic simulations. Recent advancements have led to universal MLIPs (uMLIPs) that are pre-trained on diverse datasets, providing opportunities for universal force fields and foundational machine learning models. However, their performance in extrapolating to out-of-distribution complex atomic environments remains unclear. In this study, we highlight a consistent potential energy surface (PES) softening effect in three uMLIPs: M3GNet, CHGNet, and MACE-MP-0, which is characterized by energy and force underprediction in atomic-modeling benchmarks including surfaces, defects, solid-solution energetics, ion migration barriers, phonon vibration modes, and general high-energy states. The PES softening behavior originates primarily from the systematically underpredicted PES curvature, which derives from the biased sampling of near-equilibrium atomic arrangements in uMLIP pre-training datasets. Our findings suggest that a considerable fraction of uMLIP errors are highly systematic, and can therefore be efficiently corrected. We argue for the importance of a comprehensive materials dataset with improved PES sampling for next-generation foundational MLIPs.

36 MATERIALS SCIENCE↗

Optimal invariant sets for atomistic machine learning

The representation of atomic configurations for machine learning models has led to numerous sets of descriptors. However, many descriptor sets are incomplete and/or functionally dependent. Incomplete sets cannot faithfully represent atomic environments. Yet complete constructions often suffer from a high degree of functional dependence, where some descriptors are functions of others. These redundant descriptors do not improve discrimination between atomic environments. We employ pattern recognition techniques to remove dependent descriptors to produce the smallest possible set that satisfies completeness. We apply this in two ways: First, we refine an existing description, the atomic cluster expansion. Second, we augment an incomplete construction, yielding a new message-passing neural network architecture that can recognize up to 5-body patterns. This architecture shows strong accuracy on state-of-the-art benchmarks while retaining low computational cost. Our results demonstrate the utility of this strategy to optimize descriptor sets across a range of descriptors and application datasets.

97 MATHEMATICS AND COMPUTING↗

Data mining and computational screening of Rashba-Dresselhaus splitting and optoelectronic properties in two-dimensional perovskite materials

Recent developments highlighting the promise of two-dimensional perovskites have vastly increased the compositional search space in the perovskite family. This presents a great opportunity for the realization of highly performant devices and practical challenges associated with the identification of candidate materials. High-fidelity computational screening offers great value in this regard. In this study, we carry out a multiscale computational workflow, generating a dataset of two-dimensional perovskites in the Dion-Jacobson and Ruddlesden-Popper phases. Our dataset comprises ten B-site cations, four halogens, and over 20 organic cations across over 2000 materials. We compute electronic properties, thermoelectric performance, and numerous geometric characteristics. Furthermore, we introduce a framework for the high-throughput computation of Rashba-Dresselhaus splitting. Finally, we use this dataset to train machine learning models for the accurate prediction of band gaps, candidate Rashba-Dresselhaus materials, and partial charges. The work presented herein can aid future investigations of two-dimensional perovskites with targeted applications in mind.

14 SOLAR ENERGY↗

PAH101: A GW+BSE Dataset of 101 Polycyclic Aromatic Hydrocarbon (PAH) Molecular Crystals

Abstract The excited-state properties of molecular crystals are important for applications in organic electronic devices. TheGWapproximation and Bethe-Salpeter equation (GW+BSE) is the state-of-the-art method for calculating the excited-state properties of crystalline solids with periodic boundary conditions. We present the PAH101 dataset ofGW+BSE calculations for 101 molecular crystals of polycyclic aromatic hydrocarbons (PAHs) with up to ~500 atoms in the unit cell. To the best of our knowledge, this is the firstGW+BSE dataset for molecular crystals. The data records include theGWquasiparticle band structure, the fundamental band gap, the static dielectric constant, the first singlet exciton energy (optical gap), the first triplet exciton energy, the dielectric function, and optical absorption spectra for light polarized along the three lattice vectors. The dataset can be used to (i) discover materials with desired electronic/optical properties, (ii) identify correlations between DFT andGW+BSE quantities, and (iii) train machine learned models to help in materials discovery efforts.

Science & Technology - Other Topics↗

Machine learning of 27Al NMR electric field gradient tensors for crystalline structures from DFT

NMR crystallography has emerged as a promising technique for the determination and refinement of atomic coordinates in crystal structures. The crystal structure of compounds containing quadrupolar nuclei, such as 27Al, can be improved by directly comparing solid-state NMR measurements to DFT computations of the electric field gradient (EFG) tensor. The non-negligible computational cost of these first-principles calculations limits the applicability of this method to all but the most well-defined structures. We developed a fast, low-cost machine learning model to predict EFG parameters based on local structural motifs and elemental parameters. We computed 8081 EFG tensors from 1681 27Al crystalline solids using DFT and benchmarked them against 105 experimentally measured 27Al sites. Surprisingly, simple local geometric features dominate the predictive performance of the resulting random-forest model, yielding an R2 value of 0.98 and an RMSE of 0.61 MHz for CQ, the quadrupolar coupling constant. This model accuracy should enable pre-refining future structural assignments before finally validating with first-principles calculations. Such a catalogue of 27Al NMR tensors can serve as a tool for researchers assigning complex NMR spectra influenced by the nuclear electric quadrupole interaction.

Sun, He↗

Peatland fires in Alaska will double by the end of the century

During recent summers, warm and dry conditions have increased the occurrence of wildfires and potentially peat-fires across Alaska. Limitations in resolving the fine-scale distribution of peatlands and climate observations have constrained our ability to accurately predict peat-fire dynamics. Using a new high-resolution peatland map of Alaska, we evaluated the climate and environmental controls of past and future peat-fire activity. Ensemble machine learning models identified reduced soil moisture, higher temperatures, and evapotranspiration as key predictors of annual total burned peatland area (tenfold CV R 2 = 0.62, RMSE = 221.1 km 2 ). By the end of the twenty-first century, models forced with climate datasets from representative concentration pathways (RCPs) 4.5, 6.0, and 8.5 emission scenarios project a statewide doubling of burned peatlands (increasing 61–121%), with regional increases ranging from 25–165% in polar, 61–95% in boreal, and 102–106% in maritime ecoregions. These projections indicate that wildfires will progressively encroach further into organic-rich moist and wet peaty soils, potentially amplifying soil carbon release across Alaska.

climate-change ecology↗

Image processing pipeline for AI-driven nanoparticle megalibrary characterization

Recent innovations have made it possible to produce megalibraries, millions of structurally and compositionally distinct nanoparticles on a chip. These megalibraries yield vast volumes of data that are impossible to analyze manually, necessitating the development of automated tools. In previous work, we created a binary classification machine learning model to select quality nanoparticle images for downstream analysis. In this work, we show that adding a custom image processing step before training can produce significantly higher-performing models in a fraction of the time and make them more robust to different image noise levels and microscope acquisition settings. The image processing pipeline proposed here effectively cleans raw nanoparticle images, enhances key features, and allows us to use much lower resolution images and simpler neural network model architectures. These features result in higher performance and significant cost savings. Experiments demonstrate superior performance relative to baseline, including an 18.2% improvement in recall and a 13.1% increase in accuracy. Given the high cost of downstream analysis, it is critical to minimize false positives, and our best-performing model reaches a precision of 95.9% and a weighted F-score of 95.1% on an unseen test set. Additionally, model training time is reduced from hours to less than a minute. We also show that, using this custom image processing pipeline, model performance is significantly improved at lower pixel resolutions compared to downsizing alone. We expect that adopting this pipeline for AI-driven automated nanoparticle characterization will allow researchers to rapidly and accurately analyze much greater volumes of data, thereby accelerating materials discovery.

77 NANOSCIENCE AND NANOTECHNOLOGY↗