Search NASA⌕ Search

SEARCH · Search NASA

Results for “materials informatics”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

Trust Not Verify? The Critical Need for Data Curation Standards in Materials Informatics

The importance of data curation has been recognized in multiple areas of research; however, the discussion of this important issue is only beginning to emerge in materials science. In this Perspective, we highlight the benefits of using the standardized data curation protocols in materials science and discuss current gaps in accurate and reproducible data reporting using case studies drawn from high-impact materials science papers and well-known databases such as the Crystallography Open Database (COD) and the Cambridge Structural Database (CSD). We argue that both experimental and computational materials scientists need to embrace a culture of rigorous data curation as part of modern research data management. We propose a sample data curation pipeline for materials chemistry and illustrate its use by creating two new materials chemistry databases. Here, we hope that this perspective will serve to catalyze further discussion and promote the continuous development of rigorous data curation practices within the materials science research community. We posit that adherence to best practices of data curation will promote and enhance the reliability, reproducibility, and integrity of materials research and enable the development of reliable AI and machine learning models that critically depend on the use of quality data.

Chemical structure↗

Reproducibility in materials informatics: lessons from ‘A general-purpose machine learning framework for predicting properties of inorganic materials’

The integration of machine learning techniques in materials discovery has become prominent in materials science research and has been accompanied by an increasing trend towards open data and open-source tools to propel the field. Despite the increasing usefulness and capabilities of these tools, developers neglecting to follow reproducible practices presents a significant barrier for other researchers looking to use or build upon their work. In this study, we investigate the challenges encountered while attempting to reproduce a section of the results presented in “A general-purpose machine learning framework for predicting properties of inorganic materials.” Our analysis identifies four major categories of challenges: (1) reporting software dependencies, (2) recording and sharing version logs, (3) sequential code organization, and (4) clarifying code references within the manuscript. The result is a proposed set of tangible action items for those aiming to make material informatics tools accessible to, and useful for the community.

36 MATERIALS SCIENCE↗

A Case Study of Multimodal, Multi-institutional Data Management for the Combinatorial Materials Science Community

Although the convergence of high-performance computing, automation, and machine learning has significantly altered the materials design timeline, transformative advances in functional materials and acceleration of their design will require addressing the deficiencies that currently exist in materials informatics, particularly a lack of standardized experimental data management. The challenges associated with experimental data management are especially true for combinatorial materials science, where advancements in automation of experimental workflows have produced datasets that are often too large and too complex for human reasoning. The data management challenge is further compounded by the multimodal and multi-institutional nature of these datasets, as they tend to be distributed across multiple institutions and can vary substantially in format, size, and content. Furthermore, modern materials engineering requires the tuning of not only composition but also of phase and microstructure to elucidate processing–structure–property–performance relationships. To adequately map a materials design space from such datasets, an ideal materials data infrastructure would contain data and metadata describing (i) synthesis and processing conditions, (ii) characterization results, and (iii) property and performance measurements. In this work, we present a case study for the low-barrier development of such a dashboard that enables standardized organization, analysis, and visualization of a large data lake consisting of combinatorial datasets of synthesis and processing conditions, X-ray diffraction patterns, and materials property measurements generated at several different institutions. While this dashboard was developed specifically for data-driven thermoelectric materials discovery, we envision the adaptation of this prototype to other materials applications, and, more ambitiously, future integration into an all-encompassing materials data management infrastructure.

36 MATERIALS SCIENCE↗

A high-throughput and data-driven computational framework for novel quantum materials

Two-dimensional layered materials, such as transition metal dichalcogenides (TMDs), possess an intrinsic van der Waals gap at the layer interface, allowing for remarkable tunability of the optoelectronic features via external intercalation of foreign guests such as atoms, ions, or molecules. Herein, we introduce a high-throughput, data-driven computational framework for the design of novel quantum materials derived from intercalating planar conjugated organic molecules into bilayer transition metal dichalcogenides and dioxides. By combining first-principles methods, material informatics, and machine learning, we characterize the energetic and mechanical stability of this new class of materials and identify the fifty (50) most stable hybrid materials from a vast configurational space comprising ∼105 materials, employing intercalation energy as the screening criterion.

Kastuar, Srihari M. (ORCID:0000000279001561)↗

A machine learning approach to quantify degradation of nuclear fuels and the effects of fission products

Nuclear fuel performance is critically dependent on understanding the evolution of fuel properties under operational conditions, a complex challenge driven by chemical changes and substantial radiation damage during fission. Traditionally, property evolution has been determined via empirical data collected following irradiation. However, these empirical correlations are limited in their applicability beyond the specific conditions in which they were obtained. This study explores a novel approach to address this challenge by applying materials informatics to develop a machine learning random forest (ML-RF) model that captures the effects of fission products on fuel compounds. The model predicts formation enthalpy (ΔH f ) by leveraging extensive quantum materials property data and correlating it with material descriptors such as composition, atomic and site features, and crystal lattice properties. This ML-RF model enables rapid interpolation across the compositional and structural spaces covered by the training data, thus supporting high-throughput screening and energetic ranking of candidate phases. The model demonstrates the ability to predict ΔH f with a mean absolute error (MAE) of approximately 0.1 to 0.2 eV/atom across a wide range of compounds, including key nuclear fuel systems (U-O, U-N, U-C, U-Si, and U-Mo). For example, it was used to assess shifts in stoichiometry for UO 2 (O/M) and UN (N/M) fuels, revealing their distinct tendencies in chemical potential variation and enabling preliminary convex hull analyses. Furthermore, the model provides insights into how individual fission products affect fuel properties. Results indicate that larger fission products (e.g., Nd, Pu, Ce) have a more pronounced impact on UO 2 , while lighter ones (e.g., Zr) strongly influence UN. Here, the model developed in this work can be used to support the Accelerated Fuel Qualification approach by facilitating preliminary evaluations prior to extensive materials modeling and experimentation. To this end, the trained model has been made available to the fuel community to support ongoing fuel development efforts.

Accelerated fuel qualification↗

Putting error bars on density functional theory dataset

This dataset contains submission files and raw output files from high-throughput DFT simulations to analyze the systemic errors in lattice constant, bulk moduli and formation energy predictions for a range of binary and ternary oxides using four exchange correlation functionals (LDA, PBE, PBEsol and vdW-DF-C09). This data was then used as the basis for employing materials informatics methods to predict the expected errors in the lattice constants of the studied compounds. Predicted errors were also used to better the DFT-predicted lattice parameters. Our results emphasize the link between the computed errors and the electron density and hybridization errors of a functional. In essence, these results provide “error bars” for choosing a functional for the creation of high-accuracy, high-throughput datasets as well as avenues for the development of XC functionals with enhanced performance, thereby enabling the accelerated discovery and design of new materials.

36 MATERIALS SCIENCE↗

Machine-learning-aided density functional theory calculations of stacking fault energies in steel

A combined large-scale first principles approach with machine learning and materials informatics is proposed to quickly sweep the chemistry-composition space of advanced high strength steels (AHSS). AHSS are composed of iron and key alloying elements such as aluminum and manganese. A systematic exploration of the distribution of aluminum and manganese atoms in iron is used to investigate low stacking fault energies configurations using first principles calculations. To overcome the computational cost of exploring the composition space, this process is sped up using an automated machine learning tool: DeepHyper. Here our results predict that it is energetically favorable for Al to stay away from a stacking fault, but Mn atoms do not affect the stacking fault energy and can stay in the vicinity of the fault. The distribution of Al and Mn atoms in systems containing stacking faults and the effects of their interactions on the equilibrium distribution are systematically analyzed.

36 MATERIALS SCIENCE↗

LLaMP v0.1.0

Reducing hallucination of Large Language Models (LLMs) is imperative for use in the sciences, where reliability and reproducibility are crucial. However, LLMs inherently lack long-term memory, making it a nontrivial, ad hoc, and inevitably biased task to fine-tune them on domain-specific literature and data. LLaMP is a multimodal retrieval-augmented generation (RAG) framework of hierarchical reasoning and acting (ReAct) agents that can dynamically and recursively interact with Materials Project to ground large language models on high-fidelity materials informatics.

Riebesell, Janosh [Lawrence Berkeley National Labo↗

Integrating adaptive learning with post hoc model explanation and symbolic regression to build interpretable surrogate models

Abstract We develop a materials informatics workflow to build an interpretable surrogate model for micromagnetic simulations. Our goal is to predict the energy barrier of a moving isolated skyrmion in rare-earth-free $$\hbox {Mn}_4$$ Mn 4 N. Our approach integrates adaptive learning with post hoc model explanation and symbolic regression methods. We discuss an unexplored acquisition function (information condensing active learning) within the adaptive learning loop and compare it with the known standard deviation function for efficient navigation of the search space. Model-agnostic post hoc explanation techniques then uncover trends learned by the trained model, which we then leverage to constrain the expressions used for symbolic regression. Graphical abstract

Biswas, Ankita↗

Computational modeling of grain boundary segregation: A review

Nearly all metals, alloys, ceramics, and their associated composites are polycrystalline in nature, with grain boundaries that separate well-defined crystalline regions that influence materials properties. In all but the most pure elemental systems, intentional solutes or impurities are present and can segregate to, or less commonly away from, the grain boundaries, in turn influencing boundary behavior, their stability, and associated materials properties. In some cases, grain-boundary segregation can also trigger “phase-like” structural transitions that dramatically alter the essential nature of the boundary. With the development of advanced electron microscopy techniques, researchers can directly observe grain-boundary structures and segregation with atomic precision. Despite such spatial resolution, the underlying mechanisms governing grain-boundary segregation remain difficult to characterize. As a result, computational modeling techniques such as density functional theory, molecular dynamics, mesoscale phase-field, continuum defect theory, and others are important complementary tools to experimental observations for studying grain-boundary segregation behavior. In conclusion, these computational methods offer the ability to explore the underlying formation mechanisms of grain-boundary segregation, elucidate complex segregation behavior, and provide insights into solutions to effectively controlling microstructure.

36 MATERIALS SCIENCE↗

The search for high-entropy fuel-cell catalysts using disorder descriptors

The transition to a hydrogen economy depends on efficient, affordable catalysts for fuel cells. Platinum—the industry standard for fuel-cell electrodes—is costly and scarce, highlighting the need for practical alternatives. High-entropy alloys offer vast compositional diversity and tunable properties that can mitigate these issues, yet their chemical complexity and configurational disorder have hindered rational discovery. Here, we introduce a data-driven framework that couples machine learning with first-principles disorder descriptors—including the entropy forming ability, disordered enthalpy-entropy descriptor, and electronic-structure similarity metrics to platinum—to predict alloy synthesizability and catalytic performance. These descriptors are applied for the first time in the context of fuel-cell catalyst discovery. The workflow rapidly screens more than 20 000 compositions and identifies several platinum-free candidates that are economically viable, readily scalable, and exhibit promising predicted activity. These results demonstrate that disorder descriptors are reliably predicted by machine learning models and can be effectively integrated into materials-discovery pipelines, accelerating innovation across complex compositional spaces.

fuel-cell catalysts↗

Simultaneously improving accuracy and computational cost under parametric constraints in materials property prediction tasks

Abstract Modern data mining techniques using machine learning (ML) and deep learning (DL) algorithms have been shown to excel in the regression-based task of materials property prediction using various materials representations. In an attempt to improve the predictive performance of the deep neural network model, researchers have tried to add more layers as well as develop new architectural components to create sophisticated and deep neural network models that can aid in the training process and improve the predictive ability of the final model. However, usually, these modifications require a lot of computational resources, thereby further increasing the already large model training time, which is often not feasible, thereby limiting usage for most researchers. In this paper, we study and propose a deep neural network framework for regression-based problems comprising of fully connected layers that can work with any numerical vector-based materials representations as model input. We present a novel deep regression neural network, iBRNet, with branched skip connections and multiple schedulers, which can reduce the number of parameters used to construct the model, improve the accuracy, and decrease the training time of the predictive model. We perform the model training using composition-based numerical vectors representing the elemental fractions of the respective materials and compare their performance against other traditional ML and several known DL architectures. Using multiple datasets with varying data sizes for training and testing, We show that the proposed iBRNet models outperform the state-of-the-art ML and DL models for all data sizes. We also show that the branched structure and usage of multiple schedulers lead to fewer parameters and faster model training time with better convergence than other neural networks. Scientific contribution: The combination of multiple callback functions in deep neural networks minimizes training time and maximizes accuracy in a controlled computational environment with parametric constraints for the task of materials property prediction.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗