Search NASA⌕ Search

SEARCH · Search NASA

Results for “Data-driven informatics”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

ET-AL: Entropy-targeted active learning for bias mitigation in materials data

Growing materials data and data-driven informatics drastically promote the discovery and design of materials. While there are significant advancements in data-driven models, the quality of data resources is less studied despite its huge impact on model performance. In this work, we focus on data bias arising from uneven coverage of materials families in existing knowledge. Observing different diversities among crystal systems in common materials databases, we propose an information entropy-based metric for measuring this bias. To mitigate the bias, we develop an entropy-targeted active learning (ET-AL) framework, which guides the acquisition of new data to improve the diversity of underrepresented crystal systems. We demonstrate the capability of ET-AL for bias mitigation and the resulting improvement in downstream machine learning models. This approach is broadly applicable to data-driven materials discovery, including autonomous data acquisition and dataset trimming to reduce bias, as well as data-driven informatics in other scientific domains.

36 MATERIALS SCIENCE↗

Data-driven discovery of a formation prediction rule on high-entropy ceramics

The interest in high entropy ceramics (HECs) has increased steadily due to their superior properties. However, the prediction of their formation still poses challenges for the discovery of new systems. Here, we discover a rational rule for designing single-phase high entropy metal diborides (HEBs) using data-driven approach. The machine learning (ML) model is trained on data collected via high-throughput experiments (HTEs). K nearest neighbor (KNN) model shows an experimental validation accuracy of 93.75%. By implementing interpretable ML method, we demonstrate that a mismatch of the bonds between boron and transition metals (δ B-TM ) dominates the formation of HEBs. We propose an empirical rule that HEBs favor forming a single phase when δ B-TM < 3.66; otherwise, multiphase. The rule has a high accuracy of 93.33% for new HEBs predictions. In addition, we contribute 165 high quality HEBs data in total, which can promote the development of materials informatics in HEBs. Furthermore, this data-driven strategy can be expanded to accelerate the search for new HECs, paving a pathway to design novel HECs with superior properties rapidly.

36 MATERIALS SCIENCE↗

A Digital Twin Framework Utilizing Machine Learning for Robust Predictive Maintenance: Enhancing Tire Health Monitoring

We introduce a novel digital twin (DT) framework for the predictive maintenance of long-term physical systems. Using monitoring tire health as an application, we show how the DT framework can be used to enhance automotive safety and efficiency, and how the technical challenges can be overcome using a three-step approach. First, to manage the data complexity over a long operation span, we employ data reduction techniques to concisely represent physical tires using historical performance and usage data. Relying on these data, for fast real-time prediction, we train a transformer-based model offline on our concise dataset to predict future tire health over time, represented as remaining casing potential (RCP). Based on our architecture, our model quantifies both epistemic and aleatoric uncertainties, providing reliable confidence intervals around predicted RCP. Second, to incorporate real-time data, we update the predictive model in the DT framework, ensuring its accuracy throughout its lifespan with the aid of hybrid modeling and the use of the discrepancy function. Third, to assist decision-making in predictive maintenance, we implement a tire state decision algorithm, which strategically determines the optimal timing for tire replacement based on RCP forecasted by our transformer model. This approach ensures that our DT accurately predicts system health, continually refines its digital representation, and supports predictive maintenance decisions. Furthermore, our framework effectively embodies a physical system, leveraging big data and machine learning (ML) for predictive maintenance, model updates, and decision-making.

advanced computing infrastructure↗

A data-driven framework for permeability prediction of natural porous rocks via microstructural characterization and pore-scale simulation

Understanding the microstructure–property relationships of porous media is of great practical significance, based on which macroscopic physical properties can be directly derived from measurable microstructural informatics. However, establishing reliable microstructure–property mappings in an explicit manner is difficult, due to the intricacy, stochasticity, and heterogeneity of porous microstructures. In this paper, a data-driven computational framework is presented to investigate the inherent microstructure–permeability linkage for natural porous rocks, where multiple techniques are integrated together, including microscopy imaging, stochastic reconstruction, microstructural characterization, pore-scale simulation, feature selection, and data-driven modeling. A large number of 3D digital rocks with a wide porosity range are acquired from microscopy imaging and stochastic reconstruction techniques. A broad variety of morphological descriptors are used to quantitatively characterize pore microstructures from different perspectives, and they compose the raw feature pool for feature selection. Here high-fidelity lattice Boltzmann simulations are conducted to resolve fluid flow passing through porous media, from which reliable permeability references are obtained. The optimal feature set that best represents permeability is identified through a performance-oriented feature selection process, upon which a cost-effective surrogate model is rapidly fitted to approximate the microstructure-permeability mapping via data-driven modeling. This surrogate model exhibits great advantages over empirical/analytical formulas in terms of prediction accuracy and generalization capacity, which can predict reliable permeability values spanning four orders of magnitude. Besides, feature selection also greatly enhances the interpretability of the data-driven prediction model, from which new insights into the mechanism of how microstructural characteristics determine intrinsic permeability are obtained.

58 GEOSCIENCES↗

A high-throughput and data-driven computational framework for novel quantum materials

Two-dimensional layered materials, such as transition metal dichalcogenides (TMDs), possess an intrinsic van der Waals gap at the layer interface, allowing for remarkable tunability of the optoelectronic features via external intercalation of foreign guests such as atoms, ions, or molecules. Herein, we introduce a high-throughput, data-driven computational framework for the design of novel quantum materials derived from intercalating planar conjugated organic molecules into bilayer transition metal dichalcogenides and dioxides. By combining first-principles methods, material informatics, and machine learning, we characterize the energetic and mechanical stability of this new class of materials and identify the fifty (50) most stable hybrid materials from a vast configurational space comprising ∼105 materials, employing intercalation energy as the screening criterion.

Kastuar, Srihari M. (ORCID:0000000279001561)↗

A materials-informatics based study of solid electrolytes and protective coatings for Li batteries

All-solid-state batteries with Li metal anode can address the safety issues surrounding traditional Li-ion batteries as well as the demand for higher energy densities. However, the development of solid electrolytes and protective coatings simultaneously possessing high ionic conductivity and wide electrochemical stability has proven to be a challenge. Here, we present a data-driven approach to explore the Li compound space for promising solid electrolytes and coatings. This is accomplished through the generation of a large database of battery-related materials properties of Li compounds by computing Li+ migration barriers using bond-valence-based pair potentials, and stability windows using density functional theory energies. Using this database, we implement machine learning models that can accurately predict migration barriers and electrochemical stability windows for any new Li compound. Through feature engineering, we ensure that our models are both accurate and interpretable. We perform feature importance analysis on our models to highlight materials properties that can be tuned for future design of coatings/electrolytes. Our database and informatics approach provide a valuable tool for the rapid discovery of new solid-state battery chemistries.

Solid state batteries↗

Deep learning of experimental electrochemistry for battery cathodes across diverse compositions

Artificial intelligence (AI) has emerged as a tool for discovering and optimizing novel battery materials. However, the adoption of AI in battery cathode representation and discovery is still limited due to the complexity of optimizing multiple performance properties and the scarcity of high-fidelity data. Here, we present a machine learning model (DRXNet) for battery informatics and demonstrate the application in the discovery and optimization of disordered rocksalt (DRX) cathode materials. We have compiled the electrochemistry data of DRX cathodes over the past 5 years, resulting in a dataset of more than 19,000 discharge voltage profiles on diverse chemistries spanning 14 different metal species. Learning from this extensive dataset, our DRXNet model can capture critical features in the cycling curves of DRX cathodes under various conditions. Our approach offers a data-driven solution to facilitate the rapid identification of novel cathode materials, accelerating the development of next-generation batteries for carbon neutralization.

25 ENERGY STORAGE↗

The search for high-entropy fuel-cell catalysts using disorder descriptors

The transition to a hydrogen economy depends on efficient, affordable catalysts for fuel cells. Platinum—the industry standard for fuel-cell electrodes—is costly and scarce, highlighting the need for practical alternatives. High-entropy alloys offer vast compositional diversity and tunable properties that can mitigate these issues, yet their chemical complexity and configurational disorder have hindered rational discovery. Here, we introduce a data-driven framework that couples machine learning with first-principles disorder descriptors—including the entropy forming ability, disordered enthalpy-entropy descriptor, and electronic-structure similarity metrics to platinum—to predict alloy synthesizability and catalytic performance. These descriptors are applied for the first time in the context of fuel-cell catalyst discovery. The workflow rapidly screens more than 20 000 compositions and identifies several platinum-free candidates that are economically viable, readily scalable, and exhibit promising predicted activity. These results demonstrate that disorder descriptors are reliably predicted by machine learning models and can be effectively integrated into materials-discovery pipelines, accelerating innovation across complex compositional spaces.

fuel-cell catalysts↗

A Case Study of Multimodal, Multi-institutional Data Management for the Combinatorial Materials Science Community

Although the convergence of high-performance computing, automation, and machine learning has significantly altered the materials design timeline, transformative advances in functional materials and acceleration of their design will require addressing the deficiencies that currently exist in materials informatics, particularly a lack of standardized experimental data management. The challenges associated with experimental data management are especially true for combinatorial materials science, where advancements in automation of experimental workflows have produced datasets that are often too large and too complex for human reasoning. The data management challenge is further compounded by the multimodal and multi-institutional nature of these datasets, as they tend to be distributed across multiple institutions and can vary substantially in format, size, and content. Furthermore, modern materials engineering requires the tuning of not only composition but also of phase and microstructure to elucidate processing–structure–property–performance relationships. To adequately map a materials design space from such datasets, an ideal materials data infrastructure would contain data and metadata describing (i) synthesis and processing conditions, (ii) characterization results, and (iii) property and performance measurements. In this work, we present a case study for the low-barrier development of such a dashboard that enables standardized organization, analysis, and visualization of a large data lake consisting of combinatorial datasets of synthesis and processing conditions, X-ray diffraction patterns, and materials property measurements generated at several different institutions. While this dashboard was developed specifically for data-driven thermoelectric materials discovery, we envision the adaptation of this prototype to other materials applications, and, more ambitiously, future integration into an all-encompassing materials data management infrastructure.

36 MATERIALS SCIENCE↗

Street-level temperature estimation using graph neural networks: Performance, feature embedding and interpretability

Estimating street-level air temperature is a challenging task due to the highly heterogeneous urban surfaces, canyon-like street morphology, and the diverse physical processes in the built environment. Though pioneering studies have embarked on investigations via data-driven approaches, many questions remain to be answered. Here, in this study, we leveraged an innovative framework and redefined the street-level temperature estimation problem using Graph Neural Networks (GNN) with spatial embedding techniques. The results showed that GNN models are more capable and consistent of estimating street-level temperature among tested locations, benefiting from its unique strength in handling extensive data over unstructured graph topology. In addition, we conducted in-depth analysis of feature importance to enhance the model interpretability. Among the urban features analyzed in this study, the time-variant canopy density and meter-level land use data emerge as crucial factors. Our findings highlight GNN 's high potential in capturing the complex dynamics between urban elements and their impacts on microclimate, thus offering valuable insights for comprehensive urban data collection and urban climate modeling in general. Collectively, this study also contributes to urban planning and policy by providing avenues to enhance city resilience against climate change, thereby advancing the agenda for environmental stewardship and urban sustainability.

54 ENVIRONMENTAL SCIENCES↗