Search NASA⌕ Search

SEARCH · Search NASA

Results for “Machine Learning for Data Analysis”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 199 records · Page 11

Machine-learning interatomic potentials for interfaces in all-solid-state batteries: Perspectives on training data, model selection, and validation

Interfaces play a pivotal role in dictating the performance and reliability of all-solid-state batteries (ASSBs), where complex electro-chemo-mechanical phenomena at grain boundaries (GBs) and interfaces can lead to degradation and failure. Traditional atomistic simulation methods, such as first-principles calculations and classical molecular dynamics, face limitations in modeling these interfaces due to either high computational cost or insufficient transferability to the diverse atomic environments evolving at interfaces. Machine-learning interatomic potentials (MLIPs) have emerged as a transformative approach, enabling large-scale, high-accuracy simulations of disordered and chemically complex systems by leveraging the predictability of machine learning models trained on first-principles data. Recent applications of MLIPs have demonstrated their ability to capture intricate behaviors at ASSB interfaces, including ion transport, interfacial evolution, and degradation mechanisms, with accuracy and efficiency unattainable by conventional methods. This prospective paper presents comprehensive analysis and practical guidance for MLIP development for GBs and interfaces in ASSBs, with a focus on three key pillars: data generation, model selection, and validation. Here, we review the current state of MLIP applications for GBs and interfaces in both general and ASSB-specific materials, highlighting best practices and challenges in constructing diverse and representative datasets, choosing appropriate machine learning architectures, and rigorously validating model performance. We also discuss emerging strategies and opportunities for improved reliability and efficiency of MLIPs to simulate realistic interfaces in ASSBs.

Energy - Storage↗

Control And Optimization Modular Modeling Application For Nuclear Deployment

The purpose of the COMMAND code is to provide a flexible, scalable tool for use in developing, integrating, and testing the technologies necessary for achieving autonomous operations of advanced nuclear reactors. The code enables users to efficiently implement custom simulations and experiments by combining key methods from different software modules. These modules are focused on: modeling and simulation tools, such as nuclear simulation tools used for high-fidelity modeling (e.g., Reactor Excursion and Leak Analysis Program [RELAP5-3D] and Monte Carlo N-Particle [MCNP]); machine learning and optimization tools (e.g., anomaly detection and data-driven modeling techniques); advanced control in its digital, high-performance, and supervisory control forms (e.g., proportional integral derivative (PID) control and model predictive control (MPC); and integration with hardware through industrial communication protocols. To ensure flexibility and scalability, COMMAND was designed to be both modular—the software “pieces” all inherit from generic building blocks and can be combined and connected to create complicated simulations—and high performing—designed for parallel processing, enabling simulations and experiments to take advantage of multi-core computers, servers, and nodes. The code is written in the Python programming language due to the language's popularity, active community, and open-source and cross-platform nature. Maintaining consistency with other simulation tools used within the nuclear energy community, users implement simulations and experiments through text input files, which define components, parameters, connections, etc., through lines of text. Given that COMMAND is written in Python, these input files are native Python scripts, and so use the standard Python structure and formatting. This also enables users to take advantage of Python's extensive package library to develop custom capabilities for their specific use cases.

Faber, Jacob [Idaho National Laboratory (INL), Ida↗

A 1 km soil moisture dataset over eastern CONUS generated by assimilating SMAP data into the Noah-MP land surface model

An improved fine-scale soil moisture (SM) dataset at 1 km grid spacing, covering much of the eastern continental US, was generated by assimilating 9 km Soil Moisture Active Passive (SMAP) SM data into the v4.0.1 Noah-MP land surface model. With 12 ensemble members, the assimilation was carried out using the ensemble Kalman filter algorithm within NASA's Land Information System. The SM analysis for 2016 was fully validated against in situ observations from four different networks and compared with four other existing datasets. Results indicate that this SM analysis surpasses other datasets in top-layer SM distribution, including a machine-learning-based product, despite all SM estimates being less heterogeneous than observed. The analysis of anomalous errors suggests that large similarity in intrinsic errors is likely due to overlapping data sources among the selected SM datasets. More detailed evaluations were performed over two geographic areas. The observations collected by the Atmospheric Radiation Measurement facility in Oklahoma suggest that soil temperature and surface heat fluxes are concurrently simulated with good accuracy. Investigation into the 2016 southeastern US drought response further indicates drier conditions and higher evapotranspiration estimates compared to GLEAMv4.1. Notably, large errors are associated with grids having clay soil textures, underscoring the need for refined model treatments for specific soil types to further improve SM estimates. The dataset is publicly available on Zenodo at https://doi.org/10.5281/zenodo.14370563 (Tai et al., 2024).

Tai, Sheng-Lun [Pacific Northwest National Laborat↗

Improving Trustworthiness of Data-Driven Power Grid Contingency Analysis With Bayesian Residual Graph Neural Networks

The evolving energy landscape requires novel tools to efficiently perform contingency analysis and reliability assessment of power grids, potentially in real-time. The high computational cost of traditional power flow solvers limits their applicability in practice. Machine learning (ML) surrogates such as deep neural networks (NNs) accelerate power flow solvers computations, enabling high-order contingency analysis and real-time decision-making by learning highly nonlinear functions and integrating grid topology via graph architectures. However, (graph) NNs lack predictive power away from training data and do not provide predictive confidence estimates. Here, we present a Bayesian residual graph NN that integrates knowledge from low-fidelity data via residual training and embeds granular quantification of uncertainties, improving trustworthiness critical for high-consequence decision-making. Applying Bayesian concepts to NNs is challenging due to the high-dimensionality of both the parameter space, complicating derivation of a meaningful prior, and the output space in large grid systems, requiring enhanced techniques to assess the predicted high-dimensional uncertainties. Our contributions include: (1) Deriving a prior for fully connected and graph NNs that leverages low-fidelity data to guide mean predictions and appropriately control prior predictive uncertainty. (2) Integrating this prior within an ensembling with anchoring scheme for efficient approximate posterior inference. (3) Deriving enhanced metrics to assess accuracy of both the mean and uncertainty predictions in high dimensions, appropriately accounting for correlations propagated through graph layers. The resulting Bayesian residual graph NN is tested on a contingency analysis task for 14-bus and 118-bus grids.

24 - POWER TRANSMISSION AND DISTRIBUTION↗

From Chaos to Clarity: Autonomous Materials Discovery for Extreme Environments

The pursuit of advanced functional materials for energy applications demands an understanding of their behavior under the most challenging conditions. Extreme environments, characterized by intense radiation, high temperatures, and corrosive chemistries, push materials to their limits, often revealing unexpected behaviors and degradation pathways. Traditional materials research approaches, relying on trial-and-error experimentation, are often slow and resource-intensive, ill-suited to the complexities of extreme environments. This talk will explore the transformative potential of autonomous materials science in revolutionizing our understanding of materials synthesis and degradation in extreme environments. By integrating advanced microscopy techniques, artificial intelligence, and robotic experimentation, we can accelerate the discovery and design of resilient materials for a sustainable future. The presentation will highlight recent breakthroughs in autonomous microscopy, computer vision, and machine learning, showcasing their ability to unravel complex material transformations at the atomic scale. The talk will also delve into the challenges and opportunities associated with deploying autonomous systems to probe extreme environments, emphasizing the importance of robust algorithms, real-time data analysis, and adaptive experimentation. Our ultimate goal is to empower scientists with unprecedented capabilities to explore, understand, and engineer materials that can withstand the harshest conditions, paving the way for innovations in energy, aerospace, and beyond.

artificial intelligence↗

Hybrid Quantum–Classical Graph Transformers for Efficient Sentiment Analysis

Quantum Machine Learning (QML) offers a promising paradigm that leverages quantum computing principles to develop efficient and expressive models for learning from complex and structured data. Recent advances in natural language processing (NLP) and artificial intelligence (AI) have demonstrated capabilities in understanding, generating, and reasoning over linguistic and multimodal information. In this work, we present the Quantum Graph Transformer (QGT), a hybrid quantum–classical architecture that extends graph transformer capabilities through quantum self-attention. The QGT models variable-length sentences as token graphs, where both the embedding encoding and the self-attention mechanisms are implemented using parameterized quantum circuits (PQCs), enabling efficient contextual learning with significantly fewer trainable parameters. We train QGT using both fully connected and 𝑘 -nearest-neighbor graph structures and evaluate it on five benchmark sentiment-classification datasets. Experimental results show that QGT consistently achieves higher or comparable accuracy to existing quantum NLP models and outperforms a Classical Graph Transformer (CGT) baseline with identical architecture, achieving 29.4 × fewer parameters while requiring 3–5 × fewer samples to reach comparable performance. These findings highlight the potential of graph-based quantum models as scalable and data-efficient architectures for natural language understanding.

71 CLASSICAL AND QUANTUM MECHANICS, GENERAL PHYSIC↗

Day-Ahead Probabilistic Forecasting of Net-Load and Demand Response Potentials with High Penetration of Behind-the-Meter Solar-plus-Storage

The goal of this project is to develop advanced methods for day-ahead net-load forecasting, by leveraging the state-of-the-art machine learning techniques. The developed models produce both point and probabilistic forecasts for a variety of use cases, and are versatile to work with different types of data sets. The innovation lies in the novel design of the architectures, leveraging the most recent advances in machine learning that have not been explored in power systems, accompanied by techniques in the broader artificial intelligence fields such as fuzzy systems. This project has achieved the following accomplishments: (1) preprocessing of over 10 data sets covering varying geographical regions, time horizons, and system levels, which form a robust foundation for training and evaluating forecasting models across a wide range of realistic grid scenarios; (2) development of an interactive web app that enables exploratory analysis of load and generation data, and supports better understanding of data trends, anomalies, and correlations, facilitating model development and stakeholder engagement; (3) implementation of over 10 benchmark models for point and probabilistic forecasting, which include a mix of conventional machine learning methods and state-of-the-art deep learning approaches, providing a comprehensive baseline for performance comparison and validation of the proposed models; (4) development of a fuzzy system based gradient boosting model, tailored for small (less than 3 years) data sets, which achieves a mean absolute percentage error (MAPE) of 4% for point forecasting and a 20% improvement in average pinball loss for probabilistic forecasting; (5) development of a Transformer (a state-of-the-art deep learning architecture) based neural network model, tailored for large (3 years or more) data sets, which achieves a MAPE of 2% for point forecasting and a 20% improvement in average pinball loss for probabilistic forecasting; (6) development of a methodology for quantifying DR potential, and extensions of the previous models for multi-target forecasting of net load and DR potential, which achieve a MAPE of 10% for DR potential.

24 POWER TRANSMISSION AND DISTRIBUTION↗

Quantitative Analysis and Prediction of Thermal Runaway Metrics of High-Nickel Oxide Cathodes by Machine Learning Models

The pursuit of higher energy density in lithium-ion batteries has made high-nickel (Ni) layered oxides leading cathode candidates for next-generation electric vehicles. However, their poor thermal stability, particularly at Ni contents ≥ 90%, increases the risk of cathode-initiated thermal runaway. Furthermore, we present a data-driven framework combining linear and nonlinear machine learning models to predict key thermal runaway descriptors from a high-throughput differential scanning calorimetry database. With cathode composition and state of charge (SOC) as input features, the ensemble model accurately predicts peak temperature, heat release, and peak heat flow. SHAP analysis identifies Ni content and SOC as the dominant factors controlling thermal runaway temperature, while SOC primarily governs heat release and peak heat flow. Al, Mg, and Mn improve thermal stability by strengthening metal–oxygen bonding and delaying structural transformation, whereas B mainly reduces heat release through surface passivation. Validation with a new cathode composition confirms accurate prediction of SOC-dependent thermal runaway behavior and critical SOC.

25 ENERGY STORAGE↗

Pavement condition and climatic data in southeast Texas: A dataset for evaluating flood impacts on pavement performance

Effective pavement maintenance is essential for economic stability, optimal network performance, and roadway safety. Achieving this requires thorough evaluation of pavement conditions, including structural integrity, surface roughness, and distress characteristics. Pavement performance indicators play a critical role in influencing vehicle safety and ride quality. Recent advances have emphasized the use of data-driven modeling to anticipate pavement behavior, with the goal of optimizing resource allocation and refining Maintenance and Rehabilitation (M&R) strategies through accurate condition assessment. A foundational requirement for these modeling efforts is the availability of standardized, high-quality datasets that can support robust and reproducible infrastructure analysis. This data article presents a comprehensive dataset assembled to facilitate pavement performance prediction, with a geographic focus on Southeast Texas, particularly the flood-vulnerable area of Beaumont. The dataset encompasses pavement and traffic attributes, meteorological records, flood simulation outputs, ground deformation measurements, and topographic indices, enabling detailed examination of both load-associated and non-load-associated degradation mechanisms. Data preprocessing was performed using ArcGIS Pro, Microsoft Excel, and Python to ensure consistency and usability in data-driven modeling applications, including machine learning workflows. Key contributions of this dataset include its utility in analyzing the climatic and environmental factors affecting pavement conditions, identifying critical predictive features, and enabling in-depth correlation analysis across diverse variables. By filling existing gaps in input variable selection resources, this dataset supports the development of predictive tools for estimating future maintenance demand and enhancing the resilience of pavement networks in flood-impacted areas. The resource highlights the importance of standardized datasets for advancing pavement management practices and provides a robust foundation for ongoing infrastructure performance modeling.

42 ENGINEERING↗

From Chaos to Clarity: Autonomous Materials Discovery for Extreme Environments [Slides]

The pursuit of advanced functional materials for energy applications demands an understanding of their behavior under the most challenging conditions. Extreme environments, characterized by intense radiation, high temperatures, and corrosive chemistries, push materials to their limits, often revealing unexpected behaviors and degradation pathways. Traditional materials research approaches, relying on trial-and-error experimentation, are often slow and resource-intensive, ill-suited to the complexities of extreme environments. This talk will explore the transformative potential of autonomous materials science in revolutionizing our understanding of materials synthesis and degradation in extreme environments. By integrating advanced microscopy techniques, artificial intelligence, and robotic experimentation, we can accelerate the discovery and design of resilient materials for a sustainable future. The presentation will highlight recent breakthroughs in autonomous microscopy, computer vision, and machine learning, showcasing their ability to unravel complex material transformations at the atomic scale. The talk will also delve into the challenges and opportunities associated with deploying autonomous systems to probe extreme environments, emphasizing the importance of robust algorithms, real-time data analysis, and adaptive experimentation. The ultimate goal is to empower scientists with unprecedented capabilities to explore, understand, and engineer materials that can withstand the harshest conditions, paving the way for innovations in energy, aerospace, and beyond.

14 SOLAR ENERGY↗

An integrated approach to examine fuel-cladding chemical interaction in HT9/U-10Zr metallic fast reactor fuels: Coupling machine learning with electron microscopy and local mechanical properties analysis

The metallic U-Zr nuclear fuel alloy has garnered renewed interest as a promising candidate for next-generation sodium-cooled fast reactors. Recent studies and technology assessments have identified several areas requiring improvements, enhanced knowledge, and reliable data to strengthen the U-Zr fuel design basis for qualification and commercial applications. One of the most challenging phenomena impacting this fuel system’s performance is fuel-cladding chemical interaction (FCCI). This work aimed to harvest FCCI data by examining selected HT9/U-10Zr (wt. %) fuel samples of prototypic full-length fuel pins through an integrated approach. This approach integrated scanning electron microscopy (SEM) microstructure characterization with localized mechanical properties examination to deepen understanding of FCCI phenomenon in HT9/U-10Zr fuel system. Particularly, this study focused on MFF fuel pins irradiated at Fast Flux Test Facility (FFTF), which aimed to qualify metallic fuel as a driver fuel for FFTF and to assess its viability for larger-scale fast reactors. Electron microscopy provided high confidence in detecting and distinguishing the different FCCI layers, while small-scale mechanical testing (SSMT) probed the mechanical properties of these layers. SEM examination of a MFF-2 pin 192167, with a time averaged inner cladding temperature (TICT) slightly over 500°C, revealed minimal cladding-side FCCI (cladding wastage). In contrast, significantly thicker cladding wastage comprising two distinct sublayers was observed in samples from the thermally hot MFF-3 pin 193045 and MFF-5 pin 195011 where the TICT ranged from 610-635°C. SSMT indicated complete embrittlement in the sublayer adjacent to the fuel and a tendency toward embrittlement in the other sublayer. Additionally, a new machine learning method was developed, validated, and used to quantify cladding wastage thickness. The machine learning method reliably predicted the wastage thickness across various fuel pins and sample cross-sections. Furthermore, the available cladding wastage data from HT9/U-10Zr fuel system demonstrated a strong temperature dependency. However, the dataset remains small, and ongoing research activities are essential to further understand the FCCI phenomenon and develop a reliable FCCI model for enhanced fuel performance simulation under various conditions.

36 - MATERIALS SCIENCE↗

From microbial diversity to functional potential using dimensionality reduction

The high dimensionality of microbial diversity data from ‘omics observations can be reduced using Machine Learning, with many recent studies showcasing ML utility for exploratory ecological feature finding and process prediction. Here, we compare the Self Organizing Map (SOM) dimensionality reduction method to the well-documented sample-based Principal Coordinate Analysis (PCoA) and taxa-based Weighted Gene Correlation Network Analysis (WGCNA) using near daily 16S rRNA gene amplicon sequencing data from the 2019 to 2020 MOSAiC International Arctic Drift Expedition. We then map k-means clustering outputs from each method to available metagenomes, extracting functionally distinct seasonal microbial ecotypes in the surface Arctic Ocean. Our results indicate the SOM method better represented expected seasonal transitions and identified a greater number of metabolically distinct functional groups than the more traditional PCoA ordination. Ultimately, we identified four community ecotypes with distinct taxonomic and functional cut-offs driven by seasonality, water mass, and substrate turnover, highlighting the importance of succession in functional diversity for the central Arctic Ocean. These results reinforce ML dimensionality reduction as a meaningful translator in the mining of historical amplicon datasets to address modern mechanistic questions and potentially provide ’omics informed ecotype diversity to leverage in mechanistic biogeochemical models.

Arctic Ocean↗

Tutorial: Machine-Learning-Based CREASE-2D Analysis of 2D SAXS Profiles to Characterize Anisotropic Nanostructures in Soft Materials

We present a tutorial to guide users on how to extend the Computational Reverse Engineering Analysis of Scattering Experiments-2D (CREASE-2D) framework to interpret their experimental two-dimensional small-angle scattering (SAS) data from soft materials (e.g., polymers, peptide amphiphiles, biomolecular fibrils). Unlike most traditional SAS analysis approaches, which typically rely on azimuthally averaged onedimensional (1D) profiles, CREASE-2D utilizes the complete 2D scattering profile to reveal information about anisotropy in the structure. In past applications, CREASE has provided insights into complex structural features, including the cross-sectional shapes of assembled nanostructures and dispersity in these features, which are difficult to discern with existing analytical models. While (1D- ) CREASE has been applied to SANS and SAXS data, this tutorial shares the steps for implementing CREASE-2D using an example of a dipeptide solution system, for which we have SAXS data. We present details for these steps involved in using CREASE-2D to interpret SAXS profiles: how to preprocess SAXS data, define relevant structural features, generate three-dimensional real-space structures for specific values of these features, train a machine learning (ML) surrogate model to predict scattering profiles for given structural features, and optimize these features using genetic algorithms (GA). Then, we use these steps to interpret complex 2DSAXS data collected from dipeptide solutions that, in microscopy images, exhibit nanoscale structures that could be elliptical tubes/ flat tapes/cylinders or a combination of these cross sections. Open-source codes, computational hardware, and software requirements, as well as the strengths and limitations of this protocol, are also presented. We expect researchers working with (soft) biomaterials, peptide amphiphiles, amphiphilic polymer solutions, polymer nanocomposites, and blends of particles/polymers will find this CREASE-2D method and this tutorial of use.

CREASE↗

Reduced-order model to approximate response matrices for filter stack spectrometers

We present a reduced-order model to calculate response matrices rapidly for filter stack spectrometers (FSSs). The reduced-order model allows response matrices to be built modularly from a set of pre-computed photon and electron transport and scattering calculations through various filter and detector materials. While these modular response matrices are not appropriate for high-fidelity analysis of experimental data, they encode sufficient physics to be used as a forward model in design optimization studies of FSSs, particularly for machine learning approaches that require sampling and testing a large number of FSS designs.

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND ↗

Search for squarks and gluinos in pp collisions at $\sqrt{s} = 13$ TeV and 13.6 TeV in events with $\tau$-leptons, jets and missing transverse momentum using the ATLAS detector

A search for R-parity-conserving supersymmetry in events with large missing transverse momentum, jets and at least one hadronically decaying $\tau$-lepton is presented. Both gluino and squark pair production are considered, with the cascade decay of each gluino or squark producing either a $\tau$-slepton or a $\tau$-sneutrino. Three channels are examined, requiring either exactly one hadronically decaying $\tau$-lepton and no other leptons, exactly one hadronically decaying $\tau$-lepton and at least one other lepton, or two or more hadronically decaying $\tau$-leptons. Analyses in the three channels are optimised independently and combined statistically. Two separate analysis strategies, either a cut-and-count or machine-learning approach, are used. The search uses 140 and 51.8 of pp collision data recorded by the ATLAS detector at the Large Hadron Collider during 2015–2018 at TeV and 2022–2023 at TeV, respectively. Gluino masses below 2.25 TeV and squark masses up to 1.7 TeV are excluded

Aad, G. [CNRS/IN2P3] (ORCID:0000000266654934)↗

Nonlinear Ensemble Filtering with Diffusion Models: Application to the Surface Quasigeostrophic Dynamics

The intersection between classical data assimilation methods and novel machine learning techniques has attracted significant interest in recent years. Here, we explore another promising solution in which diffusion models are used to formulate a robust nonlinear ensemble filter for sequential data assimilation. Unlike standard machine learning methods, the proposed ensemble score filter (EnSF) is completely training free and can efficiently generate a set of analysis ensemble members. Here, in this study, we apply the EnSF to a surface quasigeostrophic model and compare its performance against the popular local ensemble transform Kalman filter (LETKF), which makes Gaussian assumptions in the analysis step. Numerical tests demonstrate that EnSF maintains stable performance in the absence of localization and for a variety of experimental settings. We find that while LETKF maintains optimal performance in the case of linear observations of the entire state and a perfect model, EnSF shows improvements over LETKF when nonlinear observations are assimilated and the system is subject to unexpected model errors. A spectral decomposition of the analysis results in this nonlinear observation regime shows that the largest improvements over LETKF occur at large scales (small wavenumbers), where LETKF lacks sufficient ensemble spread. Overall, this initial application of EnSF to a geophysical model of intermediate complexity motivates further development of the algorithm for more realistic problems.

Artificial intelligence↗

Curating Carbon Storage Data for Reuse: Enabling Research and Modeling from Earth’s Surface to Subsurface

The volume of public geologic carbon storage (GCS) data resources has continued to increase in recent years as the result of an increase in funding from government, industry, and academia towards national, basin, regional and field scale studies to ensure carbon capture and storage becomes a commercially viable operation. Despite the increasing volume of data, GCS data applied towards analyses such as geologic, cost, and risk modeling continues to be multi-sourced and often disparate in nature, published across government agencies, websites, data repositories and buried in derivative reports and documents. Much of the time preparing for an analysis and derivative product development is spent collecting, aggregating, transforming and preparing input data. There have been significant efforts within the DOE National Energy Technology Laboratory’s Carbon Storage Program to optimize multi-source, multi-scale subsurface geologic data curation and aggregation to support data discovery, interoperability, and reuse. Methods include the use of artificial intelligence, machine learning, and data science techniques. This talk will discuss the workflows, best practices, and processes developed to support the aggregation and curation of data through the whole system – surface to subsurface data - that support multi-scale, multi-purpose analysis for carbon storage research.

Morkner, Paige↗

Unveiling X-ray absorption signatures of boron nitride via first-principles simulation and machine learning

Boron nitride (BN) allotropes hold great promise in many advanced applications ranging from optical and photonic devices to energy storage and battery systems to tribological components. The diverse functionalities of this material stem from BN’s highly tunable structural and electronic properties, which are governed by the versatile boron–nitrogen bonding configurations. Exploring the structural landscape of BN can unveil novel structures possessing unique properties suited for specific applications, therefore accelerating the design of next-generation advanced functional materials. In this work, we leverage boron K-edge X-ray absorption spectroscopy (XAS) as an effective probe for local structural features and chemical environments. A total of 210 BN crystal structures are generated via analogies to the extensive array of carbon allotropes, and XAS is simulated for each unique local motif within the resulting collection of structures. A mapping between structural features and spectral signatures was established by synergizing first-principle simulations with data-driven based post-analysis approaches. Specifically, we developed a neural network model that can satisfactorily predict spectra line shapes from local structural descriptors. Toward automatic spectroscopic interpretation of any new BN structures, supervised machine learning models, trained on this structure–spectrum dataset, can accurately infer local coordination environments from simulated XAS, highlighting the strength of this unique approach of combining high-fidelity first-principles simulation and machine-learning to accelerate target design of novel BN materials via rational understanding of local structure-spectrum correlations.

36 MATERIALS SCIENCE↗