Search NASASearch

SEARCH · Search NASA

Results for “Data scarcity”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 37 records · Page 2

A review on machine learning-guided design of energy materials

Abstract The development and design of energy materials are essential for improving the efficiency, sustainability, and durability of energy systems to address climate change issues. However, optimizing and developing energy materials can be challenging due to large and complex search spaces. With the advancements in computational power and algorithms over the past decade, machine learning (ML) techniques are being widely applied in various industrial and research areas for different purposes. The energy material community has increasingly leveraged ML to accelerate property predictions and design processes. This article aims to provide a comprehensive review of research in different energy material fields that employ ML techniques. It begins with foundational concepts and a broad overview of ML applications in energy material research, followed by examples of successful ML applications in energy material design. We also discuss the current challenges of ML in energy material design and our perspectives. Our viewpoint is that ML will be an integral component of energy materials research, but data scarcity, lack of tailored ML algorithms, and challenges in experimentally realizing ML-predicted candidates are major barriers that still need to be overcome.

36 MATERIALS SCIENCE

The Past, Present and Future of Structural Health Monitoring: An Overview of Three Ages

This paper presents an overview of the discipline of structural health monitoring (SHM), organised in terms of three proposed ages. The first age is delineated by the prehistory of SHM and the period where nondestructing testing methods evolved into an organised set of principles built upon physics-based models; this age ended when the model-based approaches reached an impasse in terms of their ability to properly deal with real-world problems. The second age of SHM began with a transition to data-based methods based on statistical pattern recognition, which provided a holistic approach to SHM problems for the first time. This age arguably ended when the methods foundered in situations where the necessary training data were scarce. It is argued here that the third age began with the development of population-based SHM, which has been designed to overcome the problem of data scarcity. As there is very limited space in a single article to provide a comprehensive overview, an appendix has been provided here that gives a very systematic bibliography of SHM reviews—a meta-bibliography.

60 APPLIED LIFE SCIENCES

Health Management and Prognostics for Electric Aircraft Powertrain

W and c Any air borne vehicle needs incorporating safety as key parameter of measure, and inclusion of autonomy raises the critical need for safety under autonomous operations. Management of faults and component degradation is key as complexity in autonomous operations grow over the period of time. Therefore, in addition to basic operational requirements, an autonomous electric vehicle should be able to make accurate estimates of its current system health and take the correct decisions to complete its mission successfully. Real-time safety and state-awareness tools are therefore essential for the vehicle to be able to reach its destination in a safe and successful manner. The need for safety assurance and health management capabilities is particularly relevant for aircraft electric propulsion systems, which are relatively new and with limited historical to learn. They are critical systems requiring high power density along with reliability, resilience, efficient management of weight, and operational costs. A model- based fault diagnosis and prognostics approach of complex critical systems can successfully accomplish the safety and state awareness goal for such electric propulsion systems, enabling autonomous decision making capability for safe and efficient operation. To identify critical components in the system a Qualitative Bayesian approach using FMECA is implemented. This requires the assessment of some quantities representing the state of the electric unmanned aerial systems (e-UAS), as well as look-ahead forecasts of such states during the entire flight, presented in form of safety metrics (SM). In-service data and performance data gathered from degraded components sup- ports diagnostic and prognostic methods for these systems, but this data can be difficult to obtain as weight and packaging restrictions reduce redundancy and instrumentation on-board the vehicle. Therefore, an model-based framework should be capable or operating with limited data. In addition to data scarcity, the variability of such complex critical systems re- quires the model-based framework to reason in the presence of uncertainty, such as sensor noise, and modeling imperfections. Quantification of errors and uncertainties in the measured states and quantities is therefore a fundamental step for a precise estimation of such SMs; un-modeled uncertainty may result in erroneous state assessment and un- reliable predictions of future states of e-UAVs. Typical, centralized model-based schemes suffer from inherent disadvantages such as computational complexity, single point of failure, and scalability issues, and therefore may fail in such a complex scenario. This paper presents a methodology for developing a system level diagnostics and prognostics approach using a Qualitative Bayesian FMECA approach along with a formal uncertainty management framework for an e-UAS. In this work we demonstrate the efficacy of the framework to predict effects of sub-system level degradation on vehicle operation incorporating uncertainty management to predict future behavior under different operating conditions.

Kulkarni, Chetan

PROS: An IRAF based system for analysis of x ray data

PROS is an IRAF based software package for the reduction and analysis of x-ray data. The use of a standard, portable, integrated environment provides for both multi-frequency and multi-mission analysis. The analysis of x-ray data differs from optical analysis due to the nature of the x-ray data and its acquisition during constantly varying conditions. The scarcity of data, the low signal-to-noise ratio and the large gaps in exposure time make data screening and masking an important part of the analysis. PROS was developed to support the analysis of data from the ROSAT and Einstein missions but many of the tasks have been used on data from other missions. IRAF/PROS provides a complete end-to-end system for x-ray data analysis: (1) a set of tools for importing and exporting data via FITS format -- in particular, IRAF provides a specialized event-list format, QPOE, that is compatible with its IMAGE (2-D array) format; (2) a powerful set of IRAF system capabilities for both temporal and spatial event filtering; (3) full set of imaging and graphics tasks; (4) specialized packages for scientific analysis such as spatial, spectral and timing analysis -- these consist of both general and mission specific tasks; and (5) complete system support including ftp and magnetic tape releases, electronic and conventional mail hotline support, electronic mail distribution of solutions to frequently asked questions and current known bugs. We will discuss the philosophy, architecture and development environment used by PROS to generate a portable, multimission software environment. PROS is available on all platforms that support IRAF, including Sun/Unix, VAX/VMS, HP, and Decstations. It is available on request at no charge.

Conroy, M. A.

Coincident learning for beam-based rf station fault identification using phase information at the SLAC linac coherent light source

Anomalies in radio-frequency (rf) stations can result in unplanned downtime and performance degradation in linear accelerators such as SLAC’s Linac Coherent Light Source (LCLS). Detecting these anomalies is challenging due to the complexity of accelerator systems, high data volume, and scarcity of labeled fault data. Prior work identified faults using beam-based detection, combining rf amplitude and beam position monitor data. Due to the simplicity of the rf amplitude data, classical methods are sufficient to identify faults, but the recall is constrained by the low-frequency and asynchronous characteristics of the data. In this work, we leverage high-frequency, time-synchronous rf phase data to enhance anomaly detection in the LCLS accelerator. Due to the complexity of phase data, classical methods fail, and we instead train deep neural networks within the Coincident Anomaly Detection (CoAD) framework. We find that applying CoAD to phase data detects nearly 3 times as many anomalies as when applied to amplitude data, while achieving broader coverage across rf stations. Furthermore, the rich structure of phase data enables us to cluster anomalies into distinct physical categories. Through the integration of auxiliary system status bits, we link clusters to specific fault signatures, providing additional granularity for uncovering the root cause of faults. We also investigate interpretability via Shapley values, confirming that the learned models focus on the most informative regions of the data and providing insight for cases where the model makes mistakes. This work demonstrates that phase-based anomaly detection for rf stations improves both diagnostic coverage and root cause analysis in accelerator systems and that deep neural networks are essential for effective analysis.

Accelerator Physics (physics.acc-ph)

Counter Data Paucity through Adversarial Invariance Encoding: A Case Study on Modeling Battery Thermal Runaway

Lithium-ion batteries, widely used for their durability and high energy storage, face the risk of internal short circuits leading to catastrophic thermal runaway events. These events, triggered by external stimuli like mechanical loads, pose safety concerns in applications such as electric vehicles. Detecting and understanding thermal runaway events is crucial, but physics-driven models struggle to explain the non-linear evolution of battery temperature during these events, considering factors like material composition and state-of-charge. Due to the rarity of these events and the cost of data collection, we propose a deep learning (DL) model to predict battery temperature responses during thermal runaway. The challenge lies in the scarcity of data, making traditional DL models prone to overfitting and learning low-quality representations of the complex process.Our approach introduces a novel few-shot architecture that incorporates an adversarially governed invariant encoding process. This architecture aims to distill "invariant" relationships by addressing distributional shifts in data across various battery properties, facilitating the detection of thermal runaway events. Specifically, our results demonstrate that deep learning models conditioned on these "invariant" representations outperform state-of-the-art baselines, achieving a remarkable 96.8% performance improvement in terms of the popular metric MAPE. This framework presents a promising direction for enhancing battery safety modeling, particularly in the context of rare and complex events like thermal runaway. Our code and code and dataset used for the paper are public1.

Tabassum, Anika [ORNL] (ORCID:0000000254600955)

Summary and recommendations for initial exercise prescription

The recommendations summarized herein constitute a basis on which an initial exercise prescription can be formulated. It is noteworthy that any exercise program designed currently would be an approximation. Examination of the existing space-flight data reveals a scarcity of in-flight data on which to rigorously design an exercise program. The relevant experience within the U.S. space program (with regard to long-duration space flight) is limited to the Skylab Program. Lessons learned from Skylab are relevant to the design of a Space Station exercise program, especially with regard to the total length of exercise time required, cardiovascular (CV) deconditioning/reconditioning, and bone loss. Certain observations of the U.S.S.R. exercise activities can also contribute to the formulation of an exercise prescription of Space Station. Reportedly, the U.S.S.R. uses both a bicycle ergometer and a treadmill device on long-duration missions with some degree of success. Using the third crew of Salyut 6, which was a 175-day stay, as a representative mission, the typical time dedicated to exercise varies from 2 to 3 hours per day. In addition, the cosmonauts wear an elasticized suit, called a penquin suit, for time periods ranging from 12 to 16 hours per day. This device provides a load across the axial skeleton against which the wearer must exert himself. Despite these extensive countermeasures, the effects of adaptation are not totally prevented.

Stewart, Donald F.

Evaluating the Use of Foundational Chemical Language Models in Multimodal Graph Fusion

Rapid and accurate prediction of the physicochemical properties of molecules given their structures remains a key challenge in cheminformatics. Machine learning approaches offer high-throughput options, but the optimality of inductive biases and data representations are up for debate. For example, BERT-based masked language models (MLMs) can be trained in a self-supervised way on hundreds of millions to billions of readily available SMILES strings. Another option is graph neural networks (GNNs), which can operate directly on molecular structures. Yet, generating accurate molecular geometry is computationally expensive, leading to a relative scarcity in data compared to SMILES strings. It is attractive to combine these two paradigms by pre-training an LM on a large corpus of SMILES strings and embedding these representation into a geometric graph neural network. Despite the promise of such an approach, and contrary to previous studies, we find mixed results with the combination of the LMs and GNNs on several molecule datasets. In particular, we found evidence for improvement on the FreeSolv and QM7 benchmarks, but degraded performance on the ESOL, LIPO and QM9 datasets compared to a GNN baseline.

Francel, Collin [University of Alabama]

Toward a New Generation of Agricultural System Data, Models, and Knowledge Products: State of Agricultural Systems Science

We review the current state of agricultural systems science, focusing in particular on the capabilities and limitations of agricultural systems models. We discuss the state of models relative to five different Use Cases spanning field, farm, landscape, regional, and global spatial scales and engaging questions in past, current, and future time periods. Contributions from multiple disciplines have made major advances relevant to a wide range of agricultural system model applications at various spatial and temporal scales. Although current agricultural systems models have features that are needed for the Use Cases, we found that all of them have limitations and need to be improved. We identified common limitations across all Use Cases, namely 1) a scarcity of data for developing, evaluating, and applying agricultural system models and 2) inadequate knowledge systems that effectively communicate model results to society. We argue that these limitations are greater obstacles to progress than gaps in conceptual theory or available methods for using system models. New initiatives on open data show promise for addressing the data problem, but there also needs to be a cultural change among agricultural researchers to ensure that data for addressing the range of Use Cases are available for future model improvements and applications. We conclude that multiple platforms and multiple models are needed for model applications for different purposes. The Use Cases provide a useful framework for considering capabilities and limitations of existing models and data.

Livestock models

The Stable Isotope Fractionation of Abiotic Reactions: A Benchmark in the Detection of Life

One very important tool in the analysis of biogenic, and potentially biogenic, samples is the study of their stable isotope distributions. The isotope distribution of a sample depends on the process(es) that created it. One important application of the analysis of C & N stable isotope ratios has been in the determination of whether organic matter in a sample is of biological origin or was produced abiotically. For example, the delta C-13 of organic material found embedded in phosphate grains was cited as a critical part of the evidence for life in 3.8 billion year old samples. The importance of such analysis in establishing biogenicity was highlighted again by the role this issue played in the recent debate over the validity of what had been accepted as the Earth s earliest microfossils. These kinds of analysis imply a comparison with the fractionation that one would have seen if the organic material had been produced by alternative, abiotic, pathways. Could abiotic reactions account for the same level of fractionation? Additionally, since the fractionation can vary between different abiotic reactions, understanding their fractionations can be important in distinguishing what reactions may have been significant in the formation of different abiological samples (such as extraterrestrial samples). There is however, a scarcity of data on the fractionation of carbon and nitrogen by abiotic reactions. In order to interpret properly what the stable isotope ratios of samples tell us about their biotic or abiotic nature, more needs to be known about how abiotic reactions fractionate C and N. Carbon isotope fractionations have been studied for a few abiotic processes. These studies presumed the presence of a reducing atmosphere, focusing on reactions involving spark discharge, W photolysis of reducing gas mixtures, and cyanide polymerization in the presence of ammonia. They did find that the initial products showed a depletion in I3C with values in the range of a few per mil to as low as -60 % (potentially comparable to that which accompanies the biosynthesis of organic matter). We need to understand what kind of fractionations are observed with reactions under the non-reducing or mildly reducing conditions now thought to be present on the early Earth. While nitrogen is receiving increased attention as a tool for these kinds of analyses, almost nothing is known about the isotope fractionation that one would expect for abiotic sources of fixed/reduced nitrogen. This project will measure the fixation from a series of abiotic reactions that may have been present on the early Earth (and other terrestrial planets) and produced organic material that could have ended up in the rock record. The work will look at a number of reactions, under a non- reducing, or mildly reducing, atmosphere, covering sources of prebiotic organic C & N from shock heating, to photochemistry, to hydrothermal reactions. Some reactions that we plan to study are; Shock heating of a non-reducing atmosphere to produce CO and NO (in collaboration with Chris McKay), formation of formaldehyde (and related compounds) from COY the formation of ammonia from nitrogen oxides (ultimately from NO) by ferrous iron reduction, and the hydrothermal synthesis of compounds including the hydrocarboxylation/hydrocarbonylation reaction (in collaboration with George Cody), reactions of oxalate to form hydrocarbons and other oxygenated compounds and the formation of lipids from oxalic/formic acid (in collaboration with Tom McCollom), and reactions of carbon monoxide & carbon dioxide with N2, ammonia or nitritehitrate to form hydrogen cyanide, nitriles, ammonia/amines and nitrous

Summers, David P.

Additively Printed Flexible Temperature Sensor for Wearable Applications

The flexible sensor has the capability to be mounted on any curved surfaces of applications and be used for portable devices. Additively printed sensors have received attention owing to their compact design and ability of application to non-planar surfaces. Wearable applications require capability of integration into a variety of surfaces with ability to flex, fold, twist and stretch under the stresses of daily motion. There is scarcity of data on the interaction of the process parameters with the realized performance. In addition, there is need for data focused on sensor accuracy, repeatability, and reliability. In this study, experimental analysis on function of the fabricated sensing board is conducted. The temperature sensors are made by direct write printing method with nScrypt printer. A calibration of the sensors has been conducted to confirm that resistance is well related to actual temperature and find TCR (temperature coefficient to resistance). The evolution of resistance has been correlated with the environmental temperature. The sensor hysteresis has been quantified using upswing and downswing of the environmental temperature. In addition, the effect of humidity on the temperature sensor accuracy and performance has been quantified. The effect of a polymide coat on the sensor to prevent humidity effects has also been quantified.

Pradeep Lall

PMU Data Quality and Sensor Health Monitoring

Phasor Measurement Units (PMUs) play a critical role in the evolution of the electric power industry by providing high-precision, real-time monitoring of essential power system metrics. However, effectively detecting abnormalities and critical events from PMU data is a complex task, complicated by intricate temporal patterns, a scarcity of labeled data for training algo- rithms, and constraints on online computational power. In this study, we apply TranAD, an innovative algorithm that combines transformer architectures with the refinement of adversarial learning, to both synthetic and real-world PMU datasets for developing a data quality and sensor online health monitoring platform for utilities. Our findings reveal that TranAD not only provides efficient detection and localization but also enhances the detail with which abnormalities are detected, marking a a significant step forward in the field of clean data acquisition processes for power system monitoring

deep neural network, machine learning (ML)

Machine Learning‐Assisted Microearthquake Location Workflow for Monitoring the Newberry Enhanced Geothermal System

Abstract Enhanced geothermal systems (EGS) offer a sustainable energy source but face challenges in accurately locating microearthquakes induced during reservoir stimulation. Locating these microearthquakes provides reliable feedback on the stimulation progress. Current deep learning methods for locating earthquakes require extensive data sets for training, which is problematic as detected microearthquakes are often limited. To address the scarcity of training data, we propose a practical workflow using probabilistic multilayer perceptron (PMLP) which predicts microearthquake locations from cross‐correlation time lags in waveforms. Utilizing a 3D velocity model of Newberry site derived from ambient noise interferometry, we generate numerous synthetic microearthquakes and 3D acoustic waveforms for PMLP training. Accurate synthetic tests prompt us to apply the trained network to the 2012 and 2014 stimulation field waveforms. To enhance the accuracy of source localization, we carefully handpick the P‐arrival times. Predictions on the 2012 stimulation data set show major microseismic activity at depths of 0.5–1.2 km, correlating with a known casing leakage scenario. In the 2014 data set, the majority of predictions concentrate at 2.0–2.9 km depths, consistent with results obtained from conventional physics‐based inversion, and align with the presence of natural fractures from 2.0 to 2.7 km. We validate our findings by comparing the synthetic and field picks, demonstrating a satisfactory match for the first arrivals. By combining the benefits of quick inference speeds and accurate location predictions, we demonstrate the feasibility of using realistic synthetic data set to locate microseismicity for EGS monitoring.

15 GEOTHERMAL ENERGY

Towards a liana plant functional type for vegetation models

Lianas (woody climbers) are crucial components of tropical forests and they have been increasingly recognized to have profound effects on tropical forest carbon dynamics. Despite their importance, lianas' representation in vegetation models remains limited, partly due to the complexity of liana-tree dynamics and the diversity in liana life history strategies. This paper provides a comprehensive review of advances and challenges for mechanistically representing lianas in forest ecosystem models and a proposed path towards effectively representing lianas in these models. Defining a liana plant functional type is a significant challenge because of the high morphological and physiological diversity amongst liana species, and because of their structural association with trees. Here, we identify critical liana traits that likely should contribute to establishing a liana plant functional type, along with key processes to properly represent lianas in ecosystem models. Subsequently, we discuss a variety of possible liana implementation strategies with their associated strengths, limitations, computational costs and data requirements. A fundamental redesign of the tree-centric demographic vegetation models seems appropriate to accommodate the unique growth and competition strategies of lianas. We illustrate the potential of such models with a single-site case study where we disentangle putative mechanisms of liana increasing abundance. Furthermore, we underscore the critical need for comprehensive liana demographic and functional data (including long-term, physiological, and pantropical observations) for the qualitative implementation and evaluation in the proposed modeling efforts. Currently, there is a scarcity of liana data and the data that do exist have a neotropical bias. We finally introduce a new liana functional trait database that can centralize existing liana trait data, incentivize improved data gathering and thus facilitate model development and scientific analyses.

54 ENVIRONMENTAL SCIENCES

Simulated convective systems using a cloud resolving model: Impact of large-scale temperature and moisture forcing using observations and GEOS-3 reanalysis

The GCE (Goddard Cumulus Ensemble) model, which has been developed and improved at NASA Goddard Space Flight Center over the past two decades, is considered as one of the finer and state-of-the-art CRMs (Cloud Resolving Models) in the research community. As the chosen CRM for a NASA Interdisciplinary Science (IDS) Project, GCE has recently been successfully upgraded into an MPI (Message Passing Interface) version with which great improvement has been achieved in computational efficiency, scalability, and portability. By basically using the large-scale temperature and moisture advective forcing, as well as the temperature, water vapor and wind fields obtained from TRMM (Tropical Rainfall Measuring Mission) field experiments such as SCSMEX (South China Sea Monsoon Experiment) and KWAJEX (Kwajalein Experiment), our recent 2-D and 3-D GCE simulations were able to capture detailed convective systems typical of the targeted (simulated) regions. The GEOS-3 [Goddard EOS (Earth Observing System) Version-3] reanalysis data have also been proposed and successfully implemented for usage in the proposed/performed GCE long-term simulations (i.e., aiming at producing massive simulated cloud data -- Cloud Library) in compensating the scarcity of real field experimental data in both time and space (location). Preliminary 2-D or 3-D pilot results using GEOS-3 data have generally showed good qualitative agreement (yet some quantitative difference) with the respective numerical results using the SCSMEX observations. The first objective of this paper is to ensure the GEOS-3 data quality by comparing the model results obtained from several pairs of simulations using the real observations and GEOS-3 reanalysis data. The different large-scale advective forcing obtained from these two kinds of resources (i.e., sounding observations and GEOS-3 reanalysis) has been considered as a major critical factor in producing various model results. The second objective of this paper is therefore to investigate and present such an impact of large-scale forcing on various modeled quantities (such as hydrometeors, rainfall, and etc.). A third objective is to validate the overall GCE 3-D model performance by comparing the numerical results with sounding observations, as well as available satellite retrievals.

Shie, C.-L.

Advancements and opportunities to improve bottom–up estimates of global wetland methane emissions

Wetlands are the single largest natural source of atmospheric methane (CH 4 ), contributing approximately 30% of total surface CH 4 emissions, and they have been identified as the largest source of uncertainty in the global CH 4 budget based on the most recent Global Carbon Project CH 4 report. High uncertainties in the bottom–up estimates of wetland CH 4 emissions pose significant challenges for accurately understanding their spatiotemporal variations, and for the scientific community to monitor wetland CH 4 emissions from space. In fact, there are large disagreements between bottom–up estimates versus top–down estimates inferred from inversion of atmospheric CH 4 concentrations. To address these critical gaps, we review recent development, validation, and applications of bottom–up estimates of global wetland CH 4 emissions, as well as how they are used in top–down inversions. These bottom–up estimates, using (1) empirical biogeochemical modeling (e.g. WetCHARTs: 125–208 TgCH 4 yr -1 ); (2) process-based biogeochemical modeling (e.g. WETCHIMP: 190 ± 39 TgCH 4 yr -1 ); and (3) data-driven machine learning approach (e.g. UpCH4: 146 ± 43 TgCH 4 yr -1 ). Bottom–up estimates are subject to significant uncertainties (~80 Tg CH 4 yr -1 ), and the ranges of different estimates do not overlap, further amplifying the overall uncertainty when combining multiple data products. These substantial uncertainties highlight gaps in our understanding of wetland CH 4 biogeochemistry and wetland inundation dynamics. Major tropical and arctic wetland complexes are regional hotspots of CH 4 emissions. However, the scarcity of satellite data over the tropics and northern high latitudes offer limited information for top–down inversions to improve bottom–up estimates. Recent advances in surface measurements of CH 4 fluxes (e.g. FLUXNET-CH 4 ) across a wide range of ecosystems including bogs, fens, marshes, and forest swamps provide an unprecedented opportunity to improve existing bottom–up estimates of wetland CH 4 estimates. We suggest that continuous long-term surface measurements at representative wetlands, high fidelity wetland mapping, combined with an appropriate modeling framework, will be needed to significantly improve global estimates of wetland CH 4 emissions. There is also a pressing unmet need for fine-resolution and high-precision satellite CH 4 observations directed at wetlands.

54 ENVIRONMENTAL SCIENCES

Detecting thermodynamic phase transition via explainable machine learning of photoemission spectroscopy

Identifying thermodynamic signatures of electronic phases, such as superconductivity, is challenging in low-dimensional materials due to strong fluctuations and low probing volume. Spectroscopic methods are often used to identify new bulk phases, but their main measurable quantity—electronic energy gaps—is no longer an effective order parameter in low-dimensional and fluctuating systems. Combining angle-resolved photoemission with a domain-adversarial neural network, we report a data-driven method to identify thermodynamic phase transitions solely based on single-particle spectra. We demonstrate 97.6% accuracy in cuprate superconductor Bi 2 Sr 2 CaCu 2 O 8+δ with strong superconducting fluctuations. This model notably compensates for the scarcity of experimental data by leveraging virtually inexhaustible simulated data. Further, its explainability reveals the crucial role of in-gap spectral weight in detecting phase fluctuations and thermodynamic transitions. Our work pinpoints the spectroscopic signatures of fluctuating orders and enables using spectroscopy for machine-learning-assisted material discovery for low-dimensional and strong coupling systems.

2D materials

Machine learning approaches for crystallographic classification from synthetic 2D X-ray diffraction data

Crystallographic structure identification is crucial for understanding material properties; however, current methodologies often depend on labor-intensive and time-consuming analyses of 2D X-ray diffraction (XRD) patterns. To address these limitations, this study employs synthetic 2D XRD patterns combined with deep learning (DL) techniques to enable automated and high-throughput classification of the seven crystal systems and 230 space groups. We introduce the novel Auto Diffraction Pipeline, designed to generate synthetic 2D XRD spot patterns from crystallographic information files under diverse conditions, including varying zone axes, atomic substitution, atomic depletion and mechanical loading. These conditions enhance the realism of synthetic data, mitigating the scarcity of experimental datasets and enabling the creation of large representative training sets. Convolutional neural networks were trained and validated on these synthetic datasets to classify crystallographic structures across multiple scenarios. Our results demonstrate that integrating synthetic 2D XRD patterns with DL facilitates rapid, accurate and automated crystallographic classification, promoting the wider adoption of data-driven approaches in materials science.

Shahnazari, Ayoub [Univ. of Rochester, NY (United