Search NASASearch

SEARCH · Search NASA

Results for “autoencoder”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

Autoencoder-Based Sensor Drift Detection and Mitigation for Resilient Charging Systems

This work presents an autoencoder-based approach for sensor signal reconstruction and drift detection for charging systems. The proposed strategy is implemented within a Simulink-based system framework and evaluated under multiple operating conditions. An autoencoder with 8 neurons in the bottleneck layer is adopted, achieving accurate reconstruction across 10 variables and strong agreement with the physical sensor readings under normal conditions. In the case of a sensor fault, the autoencoder reconstruction remains closer to the expected true value compared to the corrupted measurement. Furthermore, feeding the autoencoder-reconstructed signal value back into the control framework in place of the faulty sensor signal leads to improved power monitoring. These results highlight the potential of autoencoder-based virtual sensing to extend the concept of resiliency to all components of the charging system, including sensors.

Rezende Da Costa Reis Kimpara, Renata [ORNL] (ORCI

Dense autoencoders, clustering techniques, and semi-supervised learning for HPGe $γ$-spectra

Classifying high-resolution gamma spectra by their isotopic content is an essential task in nuclear forensics and other applications. Traditional analysis methods are often time-intensive, but machine learning (ML) may help analysts quickly process many spectra. Such methods tend to rely on abundant, well-labeled data for training. Historical gamma data exists in various fields but is not uniformly useful for supervised ML due to inconsistent labeling. Here, to address some of these challenges, we present a method to classify and organize unlabeled data from high-purity germanium detectors using an autoencoding neural network (autoencoder). We trained dense autoencoders to compress gamma data into latent representations that enable efficient data characterization. By clustering the encoded spectra or lower-dimensional mappings of them, we identified and removed portions of over-abundant data categories, resulting in a more balanced dataset and improved autoencoder performance. This encoding and clustering pipeline also enabled the organization of spectra into self-consistent categories. Finally, we found that encoded representations showed potential as inputs for semi-supervised learning of nuclide identification (NID) labels, achieving an average F1 score of 0.85 ± 0.03 when mapping encodings to a set of 65 isotope labels.

Autoencoders

Denoising Autoencoder for Reconstructing Sensor Observation Data and Predicting Evapotranspiration: Noisy and Missing Values Repair and Uncertainty Quantification

Abstract Machine learning (ML) methods applied in scientific research often deal with interrelated features in high‐dimensional data. Reducing data noise and redundancy is needed to increase prediction accuracy and efficiency especially when dealing with data from field sensors. We explored an unsupervised learning method, the denoising autoencoder (DAE), to extract the underlying data structure from noisy raw data in the context of predicting hydrologic quantities from multiple field sensors. These sensors have intrinsic instrumental noise and occasional malfunctions that cause missing values. Our DAE neural network reconstructed meteorological sensor data containing noise and missing values to predict evapotranspiration in a mountainous watershed. The DAE reconstructed the sensor variables with a mean coefficient of determination value of 0.77 across 15 dimensions representing individual sensors. It reduced variance and bias uncertainties compared to a classical autoencoder model. The reconstruction quality varied across dimensions depending on their cross‐correlation and alignment with the underlying data structure. Uncertainties arising from the model structure were overall higher than those resulting from data corruption. We attached the DAE structure to a downstream ET‐prediction neural network in three formats and achieved reasonably accurate ET predictions . The use of the DAE notably reduced variance uncertainty in ET prediction. However, excessive variance reduction may be accompanied by an increase in bias due to the intrinsic bias‐variance tradeoff. Our method of evaluating and reducing uncertainties in aggregated data from different sources can be used to improve predictive models, process understanding, and uncertainty quantification for better water resource management. Plain Language Summary We present a machine learning method, namely the denoising autoencoder, which reduces the effects of data noise and missing values typically present in scientific data sets collected through sensor measurements. This method selects the most relevant information from noisy raw data collected by the instruments and fills in missing values. To demonstrate the effectiveness of our method, we applied it to predict evapotranspiration, a hydrologic variable that represents the water moved from the land surface to the atmosphere through a combination of evaporation and plant water use (transpiration). We also used a random sampling technique (the Monte Carlo method) to compare the uncertainty in the predictions when using the raw and noisy data versus the reconstructed data. The denoising process produced more accurate predictions of evapotranspiration with less uncertainty. Improved predictions of evapotranspiration can lead to a better understanding and accounting of water budgets. This ML approach is broadly suitable for a wide variety of applications that involve noisy sensor data with missing values. Key Points We used a denoising autoencoder (DAE) neural network to reduce noise in meteorological and soil sensor observations by on average We used Monte Carlo sampling to estimate the bias and variance of all model outputs, including uncertainty sources from data and the model We attached the DAE component to a downstream neural network to predict ET with the variance reduced by , compared to that without the DAE

denoising autoencoder

A Blind Convolutional Deep Autoencoder for Spectral Unmixing of Hyperspectral Images Over Waterbodies

Harmful algal blooms have dangerous repercussions for biodiversity, the ecosystem, and public health. Automatic identification based on remote sensing hyperspectral image analysis provides a valuable mechanism for extracting the spectral signatures of harmful algal blooms and their respective percentage in a region of interest. This paper proposes a new model called a non-symmetrical autoencoder for spectral unmixing to perform endmember extraction and fractional abundance estimation. The model is assessed in benchmark datasets, such as Jasper Ridge and Samson. Additionally, a case study of the HSI2 image acquired by NASA over Lake Erie in 2017 is conducted for extracting optical water types. The results using the proposed model for the benchmark datasets improve unmixing performance, as indicated by the spectral angle distance compared to five baseline algorithms. Improved results were obtained for various metrics. In the Samson dataset, the proposed model outperformed other methods for water (0.060) and soil (0.025) endmember extraction. Moreover, the proposed method exhibited superior performance in terms of mean spectral angle distance compared to the other five baseline algorithms. The non-symmetrical autoencoder for the spectral unmixing approach achieved better results for abundance map estimation, with a root mean square error of 0.091 for water and 0.187 for soil, compared to the ground truth. For the Jasper Ridge dataset, the non-symmetrical autoencoder for the spectral unmixing model excelled in the tree (0.039) and road (0.068) endmember extraction and also demonstrated improved results for water abundance maps (0.1121). The proposed model can identify the presence of chlorophyll-a in waterbodies. Chlorophyll-a is an essential indicator of the presence of the different concentrations of macrophytes and cyanobacteria. The non-symmetrical autoencoder for spectral unmixing achieves a value of 0.307 for the spectral angle distance metric compared to a reference ground truth spectral signature of chlorophyll-a. The source code for the proposed model, as implemented in this manuscript, can be found at https://github.com/EstefaniaAlfaro/autoencoder_owt_spectral.git.

hyperspectral imaging

Wasserstein normalized autoencoder for anomaly detection

A novel anomaly detection algorithm is presented. The Wasserstein normalized autoencoder (WNAE) is a normalized probabilistic model that minimizes the Wasserstein distance between the learned probability distribution—a Boltzmann distribution where the energy is the reconstruction error of the autoencoder (AE)—and the distribution of the training data. This algorithm has been developed and applied to the identification of semivisible jets—conical sprays of visible standard model (SM) particles and invisible dark matter states—with the CMS experiment at the CERN LHC. Trained on jets of particles from simulated SM processes, the WNAE is shown to learn the probability distribution of the input data in a fully unsupervised fashion, such that it effectively identifies new physics jets as anomalies. The model exhibits stable, convergent training and recovers strong classification performance for a wide range of signals against the selected background process, for which a standard AE fails because of outlier reconstruction. In addition, the model improves upon standard normalized autoencoders while remaining fully agnostic to the signal. The WNAE directly tackles the problem of outlier reconstruction, a common failure mode of autoencoders in anomaly detection tasks.

Hayrapetyan, Aram [Yerevan Phys. Inst.]

Monte Carlo Dropout Uncertainty Quantification of Long Short-Term Memory Autoencoder Anomaly Detection in a Liquid Sodium Cold Trap

Advanced high-temperature fluid reactors, such as sodium-cooled fast reactors (SFRs) and molten salt–cooled reactors (MSCRs), require coolant purification systems to prevent fluid contamination and local freezing that can lead to plugging. Liquid sodium purification can be achieved with a cold trap, where the sodium temperature is reduced to a near-freezing point to precipitate out impurities. Automation of monitoring of the cold trap performance with machine learning algorithms can aid in early detection of incipient anomalies. An efficient approach to loss-of-coolant–type anomaly detection in a cold trap monitored with more than two dozen thermal-hydraulic sensors consists of a long short-term memory (LSTM) autoencoder. This work develops the uncertainty quantification of the LSTM autoencoder performance for cold trap anomaly detection using the Monte Carlo (MC) dropout method. The MC dropout methodology creates a distribution of sister distributions that all slightly differ from each other because of random neurons being turned off for testing. The variances of the sister network distributions are used to make an uncertainty interval. Our analysis shows that the uncertainty in the autoencoder performance is largest near the peak of the anomaly signal. Using the MC dropout method, we investigate the uncertainty in the anomaly detection with missing sensor inputs. This capability allows the reactor operator to evaluate resilience of the anomaly detection system and to make informed decisions about continuity of operation in the event of sensor failure.

22 GENERAL STUDIES OF NUCLEAR REACTORS

Temporally-consistent koopman autoencoders for forecasting dynamical systems

Absence of sufficiently high-quality data often poses a key challenge in data-driven modeling of high-dimensional spatio-temporal dynamical systems. Koopman Autoencoders (KAEs) harness the expressivity of deep neural networks (DNNs), the dimension reduction capabilities of autoencoders, and the spectral properties of the Koopman operator to learn a reduced-order feature space with simpler, linear dynamics. However, the effectiveness of KAEs is hindered by limited and noisy training datasets, leading to poor generalizability. To address this, we introduce the Temporally-Consistent Koopman Autoencoder (tcKAE), designed to generate accurate long-term predictions even with limited and noisy training data. This is achieved through a consistency regularization term that enforces prediction coherence across different time steps, thus enhancing the robustness and generalizability of tcKAE over existing models. We provide analytical justification for this approach based on Koopman spectral theory and empirically demonstrate tcKAE’s superior performance over state-of-the-art KAE models across a variety of test cases, including simple pendulum oscillations, kinetic plasma, and fluid flow data.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY

Acoustic sensing and autoencoder approach for abnormal gas detection in a spent nuclear fuel canister mock-up

Currently, spent nuclear fuel (SNF) from commercial nuclear power plants is stored in stainless-steel canisters for interim dry storage. To provide an inert environment, these canisters are backfilled with helium after vacuum drying. However, the helium environment may be contaminated during extended storage because of the material degradation. For example, the heavier fission gas xenon may be released from the fuel rods into the canister cavity should the fuel cladding be breached. Other gases such as air and water vapor may also be present as a result of leakage caused by chloride-induced stress corrosion cracking on the canister walls or by insufficient vacuum drying. Therefore, monitoring the gas composition can provide critical information about the health of SNF canisters. In this study, noninvasive testing was conducted on a 2/3-scaled SNF canister mock-up using acoustic sensing. Ultrasonic transducers were placed on the exterior surface of the canister to probe the gas composition. A dataset was collected by sealing the canister mock-up and introducing up to 1.53% argon or 1.29% air into the helium background gas. Three methods were used to detect changes in the gas composition: the time-of-flight (TOF) method, the differential method, and the autoencoder method. Results showed that the TOF method had sufficient resolution to detect abnormal gas concentrations of less than 1.0%. The differential method demonstrated a periodic in-phase and out-of-phase behavior between the benchmark (i.e., pure helium) and abnormal (i.e., with argon or air) state signals. The variational autoencoder (VAE) and the Wasserstein autoencoder (WAE) were trained on the benchmark data and were applied directly to the abnormal state data. It was found that both the unsupervised VAE and the WAE were able to distinguish the benchmark and abnormal states of the canister mock-up based on the reconstruction error.

11 NUCLEAR FUEL CYCLE AND FUEL MATERIALS

Variational autoencoders for at-source data reduction and anomaly detection in high energy particle detectors

Detectors in next-generation high-energy physics experiments face several daunting requirements, such as high data rates, damaging radiation exposure, and stringent constraints on power, space, and latency. To address these challenges, machine learning in readout electronics can be leveraged for smart detector designs, enabling intelligent inference and data reduction at-source. Variational autoencoders (VAEs) offer a variety of benefits for front-end readout; an on-sensor encoder can perform efficient lossy data compression while simultaneously providing a latent space representation that can be used for anomaly detection. Results are presented from low-latency and resource-efficient VAEs for front-end data processing in a futuristic silicon pixel detector. Encoder-based data compression is found to preserve good performance of off-detector analysis while significantly reducing the off-detector data rate as compared to a similarly sized data filtering approach. Furthermore, the latent space information is found to be a useful discriminator in the context of real-time sensor defect monitoring. Together, these results highlight the multifaceted utility of autoencoder-based front-end readout schemes and motivate their consideration in future detector designs.

47 OTHER INSTRUMENTATION

Physicochemical and Performance Characterization of Six Commercial Organic Solvent Nanofiltration Membranes

This work introduces a novel, gradient-free metamaterial design method based on Gaussian process regression to represent the density field of a unit cell. The dimension of the design space is determined by the covariance matrix dimension in the Gaussian process regression. We propose compressing this matrix using an autoencoder, enabling the decoder to generate the density field and effectively reduce the originally large design space to a lower-dimensional subspace. In this compressed space, we employ an active learning method, Bayesian Adaptive Direct Search (BADS), for efficient exploration of the design space. We demonstrate that for simple 2D designs aimed at maximizing unit cell stiffness, our method yields results comparable to those of standard topology optimization. Furthermore, we extend our approach to various mechanical problems, from linear elasticity to hyperelastic large deformation and elasto-plasticity under finite deformation, to 3D metamaterial design. This illustrates the method’s versatility and effectiveness across a range of applications.

Wu, Haoran

Enabling dynamic 3D coherent diffraction imaging via adaptive latent space tuning of generative autoencoders

Abstract Coherent diffraction imaging (CDI) is an advanced non-destructive 3D X-ray imaging technique for measuring a sample’s electron density. The main challenge of CDI is loss of phase information in diffraction intensity measurements, resulting in lengthy iterative reconstruction processes that can return non-unique solutions, which pose challenges for experiments attempting to track dynamic sample evolution through multiple states. As the increased brightness of fourth-generation light sources enables faster sample measurements and drives operando experiments with Bragg CDI, there is a growing need for faster reconstruction techniques that can keep pace. We have developed an adaptive generative autoencoder approach for uniquely tracking a sample’s electron density as it dynamically evolves. Our approach adaptively tunes the low-dimensional latent embedding of a generative autoencoder, enabling a computationally efficient manner to account for time-varying shifting distributions in real-time. Analytic proof of convergence is provided as well as numerical demonstration of sample tracking with noisy measurements.

97 MATHEMATICS AND COMPUTING

Hierarchical-embedding autoencoder with a predictor as efficient architecture for learning time-evolution in multi-scale turbulent flows

We introduce a scale-aware, data-driven deep learning modeling framework for accurately predicting the time evolution of multi-scale turbulent plasma and liquid flows. The approach is motivated by the idea of scale separation. Structures of vastly different length scales emerge in these systems, and interactions between these structures occur only locally. To exploit this structure, the flow state is transformed by a hierarchical, fully convolutional autoencoder, not into a single embedding layer as in conventional convolutional surrogate models, but into a series of embedding layers. A stepwise training strategy ensures that fine-scale features are encoded on a high-resolution grid, while larger structures are represented on progressively coarser layers. The time evolution predictor advances all embedding layers in sync, capturing local interactions between features at the same scale as well as between all scales. This approach enables efficient modeling of multi-scale systems since negligible interactions between distant, small-scale structures do not need to be directly modeled. Our hierarchical-embedding autoencoder with a predictor framework is evaluated on canonical examples of multi-scale turbulence: two-dimensional Kolmogorov flow and Hasegawa–Wakatani plasma turbulence. In both cases, the proposed framework significantly improves predictive accuracy relative to conventional convolutional network architectures. A significant improvement in prediction accuracy was observed for crucial statistical characteristics of the Hasegawa–Wakatani plasma as well as for individual trajectories of the Kolmogorov flow turbulence. Importantly, the model's rollout for the Hasegawa–Wakatani problem demonstrates a four-order-of-magnitude speedup compared to traditional numerical solvers.

Khrabry, Alexander I. [Princeton Univ., NJ (United

Clustering Algorithm for AM Parts using GSH and EDT with Autoencoder

SAND2025-10103O The Clustering Algorithm for AM Parts Using GSH and (EDT With Autoencoder is a software tool. It uses a clustering algorithm for additive manufacturing (AM) parts using generalized spherical harmonics (GSH) and Euclidean distance transform (EDT) with an autoencoder to quantify material microstructure. The tool offers improved sensitivity to microstructural changes compared to traditional approaches. The tool integrates multiple microstructural properties, such as grain morphology, crystallographic orientation, and material phase information, to provide a comprehensive analysis of material microstructures. Sandia National Laboratories is a multimission laboratory managed and operated by National Technology & Engineering Solutions of Sandia, LLC, a wholly owned subsidiary of Honeywell International Inc., for the U.S. Department of Energy’s National Nuclear Security Administration under contract DE-NA0003525.

Rodgers, Theron [Sandia National Lab. (SNL-CA), Li

Data and scripts from: “Denoising autoencoder for reconstructing sensor observation data and predicting evapotranspiration: noisy and missing values repair and uncertainty quantification”

This data package includes data and scripts from the manuscript “Denoising autoencoder for reconstructing sensor observation data and predicting evapotranspiration: noisy and missing values repair and uncertainty quantification”.The study addressed common challenges faced in environmental sensing and modeling, including uncertain input data, missing sensor observations, and high-dimensional datasets with interrelated but redundant variables. Point-scaled meteorological and soil sensor observations were perturbed with noises and missing values, and denoising autoencoder (DAE) neural networks were developed to reconstruct the perturbed data and further predict evapotranspiration. This study concluded that (1) the reconstruction quality of each variable depends on its cross-correlation and alignment to the underlying data structure, (2) uncertainties from the models were overall stronger than those from the data corruption, and (3) there was a tradeoff between reducing bias and reducing variance when evaluating the uncertainty of the machine learning models.This package includes:(1) Four ipython scripts (.ipynb): “DAE_train.ipynb” trains and evaluates DAE neural networks, “DAE_predict.ipynb” makes predictions from the trained DAE models, “ET_train.ipynb” trains and evaluates ET prediction neural networks, and “ET_predict.ipynb” makes predictions from trained ET models.(2) One python file (.py): “methods.py” includes all user-defined functions and python codes used in the ipython scripts.(3) A “sub_models” folder that includes five trained DAE neural networks (in pytorch format, .pt), which could be used to ingest input data before being fed to the downstream ET models in ‘ET_train.ipynb” or ‘ET_predict.ipynb’.(4) Two data files (.csv). Daily meteorological, vegetation, and soil data is in “df_data.csv”, where “df_meta.csv” contains the location and time information of “df_data.csv”. Each row (index) in “df_meta.csv” corresponds to each row in “df_data.csv”. These data files are formatted to follow the data structure requirements and be directly used in the ipython scripts, and they have been shuffled chronologically to train machine learning models. The meteorological and soil data was collected using point sensors between 2019-2023 at(4.a) Three shrub-dominated field sites in East River, Colorado (named “ph1”, “ph2” and “sg5” in “df_meta.csv”, where “ph1” and “ph2” were located at PumpHouse Hillslopes, and “sg5” was at Snodgrass Mountain meadow) and(4.b) One outdoor, mesoscale, and herbaceous-dominated experiment in Berkeley, California (named “tb” in “df_meta.csv”, short for Smartsoils Testbed at Lawrence Berkeley National Lab).- See "df_data_dd.csv" and "df_meta_dd.csv" for variable descriptions and the Methods section for additional data processing steps. See "flmd.csv" and "README.txt" for brief file descriptions.- All ipython scripts and python files are written in and require PYTHON language software.

54 ENVIRONMENTAL SCIENCES

Evaluating the Efficacy of Conditional Variational Autoencoders in Generating Synthetic Single Nuclei RNA-Seq Data for Space Biology Research

Astronauts are subject to unique stressors during spaceflight, leading to changes in their cellular function. However, neither astronauts nor model organisms respond the same to spaceflight, and research implicates a contribution of omics components in differential responses. Understanding how gene expression affects astronaut health is critical for the success of long-term space missions, prompting interest in developing personalized predictive models leveraging artificial intelligence (AI) and machine learning (ML) techniques. Developing such models requires extensive data, which is challenging to obtain and share. This study explores the use of conditional variational autoencoders (CVAEs) to synthetically generate single-nuclei RNA-seq (snRNA-seq) data. CVAEs build on standard variational autoencoders (VAEs) by conditioning data generation on covariates like sample identity and mission parameters, enhancing the relevance of generated data for specific contexts. For our work, we built two CVAEs with varying degrees of sparsity to optimize both interpretability and generative power. We train and validate models on existing snRNA-seq data collected from the brain tissue of mice subjected to spaceflight conditions and their ground control counterparts. We evaluate model performance using statistical tests and visualizations to compare synthetic data to real data. We aim to demonstrate that these prototype CVAE architectures could be used in future space biology work and that this is a method worth further exploring.

Sarah Golts

Predicting critical heat flux with uncertainty quantification and domain generalization using conditional variational autoencoders and deep neural networks

Deep generative models (DGMs) can generate synthetic data samples that closely resemble the original dataset, addressing data scarcity. In this work, we developed a conditional variational autoencoder (CVAE) to augment critical heat flux (CHF) data used for the 2006 Groeneveld lookup table. To compare with traditional methods, a fine-tuned deep neural network (DNN) regression model was evaluated on the same dataset. Both models achieved small mean absolute relative errors, with the CVAE showing more favorable results. Uncertainty quantification (UQ) was performed using repeated CVAE sampling and DNN ensembling. The DNN ensemble improved performance over the baseline, while the CVAE maintained consistent results with less variability and higher confidence. Both models achieved small errors inside and outside the training domain, with slightly larger errors outside. Altogether, the CVAE performed better than the DNN in predicting CHF and exhibited better uncertainty behavior.

22 GENERAL STUDIES OF NUCLEAR REACTORS

Predicting non-linear stress–strain response of mesostructured cellular materials using supervised autoencoder

Recent breakthroughs in advanced manufacturing capabilities have made it possible to design and print sophisticated topologies of cellular structures using diverse engineering materials such as metals, polymers, and ceramics. In these architectured materials, it is often desirable to tailor the mechanical properties by altering the unit cell topology. This necessitates an in-depth understanding of how the topology of the unit cell structure affects the macroscopic behavior of the material in both the linear and the non-linear regimes encountered under large compression. Here, we have developed a machine learning (ML) approach capable of accelerating the prediction of the stress–strain response of a polymer-based cellular structure under uniaxial confined compression. As part of generating the training data for ML, 60,000 mesostructures were generated using a relatively novel approach based on cellular automata, and their corresponding stress–strain responses were obtained from the finite element simulations. Principal component analysis (PCA) was used to reduce the dimensionality of the stress–strain curves. With only 20 principal components, PCA captured 99.89% of the variance in the stress–strain curves while reducing the dimensionality by 5X. ML using supervised autoencoder was able to successfully speed up the prediction of the non-linear stress–strain response of a unit cell by up to 4600X. The proposed method can serve as an efficient data generation tool and a rapid means for predicting the structure–property relationship through accelerated forward modeling of cellular materials under compaction, in cases where the macroscopic stress–strain response is governed by the unit-cell topology.

36 MATERIALS SCIENCE

Unsupervised atomic data mining via multi-kernel graph autoencoders for machine learning force fields

Constructing a chemically diverse dataset while avoiding sampling bias is critical to training efficient and generalizable force fields. However, in computational chemistry and materials science, many common dataset generation techniques are prone to oversampling regions of the potential energy surface. Furthermore, these regions can be difficult to identify and isolate from each other or may not align well with human intuition, making it challenging to systematically remove bias in the dataset. While traditional clustering and pruning (down-sampling) approaches can be useful for this, they can often lead to information loss or a failure to properly identify distinct regions of the potential energy surface due to difficulties associated with the high dimensionality of atomic descriptors. In this work, we introduce the Multi-kernel Edge Attention-based Graph Autoencoder (MEAGraph) model, an unsupervised approach for analyzing atomic datasets. MEAGraph combines multiple linear kernel transformations with attention-based message passing to capture geometric sensitivity and enable effective dataset pruning without relying on labels or extensive training. Demonstrated applications on niobium, tantalum, and iron datasets show that MEAGraph efficiently groups similar atomic environments, allowing for the use of basic pruning techniques for removing sampling bias. This approach provides an effective method for representation learning and clustering that can be used for data analysis, outlier detection, and dataset optimization.

Materials science