Search NASA⌕ Search

SEARCH · Search NASA

Results for “Unsupervised machine learning”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

Unveiling and Mapping Polymorphs in Fluorite Y2TiO5 Using 4D-STEM and Unsupervised Machine Learning

Y2TiO5 belongs to the Ln2TiO5 (Ln = lanthanide or Y) family of ceramic materials and exhibits a range of desirable material properties such as radiation tolerance, frustrated magnetism, and large dielectric constant. However, understanding the complex crystal structure of Y2TiO5 remains elusive, given that Y2TiO5 can adopt multiple polymorphs such as cubic, orthorhombic, and hexagonal phases within the lattice. In this work, we report a detailed structural analysis of Y2TiO5 using four-dimensional scanning transmission electron microscopy coupled with unsupervised machine learning. The pyrochlore nanodomains, characterized by the ordered arrangement of yttrium cations on the A site of their A2BO5 structure, are present within the matrix of a predominantly fluorite-structured Y2TiO5 along with a third polymorph, the hexagonal phase. The pyrochlore phase is found to form 2 nm boundary regions around hexagonal phase stacking faults, highlighting the potential influence of the hexagonal phase on the occurrence and distribution of the pyrochlore phase. Lastly, we identify a unique pyrochlore phase with asymmetric arrangement of cation ordering along a single planar direction. Our findings provide invaluable insights into the possible mechanisms stabilizing pyrochlore nanodomains within the fluorite lattice of Y2TiO5.

36 MATERIALS SCIENCE↗

Exploring Continuous Seismic Data at an Industry Facility Using Unsupervised Machine Learning

Seismic data recorded at industrial sites contain valuable information on anthropogenic activities. With advances in machine learning and computing power, new opportunities have emerged to explore the seismic wavefield in these complex environments. We applied two unsupervised machine learning algorithms to analyze continuous seismic data collected from an industrial facility in Texas, United States. The Uniform Manifold Approximation and Projection for Dimension Reduction algorithm was used to reduce the dimensionality of the data and generate 2D embeddings. Then, the Hierarchical Density-Based Spatial Clustering of Applications with Noise method was employed to automatically group these embeddings into distinct signal clusters. Our analysis of over 1400 hr (around 59 days) of continuous seismic data revealed five and seven signal clusters at two separate stations. At both stations, we identified clusters associated with background noise and vehicle traffic, with the latter’s temporal patterns aligning closely with the facility’s work schedule. Furthermore, the algorithms detected signal clusters from unknown sources and underline the ability of unsupervised machine learning for uncovering previously unrecognized patterns. Our analysis demonstrates the effectiveness of unsupervised approaches in examining continuous seismic data without requiring prior knowledge or pre-existing labels.

58 GEOSCIENCES↗

Unsupervised machine learning discovery of structural units and transformation pathways from imaging data

We show that unsupervised machine learning can be used to learn chemical transformation pathways from observational Scanning Transmission Electron Microscopy (STEM) data. To enable this analysis, we assumed the existence of atoms, a discreteness of atomic classes, and the presence of an explicit relationship between the observed STEM contrast and the presence of atomic units. With only these postulates, we developed a machine learning method leveraging a rotationally invariant variational autoencoder (VAE) that can identify the existing molecular fragments observed within a material. The approach encodes the information contained in STEM image sequences using a small number of latent variables, allowing the exploration of chemical transformation pathways by tracing the evolution of atoms in the latent space of the system. The results suggest that atomically resolved STEM data can be used to derive fundamental physical and chemical mechanisms involved, by providing encodings of the observed structures that act as bottom-up equivalents of structural order parameters. The approach also demonstrates the potential of variational (i.e., Bayesian) methods in the physical sciences and will stimulate the development of more sophisticated ways to encode physical constraints in the encoder–decoder architectures and generative physical laws and causal relationships in the latent space of VAEs.

97 MATHEMATICS AND COMPUTING↗

Identifying recharge sources and their impacts on a North Central New Mexico shallow aquifer using unsupervised machine learning

In this article, shallow aquifers are important but highly variable resources in arid to semi-arid regions. Limited shallow aquifer volume results in high sensitivity to recharge fluctuations, which can impact the local fauna and flora, and transport of contaminants in the aquifer or vadose zone. Aquifer response to external forcing (e.g., precipitation) is usually solved by estimating aquifer parameters and running physics-based models to match known fluctuations of hydraulic head. However, this technique is time and computationally expensive. Furthermore, high aquifer complexity decreases precision in physics-based models. Alternatively supervised machine learning is used to predict aquifer dynamics. However, these techniques rely on input data and struggle to interpret aquifer response for missing sources (i.e., snowpack data). To counter these problems, we propose an unsupervised machine learning technique (NMFk) to estimate the impact of different sources on aquifer recharge. NMFk is used to understand the influence of external forcing on shallow aquifer recharge in the Pajarito Plateau (Los Alamos, NM, USA). The results show how NMFk can be used to reduce the data dimension in a complex field dataset to three recharge signals that cause fluctuations within the field data. Here, the source signals are interpreted as rainfall, snowmelt, and a delayed aquifer response to the previous two signals. These results evidence how heterogeneous aquifers delimited by canyons incised into the Pajarito Plateau respond in similar ways to the source signals identified by NMFk. Furthermore, results show the importance of the local geology where faults act as sinks, and anthropogenic disturbances can facilitate infiltration amplifying the interpreted signal.

54 ENVIRONMENTAL SCIENCES↗

Fracture Networks Imaging in CO2 Injection Zones in IBDP Site: An Unsupervised Machine Learning Application with Multiple Datasets

Poster presented at the 17th International Conference on Greenhouse Gas Control Technologies GHGT-17 held in Calgary, Canada, October 20-24, 2024. This poster highlights the integration of unsupervised machine learning (ML) techniques as a transformative tool for advancing understanding of CO2 injection into reservoirs that could potentially contribute to optimizing injection strategies and reservoir management, ultimately bolstering the efficacy and sustainability of CO2 storage.

Kumar, Abhash↗

Fracture Networks Imaging in CO2 Injection Zones in IBDP Site: An Unsupervised Machine Learning Application with Multiple Datasets

This is the conference paper accompanying a poster presentation at the 17th International Conference on Greenhouse Gas Control Technologies GHGT-17 held in Calgary, Canada, October 20-24 , 2024. This work highlights the integration of unsupervised machine learning (ML) techniques as a transformative tool for advancing understanding of CO2 injection into reservoirs that could potentially contribute to optimizing injection strategies and reservoir management, ultimately bolstering the efficacy and sustainability of CO2 storage.

Kumar, Abhash↗

Predicting Dynamic-to-Static Correction Factor from Petrophysical Data and Chemostratigraphy using Unsupervised Machine Learning

Estimating static mechanical properties of stratigraphic layers is critical for optimizing subsurface engineering applications. To estimate dynamic-to-static correction factor F ds (static-to-dynamic Young’s modulus ratio) across the Caney shale interval in Oklahoma, USA, we integrated triaxial test measurements and petrophysical data, including well logs and X-ray fluorescence (XRF) using unsupervised machine learning (ML). We used a novel workflow that includes principal component analysis (PCA) to reduce data set dimensionality of well logs and XRF data sets—both separately and combined—creating three scenarios, and later applied inverse distance weighting (IDW) to derive F ds profiles for these scenarios. Furthermore, we applied K-means clustering on each scenario to predict depositional facies, and built a stiffness zonation profile through chemostratigraphic analysis of the terrigenous elements to validate the predicted F ds . The predicted F ds profile from each scenario using the PCA-IDW method was compared with the constant F ds approach from our previous study by calculating the root mean square error (RMSE). The combined data sets scenario yielded the lowest RMSE value of 0.113, while the RMSE values for the well logs and XRF scenarios were 0.131 and 0.129, respectively. In addition, the predicted F ds from the XRF scenario well-matched the stiffness zonation from the chemostratigraphic analysis that was built using the optimized K-means clustering of nine clusters for that scenario. These methods and findings offer a valuable tool for refining lithological classification and improving the F ds profile, potentially enhancing drilling and stimulation strategies for subsurface energy engineering applications.

clastic rock↗

Uncertainty Quantification in CO2 Trapping Mechanisms: A Case Study of PUNQ-S3 Reservoir Model Using Representative Geological Realizations and Unsupervised Machine Learning

Evaluating uncertainty in CO2 injection projections often requires numerous high-resolution geological realizations (GRs) which, although effective, are computationally demanding. This study proposes the use of representative geological realizations (RGRs) as an efficient approach to capture the uncertainty range of the full set while reducing computational costs. A predetermined number of RGRs is selected using an integrated unsupervised machine learning (UML) framework, which includes Euclidean distance measurement, multidimensional scaling (MDS), and a deterministic K-means (DK-means) clustering algorithm. In the context of the intricate 3D aquifer CO2 storage model, PUNQ-S3, these algorithms are utilized. The UML methodology selects five RGRs from a pool of 25 possibilities (20% of the total), taking into account the reservoir quality index (RQI) as a static parameter of the reservoir. To determine the credibility of these RGRs, their simulation results are scrutinized through the application of the Kolmogorov–Smirnov (KS) test, which analyzes the distribution of the output. In this assessment, 40 CO2 injection wells cover the entire reservoir alongside the full set. The end-point simulation results indicate that the CO2 structural, residual, and solubility trapping within the RGRs and full set follow the same distribution. Simulating five RGRs alongside the full set of 25 GRs over 200 years, involving 10 years of CO2 injection, reveals consistently similar trapping distribution patterns, with an average value of Dmax of 0.21 remaining lower than Dcritical (0.66). Using this methodology, computational expenses related to scenario testing and development planning for CO2 storage reservoirs in the presence of geological uncertainties can be substantially reduced.

Mahjour, Seyed Kourosh↗

Characterizing vertical upper ocean temperature structures in the European Arctic through unsupervised machine learning

In-situ observations of subsurface ocean temperatures are, in many regions, inconsistently distributed in time and space. These spatio-temporal inconsistencies in the observational network lead to difficulties in utilizing those observations effectively for ocean model evaluation or understanding larger-scale ocean characteristics. Model accuracy of subsurface ocean characteristics is especially important within regions that contain complex ocean structures. One such region is the European Arctic which not only contains several types of water masses with unique characteristics, but also wintertime sea ice coverage and complex bathymetry. This study presents an unsupervised neural networking technique that can be used in combination with traditional ocean model evaluation techniques to provide additional information on the accuracy of modeled vertical ocean temperature profiles. Self-organizing maps is an unsupervised machine learning technique that we apply to approximately twenty thousand Argo and CTD temperature profiles from 2012 to 2020 in the European Arctic to categorize the observed vertical ocean temperature structures in the top 150 m. The observed ocean profile categories, or neurons, defined by the self-organizing map show strong spatial and temporal dependencies. We then use the neuron weights, or the learned temperature profile structure of each neuron, to validate the spatial and temporal variability of modeled vertical temperature structures. This analysis gives us new insights about the model’s capabilities to reproduce specific vertical structures of the top-most ocean layer within different regions and seasons. Mapping modeled ocean temperature profiles onto the neuron-space of the observationally-defined self organized map highlights the potential of this method to advance our understanding of model deficiencies in that region.

54 ENVIRONMENTAL SCIENCES↗

Observing flow of He II with unsupervised machine learning

Abstract Time dependent observations of point-to-point correlations of the velocity vector field (structure functions) are necessary to model and understand fluid flow around complex objects. Using thermal gradients, we observed fluid flow by recording fluorescence of $${\text{He}}_{2}^{*}$$ He 2 ∗ excimers produced by neutron capture throughout a ~ cm 3 volume. Because the photon emitted by an excited excimer is unlikely to be recorded by the camera, the techniques of particle tracking (PTV) and particle imaging (PIV) velocimetry cannot be applied to extract information from the fluorescence of individual excimers. Therefore, we applied an unsupervised machine learning algorithm to identify light from ensembles of excimers (clusters) and then tracked the centroids of the clusters using a particle displacement determination algorithm developed for PTV.

75 CONDENSED MATTER PHYSICS, SUPERCONDUCTIVITY AND↗

Reclassification of ASFV into 7 Biotypes Using Unsupervised Machine Learning

In 2007, an outbreak of African swine fever (ASF), a deadly disease of domestic swine and wild boar caused by the African swine fever virus (ASFV), occurred in Georgia and has since spread globally. Historically, ASFV was classified into 25 different genotypes. However, a newly proposed system recategorized all ASFV isolates into 6 genotypes exclusively using the predicted protein sequences of p72. However, ASFV has a large genome that encodes between 150–200 genes, and classifications using a single gene are insufficient and misleading, as strains encoding an identical p72 often have significant mutations in other areas of the genome. We present here a new classification of ASFV based on comparisons performed considering the entire encoded proteome. A curated database consisting of the protein sequences predicted to be encoded by 220 reannotated ASFV genomes was analyzed for similarity between homologous protein sequences. Weights were applied to the protein identity matrices and averaged to generate a genome-genome identity matrix that was then analyzed by an unsupervised machine learning algorithm, DBSCAN, to separate the genomes into distinct clusters. We conclude that all available ASFV genomes can be classified into 7 distinct biotypes.

59 BASIC BIOLOGICAL SCIENCES↗

The influence of physical and algorithmic factors on simulated far-field waveforms and source–time functions of underground explosions using unsupervised machine learning

SUMMARY Characterizing explosion sources and differentiating between earthquake and underground explosions using distributed seismic networks becomes non-trivial when explosions are detonated in cavities or heterogeneous ground material. Moreover, there is little understanding of how changes in subsurface physical properties affect the far-field waveforms we record and use to infer information about the source. Simulations of underground explosions and the resultant ground motions can be a powerful tool to systematically explore how different subsurface properties affect far-field waveform features, but there are added variables that arise from how we choose to model the explosions that can confound interpretation. To assess how both subsurface properties and algorithmic choices affect the seismic wavefield and the estimated source functions, we ran a series of 2-D axisymmetric non-linear numerical explosion experiments and wave propagation simulations that explore a wide array of parameters. We then inverted the synthetic far-field waveform data using a linear inversion scheme to estimate source–time functions (STFs) for each simulation case. We applied principal component analysis (PCA), an unsupervised machine learning method, to both the far-field waveforms and STFs to identify the most important factors that control variance in the waveform data and differences between cases. For the far-field waveforms, the largest variance occurs in the shallower radial receiver channels in the 0–50 Hz frequency band. For the STFs, both peak amplitude and rise times across different frequencies contribute to the variance. We find that the ground equation of state (i.e. lithology and rheology) and the explosion emplacement conditions (i.e. tamped versus cavity) have the greatest effect on the variance of the far-field waveforms and STFs, with the ground yield strength and fracture pressure being secondary factors. Differences in the PCA results between the far-field waveforms and STFs could possibly be due to near-field non-linearities of the source that are not accounted for in the estimation of STFs and could be associated with yield strength, fracture pressure, cavity radius and cavity shape parameters. Other algorithmic parameters are found to be less important and cause less variance in both the far-field waveforms and STFs, meaning algorithmic choices in how we model explosions are less important, which is encouraging for the further use of explosion simulations to study how physical Earth properties affect seismic waveform features and estimated STFs.

58 GEOSCIENCES↗

Hunting for Polluted White Dwarfs and Other Treasures with Gaia XP Spectra and Unsupervised Machine Learning

White dwarfs (WDs) polluted by exoplanetary material provide the unprecedented opportunity to directly observe the interiors of exoplanets. However, spectroscopic surveys are often limited by brightness constraints, and WDs tend to be very faint, making detections of large populations of polluted WDs difficult. In this paper, we aim to increase considerably the number of WDs with multiple metals in their atmospheres. Using 96,134 WDs with Gaia DR3 BP/RP (XP) spectra, we constructed a 2D map using an unsupervised machine-learning technique called Uniform Manifold Approximation and Projection (UMAP) to organize the WDs into identifiable spectral regions. The polluted WDs are among the distinct spectral groups identified in our map. We have shown that this selection method could potentially increase the number of known WDs with five or more metal species in their atmospheres by an order of magnitude. Such systems are essential for characterizing exoplanet diversity and geology.

79 ASTRONOMY AND ASTROPHYSICS↗

4D-STEM Coupled with Unsupervised Machine Learning to Reveal at Large-Scale the Microstructural Evolution in Li- and Mn-Rich Cathodes

Li- and Mn-rich (LMR) layered oxides are known to exhibit a thin surface reconstruction layer, which grows during electrochemical cycling in a manner that depends on exposed crystallographic facets, cycling conditions, and electrolyte chemistry. Direct characterization of this layer has traditionally relied on high-resolution electron microscopy, which is inherently limited to small fields of view. Here, we employ four-dimensional scanning transmission electron microscopy (4D-STEM) combined with unsupervised machine-learning clustering to quantitatively map phase distributions over large areas and track their evolution in LMR cathodes during electrochemical aging. Our results show that the surface reconstruction layer consists predominantly of a rocksalt phase, whose thickness varies across different facets following activation cycling and becomes substantially thicker and more uniform during calendar aging. In contrast, a spinel-like phase is observed within the particle bulk. Large-area phase mapping and correlative high-resolution imaging reveal that this spinel-like phase preferentially nucleates at bulk crystallographic defects, including boundaries between 60°-rotated layered domains and associated mixed-phase regions, rather than exclusively at the particle surface. Our findings establish a mechanistic distinction between surface-driven rocksalt formation and bulk-defect-mediated spinel nucleation while demonstrating the unique capability of 4D-STEM to provide statistically robust, mesoscale insight into complex phase-evolution processes in LMR cathodes.

4D-STEM↗

Informed unsupervised machine learning analysis of dislocation microstructure from high-resolution differential aperture X-ray structural microscopy data

This study leverages high-resolution differential-aperture X-ray structural microscopy (DAXM) to probe the local dislocation structure in deformed 304L-stainless steel at small strain, by measuring the lattice rotation and deviatoric elastic strain with a sub-micron resolution. For a single grain in a polycrystalline specimen, the measured lattice rotation field over the measured volume exhibited a multimodal distribution while the deviatoric elastic strain showed a single-mode distribution. An unsupervised Cauchy mixture machine learning model was developed to resolve the multimodal distribution of the lattice rotation. By mapping the lattice rotation data associated with each Cauchy peak in the model back onto the measured volume, we identify contiguous regions of the crystal rotated near the average values corresponding to the peaks of the overall rotation distribution. These regions represent the grain subdivision in the microstructure. Finally, the dislocation density tensor was also computed and its norm was laid over the rotation field to detect the subgrain boundaries. This step provided a validation of the Cauchy mixture model for the analysis of the lattice rotation distribution. The current study highlights the integration of advanced X-ray microscopy techniques with data-driven analysis methods to uncover detailed microstructure scales in deformed crystals.

Machine learning; Lattice rotation; High-energy X-↗

Search for New Phenomena in Two-Body Invariant Mass Distributions Using Unsupervised Machine Learning for Anomaly Detection at s = 13 TeV with the ATLAS Detector

Searches for new resonances are performed using an unsupervised anomaly-detection technique. Events with at least one electron or muon are selected from 140 fb − 1 of p p collisions at s = 13 TeV recorded by ATLAS at the Large Hadron Collider. The approach involves training an autoencoder on data, and subsequently defining anomalous regions based on the reconstruction loss of the decoder. Studies focus on nine invariant mass spectra that contain pairs of objects consisting of one light jet or b jet and either one lepton ( e , μ ) , photon, or second light jet or b jet in the anomalous regions. No significant deviations from the background hypotheses are observed. Limits on contributions from generic Gaussian signals with various widths of the resonance mass are obtained for nine invariant masses in the anomalous regions. © 2024 CERN, for the ATLAS Collaboration 2024 CERN

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

Sensor Anomaly Detection for Nuclear Reactor Systems Utilizing Linear Regression and K-Means Unsupervised Machine Learning

Nuclear reactors and related systems are becoming increasingly complex due to advancing technologies in next-generation power reactors. This increased complexity necessitates enhanced automation and data management capabilities. To successfully realize autonomous systems, methods must be developed to handle vast volumes of data and effectively distinguish anomalous data from noise and expected data. While impressive models utilizing digital twins and similar approaches are under development, here we propose a simplified model for analyzing fundamental methods and techniques. Initially, we created a general dataset by using initial data from PCTRAN in order to represent ideal steady-state conditions. We then inserted anomalies based on prevalent sensor anomaly types (e.g., point anomalies, linear drift, and downward deviations), along with unusual anomalies such as exponential drift and upward deviations. To detect anomalies, we developed a program that employs data partitioning and linear regression to preprocess and filter the anomalous data. A K-Means machine learning (ML) method was then applied to separate and count the data within the anomalous partition. The results from all datasets—apart from exponential growth—demonstrated positive outcomes, with each returning multiple instances of greaterthan-95% accuracy. We conducted further investigations using Idaho National Laboratory’s RAVEN software to perform a sensitivity analysis on the input variables (R 2 Tolerance, Slope Tolerance, and Window Size) and found that the output variables (Accuracy and Time) were most sensitive to the Window Size. Despite the promising results published, further development is required to effectively apply these methods to nuclear systems. Nevertheless, the strengths of this approach are evident and hold promise for future applications in the field.

21 SPECIFIC NUCLEAR REACTORS AND ASSOCIATED PLANTS↗

Presentation: Sensor Anomaly Detection for Nuclear Reactor Systems Utilizing Linear Regression and K-Means Unsupervised Machine Learning: An overview of methods and results

This presentation is a culmination of work which has occurred over the course of a 10-week internship. Anomaly detection methods must be both robust enough to detect subtle anomalies yet not so sensitive to report false positives, which would result significant loss of revenue. Methods currently being developed for autonomous systems are often pursuing a Digital Twin method, which will look at the entire system and model it as a whole. This presentation, however, focuses less on direct application to an NPP, rather acting as a proof of concept for the methods developed. For the project, we look to develop methods to analyze steady-state data and report anomalies.

21 SPECIFIC NUCLEAR REACTORS AND ASSOCIATED PLANTS↗