Search NASA⌕ Search

SEARCH · Search NASA

Results for “Unsupervised machine learning”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

Unveiling and Mapping Polymorphs in Fluorite Y2TiO5 Using 4D-STEM and Unsupervised Machine Learning

Y2TiO5 belongs to the Ln2TiO5 (Ln = lanthanide or Y) family of ceramic materials and exhibits a range of desirable material properties such as radiation tolerance, frustrated magnetism, and large dielectric constant. However, understanding the complex crystal structure of Y2TiO5 remains elusive, given that Y2TiO5 can adopt multiple polymorphs such as cubic, orthorhombic, and hexagonal phases within the lattice. In this work, we report a detailed structural analysis of Y2TiO5 using four-dimensional scanning transmission electron microscopy coupled with unsupervised machine learning. The pyrochlore nanodomains, characterized by the ordered arrangement of yttrium cations on the A site of their A2BO5 structure, are present within the matrix of a predominantly fluorite-structured Y2TiO5 along with a third polymorph, the hexagonal phase. The pyrochlore phase is found to form 2 nm boundary regions around hexagonal phase stacking faults, highlighting the potential influence of the hexagonal phase on the occurrence and distribution of the pyrochlore phase. Lastly, we identify a unique pyrochlore phase with asymmetric arrangement of cation ordering along a single planar direction. Our findings provide invaluable insights into the possible mechanisms stabilizing pyrochlore nanodomains within the fluorite lattice of Y2TiO5.

36 MATERIALS SCIENCE↗

Exploring Continuous Seismic Data at an Industry Facility Using Unsupervised Machine Learning

Seismic data recorded at industrial sites contain valuable information on anthropogenic activities. With advances in machine learning and computing power, new opportunities have emerged to explore the seismic wavefield in these complex environments. We applied two unsupervised machine learning algorithms to analyze continuous seismic data collected from an industrial facility in Texas, United States. The Uniform Manifold Approximation and Projection for Dimension Reduction algorithm was used to reduce the dimensionality of the data and generate 2D embeddings. Then, the Hierarchical Density-Based Spatial Clustering of Applications with Noise method was employed to automatically group these embeddings into distinct signal clusters. Our analysis of over 1400 hr (around 59 days) of continuous seismic data revealed five and seven signal clusters at two separate stations. At both stations, we identified clusters associated with background noise and vehicle traffic, with the latter’s temporal patterns aligning closely with the facility’s work schedule. Furthermore, the algorithms detected signal clusters from unknown sources and underline the ability of unsupervised machine learning for uncovering previously unrecognized patterns. Our analysis demonstrates the effectiveness of unsupervised approaches in examining continuous seismic data without requiring prior knowledge or pre-existing labels.

58 GEOSCIENCES↗

Unsupervised machine learning discovery of structural units and transformation pathways from imaging data

We show that unsupervised machine learning can be used to learn chemical transformation pathways from observational Scanning Transmission Electron Microscopy (STEM) data. To enable this analysis, we assumed the existence of atoms, a discreteness of atomic classes, and the presence of an explicit relationship between the observed STEM contrast and the presence of atomic units. With only these postulates, we developed a machine learning method leveraging a rotationally invariant variational autoencoder (VAE) that can identify the existing molecular fragments observed within a material. The approach encodes the information contained in STEM image sequences using a small number of latent variables, allowing the exploration of chemical transformation pathways by tracing the evolution of atoms in the latent space of the system. The results suggest that atomically resolved STEM data can be used to derive fundamental physical and chemical mechanisms involved, by providing encodings of the observed structures that act as bottom-up equivalents of structural order parameters. The approach also demonstrates the potential of variational (i.e., Bayesian) methods in the physical sciences and will stimulate the development of more sophisticated ways to encode physical constraints in the encoder–decoder architectures and generative physical laws and causal relationships in the latent space of VAEs.

97 MATHEMATICS AND COMPUTING↗

Identifying recharge sources and their impacts on a North Central New Mexico shallow aquifer using unsupervised machine learning

In this article, shallow aquifers are important but highly variable resources in arid to semi-arid regions. Limited shallow aquifer volume results in high sensitivity to recharge fluctuations, which can impact the local fauna and flora, and transport of contaminants in the aquifer or vadose zone. Aquifer response to external forcing (e.g., precipitation) is usually solved by estimating aquifer parameters and running physics-based models to match known fluctuations of hydraulic head. However, this technique is time and computationally expensive. Furthermore, high aquifer complexity decreases precision in physics-based models. Alternatively supervised machine learning is used to predict aquifer dynamics. However, these techniques rely on input data and struggle to interpret aquifer response for missing sources (i.e., snowpack data). To counter these problems, we propose an unsupervised machine learning technique (NMFk) to estimate the impact of different sources on aquifer recharge. NMFk is used to understand the influence of external forcing on shallow aquifer recharge in the Pajarito Plateau (Los Alamos, NM, USA). The results show how NMFk can be used to reduce the data dimension in a complex field dataset to three recharge signals that cause fluctuations within the field data. Here, the source signals are interpreted as rainfall, snowmelt, and a delayed aquifer response to the previous two signals. These results evidence how heterogeneous aquifers delimited by canyons incised into the Pajarito Plateau respond in similar ways to the source signals identified by NMFk. Furthermore, results show the importance of the local geology where faults act as sinks, and anthropogenic disturbances can facilitate infiltration amplifying the interpreted signal.

54 ENVIRONMENTAL SCIENCES↗

AXEAP : a software package for X-ray emission data analysis using unsupervised machine learning

The Argonne X-ray Emission Analysis Package ( AXEAP ) has been developed to calibrate and process X-ray emission spectroscopy (XES) data collected with a two-dimensional (2D) position-sensitive detector. AXEAP is designed to convert a 2D XES image into an XES spectrum in real time using both calculations and unsupervised machine learning. AXEAP is capable of making this transformation at a rate similar to data collection, allowing real-time comparisons during data collection, reducing the amount of data stored from gigabyte-sized image files to kilobyte-sized text files. With a user-friendly interface, AXEAP includes data processing for non-resonant and resonant XES images from multiple edges and elements. AXEAP is written in MATLAB and can run on common operating systems, including Linux, Windows, and MacOS.

97 MATHEMATICS AND COMPUTING↗

Characterizing Drought Behavior in the Colorado River Basin Using Unsupervised Machine Learning

Drought is a pressing issue for the Colorado River Basin (CRB) due to the social and economic value of water resources in the region and the significant uncertainty of future drought under climate change. Here, we use climate simulations from various Earth System Models (ESMs) to force the Variable Infiltration Capacity hydrologic model and project multiple drought indicators for the sub-watersheds within the CRB. We apply an unsupervised machine learning (ML) based on Non-Negative Matrix Factorization using K-means clustering (NMFk) to synthesize the simulated historical, future, and change in drought indicators. The unsupervised ML approach can identify sub-watersheds where key changes to drought indicator behavior occur, including shifts in snowpack, snowmelt timing, precipitation, and evapotranspiration. While changes in future precipitation vary across ESMs, the results indicate that the Upper CRB will experience increasing evaporative demand and surface-water scarcity, with some locations experiencing a shift from a radiation-limited to a water-limited evaporation regime in the summer. Large shifts in peak runoff are observed in snowmelt-dominant sub-watersheds, with complete disappearance of the snowmelt signal for some sub-watersheds. The work demonstrates the utility of the NMFk algorithm to efficiently identify behavioral changes of drought indicators across space and time and to quickly analyze and interpret hydro climate model results.

54 ENVIRONMENTAL SCIENCES↗

Fracture Networks Imaging in CO2 Injection Zones in IBDP Site: An Unsupervised Machine Learning Application with Multiple Datasets

Poster presented at the 17th International Conference on Greenhouse Gas Control Technologies GHGT-17 held in Calgary, Canada, October 20-24, 2024. This poster highlights the integration of unsupervised machine learning (ML) techniques as a transformative tool for advancing understanding of CO2 injection into reservoirs that could potentially contribute to optimizing injection strategies and reservoir management, ultimately bolstering the efficacy and sustainability of CO2 storage.

Kumar, Abhash↗

Fracture Networks Imaging in CO2 Injection Zones in IBDP Site: An Unsupervised Machine Learning Application with Multiple Datasets

This is the conference paper accompanying a poster presentation at the 17th International Conference on Greenhouse Gas Control Technologies GHGT-17 held in Calgary, Canada, October 20-24 , 2024. This work highlights the integration of unsupervised machine learning (ML) techniques as a transformative tool for advancing understanding of CO2 injection into reservoirs that could potentially contribute to optimizing injection strategies and reservoir management, ultimately bolstering the efficacy and sustainability of CO2 storage.

Kumar, Abhash↗

Predicting Dynamic-to-Static Correction Factor from Petrophysical Data and Chemostratigraphy using Unsupervised Machine Learning

Estimating static mechanical properties of stratigraphic layers is critical for optimizing subsurface engineering applications. To estimate dynamic-to-static correction factor F ds (static-to-dynamic Young’s modulus ratio) across the Caney shale interval in Oklahoma, USA, we integrated triaxial test measurements and petrophysical data, including well logs and X-ray fluorescence (XRF) using unsupervised machine learning (ML). We used a novel workflow that includes principal component analysis (PCA) to reduce data set dimensionality of well logs and XRF data sets—both separately and combined—creating three scenarios, and later applied inverse distance weighting (IDW) to derive F ds profiles for these scenarios. Furthermore, we applied K-means clustering on each scenario to predict depositional facies, and built a stiffness zonation profile through chemostratigraphic analysis of the terrigenous elements to validate the predicted F ds . The predicted F ds profile from each scenario using the PCA-IDW method was compared with the constant F ds approach from our previous study by calculating the root mean square error (RMSE). The combined data sets scenario yielded the lowest RMSE value of 0.113, while the RMSE values for the well logs and XRF scenarios were 0.131 and 0.129, respectively. In addition, the predicted F ds from the XRF scenario well-matched the stiffness zonation from the chemostratigraphic analysis that was built using the optimized K-means clustering of nine clusters for that scenario. These methods and findings offer a valuable tool for refining lithological classification and improving the F ds profile, potentially enhancing drilling and stimulation strategies for subsurface energy engineering applications.

clastic rock↗

Uncertainty Quantification in CO2 Trapping Mechanisms: A Case Study of PUNQ-S3 Reservoir Model Using Representative Geological Realizations and Unsupervised Machine Learning

Evaluating uncertainty in CO2 injection projections often requires numerous high-resolution geological realizations (GRs) which, although effective, are computationally demanding. This study proposes the use of representative geological realizations (RGRs) as an efficient approach to capture the uncertainty range of the full set while reducing computational costs. A predetermined number of RGRs is selected using an integrated unsupervised machine learning (UML) framework, which includes Euclidean distance measurement, multidimensional scaling (MDS), and a deterministic K-means (DK-means) clustering algorithm. In the context of the intricate 3D aquifer CO2 storage model, PUNQ-S3, these algorithms are utilized. The UML methodology selects five RGRs from a pool of 25 possibilities (20% of the total), taking into account the reservoir quality index (RQI) as a static parameter of the reservoir. To determine the credibility of these RGRs, their simulation results are scrutinized through the application of the Kolmogorov–Smirnov (KS) test, which analyzes the distribution of the output. In this assessment, 40 CO2 injection wells cover the entire reservoir alongside the full set. The end-point simulation results indicate that the CO2 structural, residual, and solubility trapping within the RGRs and full set follow the same distribution. Simulating five RGRs alongside the full set of 25 GRs over 200 years, involving 10 years of CO2 injection, reveals consistently similar trapping distribution patterns, with an average value of Dmax of 0.21 remaining lower than Dcritical (0.66). Using this methodology, computational expenses related to scenario testing and development planning for CO2 storage reservoirs in the presence of geological uncertainties can be substantially reduced.

Mahjour, Seyed Kourosh↗

Harnessing interpretable and unsupervised machine learning to address big data from modern X-ray diffraction

The information content of crystalline materials becomes astronomical when collective electronic behavior and their fluctuations are taken into account. In the past decade, improvements in source brightness and detector technology at modern X-ray facilities have allowed a dramatically increased fraction of this information to be captured. Now, the primary challenge is to understand and discover scientific principles from big datasets when a comprehensive analysis is beyond human reach. We report the development of an unsupervised machine learning approach, X-ray diffraction (XRD) temperature clustering (X-TEC), that can automatically extract charge density wave order parameters and detect intraunit cell ordering and its fluctuations from a series of high-volume X-ray diffraction measurements taken at multiple temperatures. We benchmark X-TEC with diffraction data on a quasi-skutterudite family of materials, (Ca x Sr 1–x ) 3 Rh 4 Sn 13 , where a quantum critical point is observed as a function of Ca concentration. We apply X-TEC to XRD data on the pyrochlore metal, Cd 2 Re 2 O 7 , to investigate its two much-debated structural phase transitions and uncover the Goldstone mode accompanying them. We demonstrate how unprecedented atomic-scale knowledge can be gained when human researchers connect the X-TEC results to physical principles. Specifically, we extract from the X-TEC–revealed selection rules that the Cd and Re displacements are approximately equal in amplitude but out of phase. This discovery reveals a previously unknown involvement of 5d 2 Re, supporting the idea of an electronic origin to the structural order. Our approach can radically transform XRD experiments by allowing in operando data analysis and enabling researchers to refine experiments by discovering interesting regions of phase space on the fly.

36 MATERIALS SCIENCE↗

Characterizing vertical upper ocean temperature structures in the European Arctic through unsupervised machine learning

In-situ observations of subsurface ocean temperatures are, in many regions, inconsistently distributed in time and space. These spatio-temporal inconsistencies in the observational network lead to difficulties in utilizing those observations effectively for ocean model evaluation or understanding larger-scale ocean characteristics. Model accuracy of subsurface ocean characteristics is especially important within regions that contain complex ocean structures. One such region is the European Arctic which not only contains several types of water masses with unique characteristics, but also wintertime sea ice coverage and complex bathymetry. This study presents an unsupervised neural networking technique that can be used in combination with traditional ocean model evaluation techniques to provide additional information on the accuracy of modeled vertical ocean temperature profiles. Self-organizing maps is an unsupervised machine learning technique that we apply to approximately twenty thousand Argo and CTD temperature profiles from 2012 to 2020 in the European Arctic to categorize the observed vertical ocean temperature structures in the top 150 m. The observed ocean profile categories, or neurons, defined by the self-organizing map show strong spatial and temporal dependencies. We then use the neuron weights, or the learned temperature profile structure of each neuron, to validate the spatial and temporal variability of modeled vertical temperature structures. This analysis gives us new insights about the model’s capabilities to reproduce specific vertical structures of the top-most ocean layer within different regions and seasons. Mapping modeled ocean temperature profiles onto the neuron-space of the observationally-defined self organized map highlights the potential of this method to advance our understanding of model deficiencies in that region.

54 ENVIRONMENTAL SCIENCES↗

Observing flow of He II with unsupervised machine learning

Abstract Time dependent observations of point-to-point correlations of the velocity vector field (structure functions) are necessary to model and understand fluid flow around complex objects. Using thermal gradients, we observed fluid flow by recording fluorescence of $${\text{He}}_{2}^{*}$$ He 2 ∗ excimers produced by neutron capture throughout a ~ cm 3 volume. Because the photon emitted by an excited excimer is unlikely to be recorded by the camera, the techniques of particle tracking (PTV) and particle imaging (PIV) velocimetry cannot be applied to extract information from the fluorescence of individual excimers. Therefore, we applied an unsupervised machine learning algorithm to identify light from ensembles of excimers (clusters) and then tracked the centroids of the clusters using a particle displacement determination algorithm developed for PTV.

75 CONDENSED MATTER PHYSICS, SUPERCONDUCTIVITY AND↗

Reclassification of ASFV into 7 Biotypes Using Unsupervised Machine Learning

In 2007, an outbreak of African swine fever (ASF), a deadly disease of domestic swine and wild boar caused by the African swine fever virus (ASFV), occurred in Georgia and has since spread globally. Historically, ASFV was classified into 25 different genotypes. However, a newly proposed system recategorized all ASFV isolates into 6 genotypes exclusively using the predicted protein sequences of p72. However, ASFV has a large genome that encodes between 150–200 genes, and classifications using a single gene are insufficient and misleading, as strains encoding an identical p72 often have significant mutations in other areas of the genome. We present here a new classification of ASFV based on comparisons performed considering the entire encoded proteome. A curated database consisting of the protein sequences predicted to be encoded by 220 reannotated ASFV genomes was analyzed for similarity between homologous protein sequences. Weights were applied to the protein identity matrices and averaged to generate a genome-genome identity matrix that was then analyzed by an unsupervised machine learning algorithm, DBSCAN, to separate the genomes into distinct clusters. We conclude that all available ASFV genomes can be classified into 7 distinct biotypes.

59 BASIC BIOLOGICAL SCIENCES↗

The influence of physical and algorithmic factors on simulated far-field waveforms and source–time functions of underground explosions using unsupervised machine learning

SUMMARY Characterizing explosion sources and differentiating between earthquake and underground explosions using distributed seismic networks becomes non-trivial when explosions are detonated in cavities or heterogeneous ground material. Moreover, there is little understanding of how changes in subsurface physical properties affect the far-field waveforms we record and use to infer information about the source. Simulations of underground explosions and the resultant ground motions can be a powerful tool to systematically explore how different subsurface properties affect far-field waveform features, but there are added variables that arise from how we choose to model the explosions that can confound interpretation. To assess how both subsurface properties and algorithmic choices affect the seismic wavefield and the estimated source functions, we ran a series of 2-D axisymmetric non-linear numerical explosion experiments and wave propagation simulations that explore a wide array of parameters. We then inverted the synthetic far-field waveform data using a linear inversion scheme to estimate source–time functions (STFs) for each simulation case. We applied principal component analysis (PCA), an unsupervised machine learning method, to both the far-field waveforms and STFs to identify the most important factors that control variance in the waveform data and differences between cases. For the far-field waveforms, the largest variance occurs in the shallower radial receiver channels in the 0–50 Hz frequency band. For the STFs, both peak amplitude and rise times across different frequencies contribute to the variance. We find that the ground equation of state (i.e. lithology and rheology) and the explosion emplacement conditions (i.e. tamped versus cavity) have the greatest effect on the variance of the far-field waveforms and STFs, with the ground yield strength and fracture pressure being secondary factors. Differences in the PCA results between the far-field waveforms and STFs could possibly be due to near-field non-linearities of the source that are not accounted for in the estimation of STFs and could be associated with yield strength, fracture pressure, cavity radius and cavity shape parameters. Other algorithmic parameters are found to be less important and cause less variance in both the far-field waveforms and STFs, meaning algorithmic choices in how we model explosions are less important, which is encouraging for the further use of explosion simulations to study how physical Earth properties affect seismic waveform features and estimated STFs.

58 GEOSCIENCES↗

Supervised and unsupervised machine learning of structural phases of polymers adsorbed to nanowires

Here, we identify configurational phases and structural transitions in a polymer nanotube composite by means of machine learning. We employ various unsupervised dimensionality reduction methods, conventional neural networks, as well as the confusion method, an unsupervised neural-network-based approach. We find neural networks are able to reliably recognize all configurational phases that have been found previously in experiment and simulation. Furthermore, we locate the boundaries between configurational phases in a way that removes human intuition or bias. This could be done before only by relying on preconceived, ad hoc order parameters.

71 CLASSICAL AND QUANTUM MECHANICS, GENERAL PHYSIC↗

Hunting for Polluted White Dwarfs and Other Treasures with Gaia XP Spectra and Unsupervised Machine Learning

White dwarfs (WDs) polluted by exoplanetary material provide the unprecedented opportunity to directly observe the interiors of exoplanets. However, spectroscopic surveys are often limited by brightness constraints, and WDs tend to be very faint, making detections of large populations of polluted WDs difficult. In this paper, we aim to increase considerably the number of WDs with multiple metals in their atmospheres. Using 96,134 WDs with Gaia DR3 BP/RP (XP) spectra, we constructed a 2D map using an unsupervised machine-learning technique called Uniform Manifold Approximation and Projection (UMAP) to organize the WDs into identifiable spectral regions. The polluted WDs are among the distinct spectral groups identified in our map. We have shown that this selection method could potentially increase the number of known WDs with five or more metal species in their atmospheres by an order of magnitude. Such systems are essential for characterizing exoplanet diversity and geology.

79 ASTRONOMY AND ASTROPHYSICS↗

4D-STEM Coupled with Unsupervised Machine Learning to Reveal at Large-Scale the Microstructural Evolution in Li- and Mn-Rich Cathodes

Li- and Mn-rich (LMR) layered oxides are known to exhibit a thin surface reconstruction layer, which grows during electrochemical cycling in a manner that depends on exposed crystallographic facets, cycling conditions, and electrolyte chemistry. Direct characterization of this layer has traditionally relied on high-resolution electron microscopy, which is inherently limited to small fields of view. Here, we employ four-dimensional scanning transmission electron microscopy (4D-STEM) combined with unsupervised machine-learning clustering to quantitatively map phase distributions over large areas and track their evolution in LMR cathodes during electrochemical aging. Our results show that the surface reconstruction layer consists predominantly of a rocksalt phase, whose thickness varies across different facets following activation cycling and becomes substantially thicker and more uniform during calendar aging. In contrast, a spinel-like phase is observed within the particle bulk. Large-area phase mapping and correlative high-resolution imaging reveal that this spinel-like phase preferentially nucleates at bulk crystallographic defects, including boundaries between 60°-rotated layered domains and associated mixed-phase regions, rather than exclusively at the particle surface. Our findings establish a mechanistic distinction between surface-driven rocksalt formation and bulk-defect-mediated spinel nucleation while demonstrating the unique capability of 4D-STEM to provide statistically robust, mesoscale insight into complex phase-evolution processes in LMR cathodes.

4D-STEM↗