Search NASA⌕ Search

SEARCH · Search NASA

Results for “unsupervised”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

Unveiling and Mapping Polymorphs in Fluorite Y2TiO5 Using 4D-STEM and Unsupervised Machine Learning

Y2TiO5 belongs to the Ln2TiO5 (Ln = lanthanide or Y) family of ceramic materials and exhibits a range of desirable material properties such as radiation tolerance, frustrated magnetism, and large dielectric constant. However, understanding the complex crystal structure of Y2TiO5 remains elusive, given that Y2TiO5 can adopt multiple polymorphs such as cubic, orthorhombic, and hexagonal phases within the lattice. In this work, we report a detailed structural analysis of Y2TiO5 using four-dimensional scanning transmission electron microscopy coupled with unsupervised machine learning. The pyrochlore nanodomains, characterized by the ordered arrangement of yttrium cations on the A site of their A2BO5 structure, are present within the matrix of a predominantly fluorite-structured Y2TiO5 along with a third polymorph, the hexagonal phase. The pyrochlore phase is found to form 2 nm boundary regions around hexagonal phase stacking faults, highlighting the potential influence of the hexagonal phase on the occurrence and distribution of the pyrochlore phase. Lastly, we identify a unique pyrochlore phase with asymmetric arrangement of cation ordering along a single planar direction. Our findings provide invaluable insights into the possible mechanisms stabilizing pyrochlore nanodomains within the fluorite lattice of Y2TiO5.

36 MATERIALS SCIENCE↗

Application of unsupervised deep learning to image segmentation and in-situ contact angle measurements in a CO 2 -water-rock system

Rock surface wettability is a critical property that regulates multiphase flows in porous media, which can be quantified using the surface contact angle (CA). X-ray micro-computed tomography (μCT) provides an effective approach to in-situ measurements of surface CAs. However, the CA measurement accuracy depends significantly on the quality of CT image segmentation, which is the clustering of CT pixels into separate phases. Inspired by this, we developed a deep learning (DL)-based CA measurement workflow. Motivated by the recent tremendous progress in unsupervised learning techniques and aiming to avoid expensive manual data annotations, an unsupervised DL pipeline for CT image segmentation was proposed and implemented, which includes unsupervised model training and post-processing. The unsupervised model training was driven by a novel loss function constrained with feature similarity and spatial continuity and implemented by iterative forward and backward paths; the former clustered the pixel-wise feature vectors extracted by convolution neural networks, whereas the latter updated the parameters using gradient descent. An over-segmentation strategy was adopted for model training. The post-processing steps based on agglomerative hierarchical clustering (AHC) were implemented to further merge the over-segmented model output to the desired cluster number, which is intended to improve the efficiency of image segmentation. The developed unsupervised DL pipeline was compared with other commonly-used image segmentation methods using pixel-wise and physics-based evaluation metrics on a synthetic raw-image dataset, which had a known ground truth. The unsupervised DL pipeline showed the best performance. Next, the segmented images were input to an automatic CA measurement tool, and the results were validated by comparisons with manual measurements. The CA values from the manual and automatic measurements showed similar distributions and statistical properties. The automatic measurement demonstrated a wider spectrum because of the much larger number of measurement data points. The primary novelty of the unsupervised DL pipeline developed in this study lies in the novel loss function and the over-segmentation strategy associated with AHC post-processing. Finally, the workflow has been proven an efficient tool for pore-scale wettability characterization, which has a wide range of applications in fundamental studies of multiphase flows in natural porous media, which have critical implications to geological carbon sequestration, hydrocarbon energy recovery, and contaminant transport in groundwater.

42 ENGINEERING↗

General-Purpose Unsupervised Cyber Anomaly Detection via Non-Negative Tensor Factorization

Distinguishing malicious anomalous activities from unusual but benign activities is a fundamental challenge for cyber defenders. Prior studies have shown that statistical user behavior analysis yields accurate detections by learning behavior profiles from observed user activity. These unsupervised models are able to generalize to unseen types of attacks by detecting deviations from normal behavior, without knowledge of specific attack signatures. However, approaches proposed to date based on probabilistic matrix factorization are limited by the information conveyed in a two-dimensional space. Non-negative tensor factorization, on the other hand, is a powerful unsupervised machine learning method that naturally models multi-dimensional data, capturing complex and multi-faceted details of behavior profiles. Herein, our new unsupervised statistical anomaly detection methodology matches or surpasses state-of-the-art supervised learning baselines across several challenging and diverse cyber application areas, including detection of compromised user credentials, botnets, spam e-mails, and fraudulent credit card transactions.

97 MATHEMATICS AND COMPUTING↗

Seeking regularity from irregularity: unveiling the synthesis–nanomorphology relationships of heterogeneous nanomaterials using unsupervised machine learning

Nanoscale morphology of functional materials determines their chemical and physical properties. However, despite increasing use of transmission electron microscopy (TEM) to directly image nanomorphology, it remains challenging to quantify the information embedded in TEM data sets, and to use nanomorphology to link synthesis and processing conditions to properties. We develop an automated, descriptor-free analysis workflow for TEM data that utilizes convolutional neural networks and unsupervised learning to quantify and classify nanomorphology, and thereby reveal synthesis–nanomorphology relationships in three different systems. While TEM records nanomorphology readily in two-dimensional (2D) images or three-dimensional (3D) tomograms, we advance the analysis of these images by identifying and applying a universal shape fingerprint function to characterize nanomorphology. After dimensionality reduction through principal component analysis, this function then serves as the input for morphology grouping through unsupervised learning. We demonstrate the wide applicability of our workflow to both 2D and 3D TEM data sets, and to both inorganic and organic nanomaterials, including tetrahedral gold nanoparticles mixed with irregularly shaped impurities, hybrid polymer-patched gold nanoprisms, and polyamide membranes with irregular and heterogeneous 3D crumple structures. In each of these systems, unsupervised nanomorphology grouping identifies both the diversity and the similarity of the nanomaterial across different synthesis conditions, revealing how synthetic parameters guide nanomorphology development. Our work opens possibilities for enhancing synthesis of nanomaterials through artificial intelligence and for understanding and controlling complex nanomorphology, both for 2D systems and in the far less explored case of 3D structures, such as those with embedded voids or hidden interfaces.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Anatomy of Continuous Mars SEIS and Pressure Data from Unsupervised Learning

The seismic noise recorded by the Interior Exploration using Seismic Investigations, Geodesy, and Heat Transport (InSight) seismometer (Seismic Experiment for Interior Structure [SEIS]) has a strong daily quasi-periodicity and numerous transient microevents, associated mostly with an active Martian environment with wind bursts, pressure drops, in addition to thermally induced lander and instrument cracks. That noise is far from the Earth’s microseismic noise. Quantifying the importance of nonstochasticity and identifying these microevents is mandatory for improving continuous data quality and noise analysis techniques, including autocorrelation. Cataloging these events has so far been made with specific algorithms and operator’s visual inspection. We investigate here the continuous data with an unsupervised deep-learning approach built on a deep scattering network. This leads to the successful detection and clustering of these microevents as well as better determination of daily cycles associated with changes in the intensity and color of the background noise. We first provide a description of our approach, and then present the learned clusters followed by a study of their origin and associated physical phenomena. We show that the clustering is robust over several Martian days, showing distinct types of glitches that repeat at a rate of several tens per sol with stable time differences. We show that the clustering and detection efficiency for pressure drops and glitches is comparable to or better than manual or targeted detection techniques proposed to date, noticeably with an unsupervised approach. Finally, here we discuss the origin of other clusters found, especially glitch sequences with stable time offsets that might generate artifacts in autocorrelation analyses. We conclude with presenting the potential of unsupervised learning for long-term space mission operations, in particular, for geophysical and environmental observatories.

58 GEOSCIENCES↗

Exploring Continuous Seismic Data at an Industry Facility Using Unsupervised Machine Learning

Seismic data recorded at industrial sites contain valuable information on anthropogenic activities. With advances in machine learning and computing power, new opportunities have emerged to explore the seismic wavefield in these complex environments. We applied two unsupervised machine learning algorithms to analyze continuous seismic data collected from an industrial facility in Texas, United States. The Uniform Manifold Approximation and Projection for Dimension Reduction algorithm was used to reduce the dimensionality of the data and generate 2D embeddings. Then, the Hierarchical Density-Based Spatial Clustering of Applications with Noise method was employed to automatically group these embeddings into distinct signal clusters. Our analysis of over 1400 hr (around 59 days) of continuous seismic data revealed five and seven signal clusters at two separate stations. At both stations, we identified clusters associated with background noise and vehicle traffic, with the latter’s temporal patterns aligning closely with the facility’s work schedule. Furthermore, the algorithms detected signal clusters from unknown sources and underline the ability of unsupervised machine learning for uncovering previously unrecognized patterns. Our analysis demonstrates the effectiveness of unsupervised approaches in examining continuous seismic data without requiring prior knowledge or pre-existing labels.

58 GEOSCIENCES↗

Unsupervised Machine Learning for Exploratory Data Analysis of Exoplanet Transmission Spectra

Abstract Transit spectroscopy is a powerful tool for decoding the chemical compositions of the atmospheres of extrasolar planets. In this paper, we focus on unsupervised techniques for analyzing spectral data from transiting exoplanets. After cleaning and validating the data, we demonstrate methods for: (i) initial exploratory data analysis, based on summary statistics (estimates of location and variability); (ii) exploring and quantifying the existing correlations in the data; (iii) preprocessing and linearly transforming the data to its principal components; (iv) dimensionality reduction and manifold learning; (v) clustering and anomaly detection; and (vi) visualization and interpretation of the data. To illustrate the proposed unsupervised methodology, we use a well-known public benchmark data set of synthetic transit spectra. We show that there is a high degree of correlation in the spectral data, which calls for appropriate low-dimensional representations. We explore a number of different techniques for such dimensionality reduction and identify several suitable options in terms of summary statistics, principal components, etc. We uncover interesting structures in the principal component basis, namely well-defined branches corresponding to different chemical regimes of the underlying atmospheres. We demonstrate that those branches can be successfully recovered with a K-means clustering algorithm in a fully unsupervised fashion. We advocate for lower-dimensional representations of the spectroscopic data in terms of the main principal components, in order to reveal the existing structure in the data and quickly characterize the chemical class of a planet.

Matchev, Konstantin T. (ORCID:0000000341829096)↗

Unsupervised learning from three-component accelerometer data to monitor the spatiotemporal evolution of meso-scale hydraulic fractures

Enhanced geothermal systems can provide a substantial share of the global energy demand. There exist several hurdles in the engineering implementations of such geothermal systems. One such hurdle is the accurate monitoring of the fracture networks created in subsurface through hydraulic stimulation of these systems. Micro seismicity associated with the stimulation is the primary means to locate the event hypocenters for estimating the stimulated rock volume. Existing methods for location the hypocenters are restricted to only the highest amplitude impulsive signals that are simultaneously detected on several sensors. Consequently, a large portion (usually ~99%) of the measurements are left unused. In this paper, an unsupervised manifold-approximation followed by clustering of 3-component accelerometer data is used to analyze the seismicity recorded on a monitoring well. With this method, a larger portion of the measured signal is used for the monitoring of the hydraulic fracture network. We analyze the EGS Collab experiment 1 microseismic data, recorded at the Sanford Underground Research Facility, South Dakota. Using the data from a single three-component accelerometer, the polarization features viz. Azimuth, incidence, rectilinearity, and planarity are used as inputs for the unsupervised manifold approximation followed by clustering. Our study shows that density-based clusters in the projected 3D space correspond to distinct types of hydraulically fractured zones around the injection point. Finally, we show that the temporal evolution of these clusters can be used to track fracture creation and propagation.

58 GEOSCIENCES↗

A survey of unsupervised learning methods for high-dimensional uncertainty quantification in black-box-type problems

Constructing surrogate models for uncertainty quantification (UQ) on complex partial differential equations (PDEs) having inherently high-dimensional O(10 n ), n ≥ 2, stochastic inputs (e.g., forcing terms, boundary conditions, initial conditions) poses tremendous challenges. The “curse of dimensionality” can be addressed with suitable unsupervised learning techniques used as a pre-processing tool to encode inputs onto lower-dimensional subspaces while retaining its structural information and meaningful properties. In this work, we review and investigate thirteen dimension reduction methods including linear and nonlinear, spectral, blind source separation, convex and non-convex methods and utilize the resulting embeddings to construct a mapping to quantities of interest via polynomial chaos expansions (PCE). Here, we refer to the general proposed approach as manifold PCE (m-PCE), where manifold corresponds to the latent space resulting from any of the studied dimension reduction methods. To investigate the capabilities and limitations of these methods we conduct numerical tests for three physics-based systems (treated as black-boxes) having high-dimensional stochastic inputs of varying complexity modeled as both Gaussian and non-Gaussian random fields to investigate the effect of the intrinsic dimensionality of input data. We demonstrate both the advantages and limitations of the unsupervised learning methods and we conclude that a suitable m-PCE model provides a cost-effective approach compared to alternative algorithms proposed in the literature, including recently proposed expensive deep neural network-based surrogates and can be readily applied for high-dimensional UQ in stochastic PDEs.

42 ENGINEERING↗

Characterizing vertical upper ocean temperature structures in the European Arctic through unsupervised machine learning

In-situ observations of subsurface ocean temperatures are, in many regions, inconsistently distributed in time and space. These spatio-temporal inconsistencies in the observational network lead to difficulties in utilizing those observations effectively for ocean model evaluation or understanding larger-scale ocean characteristics. Model accuracy of subsurface ocean characteristics is especially important within regions that contain complex ocean structures. One such region is the European Arctic which not only contains several types of water masses with unique characteristics, but also wintertime sea ice coverage and complex bathymetry. This study presents an unsupervised neural networking technique that can be used in combination with traditional ocean model evaluation techniques to provide additional information on the accuracy of modeled vertical ocean temperature profiles. Self-organizing maps is an unsupervised machine learning technique that we apply to approximately twenty thousand Argo and CTD temperature profiles from 2012 to 2020 in the European Arctic to categorize the observed vertical ocean temperature structures in the top 150 m. The observed ocean profile categories, or neurons, defined by the self-organizing map show strong spatial and temporal dependencies. We then use the neuron weights, or the learned temperature profile structure of each neuron, to validate the spatial and temporal variability of modeled vertical temperature structures. This analysis gives us new insights about the model’s capabilities to reproduce specific vertical structures of the top-most ocean layer within different regions and seasons. Mapping modeled ocean temperature profiles onto the neuron-space of the observationally-defined self organized map highlights the potential of this method to advance our understanding of model deficiencies in that region.

54 ENVIRONMENTAL SCIENCES↗

Characterizing Drought Behavior in the Colorado River Basin Using Unsupervised Machine Learning

Drought is a pressing issue for the Colorado River Basin (CRB) due to the social and economic value of water resources in the region and the significant uncertainty of future drought under climate change. Here, we use climate simulations from various Earth System Models (ESMs) to force the Variable Infiltration Capacity hydrologic model and project multiple drought indicators for the sub-watersheds within the CRB. We apply an unsupervised machine learning (ML) based on Non-Negative Matrix Factorization using K-means clustering (NMFk) to synthesize the simulated historical, future, and change in drought indicators. The unsupervised ML approach can identify sub-watersheds where key changes to drought indicator behavior occur, including shifts in snowpack, snowmelt timing, precipitation, and evapotranspiration. While changes in future precipitation vary across ESMs, the results indicate that the Upper CRB will experience increasing evaporative demand and surface-water scarcity, with some locations experiencing a shift from a radiation-limited to a water-limited evaporation regime in the summer. Large shifts in peak runoff are observed in snowmelt-dominant sub-watersheds, with complete disappearance of the snowmelt signal for some sub-watersheds. The work demonstrates the utility of the NMFk algorithm to efficiently identify behavioral changes of drought indicators across space and time and to quickly analyze and interpret hydro climate model results.

54 ENVIRONMENTAL SCIENCES↗

Internship Final Report on the unsupervised learning sensor fusion (ULSF) approach

This paper describes a summer internship project undertaken at Sandia National Labs (SNL), both current status and future work. The project was to explore various machine learning approaches for use on turbulent flow data. Specifically, unsupervised classification of turbulent flow data was explored. First, the usage of models in this field is discussed, and several issues in the common usage of the models are identified. Solutions to these issues are then proposed, in the form of a Bayesian filtering approach which probabilistically incorporates multiple sources of data to improve confidence in a result. Several types of sensors are suggested for this method, the incorporation of which range from semi-supervised learning approaches to fully unsupervised. These approaches are then tested on several turbulent flow cases.

97 MATHEMATICS AND COMPUTING↗

Missing Wedge Completion via Unsupervised Learning with Coordinate Networks

Cryogenic electron tomography (cryoET) is a powerful tool in structural biology, enabling detailed 3D imaging of biological specimens at a resolution of nanometers. Despite its potential, cryoET faces challenges such as the missing wedge problem, which limits reconstruction quality due to incomplete data collection angles. Recently, supervised deep learning methods leveraging convolutional neural networks (CNNs) have considerably addressed this issue; however, their pretraining requirements render them susceptible to inaccuracies and artifacts, particularly when representative training data is scarce. To overcome these limitations, we introduce a proof-of-concept unsupervised learning approach using coordinate networks (CNs) that optimizes network weights directly against input projections. This eliminates the need for pretraining, reducing reconstruction runtime by 3–20× compared to supervised methods. Our in silico results show improved shape completion and reduction of missing wedge artifacts, assessed through several voxel-based image quality metrics in real space and a novel directional Fourier Shell Correlation (FSC) metric. Our study illuminates benefits and considerations of both supervised and unsupervised approaches, guiding the development of improved reconstruction strategies.

42 ENGINEERING↗

Reward Driven Workflows for Unsupervised Explainable Analysis of Phases and Ferroic Variants From Atomically Resolved Imaging Data

Rapid progress in aberration corrected electron microscopy necessitates development of robust methods for the identification of phases, ferroic variants, and other pertinent aspects of materials structure from imaging data. While unsupervised methods for clustering and classification are widely used for these tasks, their performance can be sensitive to hyperparameter selection in the analysis workflow. In this study, the effects of descriptors and hyperparameters are explored on the capability of unsupervised ML methods to distill local structural information, exemplified by the discovery of polarization and lattice distortion in Sm − dopped BiFeO 3 (BFO) thin films. It is demonstrated that a reward-driven approach can be used to optimize these key hyperparameters across the full workflow, where rewards are designed to reflect domain wall continuity and straightness, ensuring that the analysis aligns with the material's physical behavior. This approach allows the discovery of local descriptors that are best aligned with the specific physical behavior, providing insight into the fundamental physics of materials. The reward driven workflow is further extended to disentangle structural factors of variation via an optimized variational autoencoder (VAE). Lastly, the importance of well-defined rewards is explored as a quantifiable measure of the success of the workflow.

Barakati, Kamyar [University of Tennessee, Knoxvil↗

Evaluating lightweight unsupervised online IDS for masquerade attacks in CAN

Vehicular controller area networks (CANs) are susceptible to masquerade attacks by malicious adversaries. In masquerade attacks, adversaries silence a targeted ID and then send malicious frames with forged content at the expected timing of benign frames. As masquerade attacks could seriously harm vehicle functionality and are the stealthiest attacks to detect in CAN, recent work has devoted attention to compare frameworks for detecting masquerade attacks in CAN. However, most existing works report offline evaluations using CAN logs already collected using simulations that do not comply with the domain’s real-time constraints. Here we contribute to advance the state of the art by presenting a comparative evaluation of four different non-deep learning (DL)-based unsupervised online intrusion detection systems (IDS) for masquerade attacks in CAN. Our approach differs from existing comparative evaluations in that we analyze the effect of controlling streaming data conditions in a sliding window setting. In doing so, we use realistic masquerade attacks being replayed from the ROAD dataset. We show that although evaluated IDS are not effective at detecting every attack type, the method that relies on detecting changes in the hierarchical structure of clusters of time series produces the best results at the expense of higher computational overhead. We discuss limitations, open challenges, and how the evaluated methods can be used for practical unsupervised online CAN IDS for masquerade attacks.

Anomaly detection↗

Unsupervised domain adaptation for radioisotope identification in gamma spectroscopy

Training machine learning models for radioisotope identification using gamma spectroscopy remains an elusive challenge for many practical applications, largely stemming from the difficulty of acquiring and labeling large, diverse experimental datasets. Simulations can mitigate this challenge, but the accuracy of models trained on simulated data can deteriorate substantially when deployed to an out-of-distribution operational environment. In this study, we demonstrate that unsupervised domain adaptation (UDA) can improve the ability of a model trained on synthetic data to generalize to a new testing domain, provided unlabeled data from the target domain are available. Conventional supervised techniques are unable to utilize this data because the absence of isotope labels precludes defining a supervised classification loss. Instead, we first pretrain a spectral classifier using labeled synthetic data and subsequently leverage unlabeled target data to align the learned feature representations between the source and target domains. We compare a range of different UDA techniques, finding that minimizing the maximum mean discrepancy (MMD) between source and target feature vectors yields the most consistent improvement to testing scores. For instance, using a custom transformer-based neural network, we achieved a testing accuracy of $0.904 \pm 0.022$ on an experimental LaBr test set after performing unsupervised feature alignment via MMD minimization, compared to $0.754 \pm 0.014$ before alignment. Overall, our results highlight the potential of using UDA to adapt a radioisotope classifier trained on synthetic data for real-world deployment.

Lalor, Peter W.↗

Monitoring Fracture Hydromechanical Evolution in the Lab and Field Using Unsupervised Metric Learning

Fractures evolve in time through thermal‐hydraulic‐mechanical‐chemical (THMC) processes that alter their long‐range hydraulic transport properties and modify subsurface behavior and activities. The location of subsurface fractures makes it necessary to use remote sensing techniques such as passive or active seismic monitoring for fracture characterization. In this paper, we develop a machine learning approach to monitor the evolution of fracture properties using passive seismic sources in a laboratory setting and using active seismic monitoring from the Sanford Underground Research Facility in Lead, South Dakota, at a depth of 1.25 km in amphibolite rock during stimulation of natural fractures as well as during induced fracturing. The unsupervised metric learning technique applies tandem neural networks (twin (Siamese) or triplet) with contrastive loss and adaptive margins to track slowly varying systems for which class or similarity labels are not available. The approach adopts locality‐sensitive hashing to divide time‐ordered contiguous data into an arbitrary number of pseudo‐classes. Contrastive‐loss training with many hash bins generates an evolving latent‐space trajectory. This approach enables unsupervised metric learning for seismic data stacks under the condition of contiguous state sampling and slowly varying fracture properties. The displacement discontinuity theory provides a mechanistic foundation for the fracture‐dependent trajectories that are related to relaxation of fractures with time‐dependent specific stiffness responding to changes in stress or fluid saturation.

02 PETROLEUM↗