Search NASA⌕ Search

SEARCH · Search NASA

Results for “unsupervised”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

Unveiling and Mapping Polymorphs in Fluorite Y2TiO5 Using 4D-STEM and Unsupervised Machine Learning

Y2TiO5 belongs to the Ln2TiO5 (Ln = lanthanide or Y) family of ceramic materials and exhibits a range of desirable material properties such as radiation tolerance, frustrated magnetism, and large dielectric constant. However, understanding the complex crystal structure of Y2TiO5 remains elusive, given that Y2TiO5 can adopt multiple polymorphs such as cubic, orthorhombic, and hexagonal phases within the lattice. In this work, we report a detailed structural analysis of Y2TiO5 using four-dimensional scanning transmission electron microscopy coupled with unsupervised machine learning. The pyrochlore nanodomains, characterized by the ordered arrangement of yttrium cations on the A site of their A2BO5 structure, are present within the matrix of a predominantly fluorite-structured Y2TiO5 along with a third polymorph, the hexagonal phase. The pyrochlore phase is found to form 2 nm boundary regions around hexagonal phase stacking faults, highlighting the potential influence of the hexagonal phase on the occurrence and distribution of the pyrochlore phase. Lastly, we identify a unique pyrochlore phase with asymmetric arrangement of cation ordering along a single planar direction. Our findings provide invaluable insights into the possible mechanisms stabilizing pyrochlore nanodomains within the fluorite lattice of Y2TiO5.

36 MATERIALS SCIENCE↗

Application of unsupervised deep learning to image segmentation and in-situ contact angle measurements in a CO 2 -water-rock system

Rock surface wettability is a critical property that regulates multiphase flows in porous media, which can be quantified using the surface contact angle (CA). X-ray micro-computed tomography (μCT) provides an effective approach to in-situ measurements of surface CAs. However, the CA measurement accuracy depends significantly on the quality of CT image segmentation, which is the clustering of CT pixels into separate phases. Inspired by this, we developed a deep learning (DL)-based CA measurement workflow. Motivated by the recent tremendous progress in unsupervised learning techniques and aiming to avoid expensive manual data annotations, an unsupervised DL pipeline for CT image segmentation was proposed and implemented, which includes unsupervised model training and post-processing. The unsupervised model training was driven by a novel loss function constrained with feature similarity and spatial continuity and implemented by iterative forward and backward paths; the former clustered the pixel-wise feature vectors extracted by convolution neural networks, whereas the latter updated the parameters using gradient descent. An over-segmentation strategy was adopted for model training. The post-processing steps based on agglomerative hierarchical clustering (AHC) were implemented to further merge the over-segmented model output to the desired cluster number, which is intended to improve the efficiency of image segmentation. The developed unsupervised DL pipeline was compared with other commonly-used image segmentation methods using pixel-wise and physics-based evaluation metrics on a synthetic raw-image dataset, which had a known ground truth. The unsupervised DL pipeline showed the best performance. Next, the segmented images were input to an automatic CA measurement tool, and the results were validated by comparisons with manual measurements. The CA values from the manual and automatic measurements showed similar distributions and statistical properties. The automatic measurement demonstrated a wider spectrum because of the much larger number of measurement data points. The primary novelty of the unsupervised DL pipeline developed in this study lies in the novel loss function and the over-segmentation strategy associated with AHC post-processing. Finally, the workflow has been proven an efficient tool for pore-scale wettability characterization, which has a wide range of applications in fundamental studies of multiphase flows in natural porous media, which have critical implications to geological carbon sequestration, hydrocarbon energy recovery, and contaminant transport in groundwater.

42 ENGINEERING↗

General-Purpose Unsupervised Cyber Anomaly Detection via Non-Negative Tensor Factorization

Distinguishing malicious anomalous activities from unusual but benign activities is a fundamental challenge for cyber defenders. Prior studies have shown that statistical user behavior analysis yields accurate detections by learning behavior profiles from observed user activity. These unsupervised models are able to generalize to unseen types of attacks by detecting deviations from normal behavior, without knowledge of specific attack signatures. However, approaches proposed to date based on probabilistic matrix factorization are limited by the information conveyed in a two-dimensional space. Non-negative tensor factorization, on the other hand, is a powerful unsupervised machine learning method that naturally models multi-dimensional data, capturing complex and multi-faceted details of behavior profiles. Herein, our new unsupervised statistical anomaly detection methodology matches or surpasses state-of-the-art supervised learning baselines across several challenging and diverse cyber application areas, including detection of compromised user credentials, botnets, spam e-mails, and fraudulent credit card transactions.

97 MATHEMATICS AND COMPUTING↗

Seeking regularity from irregularity: unveiling the synthesis–nanomorphology relationships of heterogeneous nanomaterials using unsupervised machine learning

Nanoscale morphology of functional materials determines their chemical and physical properties. However, despite increasing use of transmission electron microscopy (TEM) to directly image nanomorphology, it remains challenging to quantify the information embedded in TEM data sets, and to use nanomorphology to link synthesis and processing conditions to properties. We develop an automated, descriptor-free analysis workflow for TEM data that utilizes convolutional neural networks and unsupervised learning to quantify and classify nanomorphology, and thereby reveal synthesis–nanomorphology relationships in three different systems. While TEM records nanomorphology readily in two-dimensional (2D) images or three-dimensional (3D) tomograms, we advance the analysis of these images by identifying and applying a universal shape fingerprint function to characterize nanomorphology. After dimensionality reduction through principal component analysis, this function then serves as the input for morphology grouping through unsupervised learning. We demonstrate the wide applicability of our workflow to both 2D and 3D TEM data sets, and to both inorganic and organic nanomaterials, including tetrahedral gold nanoparticles mixed with irregularly shaped impurities, hybrid polymer-patched gold nanoprisms, and polyamide membranes with irregular and heterogeneous 3D crumple structures. In each of these systems, unsupervised nanomorphology grouping identifies both the diversity and the similarity of the nanomaterial across different synthesis conditions, revealing how synthetic parameters guide nanomorphology development. Our work opens possibilities for enhancing synthesis of nanomaterials through artificial intelligence and for understanding and controlling complex nanomorphology, both for 2D systems and in the far less explored case of 3D structures, such as those with embedded voids or hidden interfaces.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Exploring Continuous Seismic Data at an Industry Facility Using Unsupervised Machine Learning

Seismic data recorded at industrial sites contain valuable information on anthropogenic activities. With advances in machine learning and computing power, new opportunities have emerged to explore the seismic wavefield in these complex environments. We applied two unsupervised machine learning algorithms to analyze continuous seismic data collected from an industrial facility in Texas, United States. The Uniform Manifold Approximation and Projection for Dimension Reduction algorithm was used to reduce the dimensionality of the data and generate 2D embeddings. Then, the Hierarchical Density-Based Spatial Clustering of Applications with Noise method was employed to automatically group these embeddings into distinct signal clusters. Our analysis of over 1400 hr (around 59 days) of continuous seismic data revealed five and seven signal clusters at two separate stations. At both stations, we identified clusters associated with background noise and vehicle traffic, with the latter’s temporal patterns aligning closely with the facility’s work schedule. Furthermore, the algorithms detected signal clusters from unknown sources and underline the ability of unsupervised machine learning for uncovering previously unrecognized patterns. Our analysis demonstrates the effectiveness of unsupervised approaches in examining continuous seismic data without requiring prior knowledge or pre-existing labels.

58 GEOSCIENCES↗

Unsupervised Machine Learning for Exploratory Data Analysis of Exoplanet Transmission Spectra

Abstract Transit spectroscopy is a powerful tool for decoding the chemical compositions of the atmospheres of extrasolar planets. In this paper, we focus on unsupervised techniques for analyzing spectral data from transiting exoplanets. After cleaning and validating the data, we demonstrate methods for: (i) initial exploratory data analysis, based on summary statistics (estimates of location and variability); (ii) exploring and quantifying the existing correlations in the data; (iii) preprocessing and linearly transforming the data to its principal components; (iv) dimensionality reduction and manifold learning; (v) clustering and anomaly detection; and (vi) visualization and interpretation of the data. To illustrate the proposed unsupervised methodology, we use a well-known public benchmark data set of synthetic transit spectra. We show that there is a high degree of correlation in the spectral data, which calls for appropriate low-dimensional representations. We explore a number of different techniques for such dimensionality reduction and identify several suitable options in terms of summary statistics, principal components, etc. We uncover interesting structures in the principal component basis, namely well-defined branches corresponding to different chemical regimes of the underlying atmospheres. We demonstrate that those branches can be successfully recovered with a K-means clustering algorithm in a fully unsupervised fashion. We advocate for lower-dimensional representations of the spectroscopic data in terms of the main principal components, in order to reveal the existing structure in the data and quickly characterize the chemical class of a planet.

Matchev, Konstantin T. (ORCID:0000000341829096)↗

An unsupervised classification approach for analysis of Landsat data to monitor land reclamation in Belmont county, Ohio

Two unsupervised classification procedures for analyzing Landsat data used to monitor land reclamation in a surface mining area in east central Ohio are compared for agreement with data collected from the corresponding locations on the ground. One procedure is based on a traditional unsupervised-clustering/maximum-likelihood algorithm sequence that assumes spectral groupings in the Landsat data in n-dimensional space; the other is based on a nontraditional unsupervised-clustering/canonical-transformation/clustering algorithm sequence that not only assumes spectral groupings in n-dimensional space but also includes an additional feature-extraction technique. It is found that the nontraditional procedure provides an appreciable improvement in spectral groupings and apparently increases the level of accuracy in the classification of land cover categories.

Brumfield, J. O.↗

Identification of Security related Bug Reports via Text Mining using Supervised and Unsupervised Classification

This paper is focused on automated classification of software bug reports to security and non-security related, using both supervised and unsupervised approaches. For both approaches, three types of feature vectors are used. For supervised learning, we experiment with multiple learning algorithms and training sets with different sizes. Furthermore, we propose a novel unsupervised approach based on anomaly detection. The evaluated is based on three NASA datasets. The results show that supervised classification is affected more by the learning algorithms than by feature vectors and using only 25% of the data for training provides as good results as if 90% of data are used for training. Both supervised and unsupervised learning can be used for identification of security bug reports; the former slightly outperforms the latter at the expense of labeling the testing set. In general, the performance differs across datasets, mainly due to the different amounts of security related information.

Goseva-Popstojanova, Katerina↗

Unsupervised Deep Persistent Monocular Visual Odometry and Depth Estimation in Extreme Environments

In recent years, unsupervised deep learning ap-proaches have received a significant attention to estimate depthand visual odometry (VO) from unlabelled monocular imagesequences. However, their performance is limited in challengingenvironments due to perceptual degradation, occlusions andrapid motions. Moreover, the existing unsupervised methodssuffer from the lack of scale-consistency constraints acrossframes, which causes that the VO estimators fail to providepersistent trajectories over long sequences. In this study, wepropose a unsupervised monocular deep VO framework thatpredicts 6 degrees-of-freedom pose camera motion and depthmap of the scene from unlabelled RGB image sequences.We provide detailed quantitative and qualitative evaluationsof the proposed framework on a) a challenging dataset col-lected during the DARPA Subterranean challenge1; and b)the benchmark KITTI and Cityscapes datasets. The proposedapproach outperforms both traditional and state-of-the-artunsupervised deep VO methods providing better results for bothpose estimation and depth recovery. The presented approach ispart of the solution used by the COSTAR team participatingat the DARPA Subterranean Challenge

Agha-mohammadi, Ali-akbar↗

A survey of unsupervised learning methods for high-dimensional uncertainty quantification in black-box-type problems

Constructing surrogate models for uncertainty quantification (UQ) on complex partial differential equations (PDEs) having inherently high-dimensional O(10 n ), n ≥ 2, stochastic inputs (e.g., forcing terms, boundary conditions, initial conditions) poses tremendous challenges. The “curse of dimensionality” can be addressed with suitable unsupervised learning techniques used as a pre-processing tool to encode inputs onto lower-dimensional subspaces while retaining its structural information and meaningful properties. In this work, we review and investigate thirteen dimension reduction methods including linear and nonlinear, spectral, blind source separation, convex and non-convex methods and utilize the resulting embeddings to construct a mapping to quantities of interest via polynomial chaos expansions (PCE). Here, we refer to the general proposed approach as manifold PCE (m-PCE), where manifold corresponds to the latent space resulting from any of the studied dimension reduction methods. To investigate the capabilities and limitations of these methods we conduct numerical tests for three physics-based systems (treated as black-boxes) having high-dimensional stochastic inputs of varying complexity modeled as both Gaussian and non-Gaussian random fields to investigate the effect of the intrinsic dimensionality of input data. We demonstrate both the advantages and limitations of the unsupervised learning methods and we conclude that a suitable m-PCE model provides a cost-effective approach compared to alternative algorithms proposed in the literature, including recently proposed expensive deep neural network-based surrogates and can be readily applied for high-dimensional UQ in stochastic PDEs.

42 ENGINEERING↗

Characterizing vertical upper ocean temperature structures in the European Arctic through unsupervised machine learning

In-situ observations of subsurface ocean temperatures are, in many regions, inconsistently distributed in time and space. These spatio-temporal inconsistencies in the observational network lead to difficulties in utilizing those observations effectively for ocean model evaluation or understanding larger-scale ocean characteristics. Model accuracy of subsurface ocean characteristics is especially important within regions that contain complex ocean structures. One such region is the European Arctic which not only contains several types of water masses with unique characteristics, but also wintertime sea ice coverage and complex bathymetry. This study presents an unsupervised neural networking technique that can be used in combination with traditional ocean model evaluation techniques to provide additional information on the accuracy of modeled vertical ocean temperature profiles. Self-organizing maps is an unsupervised machine learning technique that we apply to approximately twenty thousand Argo and CTD temperature profiles from 2012 to 2020 in the European Arctic to categorize the observed vertical ocean temperature structures in the top 150 m. The observed ocean profile categories, or neurons, defined by the self-organizing map show strong spatial and temporal dependencies. We then use the neuron weights, or the learned temperature profile structure of each neuron, to validate the spatial and temporal variability of modeled vertical temperature structures. This analysis gives us new insights about the model’s capabilities to reproduce specific vertical structures of the top-most ocean layer within different regions and seasons. Mapping modeled ocean temperature profiles onto the neuron-space of the observationally-defined self organized map highlights the potential of this method to advance our understanding of model deficiencies in that region.

54 ENVIRONMENTAL SCIENCES↗

Characterizing Drought Behavior in the Colorado River Basin Using Unsupervised Machine Learning

Drought is a pressing issue for the Colorado River Basin (CRB) due to the social and economic value of water resources in the region and the significant uncertainty of future drought under climate change. Here, we use climate simulations from various Earth System Models (ESMs) to force the Variable Infiltration Capacity hydrologic model and project multiple drought indicators for the sub-watersheds within the CRB. We apply an unsupervised machine learning (ML) based on Non-Negative Matrix Factorization using K-means clustering (NMFk) to synthesize the simulated historical, future, and change in drought indicators. The unsupervised ML approach can identify sub-watersheds where key changes to drought indicator behavior occur, including shifts in snowpack, snowmelt timing, precipitation, and evapotranspiration. While changes in future precipitation vary across ESMs, the results indicate that the Upper CRB will experience increasing evaporative demand and surface-water scarcity, with some locations experiencing a shift from a radiation-limited to a water-limited evaporation regime in the summer. Large shifts in peak runoff are observed in snowmelt-dominant sub-watersheds, with complete disappearance of the snowmelt signal for some sub-watersheds. The work demonstrates the utility of the NMFk algorithm to efficiently identify behavioral changes of drought indicators across space and time and to quickly analyze and interpret hydro climate model results.

54 ENVIRONMENTAL SCIENCES↗

Internship Final Report on the unsupervised learning sensor fusion (ULSF) approach

This paper describes a summer internship project undertaken at Sandia National Labs (SNL), both current status and future work. The project was to explore various machine learning approaches for use on turbulent flow data. Specifically, unsupervised classification of turbulent flow data was explored. First, the usage of models in this field is discussed, and several issues in the common usage of the models are identified. Solutions to these issues are then proposed, in the form of a Bayesian filtering approach which probabilistically incorporates multiple sources of data to improve confidence in a result. Several types of sensors are suggested for this method, the incorporation of which range from semi-supervised learning approaches to fully unsupervised. These approaches are then tested on several turbulent flow cases.

97 MATHEMATICS AND COMPUTING↗

Missing Wedge Completion via Unsupervised Learning with Coordinate Networks

Cryogenic electron tomography (cryoET) is a powerful tool in structural biology, enabling detailed 3D imaging of biological specimens at a resolution of nanometers. Despite its potential, cryoET faces challenges such as the missing wedge problem, which limits reconstruction quality due to incomplete data collection angles. Recently, supervised deep learning methods leveraging convolutional neural networks (CNNs) have considerably addressed this issue; however, their pretraining requirements render them susceptible to inaccuracies and artifacts, particularly when representative training data is scarce. To overcome these limitations, we introduce a proof-of-concept unsupervised learning approach using coordinate networks (CNs) that optimizes network weights directly against input projections. This eliminates the need for pretraining, reducing reconstruction runtime by 3–20× compared to supervised methods. Our in silico results show improved shape completion and reduction of missing wedge artifacts, assessed through several voxel-based image quality metrics in real space and a novel directional Fourier Shell Correlation (FSC) metric. Our study illuminates benefits and considerations of both supervised and unsupervised approaches, guiding the development of improved reconstruction strategies.

42 ENGINEERING↗

Reward Driven Workflows for Unsupervised Explainable Analysis of Phases and Ferroic Variants From Atomically Resolved Imaging Data

Rapid progress in aberration corrected electron microscopy necessitates development of robust methods for the identification of phases, ferroic variants, and other pertinent aspects of materials structure from imaging data. While unsupervised methods for clustering and classification are widely used for these tasks, their performance can be sensitive to hyperparameter selection in the analysis workflow. In this study, the effects of descriptors and hyperparameters are explored on the capability of unsupervised ML methods to distill local structural information, exemplified by the discovery of polarization and lattice distortion in Sm − dopped BiFeO 3 (BFO) thin films. It is demonstrated that a reward-driven approach can be used to optimize these key hyperparameters across the full workflow, where rewards are designed to reflect domain wall continuity and straightness, ensuring that the analysis aligns with the material's physical behavior. This approach allows the discovery of local descriptors that are best aligned with the specific physical behavior, providing insight into the fundamental physics of materials. The reward driven workflow is further extended to disentangle structural factors of variation via an optimized variational autoencoder (VAE). Lastly, the importance of well-defined rewards is explored as a quantifiable measure of the success of the workflow.

Barakati, Kamyar [University of Tennessee, Knoxvil↗

Evaluating lightweight unsupervised online IDS for masquerade attacks in CAN

Vehicular controller area networks (CANs) are susceptible to masquerade attacks by malicious adversaries. In masquerade attacks, adversaries silence a targeted ID and then send malicious frames with forged content at the expected timing of benign frames. As masquerade attacks could seriously harm vehicle functionality and are the stealthiest attacks to detect in CAN, recent work has devoted attention to compare frameworks for detecting masquerade attacks in CAN. However, most existing works report offline evaluations using CAN logs already collected using simulations that do not comply with the domain’s real-time constraints. Here we contribute to advance the state of the art by presenting a comparative evaluation of four different non-deep learning (DL)-based unsupervised online intrusion detection systems (IDS) for masquerade attacks in CAN. Our approach differs from existing comparative evaluations in that we analyze the effect of controlling streaming data conditions in a sliding window setting. In doing so, we use realistic masquerade attacks being replayed from the ROAD dataset. We show that although evaluated IDS are not effective at detecting every attack type, the method that relies on detecting changes in the hierarchical structure of clusters of time series produces the best results at the expense of higher computational overhead. We discuss limitations, open challenges, and how the evaluated methods can be used for practical unsupervised online CAN IDS for masquerade attacks.

Anomaly detection↗

Unsupervised domain adaptation for radioisotope identification in gamma spectroscopy

Training machine learning models for radioisotope identification using gamma spectroscopy remains an elusive challenge for many practical applications, largely stemming from the difficulty of acquiring and labeling large, diverse experimental datasets. Simulations can mitigate this challenge, but the accuracy of models trained on simulated data can deteriorate substantially when deployed to an out-of-distribution operational environment. In this study, we demonstrate that unsupervised domain adaptation (UDA) can improve the ability of a model trained on synthetic data to generalize to a new testing domain, provided unlabeled data from the target domain are available. Conventional supervised techniques are unable to utilize this data because the absence of isotope labels precludes defining a supervised classification loss. Instead, we first pretrain a spectral classifier using labeled synthetic data and subsequently leverage unlabeled target data to align the learned feature representations between the source and target domains. We compare a range of different UDA techniques, finding that minimizing the maximum mean discrepancy (MMD) between source and target feature vectors yields the most consistent improvement to testing scores. For instance, using a custom transformer-based neural network, we achieved a testing accuracy of $0.904 \pm 0.022$ on an experimental LaBr test set after performing unsupervised feature alignment via MMD minimization, compared to $0.754 \pm 0.014$ before alignment. Overall, our results highlight the potential of using UDA to adapt a radioisotope classifier trained on synthetic data for real-world deployment.

Lalor, Peter W.↗