Search NASA⌕ Search

SEARCH · Search NASA

Results for “unsupervised method”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

Object-oriented feature-tracking algorithms for SAR images of the marginal ice zone

An unsupervised method that chooses and applies the most appropriate tracking algorithm from among different sea-ice tracking algorithms is reported. In contrast to current unsupervised methods, this method chooses and applies an algorithm by partially examining a sequential image pair to draw inferences about what was examined. Based on these inferences the reported method subsequently chooses which algorithm to apply to specific areas of the image pair where that algorithm should work best.

Daida, Jason↗

An unsupervised feature extraction method for high dimensional image data compaction

A new on-line unsupervised feature extraction method for high-dimensional remotely sensed image data compaction is presented. This method can be utilized to solve the problem of data redundancy in scene representation by satellite-borne high resolution multispectral sensors. The algorithm first partitions the observation space into an exhaustive set of disjoint objects. Then, pixels that belong to an object are characterized by an object feature. Finally, the set of object features is used for data transmission and classification. The example results show that the performance with the compacted features provides a slight improvement in classification accuracy instead of any degradation. Also, the information extraction method does not need to be preceded by a data decompaction.

Ghassemian, Hassan↗

Automatic Feature Extraction from Planetary Images

With the launch of several planetary missions in the last decade, a large amount of planetary images has already been acquired and much more will be available for analysis in the coming years. The image data need to be analyzed, preferably by automatic processing techniques because of the huge amount of data. Although many automatic feature extraction methods have been proposed and utilized for Earth remote sensing images, these methods are not always applicable to planetary data that often present low contrast and uneven illumination characteristics. Different methods have already been presented for crater extraction from planetary images, but the detection of other types of planetary features has not been addressed yet. Here, we propose a new unsupervised method for the extraction of different features from the surface of the analyzed planet, based on the combination of several image processing techniques, including a watershed segmentation and the generalized Hough Transform. The method has many applications, among which image registration and can be applied to arbitrary planetary images.

Troglio, Giulia↗

Analysis of thematic mapper simulator data acquired during winter season over Pearl River, Mississippi, test site

Digital processed aircraft-acquired thematic mapping simulator (TMS) data collected during the winter season over a forested site in southern Mississippi are presented to investigate the utility of TMS data for use in forest inventories and monitoring. Analyses indicated that TMS data are capable of delineating the mixed forest land cover type to an accuracy of 92.5 % correct. The accuracies associated with river bottom forest and pine forest were 95.5 and 91.5 % correct. The accuracies associated with river bottom forest and pine forest were 95.5 and 91.5 % correct, respectively. The figures reflect the performance for products produced using the best subset of channels for each forest cover type. It was found that the choice of channels (subsets) has a significant effect on the accuracy of classification produced, and that the same channels are not the most desirable for all three forest types studied. Both supervised and unsupervised spectral signature development techniques are evaluated; the unsupervised methods proved unacceptable for the three forest types considered.

Anderson, J. E.↗

Analysis of thematic mapper simulator data collected over eastern North Dakota

The results of the analysis of aircraft-acquired thematic mapper simulator (TMS) data, collected to investigate the utility of thematic mapper data in crop area and land cover estimates, are discussed. Results of the analysis indicate that the seven-channel TMS data are capable of delineating the 13 crop types included in the study to an overall pixel classification accuracy of 80.97% correct, with relative efficiencies for four crop types examined between 1.62 and 26.61. Both supervised and unsupervised spectral signature development techniques were evaluated. The unsupervised methods proved to be inferior (based on analysis of variance) for the majority of crop types considered. Given the ground truth data set used for spectral signature development as well as evaluation of performance, it is possible to demonstrate which signature development technique would produce the highest percent correct classification for each crop type.

Anderson, J. E.↗

Improving Acoustic Models by Watching Television

Obtaining sufficient labelled training data is a persistent difficulty for speech recognition research. Although well transcribed data is expensive to produce, there is a constant stream of challenging speech data and poor transcription broadcast as closed-captioned television. We describe a reliable unsupervised method for identifying accurately transcribed sections of these broadcasts, and show how these segments can be used to train a recognition system. Starting from acoustic models trained on the Wall Street Journal database, a single iteration of our training method reduced the word error rate on an independent broadcast television news test set from 62.2% to 59.5%.

Witbrock, Michael J.↗

Unsupervised Segmentation Of Polarimetric SAR Data

Method of unsupervised segmentation of polarimetric synthetic-aperture-radar (SAR) image data into classes involves selection of classes on basis of multidimensional fuzzy clustering of logarithms of parameters of polarimetric covariance matrix. Data in each class represent parts of image wherein polarimetric SAR backscattering characteristics of terrain regarded as homogeneous. Desirable to have each class represent type of terrain, sea ice, or ocean surface distinguishable from other types via backscattering characteristics. Unsupervised classification does not require training areas, is nearly automated computerized process, and provides nonsubjective selection of image classes naturally well separated by radar.

Rignot, Eric J.↗

Unsupervised segmentation of polarimetric SAR data using the covariance matrix

A method for unsupervised segmentation of polarimetric synthetic aperture radar (SAR) data into classes of homogeneous microwave polarimetric backscatter characteristics is presented. Classes of polarimetric backscatter are selected on the basis of a multidimensional fuzzy clustering of the logarithm of the parameters composing the polarimetric covariance matrix. The clustering procedure uses both polarimetric amplitude and phase information, is adapted to the presence of image speckle, and does not require an arbitrary weighting of the different polarimetric channels; it also provides a partitioning of each data sample used for clustering into multiple clusters. Given the classes of polarimetric backscatter, the entire image is classified using a maximum a posteriori polarimetric classifier. Four-look polarimetric SAR complex data of lava flows and of sea ice acquired by the NASA/JPL airborne polarimetric radar (AIRSAR) are segmented using this technique. The results are discussed and compared with those obtained using supervised techniques.

Rignot, Eric J. M.↗

Data Mining for Anomaly Detection

The Vehicle Integrated Prognostics Reasoner (VIPR) program describes methods for enhanced diagnostics as well as a prognostic extension to current state of art Aircraft Diagnostic and Maintenance System (ADMS). VIPR introduced a new anomaly detection function for discovering previously undetected and undocumented situations, where there are clear deviations from nominal behavior. Once a baseline (nominal model of operations) is established, the detection and analysis is split between on-aircraft outlier generation and off-aircraft expert analysis to characterize and classify events that may not have been anticipated by individual system providers. Offline expert analysis is supported by data curation and data mining algorithms that can be applied in the contexts of supervised learning methods and unsupervised learning. In this report, we discuss efficient methods to implement the Kolmogorov complexity measure using compression algorithms, and run a systematic empirical analysis to determine the best compression measure. Our experiments established that the combination of the DZIP compression algorithm and CiDM distance measure provides the best results for capturing relevant properties of time series data encountered in aircraft operations. This combination was used as the basis for developing an unsupervised learning algorithm to define "nominal" flight segments using historical flight segments.

Biswas, Gautam↗

A Fast Implementation of the ISOCLUS Algorithm

Unsupervised clustering is a fundamental tool in numerous image processing and remote sensing applications. For example, unsupervised clustering is often used to obtain vegetation maps of an area of interest. This approach is useful when reliable training data are either scarce or expensive, and when relatively little a priori information about the data is available. Unsupervised clustering methods play a significant role in the pursuit of unsupervised classification. One of the most popular and widely used clustering schemes for remote sensing applications is the ISOCLUS algorithm, which is based on the ISODATA method. The algorithm is given a set of n data points (or samples) in d-dimensional space, an integer k indicating the initial number of clusters, and a number of additional parameters. The general goal is to compute a set of cluster centers in d-space. Although there is no specific optimization criterion, the algorithm is similar in spirit to the well known k-means clustering method in which the objective is to minimize the average squared distance of each point to its nearest center, called the average distortion. One significant feature of ISOCLUS over k-means is that clusters may be merged or split, and so the final number of clusters may be different from the number k supplied as part of the input. This algorithm will be described in later in this paper. The ISOCLUS algorithm can run very slowly, particularly on large data sets. Given its wide use in remote sensing, its efficient computation is an important goal. We have developed a fast implementation of the ISOCLUS algorithm. Our improvement is based on a recent acceleration to the k-means algorithm, the filtering algorithm, by Kanungo et al.. They showed that, by storing the data in a kd-tree, it was possible to significantly reduce the running time of k-means. We have adapted this method for the ISOCLUS algorithm. For technical reasons, which are explained later, it is necessary to make a minor modification to the ISOCLUS specification. We provide empirical evidence, on both synthetic and Landsat image data sets, that our algorithm's performance is essentially the same as that of ISOCLUS, but with significantly lower running times. We show that our algorithm runs from 3 to 30 times faster than a straightforward implementation of ISOCLUS. Our adaptation of the filtering algorithm involves the efficient computation of a number of cluster statistics that are needed for ISOCLUS, but not for k-means.

Memarsadeghi, Nargess↗

Sparse Superpixel Unmixing for Hyperspectral Image Analysis

Software was developed that automatically detects minerals that are present in each pixel of a hyperspectral image. An algorithm based on sparse spectral unmixing with Bayesian Positive Source Separation is used to produce mineral abundance maps from hyperspectral images. A superpixel segmentation strategy enables efficient unmixing in an interactive session. The algorithm computes statistically likely combinations of constituents based on a set of possible constituent minerals whose abundances are uncertain. A library of source spectra from laboratory experiments or previous remote observations is used. A superpixel segmentation strategy improves analysis time by orders of magnitude, permitting incorporation into an interactive user session (see figure). Mineralogical search strategies can be categorized as supervised or unsupervised. Supervised methods use a detection function, developed on previous data by hand or statistical techniques, to identify one or more specific target signals. Purely unsupervised results are not always physically meaningful, and may ignore subtle or localized mineralogy since they aim to minimize reconstruction error over the entire image. This algorithm offers advantages of both methods, providing meaningful physical interpretations and sensitivity to subtle or unexpected minerals.

Castano, Rebecca↗

Identifying Emerging Safety Threats Through Topic Modeling in the Aviation Safety Reporting System: A Covid-19 Study

The NASA Aviation Safety Reporting System (ASRS) is a voluntary, confidential aviation safety reporting system. The ASRS receives reports from pilots, air traffic controllers, flight attendants, and others involved in aviation operations. The reports are de-identified and coded by ASRS expert safety analysts, and a short descriptive synopsis is written to describe the safety issue. The de-identified reports are then disseminated to the aviation community in many ways, including via an online database, Safety Alert Bulletins, For Your Information Notices, and the CALLBACK newsletter. In this work, we consider whether we can improve the grouping, linking, and understanding of safety concerns through topic modeling. Specifically, we use topic modeling as a building block to identify emerging safety threats over time. This unsupervised approach, we argue, offers the flexibility to identify new emerging themes in this large dataset by constructing different timelines based on the content similarity of ASRS report narratives. This method's unsupervised nature improves upon related research, which is limited to pre-defined labels and therefore can not fully capture emerging safety threats. We apply our method to all ASRS reports in 2020 to assess if the generated timelines can highlight COVID-19 as it is emerging as a safety threat in incoming ASRS reports. We perform both a quantitative and qualitative evaluation of the automatically constructed timelines. The qualitative evaluation is performed by describing the evolution of top terms in the timelines, generated by our method, which we found explicitly convey the themes of COVID-19. Separately, we use a set of 1,213 COVID-19 reports from 2020 that were manually identified by ASRS analysts to quantitatively evaluate the COVID-19 reports distribution across the timelines. Our results have shown that COVID-19 emergence can be identified using the top terms that were generated by topic modeling. The top terms in topic modeling therefore can serve as a summary alternative to manually inspecting reports. Moreover, leveraging the manually identified COVID-19 reports, we found the manually identified timelines accounted for over 70% of the COVID-19 reports curated by the ASRS analysts, which demonstrates the potential of this approach for facilitating the understanding of safety concerns as they emerge and evolve. This method shows great potential to understand aerospace safety threats and other narrative- driven incident report databases.

ASRS↗

Identifying Emerging Safety Threats Through Topic Modeling in the Aviation Safety Reporting System: A COVID-19 Study

The NASA Aviation Safety Reporting System (ASRS) is a voluntary, confidential aviation safety reporting system. The ASRS receives reports from pilots, air traffic controllers, flight attendants, and others involved in aviation operations. The reports are de-identified and coded by ASRS expert safety analysts, and a short descriptive synopsis is written to describe the safety issue. The de-identified reports are then disseminated to the aviation community in many ways, including via an online database, Safety Alert Bulletins, For Your Information Notices, and the CALLBACK newsletter. In this work, we consider whether we can improve the grouping, linking, and understanding of safety concerns through topic modeling. Specifically, we use topic modeling as a building block to identify emerging safety threats over time. This unsupervised approach, we argue, offers the flexibility to identify new emerging themes in this large dataset by constructing different timelines based on the content similarity of ASRS report narratives. This method's unsupervised nature improves upon related research, which is limited to pre-defined labels and therefore can not fully capture emerging safety threats. We apply our method to all ASRS reports in 2020 to assess if the generated timelines can highlight COVID-19 as it is emerging as a safety threat in incoming ASRS reports. We perform both a quantitative and qualitative evaluation of the automatically constructed timelines. The qualitative evaluation is performed by describing the evolution of top terms in the timelines, generated by our method, which we found explicitly convey the themes of COVID-19. Separately, we use a set of 1,213 COVID-19 reports from 2020 that were manually identified by ASRS analysts to quantitatively evaluate the COVID-19 reports distribution across the timelines. Our results have shown that COVID-19 emergence can be identified using the top terms that were generated by topic modeling. The top terms in topic modeling therefore can serve as a summary alternative to manually inspecting reports. Moreover, leveraging the manually identified COVID-19 reports, we found the manually identified timelines accounted for over 70% of the COVID-19 reports curated by the ASRS analysts, which demonstrates the potential of this approach for facilitating the understanding of safety concerns as they emerge and evolve. This method shows great potential to understand aerospace safety threats and other narrative- driven incident report databases.

ASRS↗

Automated Classification of Transient Contamination in Stationary Acoustic Data

An automated procedure for the classification of transient contamination of stationary acoustic data is proposed and analyzed. The procedure requires the assumption that the stationary acoustic data of interest can be modeled as a band-limited, Gaussian random process. It also requires that the transient contamination be of higher variance than the acoustic data of interest. When these assumptions are satisfied, it is a blind separation procedure, aside from the initial input specifying how to subdivide the time series of interest. No a priori threshold criterion is required. Simulation results show that for a sufficient number of blocks, the method performs well, as long as the occasional false positive or false negative is acceptable. The effectiveness of the procedure is demonstrated with an application to experimental wind tunnel acoustic test data which are contaminated by hydrodynamic gusts.

binary classification↗

Cluster Method Analysis of K. S. C. Image

Information obtained from satellite-based systems has moved to the forefront as a method in the identification of many land cover types. Identification of different land features through remote sensing is an effective tool for regional and global assessment of geometric characteristics. Classification data acquired from remote sensing images have a wide variety of applications. In particular, analysis of remote sensing images have special applications in the classification of various types of vegetation. Results obtained from classification studies of a particular area or region serve towards a greater understanding of what parameters (ecological, temporal, etc.) affect the region being analyzed. In this paper, we make a distinction between both types of classification approaches although, focus is given to the unsupervised classification method using 1987 Thematic Mapped (TM) images of Kennedy Space Center.

Rodriguez, Joe, Jr.↗

Active Learning with Rationales for Identifying Operationally Significant Anomalies in Aviation

A major focus of the commercial aviation community is discovery of unknown safety events in flight operations data. Data-driven unsupervised anomaly detection methods are better at capturing unknown safety events compared to rule-based methods which only look for known violations. However, not all statistical anomalies that are discovered by these unsupervised anomaly detection methods are operationally significant (e.g., represent a safety concern). Subject Matter Experts (SMEs) have to spend significant time reviewing these statistical anomalies individually to identify a few operationally significant ones. In this paper we propose an active learning algorithm that incorporates SME feedback in the form of rationales to build a classifier that can distinguish between uninteresting and operationally significant anomalies. Experimental evaluation on real aviation data shows that our approach improves detection of operationally significant events by as much as 75% compared to the state-of-the-art. The learnt classifier also generalizes well to additional validation data sets.

anomaly detection↗