Search NASASearch

SEARCH · Search NASA

Results for “cluster analysis”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 73 records · Page 4

Satellite Remote Sensing to Assess Cyanobacterial Bloom Frequency Across the United States at Multiple Spatial Scales

Cyanobacterial blooms can have negative effects on human health and local ecosystems. Field monitoring of cyanobacterial blooms can be costly, but satellite remote sensing has shown utility for more efficient spatial and temporal monitoring across the United States. Here, satellite imagery was used to assess the annual frequency of surface cyanobacterial blooms, defined for each satellite pixel as the percentage of images for that pixel throughout the year exhibiting detectable cyanobacteria. Cyanobacterial frequency was assessed across 2,196 large lakes in 46 states across the continental United States (CONUS) using imagery from the European Space Agency’s Ocean and Land Colour Imager for the years 2017 through 2019. In 2019, across all satellite pixels considered, annual bloom frequency had a median value of 4% and a maximum value of 100%, the latter indicating that for those satellite pixels, a cyanobacterial bloom was detected by the satellite sensor for every satellite image considered. In addition to annual pixel-scale cyanobacterial frequency, results were summarized at the lake- and state-scales by averaging annual pixel-scale results across each lake and state. For 2019, average annual lake-scale frequencies also had a maximum value of 100%, and Oregon and Ohio had the highest average annual state-scale frequencies at 65% and 52%. Pixel-scale frequency results can assist in identifying portions of a lake that are more prone to cyanobacterial blooms, while lake- and state-scale frequency results can assist in the prioritization of sampling resources and mitigation efforts. Satellite imagery is limited by the presence of snow and ice, as imagery collected in these conditions are quality flagged and discarded. Thus, annual bloom frequencies within nine climate regions were investigated to determine whether missing data biased results in climate regions more prone to snow and ice, given that their annual summaries would be weighted toward the summer months when cyanobacterial blooms tend to occur. Results were unbiased by the time period selected in most climate regions, but a large bias was observed for the Northwest Rockies and Plains climate region. Moderate biases were observed for the Ohio Valley and the Southeast climate regions. Finally, a clustering analysis was used to identify areas of high and low cyanobacterial frequency across CONUS based on average annual lake-scale cyanobacterial frequencies for 2019. Several clusters were identified that transcended state, watershed, and eco-regional boundaries. Combined with additional data, results from the clustering analysis may offer insight regarding large-scale drivers of cyanobacterial blooms.

remote sensing

Technical support for creating an artificial intelligence system for feature extraction and experimental design

Techniques for classifying objects into groups or clases go under many different names including, most commonly, cluster analysis. Mathematically, the general problem is to find a best mapping of objects into an index set consisting of class identifiers. When an a priori grouping of objects exists, the process of deriving the classification rules from samples of classified objects is known as discrimination. When such rules are applied to objects of unknown class, the process is denoted classification. The specific problem addressed involves the group classification of a set of objects that are each associated with a series of measurements (ratio, interval, ordinal, or nominal levels of measurement). Each measurement produces one variable in a multidimensional variable space. Cluster analysis techniques are reviewed and methods for incuding geographic location, distance measures, and spatial pattern (distribution) as parameters in clustering are examined. For the case of patterning, measures of spatial autocorrelation are discussed in terms of the kind of data (nominal, ordinal, or interval scaled) to which they may be applied.

Glick, B. J.

Spatio-temporal multivariate cluster evolution analysis for detecting and tracking climate impacts

Recent years have seen a growing concern about climate change and its impacts. While Earth System Models (ESMs) can be invaluable tools for studying the impacts of climate change, the complex coupling processes encoded in ESMs and the large amounts of data produced by these models, together with the high internal variability of the Earth system, can obscure important source-to-impact relationships. Here, this paper presents a novel and efficient unsupervised data-driven approach for detecting statistically-significant impacts and tracing spatio-temporal source-impact pathways in the climate through a unique combination of ideas from anomaly detection, clustering and Natural Language Processing (NLP). Using as an exemplar the 1991 eruption of Mount Pinatubo in the Philippines, we demonstrate that the proposed approach is capable of detecting known post-eruption impacts/events. We additionally describe a methodology for extracting meaningful sequences of post-eruption impacts/events by using NLP to efficiently mine frequent multivariate cluster evolutions, which can be used to confirm or discover the chain of physical processes between a climate source and its impact(s).

Anomaly detection

Preliminary Comparisons of the Information Content and Utility of TM Versus MSS Data

Comparisons were made between subscenes from the first TM scene acquired of the Washington, D.C. area and a MSS scene acquired approximately one year earlier. Three types of analyses were conducted to compare TM and MSS data: a water body analysis, a principal components analysis and a spectral clustering analysis. The water body analysis compared the capability of the TM to the MSS for detecting small uniform targets. Of the 59 ponds located on aerial photographs 34 (58%) were detected by the TM with six commission errors (15%) and 13 (22%) were detected by the MSS with three commission errors (19%). The smallest water body detected by the TM was 16 meters; the smallest detected by the MSS was 40 meters. For the principal components analysis, means and covariance matrices were calculated for each subscene, and principal components images generated and characterized. In the spectral clustering comparison each scene was independently clustered and the clusters were assigned to informational classes. The preliminary comparison indicated that TM data provides enhancements over MSS in terms of (1) small target detection and (2) data dimensionality (even with 4-band data). The extra dimension, partially resultant from TM band 1, appears useful for built-up/non-built-up area separation.

Markham, B. L.

Functional Groups Based on Leaf Physiology: Are they Spatially and Temporally Robust?

The functional grouping hypothesis, which suggests that complexity in ecosystem function can be simplified by grouping species with similar responses, was tested in the Florida scrub habitat. Functional groups were identified based on how species in fire maintained Florida scrub regulate exchange of carbon and water with the atmosphere as indicated by both instantaneous gas exchange measurements and integrated measures of function (%N, delta C-13, delta N-15, C-N ratio). Using cluster analysis, five distinct physiologically-based functional groups were identified in the fire maintained scrub. These functional groups were tested to determine if they were robust spatially, temporally, and with management regime. Analysis of Similarities (ANOSIM), a non-parametric multivariate analysis, indicated that these five physiologically-based groupings were not altered by plot differences (R = -0.115, p = 0.893) or by the three different management regimes; prescribed burn, mechanically treated and burn, and fire-suppressed (R = 0.018, p = 0.349). The physiological groupings also remained robust between the two climatically different years 1999 and 2000 (R = -0.027, p = 0.725). Easy-to-measure morphological characteristics indicating functional groups would be more practical for scaling and modeling ecosystem processes than detailed gas-exchange measurements, therefore we tested a variety of morphological characteristics as functional indicators. A combination of non-parametric multivariate techniques (Hierarchical cluster analysis, non-metric Multi-Dimensional Scaling, and ANOSIM) were used to compare the ability of life form, leaf thickness, and specific leaf area classifications to identify the physiologically-based functional groups. Life form classifications (ANOSIM; R = 0.629, p 0.001) were able to depict the physiological groupings more adequately than either specific leaf area (ANOSIM; R = 0.426, p = 0.001) or leaf thickness (ANOSIM; R 0.344, p 0.001). The ability of life forms to depict the physiological groupings was improved by separating the parasitic Ximenia americana from the shrub category (ANOSIM; R = 0.794, p = 0.001). Therefore, a life form classification including parasites was determined to be a good indicator of the physiological processes of scrub species, and would be a useful method of grouping for scaling physiological processes to the ecosystem level.

Foster, Tammy E.

Detecting Living-off-the-land Attacks Using K-means And Graph Convolutional Networks

The code ingests Zeek logs derived from network packet captures and goes through data preprocessing before it gets passed into a K-Means model that labels each device as either a client or server. Graph Convolutional Network (GCN) model is used to obtain the embeddings to represent the features in lower dimension. Last, K-means cluster analysis is used to cluster the embeddings for each class.

Quach, Anna [Idaho National Laboratory (INL), Idah

Applications of Fuzzy Set Theory to Satellite Soundings

The introduction of an appropriate fuzzy setting for satellite soundings and its application to clustering methods via unimodal fuzzy sets in the future is proposed. Methods of hard clustering analysis and fuzzy partitioned clustering were applied on simulated data with very encouraging results. The proposed clustering technique is discussed. The notion of a unimodal fuzzy set was chosen to represent the partition of a data set for two reasons: (1) it detects all the locations in the vector space where highly concentrated clusters of points exist; and (2) the notion is general enough to represent clusters that exhibit quite general distributions of points. The technique detects all of the existing unimodal fuzzy sets and realizes the maximum separation among them. It is economical in memory space and computational time requirements and also detects groups that are fairly generally distributed in the feature space.

Munteanu, M. J.

Dynamics of cD clusters of galaxies. II: Analysis of seven Abell clusters

We have investigated the dynamics of the seven Abell clusters A193, A399, A401, A1795, A1809, A2063, and A2124, based on redshift data reported previously by us (Hill & Oegerle, (1993)). These papers present the initial results of a survey of cD cluster kinematics, with an emphasis on studying the nature of peculiar velocity cD galaxies and their parent clusters. In the current sample, we find no evidence for significant peculiar cD velocities, with respect to the global velocity distribution. However, the cD in A2063 has a significant (3 sigma) peculiar velocity with respect to galaxies in the inner 1.5 Mpc/h, which is likely due to the merger of a subcluster with A2063. We also find significant evidence for subclustering in A1795, and a marginally peculiar cD velocity with respect to galaxies within approximately 200 kpc/h of the cD. The available x-ray, optical, and galaxy redshift data strongly suggest that a subcluster has merged with A1795. We propose that the subclusters which merged with A1795 and A2063 were relatively small, with shallow potential wells, so that the cooling flows in these clusters were not disrupted. Two-body gravitational models of the A399/401 and A2063/MKW3S systems indicate that A399/401 is a bound pair with a total virial mass of approximately 4 x 10(exp 15) solar mass/h, while A2063 and MKW3S are very unlikely to be bound.

Oegerle, William R.

Cluster Method Analysis of K. S. C. Image

Information obtained from satellite-based systems has moved to the forefront as a method in the identification of many land cover types. Identification of different land features through remote sensing is an effective tool for regional and global assessment of geometric characteristics. Classification data acquired from remote sensing images have a wide variety of applications. In particular, analysis of remote sensing images have special applications in the classification of various types of vegetation. Results obtained from classification studies of a particular area or region serve towards a greater understanding of what parameters (ecological, temporal, etc.) affect the region being analyzed. In this paper, we make a distinction between both types of classification approaches although, focus is given to the unsupervised classification method using 1987 Thematic Mapped (TM) images of Kennedy Space Center.

Rodriguez, Joe, Jr.

Orthopyroxenes as recorders of diogenite petrogenesis: Major and minor element systematics

As a part of our research to better understand magmatic processes in the Eucrite Parent Body, we have initiated an ambitious program of study of major, minor and trace elements in orthopyroxene from diogenites. This paper reports preliminary results for major and minor elements in orthopyroxenes for a suite of 13 diogenites: Aioun El Atrouss, ALH 84001, ALH A 77256, EET 87530, Ellemeet, Garland, Ibbenburen, Johnstown, Manegoan, peckelsheim, Roda, Shalka, and Tatahouine. A companion paper by Shearer et al. reports new trace element data for ALH 84001, ALH A 77256, Ibbenburen, and Tatahouine. We have presently collected over 800 high quality pyroxene microprobe analyses for Si, Al, Ca, Na, Mn, Fe, Mg, Cr, and Ti. The chemical systematics observed for these orthopyroxenes reflect original magmatic mineral/melt partitioning plus later trapped liquid/mineral equilibration, subsolids, exsolution, and mineral/mineral metamorphic reactions. We have therefore avoided, at this point, any attempt to use statistical analysis to group (e.g. factor or cluster analysis) these orthopyroxenes chemically.

Papike, J. J.

Analysis of SETI data collected in the parasitic mode

A system for performing SETI observations continuously as part of non-SETI observations at a radio observatory is presented, and the analysis of results obtained by the parasitic system is discussed. The system, designated SERENDIP, is a real-time microprocessor-controlled spectrum analyzer with algorithms for performing statistical operations and recording those spectra with characteristics presumed to be typical of intelligent rather than astrophysical origin. Programs also exist for the post-acquisition analysis of signal autocorrelation, power spectra, time behavior and positional coordinates, and for a generalized cluster analysis to detect clusters of signal detections. In a recent run of 35 days, the SERENDIP system detected 4000 narrowband spectra exceeding a preset threshold, of which 98% were determined to be of instrumental origin. The remaining class of detections is also believed to be instrumental, although not as well organized as the first signals, and means are currently being sought for eliminating them.

Bowyer, S.

The composite sequential clustering technique for analysis of multispectral scanner data

The clustering technique consists of two parts: (1) a sequential statistical clustering which is essentially a sequential variance analysis, and (2) a generalized K-means clustering. In this composite clustering technique, the output of (1) is a set of initial clusters which are input to (2) for further improvement by an iterative scheme. This unsupervised composite technique was employed for automatic classification of two sets of remote multispectral earth resource observations. The classification accuracy by the unsupervised technique is found to be comparable to that by traditional supervised maximum likelihood classification techniques. The mathematical algorithms for the composite sequential clustering program and a detailed computer program description with job setup are given.

Su, M. Y.

Oleaginous Yeast Biology Elucidated With Comparative Transcriptomics

ABSTRACT Extremophilic yeasts have favorable metabolic and tolerance traits for biomanufacturing‐ like lipid biosynthesis, flavinogenesis, and halotolerance – yet the connection between these favorable phenotypes and strain genotype is not well understood. To this end, this study compares the phenotypes and gene expression patterns of biotechnologically relevant yeasts Yarrowia lipolytica , Debaryomyces hansenii , and Debaryomyces subglobosus grown under nitrogen starvation, iron starvation, and salt stress. To analyze the large data set across species and conditions, two approaches were used: a “network‐first” approach where a generalized metabolic network serves as a scaffold for mapping genes and a “cluster‐first” approach where unsupervised machine learning co‐expression analysis clusters genes. Both approaches provide insight into strain behavior. The network‐first approach corroborates that Yarrowia upregulates lipid biosynthesis during nitrogen starvation and provides new evidence that riboflavin overproduction in Debaryomyces yeasts is overflow metabolism that is routed to flavin cofactor production under salt stress. The cluster‐first approach does not rely on annotation; therefore, the coexpression analysis can identify known and novel genes involved in stress responses, mainly transcription factors and transporters. Therefore, this work links the genotype to the phenotype of biotechnologically relevant yeasts and demonstrates the utility of complementary computational approaches to gain insight from transcriptomics data across species and conditions.

Weintraub, Sarah J. [Department of Bioinformatics

Global Weather States and Their Properties from Passive and Active Satellite Cloud Retrievals

In this study, the authors apply a clustering algorithm to International Satellite Cloud Climatology Project (ISCCP) cloud optical thickness-cloud top pressure histograms in order to derive weather states (WSs) for the global domain. The cloud property distribution within each WS is examined and the geographical variability of each WS is mapped. Once the global WSs are derived, a combination of CloudSat and Cloud-Aerosol Lidar and Infrared Pathfinder Satellite Observations (CALIPSO) vertical cloud structure retrievals is used to derive the vertical distribution of the cloud field within each WS. Finally, the dynamic environment and the radiative signature of the WSs are derived and their variability is examined. The cluster analysis produces a comprehensive description of global atmospheric conditions through the derivation of 11 WSs, each representing a distinct cloud structure characterized by the horizontal distribution of cloud optical depth and cloud top pressure. Matching those distinct WSs with cloud vertical profiles derived from CloudSat and CALIPSO retrievals shows that the ISCCP WSs exhibit unique distributions of vertical layering that correspond well to the horizontal structure of cloud properties. Matching the derived WSs with vertical velocity measurements shows a normal progression in dynamic regime when moving from the most convective to the least convective WS. Time trend analysis of the WSs shows a sharp increase of the fair-weather WS in the 1990s and a flattening of that increase in the 2000s. The fact that the fair-weather WS is the one with the lowest cloud radiative cooling capability implies that this behavior has contributed excess radiative warming to the global radiative budget during the 1990s.

histograms

Enhancing Cluster Identification in Atom Probe Tomography Data Using Transfer Learning

Atom Probe Tomography (APT) is a powerful technique for visualizing the atomic-scale distribution of solutes in materials, but quantitative cluster analysis of APT datasets remains a challenge due to the need for subjective parameter selection in clustering algorithms. While distance-based and density-based methods such as HDBSCAN are widely used, their performance is highly sensitive to user-defined parameters, which undermines reproducibility and accuracy. This study proposes an image-based, deep learning-aided workflow for automating parameter selection and cluster detection in APT data analysis. By projecting 3D APT point clouds onto 2D planes, we leverage pretrained convolutional neural networks (ConvNeXt-Tiny and ResNet-50) through transfer learning to predict the number of clusters present in synthetic datasets. The output is used to guide K-means clustering and estimate HDBSCAN parameters, specifically minimum cluster size and minimum sample points. This approach reduces reliance on manual parameter tuning, improving consistency and scalability. The methodology demonstrates the feasibility of using image-based deep learning for interpreting complex spatial patterns in APT data, enabling faster and more objective analysis. The complete workflow and code are made publicly available to support reproducibility and future research.

Density-based clustering

The NASA Analogy Software Cost Model: A Web-Based Cost Analysis Tool

This paper provides an overview of the many new features and algorithm updates in the release of the NASA Analogy Software Cost Tool (ASCoT). ASCoT is a web-based tool that provides a suite of estimation tools to support early lifecycle NASA Flight Software analysis. ASCoT employs advanced statistical methods such as Cluster Analysis to provide an analogy based estimate of software delivered lines of code and development effort, a regression based Cost Estimating Relationships (CER) model that estimates cost (dollars), and a COCOMO II based estimate. The ASCoT algorithms are designed to primarily work with system level inputs such as mission type (earth orbiter vs. planetary vs. rover), the number of instruments, and total mission cost. This allows the user to supply a minimal number of mission-level parameters which are better understood early in the life-cycle, rather than a large number of complex inputs.

Hihn, Jairus

A CLIPS expert system for clinical flow cytometry data analysis

An expert system is being developed using CLIPS to assist clinicians in the analysis of multivariate flow cytometry data from cancer patients. Cluster analysis is used to find subpopulations representing various cell types in multiple datasets each consisting of four to five measurements on each of 5000 cells. CLIPS facts are derived from results of the clustering. CLIPS rules are based on the expertise of Drs. Stewart, Duque, and Braylan. The rules incorporate certainty factors based on case histories.

Salzman, G. C.

A nonparametric clustering technique which estimates the number of clusters

In applications of cluster analysis, one usually needs to determine the number of clusters, K, and the assignment of observations to each cluster. A clustering technique based on recursive application of a multivariate test of bimodality which automatically estimates both K and the cluster assignments is presented.

Ramey, D. B.