Search NASA⌕ Search

SEARCH · Search NASA

Results for “data analysis cluster”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

Scalable, In-situ Data Clustering Data Analysis for Extreme Scale Scientific Computing (Final Report)

The objective of this project is to address challenges in the design and development of scalable in-situ data clustering and analytics algorithms and software. Our goal is to develop parallel software consisting of a set of spatio-temporal data clustering and anomaly detection functions, both of which are very important for large-scale analysis and have wide applicability for in-situ runs as well as post-processing analysis. Our design principles for in-situ analysis consider the following: (1) identify parts of the computation can be done close to the data within the nodes, while it is still in memory; (2) extract analysis components can (and should) be performed in remote staging and analysis nodes; (3) develop error-bound approximation methods for applications tolerable for small errors; (4) identify the type of derived distributions and statistics, for spatio-temporal data, that can be kept locally in order to both accelerate computations and meet energy constraints in subsequent iterations and phases; (5) use a self-describing data format so that data can be consistent and understood among local storage (memory and SSDs) and at staging and analysis nodes, thereby providing portability and flexibility; (6) develop service-oriented functions that can schedule in-situ and post-hoc analysis tasks based on the dynamic requirements of applications. Our development focus is to produce the parallel data analysis software/library that will be scalable, reusable, extensible, and generic for applications in different disciplines. The software will be able to run in-situ with the simulations as well as post-hoc analysis. This approach will satisfy many synergistic requirements for data intensive applications executed on data coming from instruments and experiments. In particular, the proposed multilevel approach is directly applicable to perform design tradeoffs for running part of the algorithms near the instruments and the rest on remote (analysis) systems.

97 MATHEMATICS AND COMPUTING↗

Adaptive on-line classification of multi-spectral scanner data

A possible solution to the analysis of the massive amounts of multi-spectral scanner data from the Earth Resource Technical Satellite (ERTS) program is proposed. This solution is offered as an adaptive on-line classification scheme. The classifier is described as well as its controller which is based on ground truth data. Cluster analysis is presented as an alternative approach to the ground truth data. Adaptive feature selection is discussed and possible mini-computer implementations are offered.

Fromm, F. R.↗

ICAP - An Interactive Cluster Analysis Procedure for analyzing remotely sensed data

An Interactive Cluster Analysis Procedure (ICAP) was developed to derive classifier training statistics from remotely sensed data. ICAP differs from conventional clustering algorithms by allowing the analyst to optimize the cluster configuration by inspection, rather than by manipulating process parameters. Control of the clustering process alternates between the algorithm, which creates new centroids and forms clusters, and the analyst, who can evaluate and elect to modify the cluster structure. Clusters can be deleted, or lumped together pairwise, or new centroids can be added. A summary of the cluster statistics can be requested to facilitate cluster manipulation. The principal advantage of this approach is that it allows prior information (when available) to be used directly in the analysis, since the analyst interacts with ICAP in a straightforward manner, using basic terms with which he is more likely to be familiar. Results from testing ICAP showed that an informed use of ICAP can improve classification, as compared to an existing cluster analysis procedure.

Wharton, S. W.↗

ICAP: An Interactive Cluster Analysis Procedure for analyzing remotely sensed data

An Interactive Cluster Analysis Procedure (ICAP) was developed to derive classifier training statistics from remotely sensed data. The algorithm interfaces the rapid numerical processing capacity of a computer with the human ability to integrate qualitative information. Control of the clustering process alternates between the algorithm, which creates new centroids and forms clusters and the analyst, who evaluate and elect to modify the cluster structure. Clusters can be deleted or lumped pairwise, or new centroids can be added. A summary of the cluster statistics can be requested to facilitate cluster manipulation. The ICAP was implemented in APL (A Programming Language), an interactive computer language. The flexibility of the algorithm was evaluated using data from different LANDSAT scenes to simulate two situations: one in which the analyst is assumed to have no prior knowledge about the data and wishes to have the clusters formed more or less automatically; and the other in which the analyst is assumed to have some knowledge about the data structure and wishes to use that information to closely supervise the clustering process. For comparison, an existing clustering method was also applied to the two data sets.

Wharton, S. W.↗

NCUBE - A clustering algorithm based on a discretized data space

Cluster analysis involves the unsupervised grouping of data. The process provides an automatic procedure for generating known training samples for pattern classification. NCUBE, the clustering algorithm presented, is based upon the concept of imposing a gridwork on the data space. The NCUBE computer implementation of this concept provides an easily derived form of piecewise linear discrimination. This piecewise linear discrimination permits the separation of some types of data groups that are not linearly separable.

Eigen, D. J.↗

The composite sequential clustering technique for analysis of multispectral scanner data

The clustering technique consists of two parts: (1) a sequential statistical clustering which is essentially a sequential variance analysis, and (2) a generalized K-means clustering. In this composite clustering technique, the output of (1) is a set of initial clusters which are input to (2) for further improvement by an iterative scheme. This unsupervised composite technique was employed for automatic classification of two sets of remote multispectral earth resource observations. The classification accuracy by the unsupervised technique is found to be comparable to that by traditional supervised maximum likelihood classification techniques. The mathematical algorithms for the composite sequential clustering program and a detailed computer program description with job setup are given.

Su, M. Y.↗

Analysis of RXTE data on Clusters of Galaxies

This grant provided support for the reduction, analysis and interpretation of of hard X-ray (HXR, for short) observations of the cluster of galaxies RXJO658--5557 scheduled for the week of August 23, 2002 under the RXTE Cycle 7 program (PI Vahe Petrosian, Obs. ID 70165). The goal of the observation was to search for and characterize the shape of the HXR component beyond the well established thermal soft X-ray (SXR) component. Such hard components have been detected in several nearby clusters. distant cluster would provide information on the characteristics of this radiation at a different epoch in the evolution of the imiverse and shed light on its origin. We (Petrosian, 2001) have argued that thermal bremsstrahlung, as proposed earlier, cannot be the mechanism for the production of the HXRs and that the most likely mechanism is Compton upscattering of the cosmic microwave radiation by relativistic electrons which are known to be present in the clusters and be responsible for the observed radio emission. Based on this picture we estimated that this cluster, in spite of its relatively large distance, will have HXR signal comparable to the other nearby ones. The planned observation of a relatively The proposed RXTE observations were carried out and the data have been analyzed. We detect a hard X-ray tail in the spectrum of this cluster with a flux very nearly equal to our predicted value. This has strengthen the case for the Compton scattering model. We intend the data obtained via this observation to be a part of a larger data set. We have identified other clusters of galaxies (in archival RXTE and other instrument data sets) with sufficiently high quality data where we can search for and measure (or at least put meaningful limits) on the strength of the hard component. With these studies we expect to clarify the mechanism for acceleration of particles in the intercluster medium and provide guidance for future observations of this intriguing phenomenon by instrument on GLAST. The details of the nonthermal particle population has important implications for the theories of cluster formation, mergers and evolution. The results of this work were first presented at the High Energy Division meeting of the American astronomical Society at Mt. Tremblene, Canada (Petrosian et al. 2003). and in an invited review talk at the General Assembly of the International Astronomical Union at Sydney, Australia (Petrosian, 2003). A paper describe the observations, the data analysis and its implication is being prepared for publication in the Astrophysical Journal.

Petrosian, Vahe↗

Clustering High-dimensional Toxicogenomics Data with Rare Signals

Toxicogenomics studies the gene and protein activities to drug treatments or toxic exposures. As the drugs and genes are numerous, toxicogenomics data are naturally high dimensional, with dimension sizes up to millions. In addition, the distribution of toxicogenomics data is oftentimes skewed, and they contain rare but important signals representing a cell or organism’s response to toxicity. The combination of high dimension and extremely skewed distribution of toxicogenomics data makes clustering analysis extremely challenging.We present our study of clustering toxicogenomics data using classical approaches such as principal component analysis as well as deep learning approaches such as auto-encoders. Our experiments show that these approaches fail to preserve rare signals and produce high-quality clusters. We then explore augmenting matrix factorization with deep learning techniques such as attention mechanism to produce latent representations for clustering. Our technique is able to better preserve rare signals after dimensionality reduction than prior approaches. Furthermore, we combine our augmented matrix factorization with a mechanism similar to autoencoder to balance separable clusters and low regeneration errors. Our experiments demonstrate better clustering with our proposed approach.

Cong, Guojing↗

Supporting data for climatic clustering and longitudinal analysis with impacts on food, bioenergy, and pandemics

This data supports the conclusions found in climatic clustering and longitudinal analysis with impacts on food, bioenergy, and pandemics. Included here are (i) the binarized geolocation vectors used for exhaustive vector comparisons, (ii) the resulting climatic networks, (iii) the results of applying Markov clustering to the climatic networks, and (iv) the results of applying Correlation-of-Correlations (cor-cor) to the climatic networks. The set of binarized geolocation vectors that are used as inputs for the Combinatorial Metrics library (CoMet) are of the form comet-UUUUUxVVVVV-XXXX-YYYY.shuffled.tped where UUUUU is the number of vectors, VVVVV is the length of each vector, XXXX is the starting year, and YYYY is the ending year. Each line corresponds to a geolocation vector of binary elements A (i.e., 0) and T (i.e., 1). The set of climatic networks that are used for downstream network analysis are of the form network-U-way-XXXX-YYYY.parsed.txt where U is the order of the comparison (2-way or 3-way), XXXX is the starting year, and YYYY is the ending year. Each line corresponds to an edge linking two geolocations (defined by latitude and longitude) with its corresponding edge weight (i.e., DUO score). The set of cluster results are of the form clusters-U-way-XXXX-YYYY-thresh-VVVV-inflation-WWW.clustered.txt where U is the order of the comparison (2-way or 3-way), XXXX is the starting year, YYYY is the ending year, VVVV is the similarity threshold, and WWW is the Markov clustering inflation rate. Each line corresponds to a single cluster and is composed of a number of corresponding geolocations (defined by latitude and longitude). The set of cor-cor results are of the form corcor-U-way-XXXX-YYYY.cumulative.txt where U is the order of the comparison (2-way or 3-way), XXXX is the starting year, and YYYY is the ending year. Each line corresponds to a single geolocation with it's corresponding cor-cor value.

54 ENVIRONMENTAL SCIENCES↗

A method of using cluster analysis to study statistical dependence in multivariate data

A technique is presented that uses both cluster analysis and a Monte Carlo significance test of clusters to discover associations between variables in multidimensional data. The method is applied to an example of a noisy function in three-dimensional space, to a sample from a mixture of three bivariate normal distributions, and to the well-known Fisher's Iris data.

Borucki, W. J.↗

Combinatorial Exploration and Mapping of Phase Transformation in a Ni–Ti–Co Thin Film Library

Combinatorial synthesis and high-throughput characterization of a Ni–Ti–Co thin film materials library are reported for exploration of reversible martensitic transformation. The library was prepared by magnetron co-sputtering, annealed in vacuum at 500 °C without atmospheric exposure, and evaluated for shape memory behavior as an indicator of transformation. Composition, structure, and transformation behavior of the 177 pads in the library were characterized using high-throughput wavelength dispersive spectroscopy (WDS), X-ray photoelectron spectroscopy (XPS), X-ray diffraction (XRD), and four-point probe temperature-dependent resistance (R(T)) measurements. A new, expanded composition space having phase transformation with low thermal hysteresis and Co > 10 at. % is found. Unsupervised machine learning methods of hierarchical clustering were employed to streamline data processing of the large XRD and XPS data sets. Through cluster analysis of XRD data, we identified and mapped the constituent structural phases. Finally, composition–structure–property maps for the ternary system are made to correlate the functional properties to the local microstructure and composition of the Ni–Ti–Co thin film library.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Determining the Number of Clusters in a Data Set Without Graphical Interpretation

Cluster analysis is a data mining technique that is meant ot simplify the process of classifying data points. The basic clustering process requires an input of data points and the number of clusters wanted. The clustering algorithm will then pick starting C points for the clusters, which can be either random spatial points or random data points. It then assigns each data point to the nearest C point where "nearest usually means Euclidean distance, but some algorithms use another criterion. The next step is determining whether the clustering arrangement this found is within a certain tolerance. If it falls within this tolerance, the process ends. Otherwise the C points are adjusted based on how many data points are in each cluster, and the steps repeat until the algorithm converges,

Aguirre, Nathan S.↗

Differentially Private K -Means Clustering Applied to Meter Data Analysis and Synthesis

The proliferation of smart meters has resulted in a large amount of data being generated. It is increasingly apparent that methods are required for allowing a variety of stakeholders to leverage the data in a manner that preserves the privacy of the consumers. The sector is scrambling to define policies, such as the so called ‘15/15 rule’, to respond to the need. However, the current policies fail to adequately guarantee privacy. Here, in this paper, we address the problem of allowing third parties to apply K-means clustering, obtaining customer labels and centroids for a set of load time series by applying the framework of differential privacy. We leverage the method to design an algorithm that generates differentially private synthetic load data consistent with the labeled data. We test our algorithm’s utility by answering summary statistics such as average daily load profiles for a 2-dimensional synthetic dataset and a real-world power load dataset.

24 POWER TRANSMISSION AND DISTRIBUTION↗

Editing ERTS-1 data to exclude land aids cluster analysis of water targets

The author has identified the following significant results. It has been determined that an increase in the number of spectrally distinct coastal water types is achieved when data values over the adjacent land areas are excluded from the processing routine. This finding resulted from an automatic clustering analysis of ERTS-1 system corrected MSS scene 1002-18134 of 25 July 1972 over Monterey Bay, California. When the entire study area data set was submitted to the clustering only two distinct water classes were extracted. However, when the land area data points were removed from the data set and resubmitted to the clustering routine, four distinct groupings of water features were identified. Additionally, unlike the previous separation, the four types could be correlated to features observable in the associated ERTS-1 imagery. This exercise demonstrates that by proper selection of data submitted to the processing routine, based upon the specific application of study, additional information may be extracted from the ERTS-1 MSS data.

Erb, R. B.↗

Automated clustering-based workload characterization

The demands placed on the mass storage systems at various federal agencies and national laboratories are continuously increasing in intensity. This forces system managers to constantly monitor the system, evaluate the demand placed on it, and tune it appropriately using either heuristics based on experience or analytic models. Performance models require an accurate workload characterization. This can be a laborious and time consuming process. It became evident from our experience that a tool is necessary to automate the workload characterization process. This paper presents the design and discusses the implementation of a tool for workload characterization of mass storage systems. The main features of the tool discussed here are: (1)Automatic support for peak-period determination. Histograms of system activity are generated and presented to the user for peak-period determination; (2) Automatic clustering analysis. The data collected from the mass storage system logs is clustered using clustering algorithms and tightness measures to limit the number of generated clusters; (3) Reporting of varied file statistics. The tool computes several statistics on file sizes such as average, standard deviation, minimum, maximum, frequency, as well as average transfer time. These statistics are given on a per cluster basis; (4) Portability. The tool can easily be used to characterize the workload in mass storage systems of different vendors. The user needs to specify through a simple log description language how the a specific log should be interpreted. The rest of this paper is organized as follows. Section two presents basic concepts in workload characterization as they apply to mass storage systems. Section three describes clustering algorithms and tightness measures. The following section presents the architecture of the tool. Section five presents some results of workload characterization using the tool.Finally, section six presents some concluding remarks.

Pentakalos, Odysseas I.↗

Compact near-IR and mid-IR cavity ring down spectroscopy device

This invention relates to a compact cavity ring down spectrometer for detection and measurement of trace species in a sample gas using a tunable solid-state continuous-wave mid-infrared PPLN OPO laser or a tunable low-power solid-state continuous wave near-infrared diode laser with an algorithm for reducing the periodic noise in the voltage decay signal which subjects the data to cluster analysis or by averaging of the interquartile range of the data.

Miller, J. Houston↗

Crowd cluster data in the USA for analysis of human response to COVID-19 events and policies

We provide data on daily social contact intensity of clusters of people at different types of Points of Interest (POI) by zip code in Florida and California. This data is obtained by aggregating fine-scaled details of interactions of people at the spatial resolution of 10 m, which is then normalized as a social contact index. We also provide the distribution of cluster sizes and average time spent in a cluster by POI type. This data will help researchers perform fine-scaled, privacy-preserving analysis of human interaction patterns to understand the drivers of the COVID-19 epidemic spread and mitigation. Current mobility datasets either provide coarse-level metrics of social distancing, such as radius of gyration at the county or province level, or traffic at a finer scale, neither of which is a direct measure of contacts between people. We use anonymized, de-identified, and privacy-enhanced location-based services (LBS) data from opted-in cell phone apps, suitably reweighted to correct for geographic heterogeneities, and identify clusters of people at non-sensitive public areas to estimate fine-scaled contacts.

60 APPLIED LIFE SCIENCES↗

Clustering with Missing Values: No Imputation Required

Clustering algorithms can identify groups in large data sets, such as star catalogs and hyperspectral images. In general, clustering methods cannot analyze items that have missing data values. Common solutions either fill in the missing values (imputation) or ignore the missing data (marginalization). Imputed values are treated as just as reliable as the truly observed data, but they are only as good as the assumptions used to create them. In contrast, we present a method for encoding partially observed features as a set of supplemental soft constraints and introduce the KSC algorithm, which incorporates constraints into the clustering process. In experiments on artificial data and data from the Sloan Digital Sky Survey, we show that soft constraints are an effective way to enable clustering with missing values.

constraints↗