Search NASA⌕ Search

SEARCH · Search NASA

Results for “data analysis cluster”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

Scalable, In-situ Data Clustering Data Analysis for Extreme Scale Scientific Computing (Final Report)

The objective of this project is to address challenges in the design and development of scalable in-situ data clustering and analytics algorithms and software. Our goal is to develop parallel software consisting of a set of spatio-temporal data clustering and anomaly detection functions, both of which are very important for large-scale analysis and have wide applicability for in-situ runs as well as post-processing analysis. Our design principles for in-situ analysis consider the following: (1) identify parts of the computation can be done close to the data within the nodes, while it is still in memory; (2) extract analysis components can (and should) be performed in remote staging and analysis nodes; (3) develop error-bound approximation methods for applications tolerable for small errors; (4) identify the type of derived distributions and statistics, for spatio-temporal data, that can be kept locally in order to both accelerate computations and meet energy constraints in subsequent iterations and phases; (5) use a self-describing data format so that data can be consistent and understood among local storage (memory and SSDs) and at staging and analysis nodes, thereby providing portability and flexibility; (6) develop service-oriented functions that can schedule in-situ and post-hoc analysis tasks based on the dynamic requirements of applications. Our development focus is to produce the parallel data analysis software/library that will be scalable, reusable, extensible, and generic for applications in different disciplines. The software will be able to run in-situ with the simulations as well as post-hoc analysis. This approach will satisfy many synergistic requirements for data intensive applications executed on data coming from instruments and experiments. In particular, the proposed multilevel approach is directly applicable to perform design tradeoffs for running part of the algorithms near the instruments and the rest on remote (analysis) systems.

97 MATHEMATICS AND COMPUTING↗

Clustering High-dimensional Toxicogenomics Data with Rare Signals

Toxicogenomics studies the gene and protein activities to drug treatments or toxic exposures. As the drugs and genes are numerous, toxicogenomics data are naturally high dimensional, with dimension sizes up to millions. In addition, the distribution of toxicogenomics data is oftentimes skewed, and they contain rare but important signals representing a cell or organism’s response to toxicity. The combination of high dimension and extremely skewed distribution of toxicogenomics data makes clustering analysis extremely challenging.We present our study of clustering toxicogenomics data using classical approaches such as principal component analysis as well as deep learning approaches such as auto-encoders. Our experiments show that these approaches fail to preserve rare signals and produce high-quality clusters. We then explore augmenting matrix factorization with deep learning techniques such as attention mechanism to produce latent representations for clustering. Our technique is able to better preserve rare signals after dimensionality reduction than prior approaches. Furthermore, we combine our augmented matrix factorization with a mechanism similar to autoencoder to balance separable clusters and low regeneration errors. Our experiments demonstrate better clustering with our proposed approach.

Cong, Guojing↗

Supporting data for climatic clustering and longitudinal analysis with impacts on food, bioenergy, and pandemics

This data supports the conclusions found in climatic clustering and longitudinal analysis with impacts on food, bioenergy, and pandemics. Included here are (i) the binarized geolocation vectors used for exhaustive vector comparisons, (ii) the resulting climatic networks, (iii) the results of applying Markov clustering to the climatic networks, and (iv) the results of applying Correlation-of-Correlations (cor-cor) to the climatic networks. The set of binarized geolocation vectors that are used as inputs for the Combinatorial Metrics library (CoMet) are of the form comet-UUUUUxVVVVV-XXXX-YYYY.shuffled.tped where UUUUU is the number of vectors, VVVVV is the length of each vector, XXXX is the starting year, and YYYY is the ending year. Each line corresponds to a geolocation vector of binary elements A (i.e., 0) and T (i.e., 1). The set of climatic networks that are used for downstream network analysis are of the form network-U-way-XXXX-YYYY.parsed.txt where U is the order of the comparison (2-way or 3-way), XXXX is the starting year, and YYYY is the ending year. Each line corresponds to an edge linking two geolocations (defined by latitude and longitude) with its corresponding edge weight (i.e., DUO score). The set of cluster results are of the form clusters-U-way-XXXX-YYYY-thresh-VVVV-inflation-WWW.clustered.txt where U is the order of the comparison (2-way or 3-way), XXXX is the starting year, YYYY is the ending year, VVVV is the similarity threshold, and WWW is the Markov clustering inflation rate. Each line corresponds to a single cluster and is composed of a number of corresponding geolocations (defined by latitude and longitude). The set of cor-cor results are of the form corcor-U-way-XXXX-YYYY.cumulative.txt where U is the order of the comparison (2-way or 3-way), XXXX is the starting year, and YYYY is the ending year. Each line corresponds to a single geolocation with it's corresponding cor-cor value.

54 ENVIRONMENTAL SCIENCES↗

Combinatorial Exploration and Mapping of Phase Transformation in a Ni–Ti–Co Thin Film Library

Combinatorial synthesis and high-throughput characterization of a Ni–Ti–Co thin film materials library are reported for exploration of reversible martensitic transformation. The library was prepared by magnetron co-sputtering, annealed in vacuum at 500 °C without atmospheric exposure, and evaluated for shape memory behavior as an indicator of transformation. Composition, structure, and transformation behavior of the 177 pads in the library were characterized using high-throughput wavelength dispersive spectroscopy (WDS), X-ray photoelectron spectroscopy (XPS), X-ray diffraction (XRD), and four-point probe temperature-dependent resistance (R(T)) measurements. A new, expanded composition space having phase transformation with low thermal hysteresis and Co > 10 at. % is found. Unsupervised machine learning methods of hierarchical clustering were employed to streamline data processing of the large XRD and XPS data sets. Through cluster analysis of XRD data, we identified and mapped the constituent structural phases. Finally, composition–structure–property maps for the ternary system are made to correlate the functional properties to the local microstructure and composition of the Ni–Ti–Co thin film library.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Differentially Private K -Means Clustering Applied to Meter Data Analysis and Synthesis

The proliferation of smart meters has resulted in a large amount of data being generated. It is increasingly apparent that methods are required for allowing a variety of stakeholders to leverage the data in a manner that preserves the privacy of the consumers. The sector is scrambling to define policies, such as the so called ‘15/15 rule’, to respond to the need. However, the current policies fail to adequately guarantee privacy. Here, in this paper, we address the problem of allowing third parties to apply K-means clustering, obtaining customer labels and centroids for a set of load time series by applying the framework of differential privacy. We leverage the method to design an algorithm that generates differentially private synthetic load data consistent with the labeled data. We test our algorithm’s utility by answering summary statistics such as average daily load profiles for a 2-dimensional synthetic dataset and a real-world power load dataset.

24 POWER TRANSMISSION AND DISTRIBUTION↗

Crowd cluster data in the USA for analysis of human response to COVID-19 events and policies

We provide data on daily social contact intensity of clusters of people at different types of Points of Interest (POI) by zip code in Florida and California. This data is obtained by aggregating fine-scaled details of interactions of people at the spatial resolution of 10 m, which is then normalized as a social contact index. We also provide the distribution of cluster sizes and average time spent in a cluster by POI type. This data will help researchers perform fine-scaled, privacy-preserving analysis of human interaction patterns to understand the drivers of the COVID-19 epidemic spread and mitigation. Current mobility datasets either provide coarse-level metrics of social distancing, such as radius of gyration at the county or province level, or traffic at a finer scale, neither of which is a direct measure of contacts between people. We use anonymized, de-identified, and privacy-enhanced location-based services (LBS) data from opted-in cell phone apps, suitably reweighted to correct for geographic heterogeneities, and identify clusters of people at non-sensitive public areas to estimate fine-scaled contacts.

60 APPLIED LIFE SCIENCES↗

Cluster Analysis of Combined EDS and EBSD Data to Solve Ambiguous Phase Identifications

A common problem in analytical scanning electron microscopy (SEM) using electron backscatter diffraction (EBSD) is the differentiation of phases with distinct chemistry but the same or very similar crystal structure. X-ray energy dispersive spectroscopy (EDS) is useful to help differentiate these phases of similar crystal structures but different elemental makeups. However, open, automated, and unbiased methods of differentiating phases of similar EBSD responses based on their EDS response are lacking. This paper describes a simple data analytics-based method, using a combination of singular value decomposition and cluster analysis, to merge simultaneously acquired EDS + EBSD information and automatically determine phases from both their crystal and elemental data. I use hexagonal TiB 2 ceramic contaminated with multiple crystallographically ambiguous but chemically distinct cubic phases to illustrate the method. Code, in the form of a Python 3 Jupyter Notebook, and the necessary data to replicate the analysis are provided as Supplementary material.

47 OTHER INSTRUMENTATION↗

AutoEnRichness: A hybrid empirical and analytical approach for estimating the richness of galaxy clusters

ABSTRACT We introduce AutoEnRichness, a hybrid approach that combines empirical and analytical strategies to determine the richness of galaxy clusters (in the redshift range of 0.1 ≤ z ≤ 0.35) using photometry data from the Sloan Digital Sky Survey Data Release 16, where cluster richness can be used as a proxy for cluster mass. In order to reliably estimate cluster richness, it is vital that the background subtraction is as accurate as possible when distinguishing cluster and field galaxies to mitigate severe contamination. AutoEnRichness is comprised of a multistage machine learning algorithm that performs background subtraction of interloping field galaxies along the cluster line of sight and a conventional luminosity distribution fitting approach that estimates cluster richness based only on the number of galaxies within a magnitude range and search area. In this proof-of-concept study, we obtain a balanced accuracy of 83.20 per cent when distinguishing between cluster and field galaxies as well as a median absolute percentage error of 33.50 per cent between our estimated cluster richnesses and known cluster richnesses within r200. In the future, we aim for AutoEnRichness to be applied on upcoming large-scale optical surveys, such as the Legacy Survey of Space and Time and Euclid, to estimate the richness of a large sample of galaxy groups and clusters from across the halo mass function. This would advance our overall understanding of galaxy evolution within overdense environments as well as enable cosmological parameters to be further constrained.

79 ASTRONOMY AND ASTROPHYSICS↗

Estimating cluster masses from SDSS multiband images with transfer learning

ABSTRACT The total masses of galaxy clusters characterize many aspects of astrophysics and the underlying cosmology. It is crucial to obtain reliable and accurate mass estimates for numerous galaxy clusters over a wide range of redshifts and mass scales. We present a transfer-learning approach to estimate cluster masses using the ugriz-band images in the SDSS Data Release 12. The target masses are derived from X-ray or SZ measurements that are only available for a small subset of the clusters. We designed a semisupervised deep learning model consisting of two convolutional neural networks. In the first network, a feature extractor is trained to classify the SDSS photometric bands. The second network takes the previously trained features as inputs to estimate their total masses. The training and testing processes in this work depend purely on real observational data. Our algorithm reaches a mean absolute error (MAE) of 0.232 dex on average and 0.214 dex for the best fold. The performance is comparable to that given by redMaPPer, 0.192 dex. We have further applied a joint integrated gradient and class activation mapping method to interpret such a two-step neural network. The performance of our algorithm is likely to improve as the size of training data set increases. This proof-of-concept experiment demonstrates the potential of deep learning in maximizing the scientific return of the current and future large cluster surveys.

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND ↗

Cosmological constraints from DES Y1 cluster abundances and SPT multiwavelength data

We perform a joint analysis of the counts of redMaPPer clusters selected from the Dark Energy Survey (DES) year 1 data and multiwavelength follow-up data collected within the 2500 deg2 South Pole Telescope (SPT) Sunyaev-Zel’dovich (SZ) survey. The SPT follow-up data, calibrating the richness-mass relation of the optically selected redMaPPer catalog, enable the cosmological exploitation of the DES cluster abundance data. To explore possible systematics related to the modeling of projection effects, we consider two calibrations of the observational scatter on richness estimates: a simple Gaussian model which account only for the background contamination (BKG), and a model which further includes contamination and incompleteness due to projection effects (PRJ). Assuming either a ΛCDM+∑mν or wCDM+∑mν cosmology, and for both scatter models, we derive cosmological constraints consistent with multiple cosmological probes of the low and high redshift Universe, and in particular with the SPT cluster abundance data. This result demonstrates that the DES Y1 and SPT cluster counts provide consistent cosmological constraints, if the same mass calibration data set is adopted. It thus supports the conclusion of the DES Y1 cluster cosmology analysis which interprets the tension observed with other cosmological probes in terms of systematics affecting the stacked weak lensing analysis of optically selected low–richness clusters. Finally, we analyze the first combined optically SZ selected cluster catalog obtained by including the SPT sample above the maximum redshift probed by the DES Y1 redMaPPer sample (z=0.65). Besides providing a mild improvement of the cosmological constraints, this data combination serves as a stricter test of our scatter models: the PRJ model, providing scaling relations consistent between the two abundance and multiwavelength follow-up data, is favored over the BKG model.

79 ASTRONOMY AND ASTROPHYSICS↗

A Parameter-masked Mock Data Challenge for Beyond-two-point Galaxy Clustering Statistics

The past few years have seen the emergence of a wide array of novel techniques for analyzing high-precision data from upcoming galaxy surveys, which aim to extend the statistical analysis of galaxy clustering data beyond the linear regime and the canonical two-point (2pt) statistics. We test and benchmark some of these new techniques in a community data challenge named “Beyond-2pt,” initiated during the Aspen 2022 Summer Program “Large-Scale Structure Cosmology beyond 2-Point Statistics,” whose first round of results we present here. The challenge data set consists of high-precision mock galaxy catalogs for clustering in real space, in redshift space, and on a light cone. Participants in the challenge have developed end-to-end pipelines to analyze mock catalogs and extract unknown (“masked”) cosmological parameters of the underlying ΛCDM models with their methods. The methods represented are density-split clustering, nearest neighbor statistics, BACCO power spectrum emulator, void statistics, LEFTfield field-level inference using effective field theory (EFT), and joint power spectrum and bispectrum analyses using both EFT and simulation-based inference. In this work, we review the results of the challenge, focusing on problems solved, lessons learned, and future research needed to perfect the emerging beyond-2pt approaches. The unbiased parameter recovery demonstrated in this challenge by multiple statistics and the associated modeling and inference frameworks supports the credibility of cosmology constraints from these methods. The challenge data set is publicly available, and we welcome future submissions from methods that are not yet represented.

Krause, Elisabeth [Univ. of Arizona, Tucson, AZ (U↗

SITCOMTN-161: PSF assessment in the field of Abell 360 and shapeHSM shear profile using LSSTComCam data

The Rubin LSSTComCam on-sky campaign performed at the end of 2024 provided observations of the Abell 360 galaxy cluster; these data allow a preliminary study of cluster weak lensing analysis using Rubin Data Preview 1 (DP1) data. Among all the steps required for such analyses, accurate modeling of the PSF is essential. This work uses several diagnostics, mostly based on the residuals between the second moments of stars and the PSF model, to characterize the accuracy of the PSF modeling in the A360 field. We find the level of the residuals to be sufficiently low not to hinder the measurement of the tangential shear profile around A360. With a simple source selection process, we demonstrate that outputs of the LSST Science Pipelines can be used to detect the tangential shear profile in Abell 360 at the 3.6σ level, and our analysis indicates that contamination from PSF modeling systematics is negligible.

Dell'Antonio, Ian [Brown University]↗

BiG-SLiCE 2 v1.0.0

BiG-SLiCE was originally an open source Python-based command line bioinformatics software that offers a highly scalable clustering analysis on biosynthetic gene clusters (BGC) data. It allows a simultaneous analysis of millions of BGCs, exceeding the capability of other existing tools (around one hundred thousands). As a tradeoff, the clustering accuracy is relatively lower and sometimes fall short in corner cases and specific BGC classes such as the RiPPs (Ribosomally-translated, Post-translationally modified Peptides). In BiG-SLiCE V2 (developed in LBNL), the clustering algorithm has been significantly improved to deliver a much accurate result even for RiPPs and other previous corner case classes. Moreover, the speed of the overall pipeline has been improved by 50-100%. Finally, additional features were implemented to support downstream analyses of BiG-SLiCE results, such as customized tabular (TSV/CSV) and columnar (Parquet) outputs.

Kautsar, Satria↗

Web-based wide-area monitoring platform for ringdown and clustering analytics in power systems

This paper introduces an open-source research platform for monitoring the Mexican interconnected power grid, allowing real-time processing and information extraction of the grid’s dynamic condition. Moreover, the platform is a Python-based development that embeds different ringdown and clustering analytics tools. In the case of ringdown analysis, the modal information can be extracted using some of the most known algorithms, i.e., Prony analysis, eigensystem realization algorithm (ERA), and matrix pencil (MP). For clustering analysis, the coherent behaviour of generator and non-generator buses is provided by applying recent state-of-the-art techniques such as affinity propagation, K-means, hierarchical agglomerative clustering, and typicality data analysis. The results of up to 93 PMUs show that this open-source platform suits researchers’ and engineers’ power system dynamic analysis requirements.

Clustering↗

Estimating Cosmological Constraints from Galaxy Cluster Abundance using Simulation-Based Inference

Inferring the values and uncertainties of cosmological parameters in a cosmology model is of paramount importance for modern cosmic observations. In this paper, we use the simulation-based inference (SBI) approach to estimate cosmological constraints from a simplified galaxy cluster observation analysis. Using data generated from the Quijote simulation suite and analytical models, we train a machine learning algorithm to learn the probability function between cosmological parameters and the possible galaxy cluster observables. The posterior distribution of the cosmological parameters at a given observation is then obtained by sampling the predictions from the trained algorithm. Our results show that the SBI method can successfully recover the truth values of the cosmological parameters within the 2σ limit for this simplified galaxy cluster analysis, and acquires similar posterior constraints obtained with a likelihood-based Markov Chain Monte Carlo method, the current state-of the-art method used in similar cosmological studies.

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND ↗

Enhancing Cluster Identification in Atom Probe Tomography Data Using Transfer Learning

Atom Probe Tomography (APT) is a powerful technique for visualizing the atomic-scale distribution of solutes in materials, but quantitative cluster analysis of APT datasets remains a challenge due to the need for subjective parameter selection in clustering algorithms. While distance-based and density-based methods such as HDBSCAN are widely used, their performance is highly sensitive to user-defined parameters, which undermines reproducibility and accuracy. This study proposes an image-based, deep learning-aided workflow for automating parameter selection and cluster detection in APT data analysis. By projecting 3D APT point clouds onto 2D planes, we leverage pretrained convolutional neural networks (ConvNeXt-Tiny and ResNet-50) through transfer learning to predict the number of clusters present in synthetic datasets. The output is used to guide K-means clustering and estimate HDBSCAN parameters, specifically minimum cluster size and minimum sample points. This approach reduces reliance on manual parameter tuning, improving consistency and scalability. The methodology demonstrates the feasibility of using image-based deep learning for interpreting complex spatial patterns in APT data, enabling faster and more objective analysis. The complete workflow and code are made publicly available to support reproducibility and future research.

Density-based clustering↗

Machine-learning identification of the variability of mean velocity and turbulence intensity for wakes generated by onshore wind turbines: Cluster analysis of wind LiDAR measurements

Light detection and ranging (LiDAR) measurements of isolated wakes generated by wind turbines installed at an onshore wind farm are leveraged to characterize the variability of the wake mean velocity and turbulence intensity during typical operations, which encompass a breadth of atmospheric stability regimes and rotor thrust coefficients. The LiDAR measurements are clustered through the k-means algorithm, which enables identifying the most representative realizations of wind turbine wakes while avoiding the imposition of thresholds for the various wind and turbine parameters. Considering the large number of LiDAR samples collected to probe the wake velocity field, the dimensionality of the experimental dataset is reduced by projecting the LiDAR data on an intelligently truncated basis obtained with the proper orthogonal decomposition (POD). The coefficients of only five physics-informed POD modes are then injected in the k-means algorithm for clustering the LiDAR dataset. The analysis of the clustered LiDAR data and the associated supervisory control and data acquisition and meteorological data enables the study of the variability of the wake velocity deficit, wake extent, and wake-added turbulence intensity for different thrust coefficients of the turbine rotor and regimes of atmospheric stability. Furthermore, the cluster analysis of the LiDAR data allows for the identification of systematic off-design operations with a certain yaw misalignment of the turbine rotor with the mean wind direction.

17 WIND ENERGY↗