Search NASA⌕ Search

SEARCH · Search NASA

Results for “Cluster algorithms”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

Clustering Algorithm for AM Parts using GSH and EDT with Autoencoder

SAND2025-10103O The Clustering Algorithm for AM Parts Using GSH and (EDT With Autoencoder is a software tool. It uses a clustering algorithm for additive manufacturing (AM) parts using generalized spherical harmonics (GSH) and Euclidean distance transform (EDT) with an autoencoder to quantify material microstructure. The tool offers improved sensitivity to microstructural changes compared to traditional approaches. The tool integrates multiple microstructural properties, such as grain morphology, crystallographic orientation, and material phase information, to provide a comprehensive analysis of material microstructures. Sandia National Laboratories is a multimission laboratory managed and operated by National Technology & Engineering Solutions of Sandia, LLC, a wholly owned subsidiary of Honeywell International Inc., for the U.S. Department of Energy’s National Nuclear Security Administration under contract DE-NA0003525.

Rodgers, Theron [Sandia National Lab. (SNL-CA), Li↗

Minijet clustering algorithm using transverse-momentum seeds in high-energy nuclear collisions

We propose an algorithm to detect mini-jet clusters in high-energy nuclear collisions, by selecting a high-transverse-momentum (pT) particle as a seed and assigning a clustering radius (R) in the pseudorapidity and azimuthal-angle space. Our PYTHIA simulations for p+p collisions show that a scheme with a seeding p T of around 0.5 GeV/c and R of approximately 0.6 satisfactorily identifies mini-jet clusters. The correlation between clusters obtained in PYTHIA calculations using the algorithm exhibits the proper behavior of hard-scattering-like processes, suggesting its usefulness in isolating mini-jet-like clusters from non-hard-scattering soft processes when applied to actual nuclear-collision data, thereby allowing a closer examination of both the mini-jet and the soft mechanisms.

73 NUCLEAR PHYSICS AND RADIATION PHYSICS↗

Optimal Electrification Using Renewable Energies: Microgrid Installation Model with Combined Mixture k-Means Clustering Algorithm, Mixed Integer Linear Programming, and Onsset Method

Optimal planning and design of microgrids are priorities in the electrification of off-grid areas. Indeed, in one of the Sustainable Development Goals (SDG 7), the UN recommends universal access to electricity for all at the lowest cost. Several optimization methods with different strategies have been proposed in the literature as ways to achieve this goal. This paper proposes a microgrid installation and planning model based on a combination of several techniques. The programming language Python 3.10 was used in conjunction with machine learning techniques such as unsupervised learning based on K-means clustering and deterministic optimization methods based on mixed linear programming. These methods were complemented by the open-source spatial method for optimal electrification planning: onsset. Four levels of study were carried out. The first level consisted of simulating the model obtained with a cluster, which is considered based on the elbow and k-means clustering method as a case study. The second level involved sizing the microgrid with a capacity of 40 kW and optimizing all the resources available on site. The example of the different resources in the Togo case was considered. At the third level, the work consisted of proposing an optimal connection model for the microgrid based on voltage stability constraints and considering, above all, the capacity limit of the source substation. Finally, the fourth level involved a planning study of electrification strategies based mainly on microgrids according to the study scenario. The results of the first level of study enabled us to obtain an optimal location for the centroid of the cluster under consideration, according to the different load positions of this cluster. Then, the results of the second level of study were used to highlight the optimal resources obtained and proposed by the optimization model formulated based on the various technology costs, such as investment, maintenance, and operating costs, which were based on the technical limits of the various technologies. In these results, solar systems account for 80% of the maximum load considered, compared to 7.5% for wind systems and 12.5% for battery systems. Next, an optimal microgrid connection model was proposed based on the constraints of a voltage stability limit estimated to be 10% of the maximum voltage drop. The results obtained for the third level of study enabled us to present selective results for load nodes in relation to the source station node. Finally, the last results made it possible to plan electrification using different network technologies and systems in the short and long term. The case study of Togo was taken into account. The various results obtained from the different techniques provide the necessary leads for a feasibility study for optimal electrification of off-grid areas using microgrid systems.

24 POWER TRANSMISSION AND DISTRIBUTION↗

DONKEY: A Flexible and Accurate Algorithm for Clustering

We propose an accurate clustering algorithm suitable for the varied and multidimensional data sets that correspond to temporal snapshots from on-the-fly nonadiabatic trajectory-based simulations of photoexcited dynamics. The algorithm approximates the underlying probability density function using variable kernel density estimation, with local maxima corresponding to cluster centers. Each data point is then assigned to one of the maxima by employing a maximization procedure. Finally, clusters artificially separated by minor fluctuations in the probability density are merged. The algorithm does not require parameter tuning, which ensures flexibility and reduces the risk of bias. It is tested on several synthetic data sets, where it consistently outperforms conventional clustering algorithms. As a final example, the algorithm is applied to the excited dynamics of the norbornadiene ⇌ quadricyclane (C 7 H 8 ) molecular photoswitch, demonstrating how distinct reaction pathways can be identified.

algorithms↗

The CluMPR galaxy cluster-finding algorithm and DESI legacy survey galaxy cluster catalogue

ABSTRACT Galaxy clusters enable unique opportunities to study cosmology, dark matter, galaxy evolution, and strongly lensed transients. We here present a new cluster-finding algorithm, CluMPR (Clusters from Masses and Photometric Redshifts), that exploits photometric redshifts (photo-z’s) as well as photometric stellar mass measurements. CluMPR uses a 2D binary search tree to search for overdensities of massive galaxies with similar redshifts on the sky and then probabilistically assigns cluster membership by accounting for photo-z uncertainties. We leverage the deep DESI Legacy Survey grzW1W2 imaging over one-third of the sky to create a catalogue of $\sim 300\, 000$ galaxy cluster candidates out to z = 1, including tabulations of member galaxies and estimates of each cluster’s total stellar mass. Compared to other methods, CluMPR is particularly effective at identifying clusters at the high end of the redshift range considered (z = 0.75–1), with minimal contamination from low-mass groups. These characteristics make it ideal for identifying strongly lensed high-redshift supernovae and quasars that are powerful probes of cosmology, dark matter, and stellar astrophysics. As an example application of this cluster catalogue, we present a catalogue of candidate wide-angle strongly lensed quasars in Appendix C. The nine best candidates identified from this sample include two known lensed quasar systems and a possible changing-look lensed QSO with SDSS spectroscopy. All code and catalogues produced in this work are publicly available (see Data Availability).

79 ASTRONOMY AND ASTROPHYSICS↗

sOPTICS: a modified density-based algorithm for identifying galaxy groups/clusters and brightest cluster galaxies

A direct approach to studying the galaxy–halo connection is to analyse groups and clusters of galaxies that trace the underlying dark matter haloes, emphasizing the importance of identifying galaxy clusters and their associated brightest cluster galaxies (BCGs). In this work, we test and propose a robust density-based clustering algorithm that outperforms the traditional Friends-of-Friends (FoF) algorithm in the currently available galaxy group/cluster catalogues. Our new approach is a modified version of the Ordering Points To Identify the Clustering Structure (OPTICS) algorithm, which accounts for line-of-sight positional uncertainties due to redshift space distortions by incorporating a scaling factor, and is thereby referred to as sOPTICS. When tested on both a galaxy group catalogue based on semi-analytic galaxy formation simulations and observational data, our algorithm demonstrated robustness to outliers and relative insensitivity to hyperparameter choices. In total, we compared the results of eight clustering algorithms. The proposed density-based clustering method, sOPTICS, outperforms FoF in accurately identifying giant galaxy clusters and their associated BCGs in various environments with higher purity and recovery rate, also successfully recovering 115 BCGs out of 118 reliable BCGs from a large galaxy sample. Furthermore, when applied to an independent observational catalogue without extensive re-tuning, sOPTICS maintains high recovery efficiency, confirming its flexibility and effectiveness for large-scale astronomical surveys.

79 ASTRONOMY AND ASTROPHYSICS↗

Combined Machine Learning and Molecular Dynamics Reveal Two States of Hydration of a Single Functional Group of Cationic Polymeric Brushes

The state of hydration of a macromolecular system regulates a plethora of different properties of such a system. In this article, we develop a novel machine learning (ML) approach, based on the unsupervised clustering algorithm, for probing the hydration behavior of the {N(CH 3 ) 3 } + functional group of the PMETAC [Poly(2-(methacryloyloxy)ethyl trimethylammonium chloride] polyelectrolyte (PE) brush system. The PE brushes and the brush-supported water molecules and counterions (chloride ions) are first described using all-atom molecular dynamics (MD) simulations. The simulation data is subsequently used in our ML framework to identify that (1) the {N(CH 3 ) 3 } + functional groups of the PMETAC brushes have two distinct hydration states with one state (state 1) being characterized by less structured water molecules and the other state (state 2) being characterized by more structured water molecules and (2) an enhancement in the brush grafting density leads to the progressive dissapparenace of state 2. An increase in the grafting density increases the number of chloride counterions in a given volume around the {N(CH 3 ) 3 } + functional group and increases the number of shared water molecules between the {N(CH 3 ) 3 } + and Cl - . The chloride counterions are associated with a hydration layer with much less structured water molecules. Therefore, with an increase in the grafting density, an increase in the percentage of shared water molecules leads to the prevalence of the hydration state [of the {N(CH 3 ) 3 } + moiety] with less structured water molecules. Finally, we explain how the present findings are commensurate with two key previous related results, namely a significantly large chloride ion mobility inside the PMETAC brush layer and the {N(CH 3 ) 3 } + -Cl - average distance remaining independent of the PMETAC brush grafting density. Furthermore, we anticipate that the combined ML-MD-simulation approach proposed in this study can be adapted to probe other soft matter systems to reveal new insights of the underlying mechanisms of emergent phenomenon.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

GPU acceleration of Swendsen–Wang dynamics

When simulating a lattice system near its critical temperature, local algorithms for modeling the system’s evolution can introduce very large autocorrelation times into sampled data. Here, this critical slowing down places restrictions on the analysis that can be completed in a timely manner of the behavior of systems around the critical point. Because it is often desirable to study such systems around this point, a new algorithm must be introduced. Therefore, we turn to cluster algorithms, such as the Swendsen–Wang algorithm and the Wolff clustering algorithm. They incorporate global updates which generate new lattice configurations with little correlation to previous states, even near the critical point. We look to accelerate the rate at which these algorithm are capable of running by implementing and benchmarking a parallel implementation of each algorithm designed to run on GPUs under NVIDIA’s CUDA framework. A 17 and 90 fold increase in the computational rate was, respectively, experienced when measured against the equivalent algorithm implemented in serial code.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

Possibilities and Limitations of Kinematically Identifying Stars from Accreted Ultra-faint Dwarf Galaxies

Abstract The Milky Way has accreted many ultra-faint dwarf galaxies (UFDs), and stars from these galaxies can be found throughout our Galaxy today. Studying these stars provides insight into galaxy formation and early chemical enrichment, but identifying them is difficult. Clustering stellar dynamics in 4D phase space ( E , L z , J r , J z ) is one method of identifying accreted structure that is currently being utilized in the search for accreted UFDs. We produce 32 simulated stellar halos using particle tagging with the Caterpillar simulation suite and thoroughly test the abilities of different clustering algorithms to recover tidally disrupted UFD remnants. We perform over 10,000 clustering runs, testing seven clustering algorithms, roughly twenty hyperparameter choices per algorithm, and six different types of data sets each with up to 32 simulated samples. Of the seven algorithms, HDBSCAN most consistently balances UFD recovery rates and cluster realness rates. We find that, even in highly idealized cases, the vast majority of clusters found by clustering algorithms do not correspond to real accreted UFD remnants and we can generally only recover 6% of UFDs remnants at best. These results focus exclusively on groups of stars from UFDs, which have weak dynamic signatures compared to the background of other stars. The recoverable UFD remnants are those that accreted recently, z accretion ≲ 0.5. Based on these results, we make recommendations to help guide the search for dynamically linked clusters of UFD stars in observational data. We find that real clusters generally have higher median energy and J r , providing a way to help identify real versus fake clusters. We also recommend incorporating chemical tagging as a way to improve clustering results.

79 ASTRONOMY AND ASTROPHYSICS↗

Scalable edge clustering of dynamic graphs via weighted line graphs

Timestamped relational datasets consisting of records (or connections) between pairs of entities are ubiquitous in network science. For applications like peer-to-peer communication, email, various social network interactions, and computer network security, it is useful to organize these records into groups based on how and when they are occurring. Weighted line graphs offer a natural way to model how records are related in such datasets but for large real-world graph topologies, building and utilizing the line graph is prohibitively expensive. Here, we present the framework to cluster the edges of a dynamic graph via the associated line graph that contains two major contributions. The first is a method to work with the line graph implicitly and the second is a distributed scale implementation of an agglomerative hierarchical graph clustering algorithm. We outline a novel hierarchical dynamic graph edge clustering approach that efficiently breaks massive relational datasets into small sets of edges containing events at various timescales. This is in stark contrast to traditional graph clustering algorithms that prioritize highly connected (clique-like) community structures. Our approach relies on constructing a sufficient subgraph of a weighted line graph and applying a hierarchical agglomerative clustering. This approach is related to scalable techniques from spatial clustering, nonlinear-dimension reduction, topological data analysis, and draws particular inspiration from HDBSCAN. As an edge clustering, this method yields an overlapping node clustering. Our algorithm is parallelizable and we demonstrate efficient clustering of a billion-scale, real-world dynamic graph into small edge sets that correlate in topology and time. The entire clustering process for a graph with tens of billions of edges takes just a few minutes of run time on 256 nodes of a distributed compute environment. We argue how the output of the edge clustering is useful for a multitude of data visualization and powerful machine learning tasks, both involving the original massive dynamic graph data and metadata associated with the nodes and edges. Finally, we describe how this approach can be extended to dynamic hypergraphs and dynamic graphs/hypergraphs with unstructured data living on vertices and edges.

Data Analysis↗

Integrated Land Suitability Assessment for Depots Siting in a Sustainable Biomass Supply Chain

A sustainable biomass supply chain would require not only an effective and fluid transportation system with a reduced carbon footprint and costs, but also good soil characteristics ensuring durable biomass feedstock presence. Unlike existing approaches that fail to account for ecological factors, this work integrates ecological as well as economic factors for developing sustainable supply chain development. For feedstock to be sustainably supplied, it necessitates adequate environmental conditions, which need to be captured in supply chain analysis. Using geospatial data and heuristics, we present an integrated framework that models biomass production suitability, capturing the economic aspect via transportation network analysis and the environmental aspect via ecological indicators. Production suitability is estimated using scores, considering both ecological factors and road transportation networks. These factors include land cover/crop rotation, slope, soil properties (productivity, soil texture, and erodibility factor) and water availability. This scoring determines the spatial distribution of depots with priority to fields scoring the highest. Two methods for depot selection are presented using graph theory and a clustering algorithm to benefit from contextualized insights from both and potentially gain a more comprehensive understanding of biomass supply chain designs. Graph theory, via the clustering coefficient, helps determine dense areas in the network and indicate the most appropriate location for a depot. Clustering algorithm, via K-means, helps form clusters and determine the depot location at the center of these clusters. An application of this innovative concept is performed on a case study in the US South Atlantic, in the Piedmont region, determining distance traveled and depot locations, with implications on supply chain design. The findings from this study show that a more decentralized depot-based supply chain design with 3depots, obtained using the graph theory method, can be more economical and environmentally friendly compared to a design obtained from the clustering algorithm method with 2 depots. In the former, the distance from fields to depots totals 801,031,476 miles, while in the latter, it adds up to 1,037,606,072 miles, which represents about 30% more distance covered for feedstock transportation.

29 ENERGY PLANNING, POLICY, AND ECONOMY↗

Metric DBSCAN

SAND2025-11725O Metric DBSCAN is an implementation of the popular DBSCAN clustering algorithm that works in general metric spaces. DBSCAN is a clustering algorithm, a fundamental building block in machine learning. It takes a set of objects and, given some notion of distance, identifies coherent groups of objects. With Metric DBSCAN, users can provide an arbitrary function to compute distance. Nearly all existing implementations of DBSCAN restrict distance to one of a few formulations. Metric DBScan accomplishes this cleanly and efficiently. The Python source code is on Github. Sandia National Laboratories is a multimission laboratory managed and operated by National Technology & Engineering Solutions of Sandia, LLC, a wholly owned subsidiary of Honeywell International Inc., for the U.S. Department of Energy’s National Nuclear Security Administration under contract DE-NA0003525.

Dalbey, Keith↗

Robust clustering of the local Milky Way stellar kinematic substructures with Gaia eDR3

Understanding local stellar kinematic substructures in the solar neighbourhood helps build a complete picture of the formation of the Milky Way, as well as an empirical phase space distribution of dark matter that would inform detection experiments. We apply the clustering algorithm HDBSCAN on the Gaia early third data release to identify a list of stable clusters in velocity space and action-angle space by taking into account the measurement uncertainties and studying the stability of the clustering results. We find 1405 (497) stars in 23 (6) robust clusters in velocity space (action-angle space) that are consistently not associated with noise. We discuss the kinematic properties of these structures and study whether many of the small clusters belong to a similar larger cluster based on their chemical abundances. They are attributed to the known structures: the Gaia Sausage-Enceladus, the Helmi Stream, and globular cluster NGC 3201 are found in both spaces, while NGC 104 and the thick disc (Sequoia) are identified in velocity space (action-angle space). Although we do not identify any new structures, we find that the HDBSCAN member selection of already known structures is unstable to input kinematics of the stars when resampled within their uncertainties. We therefore present the stable subset of local kinematic structures, which are consistently identified by the clustering algorithm, and emphasize the need to take into account error propagation during both the manual and automated identification of stellar structures, both for existing ones as well as future discoveries.

79 ASTRONOMY AND ASTROPHYSICS↗

VoroClust

SAND2025-11465O VoroClust, also known as Voronoi Clustering, is a fast, density-based unsupervised clustering algorithm applicable to high-resolution and high-dimensional data. It operates as quickly as distance-based clustering methods while effectively capturing complex regional geometries, matching the performance of current density-based methods. VoroClust employs a data-centered sphere cover to reduce computational demands while preserving data topology. It propagates clusters outward from local density peaks. Although supervised machine learning is powerful for applications like image classification and segmentation, it requires comprehensive, consistent datasets, which many applications lack. Unsupervised clustering algorithms analyze the structure of each dataset rather than relying on similarities with other examples, making them well-suited for practical applications with insufficient or inappropriate data for supervised learning. Sandia National Laboratories is a multimission laboratory managed and operated by National Technology & Engineering Solutions of Sandia, LLC, a wholly owned subsidiary of Honeywell International Inc., for the U.S. Department of Energy’s National Nuclear Security Administration under contract DE-NA0003525.

Ebeida, Mohamed [Sandia National Lab. (SNL-CA), Li↗

VoroClust: Scalable Clustering for Remote Sensing

Although supervised machine learning provides a powerful framework for image classification and segmentation, it requires comprehensive consistent datasets, which are not available for many remote-sensing applications. Remote-sensing datasets are expensive to collect, and each is acquired under different environmental conditions or with significant variations in system operating parameters. Unsupervised clustering algorithms analyze the structure of each dataset independently, rather than drawing on similarities with existing “training” examples, and are thus well suited for practical remote-sensing applications. We introduce VoroClust, a fast density-based unsupervised clustering algorithm applicable to high-resolution and high-dimensional data. VoroClust runs as fast as distance-based clustering methods, while capturing complex regional geometries at least as well as current-density-based methods. It uses a data-centered sphere cover to reduce computational demands, while still capturing data topology. It then propagates clusters outward from local peaks in density. We show that VoroClust provides fast state-of-the-art clustering for both high-resolution polarimetric synthetic aperture radar and high-dimensional hyperspectral imaging datasets.

42 ENGINEERING↗