Search NASA⌕ Search

SEARCH · Search NASA

Results for “Cluster analysis”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 91 records · Page 5

Systematic study of projection biases in the weak lensing analysis of cosmic shear and the combination of galaxy clustering and galaxy-galaxy lensing

This paper presents the results of a systematic study of projection biases in the weak lensing analysis of cosmic shear and the combination of galaxy clustering and galaxy-galaxy lensing using data collected during the first year of running the Dark Energy Survey experiment. The study uses Lambda cold dark matter ( Λ CDM ) as the cosmological model and two-point correlation functions for the weak lensing (WL) analysis. The results in this paper show that, independent of the WL analysis, projection biases of more than 1 σ exist and are a function of the position of the true values of the parameters h , n s , Ω b , and Ω ν h 2 with respect to their prior probabilities. For cosmic shear, and the combination of galaxy clustering and galaxy-galaxy lensing, this study shows that the coverage probability of the 68.27% credible intervals ranges from as high as 93% to as low as 16% and that these credible intervals are inflated, on average, by 29% for cosmic shear and 20% for the combination of galaxy clustering and galaxy-galaxy lensing. The results of the study also show that, in six out of nine tested cases, the reduction in error bars obtained by transforming credible intervals into confidence intervals is equivalent to an increase in the amount of data by a factor of 3.

79 ASTRONOMY AND ASTROPHYSICS↗

A Parameter-masked Mock Data Challenge for Beyond-two-point Galaxy Clustering Statistics

The past few years have seen the emergence of a wide array of novel techniques for analyzing high-precision data from upcoming galaxy surveys, which aim to extend the statistical analysis of galaxy clustering data beyond the linear regime and the canonical two-point (2pt) statistics. We test and benchmark some of these new techniques in a community data challenge named “Beyond-2pt,” initiated during the Aspen 2022 Summer Program “Large-Scale Structure Cosmology beyond 2-Point Statistics,” whose first round of results we present here. The challenge data set consists of high-precision mock galaxy catalogs for clustering in real space, in redshift space, and on a light cone. Participants in the challenge have developed end-to-end pipelines to analyze mock catalogs and extract unknown (“masked”) cosmological parameters of the underlying ΛCDM models with their methods. The methods represented are density-split clustering, nearest neighbor statistics, BACCO power spectrum emulator, void statistics, LEFTfield field-level inference using effective field theory (EFT), and joint power spectrum and bispectrum analyses using both EFT and simulation-based inference. In this work, we review the results of the challenge, focusing on problems solved, lessons learned, and future research needed to perfect the emerging beyond-2pt approaches. The unbiased parameter recovery demonstrated in this challenge by multiple statistics and the associated modeling and inference frameworks supports the credibility of cosmology constraints from these methods. The challenge data set is publicly available, and we welcome future submissions from methods that are not yet represented.

Krause, Elisabeth [Univ. of Arizona, Tucson, AZ (U↗

Cluster-Graph Fingerprinting: A Framework for Quantitative Analysis of Machine-Learned Interatomic Model Training and Simulation Data

Machine-learned interatomic models represent a significant advancement in simulation methods, extending the predictive ability of first-principles methods to previously inaccessible length and time scales. However, the data-driven nature of these models can lead to difficult-to-detect errors that can compromise prediction accuracy. To address this challenge, we introduce a novel fingerprinting approach based on the Chebyshev Interaction Model for Efficient Simulation (ChIMES) ML-IAM graph-based descriptor. Our strategy enables efficient and statistically rigorous analysis of system configurations used in ML-IAM training and those generated by their application, e.g., in molecular dynamics simulations. We demonstrate that these fingerprints can effectively assess novelty of a configuration relative to an existing data set and determine dissimilarity among individual configurations, which are two key tasks in workflows for active learning-based ML-IAM training, data set curation, and on-the-fly uncertainty quantification.

36 MATERIALS SCIENCE↗

Ion Clusters Reveal the Sources, Impacts, and Drivers of Freshwater Salinization

Population growth, land use change, climate change, and natural resource extraction are driving the salinization of freshwater resources worldwide. Reversing these trends will require data-centric approaches that identify salt sources, environmental drivers, and ecosystem responses. In this study, we applied principal component analysis and hierarchical clustering to identify ion covariance patterns, or “ion clusters,” in Broad Run, an urban stream in the Mid-Atlantic United States. These clusters correspond to distinct hydrologic regimes and reveal specific salinization risks: (1) phosphorus pollution mobilized during summer storms (Cluster 1); (2) elevated concentrations of sulfate and bicarbonate during baseflow (Cluster 2), likely reflecting groundwater discharge; and (3) elevated specific conductance and sodium, chloride, and potassium ion concentrations during snowmelt and rain-on-snow events (Cluster 3), driven by deicer and anti-icer wash-off. These ion fingerprints offer a transferable framework for diagnosing salt sources, assessing ecological risk, and identifying management targets. Our findings underscore the need for next-generation stormwater infrastructure and smart growth policies to protect aquatic life in rapidly urbanizing watersheds.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Ion Clusters in Multicomponent Solutions Determined from X-ray PDF and SAXS Analysis: The NaNO 2 –NaOH–H 2 O System

Here, this study explores ion cluster formation in the NaNO 2 –NaOH–H 2 O system to understand how ion cluster formation is influenced by the composition in multicomponent mixtures. X-ray pair distribution function (PDF) and small-angle X-ray scattering (SAXS) identified complex ion clusters in concentrated NaNO 2 and NaOH solutions as well as their mixtures. PDF analysis showed that the Na–O distance depended primarily on total Na + concentrations rather than anion composition, whereas the nitrite-water oxygen distance stayed the same regardless of concentration or composition. This result indicates that the nitrite ion hydration was relatively independent of composition. SAXS confirms local fluctuations and coherent clusters with notable size differences between NaOH and NaNO 2 solutions. SAXS analysis of mixed solutions shows that their clusters are an average of the individual solutions, indicating mixed clusters.

Reynolds, Jacob G. [Central Plateau Cleanup Compan↗

Constraints on Nonthermal Pressure at Galaxy Cluster Outskirts from a Joint SPT and XMM-Newton Analysis

We present joint South Pole Telescope and XMM-Newton observations of eight massive galaxy clusters (0.8–2 × 10$^{15}$ M$_{⊙}$) spanning a redshift range of 0.16–0.35. Employing a novel Sunyaev–Zel’dovich + X-ray fitting technique, we effectively constrain the thermodynamic properties of these clusters out to the virial radius. The resulting best-fit electron density, deprojected temperature, and deprojected pressure profiles are in good agreement with previous observations of massive clusters. For the majority of the cluster sample (five out of eight clusters), the entropy profiles exhibit a self-similar behavior near the virial radius. We further derive hydrostatic mass, gas mass, and gas fraction profiles for all clusters up to the virial radius. Comparing the enclosed gas fraction profiles with the universal gas fraction profile, we obtain nonthermal pressure fraction profiles for our cluster sample at >0.5R$_{500}$, demonstrating a steeper increase between R$_{500}$ and R$_{200}$ that is consistent with the hydrodynamical simulations. Our analysis yields nonthermal pressure fraction ranges of 8%–28% (median: 15% ± 11%) at R$_{500}$ and 21%–35% (median: 27% ± 12%) at R$_{200}$. Notably, weak-lensing mass measurements are available for only four clusters in our sample, and our recovered total cluster masses, after accounting for nonthermal pressure, are consistent with these measurements.

79 ASTRONOMY AND ASTROPHYSICS↗

Dark Energy Survey Year 6 Results: Redshift Calibration of the MagLim++ Lens Sample

In this work, we derive and calibrate the redshift distribution of the MagLim++ lens galaxy sample used in the Dark Energy Survey Year 6 (DES Y6) 3x2pt cosmology analysis. The 3x2pt analysis combines galaxy clustering from the lens galaxy sample and weak gravitational lensing. The redshift distributions are inferred using the SOMPZ method - a Self-Organizing Map framework that combines deep-field multi-band photometry, wide-field data, and a synthetic source injection (Balrog) catalog. Key improvements over the DES Year 3 (Y3) calibration include a noise-weighted SOM metric, an expanded Balrog catalogue, and an improved scheme for propagating systematic uncertainties, which allows us to generate O($10^8$) redshift realizations that collectively span the dominant sources of uncertainty. These realizations are then combined with independent clustering-redshift measurements via importance sampling. The resulting calibration achieves typical uncertainties on the mean redshift of 1-2%, corresponding to a 20-30% average reduction relative to DES Y3. We compress the $n(z)$ uncertainties into a small number of orthogonal modes for use in cosmological inference. Marginalizing over these modes leads to only a minor degradation in cosmological constraints. This analysis establishes the MagLim++ sample as a robust lens sample for precision cosmology with DES Y6 and provides a scalable framework for future surveys.

Giannini, G. [Chicago U., Astron. Astrophys. Ctr.;↗

Dark Energy Survey Year 6 results: Redshift calibration of the MagLim++ lens sample

In this work, we derive and calibrate the redshift distribution of the MagLim++ lens galaxy sample used in the Dark Energy Survey Year 6 (DES Y6) 3 x 2pt cosmology analysis. The 3 x 2pt analysis combines galaxy clustering from the lens galaxy sample and weak gravitational lensing. The redshift distributions are inferred using the SOMPZ method - a Self-Organizing Map framework that combines deep-field multi-band photometry, wide-field data, and a synthetic source injection ( B alrog) catalog. Key improvements over the DES Year 3 (Y3) calibration include a noise-weighted SOM metric, an expanded Balrog catalogue, and an improved scheme for propagating systematic uncertainties, which allows us to generate O(10 8 ) redshift realizations that collectively span the dominant sources of uncertainty. These realizations are then combined with independent clustering-redshift measurements via importance sampling. The resulting calibration achieves typical uncertainties on the mean redshift of 1-2%, corresponding to a 20-30% average reduction relative to DES Y3. We compress the n(z) uncertainties into a small number of orthogonal modes for use in cosmological inference. Marginalizing over these modes leads to only a minor degradation in cosmological constraints. Here, this analysis establishes the MagLim++ sample as a robust lens sample for precision cosmology with DES Y6 and provides a scalable framework for future surveys.

dark energy↗

The z ∼ 1.03 Merging Cluster SPT-CL J0356–5337: New Strong Lensing Analysis with HST and MUSE

We present a strong lensing analysis and reconstruct the mass distribution of SPT-CL J0356−5337, a galaxy cluster at redshift z = 1.034. Our model supersedes previous models by making use of new multiband Hubble Space Telescope data and Multi-Unit Spectroscopic Explorer (MUSE) spectroscopy. We identify two additional lensed galaxies to inform a more well-constrained model using 12 sets of multiple images in five separate lensed sources. The three previously known sources were spectroscopically confirmed by G. Mahler et al. at redshifts of z = 2.363, z = 2.364, and z = 3.048. We measured the spectroscopic redshifts of two of the newly discovered arcs using MUSE data, at z = 3.0205 and z = 5.3288. We increase the number of cluster member galaxies by a factor of 3 compared to previous work. We also report the detection of extended Lyα emission from several background galaxies. We measure the total projected mass density of the two major subcluster components, one dominated by the brightest cluster galaxy and the other by a compact group of luminous red galaxies. We find ${M}_{{\rm{B}}{\rm{C}}{\rm{G}}}(\lt 80\,\rm{kpc})=3.9{3}_{-0.14}^{+0.21}\times 1{0}^{13}$ M⊙ and ${M}_{{\rm{LRG}}}(\lt 80\,\rm{kpc})=2.9{2}_{-0.23}^{+0.16}\times 1{0}^{13}$ M⊙, yielding a mass ratio of $1.3{5}_{-0.08}^{+0.16}$ . The strong lensing constraints offer a robust estimate of the projected mass density regardless of modeling assumptions; allowing more substructure in this line of sight does not change the results or conclusions. Our results corroborate the conclusion that SPT-CL J0356−5337 is dominated by two mass components and is likely undergoing a major merger on the plane of the sky.

79 ASTRONOMY AND ASTROPHYSICS↗

Reference Site Conditions for Floating Wind Arrays in the United States

Floating offshore wind farm design is highly site-specific, requiring detailed information about the specific conditions of a project area for realistic design studies. Unfortunately, publicly available site condition data for potential floating offshore wind project sites in the United States is scarce. To support U.S. offshore wind research, we developed reference site condition datasets, including metocean and seabed information, for four potential floating wind project areas in the U.S.: Humboldt Bay, Morro Bay, the Gulf of Maine, and the Gulf of Mexico. These datasets were compiled using publicly available data. Our metocean analysis, covering wind, waves, and surface currents, utilized measurement data from 2000 to 2020. Sources included the National Renewable Energy Laboratory’s National Offshore Wind Dataset for wind data, National Data Buoy Center buoys for wave data, and the High Frequency Radar Network for surface currents. These data were integrated into hourly time series used to compute extreme return periods up to 500 years, monthly statistics, and joint probability clusters for fatigue analysis. Soil conditions were evaluated using the usSEABED database and bathymetry grids were interpolated from the NCEI Digital Elevation Model Global Mosaic. In addition to providing curated reference site condition datasets for four U.S. areas, our assessment highlights the need for more publicly available metocean and soil condition data.

17 WIND ENERGY↗

Reference Site Condition Datasets for Floating Wind Arrays in the United States

Floating offshore wind farm design is highly site-specific, requiring detailed information about the specific conditions of a project area for realistic design studies. Unfortunately, publicly available site condition data for potential floating offshore wind project sites in the United States is scarce. To support U.S. offshore wind research, we developed reference site condition datasets, including metocean and seabed information, for four potential floating wind project areas in the U.S.: Humboldt Bay, Morro Bay, the Gulf of Maine, and the Gulf of Mexico. These datasets were compiled using publicly available data. Our metocean analysis, covering wind, waves, and surface currents, utilized measurement data from 2000 to 2020. Sources included the National Renewable Energy Laboratory’s National Offshore Wind Dataset for wind data, National Data Buoy Center buoys for wave data, and the High Frequency Radar Network for surface currents. These data were integrated into hourly time series used to compute extreme return periods up to 500 years, monthly statistics, and joint probability clusters for fatigue analysis. Soil conditions were evaluated using the usSEABED database and bathymetry grids were interpolated from the NCEI Digital Elevation Model Global Mosaic. Further information on the datasets and how they were created can be found in: Biglu, M., M. Hall, E. Lozon, S. Housner. 2024. Reference Site Conditions for Floating Wind Arrays in the United States. Golden, CO: National Renewable Energy Laboratory (NREL). NREL/TP-5000-89897. The data are also available at: https://github.com/FloatingArrayDesign/SiteConditions The content of each dataset is as follows: _NOW23_wind.txt: Hourly NOW-23 wind data up to a height of 400 meter. _metocean_1hr.txt: Hourly time series including wind, wave, surface current and temperature data. _Summary.xlsx: Metocean data, including extreme values, joint probability distributions and monthly statistics. _usSEABED_soil.csv: Extract of the usSEABED database for this specific site. _bathymetry_200m.txt (and 500m, 1000m): Gridded seabed depth data.

16 TIDAL AND WAVE POWER↗

Masses of Sunyaev-Zel’dovich galaxy clusters detected by the Atacama Cosmology Telescope: Stacked lensing measurements with Subaru HSC year 3 data

We present a stacked lensing analysis of 96 galaxy clusters selected by the thermal Sunyaev-Zel’dovich (SZ) effect in maps of the cosmic microwave background (CMB). We select foreground galaxy clusters with a 5 σ -level SZ threshold in CMB observations from the Atacama Cosmology Telescope, while we define background source galaxies for the lensing analysis with secure photometric redshift cuts in Year 3 data of the Subaru Hyper Suprime Cam survey. We detect the stacked lensing signal in the range of 0.1 < R [ h - 1 Mpc ] < 100 in each of three cluster redshift bins, 0.092 < z ≤ 0.445 , 0.445 < z ≤ 0.695 , and 0.695 < z ≤ 1.180 , with 32 galaxy clusters in each bin. The cumulative signal-to-noise ratios of the lensing signal are 14.6, 12.0, and 6.6, respectively. Using a halo-based forward model, we then constrain statistical relationships between the mass inferred from the SZ observation (i.e. SZ mass) and the total mass derived from our stacked lensing measurements. At the average SZ mass in the cluster sample ( 2.1 - 2.4 × 10 14 h - 1 M ⊙ ), our likelihood analysis shows that the average total mass differs from the SZ counterpart by a factor of 1.3 ± 0.2 , 1.6 ± 0.2 , and 1.6 ± 0.3 (68%) in the aforementioned redshift ranges, respectively. Our limits are consistent with previous lensing measurements, and we find that the cluster modeling choices can introduce a 1 σ -level difference in our parameter inferences.

79 ASTRONOMY AND ASTROPHYSICS↗

Using Apptainer in a Pilot-based Distributed Workload

GlideinWMS is a pilot and pressure-based workload manager for distributed scientific computing. Many experiments like CMS and Fermilab’s Neutrino experiments use it to provision elastic clusters for their analysis and simulations, split into close to a million concurrent jobs. Most user jobs require containers, and the pilots use Apptainer to set up the desired platform. For the pilots that run as regular batch jobs, Apptainer is safer, lighter, and easier to use than other containerization solutions. Many images used by the pilots are expanded SIF images distributed via the CernVM-FS: this combination is very efficient. At Fermilab, for example, we store on GitHub Dockerfiles that mimic the platform in the worker nodes of local clusters. GitHub workflows build and push the images to Docker Hub, and a service periodically pulls and converts them to the expanded SIF images in the CernVM-FS, so the scientists can find a familiar environment everywhere. Apptainer has also been used to run services inside the pilot jobs, like benchmarks that characterize the worker node being used, or a Triton Inference Server that allows sharing a GPU with all the jobs that run in parallel on a node.

Mambelli, Marco [Fermilab] (ORCID:0000000294892681↗

Structural features of xylan dictate reactivity and functionalization potential for bio-based materials

Plant-based materials have the potential to replace some petroleum-based products, offering compostability and biodegradability as critical advantages. Xylan-rich biomass sources are gaining recognition due to their abundance and underutilization in current industrial applications. Research of potential xylan applications has been complicated by the complex and heterogeneous structure that varies for different xylan feedstocks. Acylation is a broadly used reaction in functionalization of polysaccharides at an industrial scale. However, the efficiency of this reaction varies with the xylan source. To optimize xylan valorization, a systematic understanding of structure–reactivity relationships is essential. This study explores, characterizes, and compares various xylan feedstocks in the acylation process. Xylan feedstocks were analyzed for their chemical composition, degree of polymerization, branching, solubility, and presence of impurities. These features were correlated with xylan glycotypes’ reactivity toward functionalization with succinic anhydride in an optimized DMSO/KOH condition, achieving carboxyl contents of up to 1.46. We used principal component analysis and hierarchical clustering to identify key structural features of xylan that promote its reactivity. Our findings reveal that xylans with higher xylose content and lower degrees of branching exhibit enhanced reactivity, achieving higher carboxyl content and yields. Structural analyses confirmed successful modification, and light scattering analyses showed dramatic changes in the solution properties. Succinylation improves the solubility and film-forming properties of native xylans. This study shows key structure–reactivity relationships in xylan succinylation, establishing that low branching, high xylose content, and reduced lignin impurity enhance chemical functionalization. The results offer a framework for selecting optimal biomass feedstocks and support future efforts in genetic and synthetic biology to design plants with tunable xylan architectures. These findings advance the hemicellulose valorization for applications in coatings and packaging.

Acylation↗

Investigation of acoustic waves under subsurface conditions to improve the predictions of rock mechanical properties and natural fracture characteristics

Mechanical properties and natural fracture characteristics are critical to investigate for subsurface engineering applications, including carbon storage, well drilling, and stimulation, as they govern rock stability, fluid flow, and mechanical behavior under stress. This dissertation integrates experimental and machine learning approaches to enhance the prediction and understanding of these properties by analyzing acoustic wave behavior under varied subsurface conditions. First, the influence of temperature, pore pressure, and supercritical CO2 (scCO2) saturation on poroelastic properties is examined using Gray Berea sandstone samples. The results show that temperature and pore pressure significantly affect the bulk modulus and Biot’s coefficient, while scCO2 saturation impacts rock compressibility, informing strategies for effective geological carbon storage. The study extends this understanding by experimentally evaluating the impact of reservoir depletion on the dynamic mechanical properties of the emerging Caney shale in South Oklahoma with the employment of unsupervised machine learning to predict static mechanical properties across the Caney shale. Integrating petrophysical data and chemostratigraphy, the workflow—featuring K-means clustering, principal component analysis (PCA), and inverse distance weighting (IDW)—improves stratigraphic characterization and the estimation of static-to-dynamic modulus ratios, which is vital for optimizing drilling and stimulation strategies. Finally, the work explores how natural fracture characteristics in shale influence acoustic waveforms and shear wave splitting (SWS) analysis. Experimental data on fractured samples under different stress and temperature conditions, combined with machine learning models such as K-nearest neighbors (KNN) and extreme gradient boosting (XGBoost), reveal key fracture properties impacting SWS and wave propagation. Together, these studies provide a comprehensive framework for linking acoustic wave behavior with rock properties, advancing the methods for monitoring and predicting geomechanical changes. The insights offered valuable implications for safer, more efficient CO2 injection, hydrocarbon extraction, and subsurface management.

Elkholy, Sherif↗

Lyman-$α$ forest holography: 3D predictions from 1D measurements

Cosmological analyses of Lyman-$α$ forest clustering rely on either one-dimensional correlations along individual sightlines or three-dimensional correlations between different sightlines. Because these observables probe the matter distribution on very different scales, they have traditionally been analyzed independently. In this work, we bridge this gap using ForestFlow, an emulator trained on a suite of cosmological hydrodynamical simulations that provides a unified description of Lyman-$α$ forest clustering from linear to nonlinear scales. This framework enables us to determine the range of three-dimensional clustering models compatible with the DESI one-dimensional flux power spectrum ($P_{\rm 1D}$). The resulting predictions successfully reproduce the large-scale clustering measured by the DESI BAO analysis and provide physically motivated priors on nonlinear clustering that are used in a companion paper presenting the full-shape analysis of the DESI DR2 Lyman-$α$ forest. We validate our methodology using the large-volume, high-resolution hydrodynamical simulation ACCEL-2, demonstrating excellent agreement across the full range of scales considered. Finally, we combine constraints from the $P_{\rm 1D}$ and BAO analyses on the parameter combinations $b_δσ_8$ and $b_ηf σ_8$, finding that the two probes provide comparable constraining power while exhibiting complementary parameter degeneracies. Our results establish a direct connection between one- and three-dimensional Lyman-$α$ forest measurements through ForestFlow, an approach we term Lyman-$α$ holography by analogy with the reconstruction of higher-dimensional structure from lower-dimensional information.

Chaves-Montero, J. [Barcelona, IFAE] (ORCID:000000↗

LoVoCCS. II. Weak Lensing Mass Distributions, Red-sequence Galaxy Distributions, and Their Alignment with the Brightest Cluster Galaxy in 58 Nearby X-Ray-luminous Galaxy Clusters

The Local Volume Complete Cluster Survey is an ongoing program to observe nearly a hundred low-redshift X-ray-luminous galaxy clusters (redshifts 0.03 < z < 0.12 and X-ray luminosities in the 0.1–2.4 keV band L X500c > 10 44 erg s −1 ) with the Dark Energy Camera, capturing data in the u, g, r, i, z bands with a 5σ point source depth of approximately 25th–26th AB magnitudes. Here, we map the aperture masses in 58 galaxy cluster fields using weak gravitational lensing. These clusters span a variety of dynamical states, from nearly relaxed to merging systems, and approximately half of them have not been subject to detailed weak lensing analysis before. In each cluster field, we analyze the alignment between the 2D mass distribution described by the aperture mass map, the 2D red-sequence (RS) galaxy distribution, and the brightest cluster galaxy (BCG). We find that the orientations of the BCG and the RS distribution are strongly aligned throughout the interiors of the clusters: the median misalignment angle is 19° within 2 Mpc. We also observe the alignment between the orientations of the RS distribution and the overall cluster mass distribution (by a median difference of 32° within 1 Mpc), although this is constrained by galaxy shape noise and the limitations of our cluster sample size. These types of alignment suggest long-term dynamical evolution within the clusters over cosmic timescales.

79 ASTRONOMY AND ASTROPHYSICS↗

Scalable edge clustering of dynamic graphs via weighted line graphs

Timestamped relational datasets consisting of records (or connections) between pairs of entities are ubiquitous in network science. For applications like peer-to-peer communication, email, various social network interactions, and computer network security, it is useful to organize these records into groups based on how and when they are occurring. Weighted line graphs offer a natural way to model how records are related in such datasets but for large real-world graph topologies, building and utilizing the line graph is prohibitively expensive. Here, we present the framework to cluster the edges of a dynamic graph via the associated line graph that contains two major contributions. The first is a method to work with the line graph implicitly and the second is a distributed scale implementation of an agglomerative hierarchical graph clustering algorithm. We outline a novel hierarchical dynamic graph edge clustering approach that efficiently breaks massive relational datasets into small sets of edges containing events at various timescales. This is in stark contrast to traditional graph clustering algorithms that prioritize highly connected (clique-like) community structures. Our approach relies on constructing a sufficient subgraph of a weighted line graph and applying a hierarchical agglomerative clustering. This approach is related to scalable techniques from spatial clustering, nonlinear-dimension reduction, topological data analysis, and draws particular inspiration from HDBSCAN. As an edge clustering, this method yields an overlapping node clustering. Our algorithm is parallelizable and we demonstrate efficient clustering of a billion-scale, real-world dynamic graph into small edge sets that correlate in topology and time. The entire clustering process for a graph with tens of billions of edges takes just a few minutes of run time on 256 nodes of a distributed compute environment. We argue how the output of the edge clustering is useful for a multitude of data visualization and powerful machine learning tasks, both involving the original massive dynamic graph data and metadata associated with the nodes and edges. Finally, we describe how this approach can be extended to dynamic hypergraphs and dynamic graphs/hypergraphs with unstructured data living on vertices and edges.

Data Analysis↗