Search NASA⌕ Search

SEARCH · Search NASA

Results for “clustering statistics”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

The Aemulus Project. VI. Emulation of Beyond-standard Galaxy Clustering Statistics to Improve Cosmological Constraints

Abstract There is untapped cosmological information in galaxy redshift surveys in the nonlinear regime. In this work, we use the Aemulus suite of cosmological N -body simulations to construct Gaussian process emulators of galaxy clustering statistics at small scales (0.1–50 h −1 Mpc) in order to constrain cosmological and galaxy bias parameters. In addition to standard statistics—the projected correlation function w p ( r p ), the redshift-space monopole of the correlation function ξ 0 ( s ), and the quadrupole ξ 2 ( s )—we emulate statistics that include information about the local environment, namely the underdensity probability function P U ( s ) and the density-marked correlation function M ( s ). This extends the model of Aemulus III for redshift-space distortions by including new statistics sensitive to galaxy assembly bias. In recovery tests, we find that the beyond-standard statistics significantly increase the constraining power on cosmological parameters of interest: including P U ( s ) and M ( s ) improves the precision of our constraints on Ω m by 27%, σ 8 by 19%, and the growth of structure parameter, f σ 8 , by 12% compared to standard statistics. We additionally find that scales below ∼6 h −1 Mpc contain as much information as larger scales. The density-sensitive statistics also contribute to constraining halo occupation distribution parameters and a flexible environment-dependent assembly bias model, which is important for extracting the small-scale cosmological information as well as understanding the galaxy–halo connection. This analysis demonstrates the potential of emulating beyond-standard clustering statistics at small scales to constrain the growth of structure as a test of cosmic acceleration.

79 ASTRONOMY AND ASTROPHYSICS↗

A Parameter-masked Mock Data Challenge for Beyond-two-point Galaxy Clustering Statistics

The past few years have seen the emergence of a wide array of novel techniques for analyzing high-precision data from upcoming galaxy surveys, which aim to extend the statistical analysis of galaxy clustering data beyond the linear regime and the canonical two-point (2pt) statistics. We test and benchmark some of these new techniques in a community data challenge named “Beyond-2pt,” initiated during the Aspen 2022 Summer Program “Large-Scale Structure Cosmology beyond 2-Point Statistics,” whose first round of results we present here. The challenge data set consists of high-precision mock galaxy catalogs for clustering in real space, in redshift space, and on a light cone. Participants in the challenge have developed end-to-end pipelines to analyze mock catalogs and extract unknown (“masked”) cosmological parameters of the underlying ΛCDM models with their methods. The methods represented are density-split clustering, nearest neighbor statistics, BACCO power spectrum emulator, void statistics, LEFTfield field-level inference using effective field theory (EFT), and joint power spectrum and bispectrum analyses using both EFT and simulation-based inference. In this work, we review the results of the challenge, focusing on problems solved, lessons learned, and future research needed to perfect the emerging beyond-2pt approaches. The unbiased parameter recovery demonstrated in this challenge by multiple statistics and the associated modeling and inference frameworks supports the credibility of cosmology constraints from these methods. The challenge data set is publicly available, and we welcome future submissions from methods that are not yet represented.

Krause, Elisabeth [Univ. of Arizona, Tucson, AZ (U↗

Full forward model of galaxy clustering statistics with AbacusSummit light cones

ABSTRACT Novel summary statistics beyond the standard 2-point correlation function (2PCF) are necessary to capture the full astrophysical and cosmological information from the small-scale (r < 30h−1Mpc) galaxy clustering. However, the analysis of beyond-2PCF statistics on small scales is challenging because we lack the appropriate treatment of observational systematics for arbitrary summary statistics of the galaxy field. In this paper, we develop a full forward modelling pipeline for a wide range of summary statistics using the large high-fidelity AbacusSummit light cones that account for many systematic effects as well as remain flexible and computationally efficient to enable posterior sampling. We apply our forward model approach to a fully realistic mock galaxy catalog and demonstrate that we can recover unbiased constraints on the underlying galaxy–halo connection model using two separate summary statistics: the standard 2PCF and the novel k-th nearest neighbour (kNN) statistics, which are sensitive to correlation functions of all orders. We will demonstrate its strong constraining power on extended galaxy–halo connection models and cosmology in follow up papers. We expect this to become a powerful approach when applying to upcoming surveys such as DESI where we can leverage a multitude of summary statistics across a wide redshift range to maximally extract information from the non-linear scales.

79 ASTRONOMY AND ASTROPHYSICS↗

DESI 2024 II: sample definitions, characteristics, and two-point clustering statistics

We present the samples of galaxies and quasars used for DESI 2024 cosmological analyses, drawn from the DESI Data Release 1 (DR1). We describe the construction of largescale structure (LSS) catalogs from these samples, which include matched sets of synthetic reference ‘randoms’ and weights that account for variations in the observed density of the samples due to experimental design and varying instrument performance. We detail how we correct for variations in observational completeness, the input ‘target’ densities due to imaging systematics, and the ability to confidently measure redshifts from DESI spectra. We then summarize how remaining uncertainties in the corrections can be translated to systematic uncertainties for particular analyses. We describe the weights added to maximize the signalto-noise of DESI DR1 2-point clustering measurements. We detail measurement pipelines applied to the LSS catalogs that obtain 2-point clustering measurements in configuration and Fourier space. The resulting 2-point measurements depend on window functions and normalization constraints particular to each sample, and we present the corrections required to match models to the data. We compare the configuration- and Fourier-space 2-point clustering of the data samples to that recovered from simulations of DESI DR1 and find they are, generally, in statistical agreement to within 2% in the inferred real-space over-density field. The LSS catalogs, 2-point measurements, and their covariance matrices will be released publicly with DESI DR1.

79 ASTRONOMY AND ASTROPHYSICS↗

The DESI One-Percent Survey: Modelling the clustering and halo occupation of all four DESI tracers with U CHUU

We present results from a set of mock lightcones for the DESI One-Percent Survey, created from the UCHUU simulation. This 8 h −3 Gpc 3 N-body simulation comprises 2.1 trillion particles and provides high-resolution dark matter (sub)haloes in the framework of the Planck-based ΛCDM cosmology. Employing the subhalo abundance matching (SHAM) technique, we populated the UCHUU (sub)haloes with all four DESI tracers – Bright Galaxy Survey (BGS), luminous red galaxies (LRGs), emission line galaxies (ELGs), and quasars (QSOs) – to z = 2.1. Our method accounts for redshift evolution as well as the clustering dependence on luminosity and stellar mass. The two-point clustering statistics of the DESI One-Percent Survey generally agree with predictions from UCHUU across scales ranging from 0.3 h −1 Mpc to 100 h −1 Mpc for the BGS and across scales ranging from 5 h −1 Mpc to 100 h −1 Mpc for the other tracers. We observed some differences in clustering statistics that can be attributed to incompleteness of the massive end of the stellar mass function of LRGs, our use of a simplified galaxy-halo connection model for ELGs and QSOs, and cosmic variance. We find that at the high precision of UCHUU, the shape of the halo occupation distribution (HOD) of the BGS and LRG samples is smaller bias values, likely due to cosmic variance. The bias dependence on absolute magnitude, stellar mass, and redshift aligns with that of previous surveys. These results provide DESI with tools to generate high-fidelity lightcones for the remainder of the survey and enhance our understanding of the galaxy-halo connection.

cosmology↗

First Constraints on Growth Rate from Redshift-space Ellipticity Correlations of SDSS Galaxies at 0.16 < z < 0.70

We report the first constraints on the growth rate of the universe, f(z)σ 8 (z), with intrinsic alignments (IAs) of galaxies. We measure the galaxy density-intrinsic ellipticity cross-correlation and intrinsic ellipticity autocorrelation functions over 0.16 < z < 0.7 from luminous red galaxies (LRGs) and LOWZ and CMASS galaxy samples in the Sloan Digital Sky Survey (SDSS) and SDSS-III BOSS survey. We detect clear anisotropic signals of IA due to redshift-space distortions. By combining measured IA statistics with the conventional galaxy clustering statistics, we obtain tighter constraints on the growth rate. The improvement is particularly prominent for the LRG, which is the brightest galaxy sample and known to be strongly aligned with underlying dark matter distribution; using the measurements on scales above 10 h -1 Mpc, we obtain $f{\sigma }_{8}={0.5196}_{-0.0354}^{+0.0352}$ (68% confidence level) from the clustering-only analysis and $f{\sigma }_{8}={0.5322}_{-0.0291}^{+0.0293}$ with clustering and IA, meaning 19% improvement. The constraint is in good agreement with the prediction of general relativity, f σ 8 = 0.4937 at z = 0.34. For LOWZ and CMASS samples, the improvement of constraints on f σ 8 is found to be 10% and 3.5%, respectively. Our results indicate that the contribution from IA statistics for cosmological constraints can be further enhanced by carefully selecting galaxies for a shape sample.

79 ASTRONOMY AND ASTROPHYSICS↗

High-precision Galaxy Clustering Predictions from Small-volume Hydrodynamical Simulations via Control Variates

Abstract Cosmological simulations of galaxy formation are an invaluable tool for understanding galaxy formation and its impact on cosmological parameter inference from large-scale structures. However, their high computational cost is a significant obstacle for running simulations that probe cosmological volumes comparable to those analyzed by contemporary large-scale structure experiments. In this work, we explore the possibility of obtaining high-precision galaxy clustering predictions from small-volume hydrodynamical simulations such as MillenniumTNG and FLAMINGO via control variates. In this approach, the hydrodynamical full-physics simulation is paired with a matched low-resolution gravity-only simulation. By learning the galaxy–halo connection from the hydrodynamical simulation and applying it to the gravity-only counterpart, one obtains a galaxy population that closely mimics the one in the more expensive simulation. One can then construct an estimator of galaxy clustering that combines the clustering amplitudes in the small-volume hydrodynamical and gravity-only simulations with clustering amplitudes in a large-volume gravity-only simulation. Depending on the galaxy sample, clustering statistic, and scale, this galaxy clustering estimator can have an effective volume of up to around 100 times the volume of the original hydrodynamical simulation in the nonlinear regime. With this approach, we can construct galaxy clustering predictions from existing simulations that are precise enough for mock analyses of next-generation large-scale structure surveys such as the Dark Energy Spectroscopic Instrument and the Legacy Survey of Space and Time.

Doytcheva, Alexandra (ORCID:0009000111254888)↗

Dark Energy Survey Year 3 Results: clustering redshifts – calibration of the weak lensing source redshift distributions with redMaGiC and BOSS/eBOSS

ABSTRACT We present the calibration of the Dark Energy Survey Year 3 (DES Y3) weak lensing (WL) source galaxy redshift distributions n(z) from clustering measurements. In particular, we cross-correlate the WL source galaxies sample with redMaGiC galaxies (luminous red galaxies with secure photometric redshifts) and a spectroscopic sample from BOSS/eBOSS to estimate the redshift distribution of the DES sources sample. Two distinct methods for using the clustering statistics are described. The first uses the clustering information independently to estimate the mean redshift of the source galaxies within a redshift window, as done in the DES Y1 analysis. The second method establishes a likelihood of the clustering data as a function of n(z), which can be incorporated into schemes for generating samples of n(z) subject to combined clustering and photometric constraints. Both methods incorporate marginalization over various astrophysical systematics, including magnification and redshift-dependent galaxy-matter bias. We characterize the uncertainties of the methods in simulations; the first method recovers the mean z of tomographic bins to RMS (precision) of ∼0.014. Use of the second method is shown to vastly improve the accuracy of the shape of n(z) derived from photometric data. The two methods are then applied to the DES Y3 data.

79 ASTRONOMY AND ASTROPHYSICS↗

ADDGALS: Simulated Sky Catalogs for Wide Field Galaxy Surveys

Abstract We present a method for creating simulated galaxy catalogs with realistic galaxy luminosities, broadband colors, and projected clustering over large cosmic volumes. The technique, denoted Addgals (Adding Density Dependent GAlaxies to Lightcone Simulations), uses an empirical approach to place galaxies within lightcone outputs of cosmological simulations. It can be applied to significantly lower-resolution simulations than those required for commonly used methods such as halo occupation distributions, subhalo abundance matching, and semi-analytic models, while still accurately reproducing projected galaxy clustering statistics down to scales of r ∼ 100 h −1 kpc . We show that Addgals catalogs reproduce several statistical properties of the galaxy distribution as measured by the Sloan Digital Sky Survey (SDSS) main galaxy sample, including galaxy number densities, observed magnitude and color distributions, as well as luminosity- and color-dependent clustering. We also compare to cluster–galaxy cross correlations, where we find significant discrepancies with measurements from SDSS that are likely linked to artificial subhalo disruption in the simulations. Applications of this model to simulations of deep wide-area photometric surveys, including modeling weak-lensing statistics, photometric redshifts, and galaxy cluster finding, are presented in DeRose et al., and an application to a full cosmology analysis of Dark Energy Survey (DES) Year 3 like data is presented in DeRose et al. We plan to publicly release a 10,313 square degree catalog constructed using Addgals with magnitudes appropriate for several existing and planned surveys, including SDSS, DES, VISTA, Wide-field Infrared Survey Explorer, and Rubin Observatory’s Legacy Survey of Space and Time.

79 ASTRONOMY AND ASTROPHYSICS↗

DESI DR2 Reference Mocks: Clustering results from UCHUU ELGs and QSOs

High-redshift galaxy clustering provides a powerful probe of the growth of structure, testing models of dark matter, dark energy, and galaxy formation during the epoch when the Universe was rapidly evolving. Emission line galaxies (ELGs) and quasars (QSOs) are used as tracers of dark matter by the Dark Energy Spectroscopic Instrument (DESI) to probe this redshift regime. We present results from ELG and QSO mock catalogs created from the Uchuu N-body simulation and tuned to DESI Data Release 2 (DR2) clustering. Employing a modified subhalo abundance matching (SHAM) technique, we populate Uchuu halos and subhalos with QSOs between 0.8 < z < 2.1. For ELGs, we modify this method to select satellite galaxies with low velocities relative to their associated central halos, and populate a separate set of Uchuu halos and subhalos with ELGs between 0.8 < z < 1.6. In this paper, we reproduce the redshift evolution of number density and clustering statistics across the fitted range of scales. We also measure the large-scale clustering bias of both the data and mock samples. These results improve simulated lightcone construction from cosmological models and enhance our understanding of the galaxy-halo connection.

Vaisakh, R. [Southern Methodist U.] (ORCID:0009000↗

Galaxy-multiplet clustering from DESI DR2

We present an efficient estimator for higher-order galaxy clustering using small groups of nearby galaxies, or multiplets. Using the Luminous Red Galaxy (LRG) sample from the Dark Energy Spectroscopic Instrument (DESI) Data Release 2, we identify galaxy multiplets as discrete objects and measure their cross-correlations with the general galaxy field. Our results show that the multiplets exhibit stronger clustering bias as they trace more massive dark matter halos than individual galaxies. When comparing the observed clustering statistics with the mock catalogs generated from the N-body simulation AbacusSummit, we find that the mocks underpredict multiplet clustering despite reproducing the galaxy two-point auto-correlation reasonably well. This discrepancy indicates that the standard Halo Occupation Distribution (HOD) model is insufficient to describe the properties of galaxy multiplets, revealing the greater constraining power of this higher-order statistic on galaxy-halo connection and the possibility that multiplets are specific to additional assembly bias. We demonstrate that incorporating secondary biases into the HOD model improves agreement with the observed multiplet statistics, specifically by allowing galaxies to preferentially occupy halos in denser environments. Our results highlight the potential of utilizing multiplet clustering, beyond traditional two-point correlation measurements, to break degeneracies in models describing the galaxy-dark matter connection.

cosmology↗

Significant DBSCAN+: Statistically Robust Density-based Clustering

Cluster detection is important and widely used in a variety of applications, including public health, public safety, transportation, and so on. Given a collection of data points, we aim to detect density-connected spatial clusters with varying geometric shapes and densities, under the constraint that the clusters are statistically significant. The problem is challenging, because many societal applications and domain science studies have low tolerance for spurious results, and clusters may have arbitrary shapes and varying densities. As a classical topic in data mining and learning, a myriad of techniques have been developed to detect clusters with both varying shapes and densities (e.g., density-based, hierarchical, spectral, or deep clustering methods). However, the vast majority of these techniques do not consider statistical rigor and are susceptible to detecting spurious clusters formed as a result of natural randomness. On the other hand, scan statistic approaches explicitly control the rate of spurious results, but they typically assume a single “hotspot” of over-density and many rely on further assumptions such as a tessellated input space. To unite the strengths of both lines of work, we propose a statistically robust formulation of a multi-scale DBSCAN, namely Significant DBSCAN+, to identify significant clusters that are density connected. As we will show, incorporation of statistical rigor is a powerful mechanism that allows the new Significant DBSCAN+ to outperform state-of-the-art clustering techniques in various scenarios. We also propose computational enhancements to speed-up the proposed approach. Experiment results show that Significant DBSCAN+ can simultaneously improve the success rate of true cluster detection (e.g., 10–20% increases in absolute F1 scores) and substantially reduce the rate of spurious results (e.g., from thousands/hundreds of spurious detections to none or just a few across 100 datasets), and the acceleration methods can improve the efficiency for both clustered and non-clustered data.

Computer Science↗

Production of alternate realizations of DESI fiber assignment for unbiased clustering measurement in data and simulations

A critical requirement of spectroscopic large scale structure analyses is correcting for selection of which galaxies to observe from an isotropic target list. This selection is often limited by the hardware used to perform the survey which will impose angular constraints of simultaneously observable targets, requiring multiple passes to observe all of them. In SDSS this manifested solely as the collision of physical fibers and plugs placed in plates. In DESI, there is the additional constraint of the robotic positioner which controls each fiber being limited to a finite patrol radius. A number of approximate methods have previously been proposed to correct the galaxy clustering statistics for these effects, but these generally fail on small scales. To accurately correct the clustering we need to upweight pairs of galaxies based on the inverse probability that those pairs would be observed (Bianchi & Percival 2017). This paper details an implementation of that method to correct the Dark Energy Spectroscopic Instrument (DESI) survey for incompleteness. To calculate the required probabilities, we need a set of alternate realizations of DESI where we vary the relative priority of otherwise identical targets. These realizations take the form of alternate Merged Target Ledgers (AMTL), the files that link DESI observations and targets. We present the method used to generate these alternate realizations and how they are tracked forward in time using the real observational record and hardware status, propagating the survey as though the alternate orderings had been adopted. We detail the first applications of this method to the DESI One-Percent Survey (SV3) and the DESI year 1 data. We include evaluations of the pipeline outputs, estimation of survey completeness from this and other methods, and validation of the method using mock galaxy catalogs.

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND ↗

The impact of anisotropic redshift distributions on angular clustering

A leading way to constrain physical theories from cosmological observations is to test their predictions for the angular clustering statistics of matter tracers, a technique that is set to become ever more central with the next generation of large imaging surveys. Interpretation of this clustering requires knowledge of the projection kernel, or the redshift distribution of the sources, and the typical assumption is an isotropic redshift distribution for the objects. However, variations in the kernel are expected across the survey footprint due to photometric variations and residual observational systematic effects. Here, we develop the formalism for anisotropic projection and present several limiting cases that elucidate the key aspects. We quantify the impact of anisotropies in the redshift distribution on a general class of angular two-point statistics. In particular, we identify a mode-coupling effect that can add power to auto-correlations, including galaxy clustering and cosmic shear, and remove it from certain cross-correlations. If the projection anisotropy is primarily at large scales, the mode-coupling depends upon its variance as a function of redshift; furthermore, it is often of similar shape to the signal. In contrast, the cross-correlation of a field whose selection function is anisotropic with another one featuring no such variations — such as CMB lensing — is immune to these effects. We discuss explicitly several special cases of the general formalism including galaxy clustering, galaxy-galaxy lensing, cosmic shear and cross-correlations with CMB lensing, and publicly release a code to compute the biases.

79 ASTRONOMY AND ASTROPHYSICS↗

Covariance matrices for variance-suppressed simulations

ABSTRACT Cosmological N-body simulations provide numerical predictions of the structure of the Universe against which to compare data from ongoing and future surveys, but the growing volume of the Universe mapped by surveys requires correspondingly lower statistical uncertainties in simulations, usually achieved by increasing simulation sizes at the expense of computational power. It was recently proposed to reduce simulation variance without incurring additional computational costs by adopting fixed-amplitude initial conditions. This method has been demonstrated not to introduce bias in various statistics, including the two-point statistics of galaxy samples typically used for extracting cosmological parameters from galaxy redshift survey data, but requires us to revisit current methods for estimating covariance matrices of clustering statistics for simulations. In this work, we find that it is not trivial to construct covariance matrices analytically for fixed-amplitude simulations, but we demonstrate that ezmock (Effective Zel’dovich approximation mock catalogue), the most efficient method for constructing mock catalogues with accurate two- and three-point statistics, provides reasonable covariance matrix estimates for such simulations. We further examine how the variance suppression obtained by amplitude-fixing depends on three-point clustering, small-scale clustering, and galaxy bias, and propose intuitive explanations for the effects we observe based on the ezmock bias model.

79 ASTRONOMY AND ASTROPHYSICS↗

Precision redshift-space galaxy power spectra using Zel'dovich control variates

Numerical simulations in cosmology require trade-offs between volume, resolution and run-time that limit the volume of the Universe that can be simulated, leading to sample variance in predictions of ensemble-average quantities such as the power spectrum or correlation function(s). Sample variance is particularly acute at large scales, which is also where analytic techniques can be highly reliable. This provides an opportunity to combine analytic and numerical techniques in a principled way to improve the dynamic range and reliability of predictions for clustering statistics. In this paper we extend the technique of Zel'dovich control variates, previously demonstrated for 2-point functions in real space, to reduce the sample variance in measurements of 2-point statistics of biased tracers in redshift space. We demonstrate that with this technique, we can reduce the sample variance of these statistics down to their shot-noise limit out to k ~ 0.2 h Mpc -1 . This allows a better matching with perturbative models and improved predictions for the clustering of e.g. quasars, galaxies and neutral Hydrogen measured in spectroscopic redshift surveys at very modest computational expense. We discuss the implementation of ZCV, give some examples and provide forecasts for the efficacy of the method under various conditions.

79 ASTRONOMY AND ASTROPHYSICS↗

Brightest cluster galaxies are statistically special from z = 0.3 to z = 1

ABSTRACT We study brightest cluster galaxies (BCGs) in ∼5000 galaxy clusters from the Hyper Suprime-Cam (HSC) Subaru Strategic Program. The sample is selected over an area of 830 deg2 and is uniformly distributed in redshift over the range of z = 0.3−1.0. The clusters have stellar masses in the range of 1011.8−1012.9M⊙. We compare the stellar mass of the BCGs in each cluster to what we would expect if their masses were drawn from the mass distribution of the other member galaxies of the clusters. The BCGs are found to be ‘special’, in the sense that they are not consistent with being a statistical extreme of the mass distribution of other cluster galaxies. This result is robust over the full range of cluster stellar masses and redshifts in the sample, indicating that BCGs are special up to a redshift of z = 1.0. However, BCGs with a large separation from the centre of the cluster are found to be consistent with being statistical extremes of the cluster member mass distribution. We discuss the implications of these findings for BCG formation scenarios.

Dalal, Roohi (ORCID:0000000279989899)↗