Search NASA⌕ Search

SEARCH · Search NASA

Results for “clustering statistics”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

The Aemulus Project. VI. Emulation of Beyond-standard Galaxy Clustering Statistics to Improve Cosmological Constraints

Abstract There is untapped cosmological information in galaxy redshift surveys in the nonlinear regime. In this work, we use the Aemulus suite of cosmological N -body simulations to construct Gaussian process emulators of galaxy clustering statistics at small scales (0.1–50 h −1 Mpc) in order to constrain cosmological and galaxy bias parameters. In addition to standard statistics—the projected correlation function w p ( r p ), the redshift-space monopole of the correlation function ξ 0 ( s ), and the quadrupole ξ 2 ( s )—we emulate statistics that include information about the local environment, namely the underdensity probability function P U ( s ) and the density-marked correlation function M ( s ). This extends the model of Aemulus III for redshift-space distortions by including new statistics sensitive to galaxy assembly bias. In recovery tests, we find that the beyond-standard statistics significantly increase the constraining power on cosmological parameters of interest: including P U ( s ) and M ( s ) improves the precision of our constraints on Ω m by 27%, σ 8 by 19%, and the growth of structure parameter, f σ 8 , by 12% compared to standard statistics. We additionally find that scales below ∼6 h −1 Mpc contain as much information as larger scales. The density-sensitive statistics also contribute to constraining halo occupation distribution parameters and a flexible environment-dependent assembly bias model, which is important for extracting the small-scale cosmological information as well as understanding the galaxy–halo connection. This analysis demonstrates the potential of emulating beyond-standard clustering statistics at small scales to constrain the growth of structure as a test of cosmic acceleration.

79 ASTRONOMY AND ASTROPHYSICS↗

A Parameter-masked Mock Data Challenge for Beyond-two-point Galaxy Clustering Statistics

The past few years have seen the emergence of a wide array of novel techniques for analyzing high-precision data from upcoming galaxy surveys, which aim to extend the statistical analysis of galaxy clustering data beyond the linear regime and the canonical two-point (2pt) statistics. We test and benchmark some of these new techniques in a community data challenge named “Beyond-2pt,” initiated during the Aspen 2022 Summer Program “Large-Scale Structure Cosmology beyond 2-Point Statistics,” whose first round of results we present here. The challenge data set consists of high-precision mock galaxy catalogs for clustering in real space, in redshift space, and on a light cone. Participants in the challenge have developed end-to-end pipelines to analyze mock catalogs and extract unknown (“masked”) cosmological parameters of the underlying ΛCDM models with their methods. The methods represented are density-split clustering, nearest neighbor statistics, BACCO power spectrum emulator, void statistics, LEFTfield field-level inference using effective field theory (EFT), and joint power spectrum and bispectrum analyses using both EFT and simulation-based inference. In this work, we review the results of the challenge, focusing on problems solved, lessons learned, and future research needed to perfect the emerging beyond-2pt approaches. The unbiased parameter recovery demonstrated in this challenge by multiple statistics and the associated modeling and inference frameworks supports the credibility of cosmology constraints from these methods. The challenge data set is publicly available, and we welcome future submissions from methods that are not yet represented.

Krause, Elisabeth [Univ. of Arizona, Tucson, AZ (U↗

Performance analysis of the Alliant FX/8 multiprocessor using statistical clustering

Results for two distinct, real, scientific workloads executed on an Alliant FX/8 are discussed. A combination of user concurrency and system overhead measurements was taken for both workloads. Preliminary analysis shows that the first sampled workload is comprised of consistently high user concurrency, low system overhead, and little paging. The second sample has much less user concurrency, but significant paging and system overhead. Statistical cluster analysis is used to extract a state transition model to jointly characterize user concurrency and system overhead. A skewness factor is introduced and used to bring out the effects of unbalanced clustering when determining states with important transitions. The results from the models show that during the collection of the first sample, the system was operating in states of high user concurrency approximately 75 percent of the time. The second workload sample shows the system in high user concurrency states only 26 percent of the time. In addition, it is ascertained that high system overhead is usually accompanied by low user concurrency. The analysis also shows a high predictability of system behavior for both workloads.

Dimpsey, Robert Tod↗

Full forward model of galaxy clustering statistics with AbacusSummit light cones

ABSTRACT Novel summary statistics beyond the standard 2-point correlation function (2PCF) are necessary to capture the full astrophysical and cosmological information from the small-scale (r < 30h−1Mpc) galaxy clustering. However, the analysis of beyond-2PCF statistics on small scales is challenging because we lack the appropriate treatment of observational systematics for arbitrary summary statistics of the galaxy field. In this paper, we develop a full forward modelling pipeline for a wide range of summary statistics using the large high-fidelity AbacusSummit light cones that account for many systematic effects as well as remain flexible and computationally efficient to enable posterior sampling. We apply our forward model approach to a fully realistic mock galaxy catalog and demonstrate that we can recover unbiased constraints on the underlying galaxy–halo connection model using two separate summary statistics: the standard 2PCF and the novel k-th nearest neighbour (kNN) statistics, which are sensitive to correlation functions of all orders. We will demonstrate its strong constraining power on extended galaxy–halo connection models and cosmology in follow up papers. We expect this to become a powerful approach when applying to upcoming surveys such as DESI where we can leverage a multitude of summary statistics across a wide redshift range to maximally extract information from the non-linear scales.

79 ASTRONOMY AND ASTROPHYSICS↗

Discovery of Activities via Statistical Clustering of Fixation Patterns

Human behavior often consists of a series of distinct activities, each characterized by a unique pattern of interaction with the visual environment. This is true even in a restricted domain, such as a piloting an aircraft, where activities with distinct visual signatures might be things like communicating, navigating, and monitoring. We propose a novel analysis method for gaze-tracking data, to perform blind discovery of these hypothetical activities. The method is in some respects similar to recurrence analysis, but here we compare not individual fixations, but groups of fixations aggregated over a fixed time interval. The duration of this interval is a parameter that we will refer to as delta. We assume that the environment has been divided into a set of N different areas-of-interest (AOIs). For a given interval of time of duration delta, we compute the proportion of time spent fixating each AOI, resulting in an N-dimensional vector. These proportions can be converted to integer counts by multiplying by delta divided by the average fixation duration (another parameter that we fix at 280 milliseconds). We compare different intervals by computing the chi-square statistic. The p-value associated with the statistic is the likelihood of observing the data under the hypothesis that the data in the two intervals were generated by a single process with a single set of probabilities governing the fixation of each AOI. The method has been applied to approximately 100 hours of eye movement data collected from pilots in a high-fidelity B747 flight simulator, and the results have been compared to synthetic data in which the each activity is represented as first-order Markov process with random probabilities assigned to the AOIs. Randomly-generated synthetic activities can require thousands of fixations to be discriminated with statistical significance, while the human data can be clustered using averaging windows of some 10's of seconds, suggesting that the actual activities are much more narrowly focused than random Markov models.

activity analysis↗

DESI 2024 II: sample definitions, characteristics, and two-point clustering statistics

We present the samples of galaxies and quasars used for DESI 2024 cosmological analyses, drawn from the DESI Data Release 1 (DR1). We describe the construction of largescale structure (LSS) catalogs from these samples, which include matched sets of synthetic reference ‘randoms’ and weights that account for variations in the observed density of the samples due to experimental design and varying instrument performance. We detail how we correct for variations in observational completeness, the input ‘target’ densities due to imaging systematics, and the ability to confidently measure redshifts from DESI spectra. We then summarize how remaining uncertainties in the corrections can be translated to systematic uncertainties for particular analyses. We describe the weights added to maximize the signalto-noise of DESI DR1 2-point clustering measurements. We detail measurement pipelines applied to the LSS catalogs that obtain 2-point clustering measurements in configuration and Fourier space. The resulting 2-point measurements depend on window functions and normalization constraints particular to each sample, and we present the corrections required to match models to the data. We compare the configuration- and Fourier-space 2-point clustering of the data samples to that recovered from simulations of DESI DR1 and find they are, generally, in statistical agreement to within 2% in the inferred real-space over-density field. The LSS catalogs, 2-point measurements, and their covariance matrices will be released publicly with DESI DR1.

79 ASTRONOMY AND ASTROPHYSICS↗

Discovery of Activities via Statistical Clustering of Fixation Patterns

Human behavior often consists of a series of distinct activities, each characterized by a unique pattern of interaction with the visual environment. This is true even in a restricted domain, such as a pilot flying an airplane; in this case, activities with distinct visual signatures might be things like communicating, navigating, monitoring, etc. We propose a novel analysis method for gaze-tracking data, to perform blind discovery of these hypothetical activities. We compare, not individual fixations, but groups of fixations aggregated over a fixed time interval (Tau). We assume that the environment has been divided into a finite set of discrete areas-of-interest (AOIs). For a given time interval, we compute the proportion of time spent fixating each AOI, resulting in an N-dimensional vector, where N is the number of AOIs. These proportions can be converted to integer counts by multiplying by Tau divided by the average fixation duration, a parameter that we fix at 283 milliseconds. We compare different intervals by computing the chi-squared statistic. The p-value associated with the statistic is the likelihood of observing the data under the hypothesis that the data in the two intervals were generated by a single process with a single set of probabilities governing the fixation of each AOI. We cluster the intervals, first by merging adjacent intervals that are sufficiently similar, optionally shifting the boundary between non-merged intervals to maximize the difference. Then we compare and cluster non-adjacent intervals. The method is evaluated using synthetic data generated by a hand-crafted set of activities. While the method generally finds more activities than put into the simulation, we have obtained agreement as high as 80 percent between the inferred activity labels and ground truth.

Eye Movements↗

Discovery of Activities via Statistical Clustering of Fixation Patterns

Human behavior often consists of a series of distinct activities, each characterized by a unique signature of visual behavior. This is true even in a restricted domain, such as piloting an aircraft, where patterns of visual signatures might represent activities like communicating, navigating, and monitoring. We propose a novel analysis method for gaze-tracking data, to perform blind discovery of these activities based on their behavioral signatures. The method is in some respects similar to recurrence analysis, but here we compare not individual fixations, but groups of fixations aggregated over a fixed time interval. The duration of this interval is a parameter that we will refer to as τ. We assume that the environment has been divided into a set of N different areas-of-interest (AOIs). For a given interval of time of duration τ, we compute the proportion of time spent fixating each AOI, resulting in an N-dimensional vector. These proportions can be converted to counts by multiplying by τ divided by the average fixation duration (another parameter that we fix at 280 milliseconds). We compare different intervals by computing the chi-square statistic. The p-value associated with the statistic is the likelihood of observing the data under the hypothesis that the data in the two intervals were generated by a single process with a single set of probabilities governing the fixation of each AOI. We have investigated the method using a set of 10 synthetic "activities," that sample 4 AOIs. Four of these activities visit 3 of the 4 AOIs, with equal probability; as there are four different ways to leave-one- out, there are four such activities. Similarly, there are six different activities that leave-two-out. Sequences of simulated behavior were generated by running each activity for 40 seconds, in sequence, for a total of 6.7 minutes. The figure to the right shows the matrix of chi-square statistics, using a value of 2.8 seconds for τ, corresponding to 10 fixations. Low values (dark) indicate poor evidence for activity differences, while high values (bright) indicate strong evidence. The dark squares along the main diagonal each correspond to the forty second intervals in which the activity was held constant; the 4x4 block at the lower left corresponds to the four leave-one-out activities, while the 6x6 block in the upper right corresponds to the leave-two-out activities. (The anti-diagonal pattern of white squares indicates those activity pairs that share no AOIs.) The chi-square values can be binarized by choosing a particular significance level; we are interested in grouping bins that represent the same activity, effectively accepting the null hypothesis. Therefore, we may adopt a relatively lax criterion; for example, choosing a p-value of 0.2 means that two behaviors that have only a 1-in-5 chance of being produced by a single activity might nevertheless be clustered together. We have explored several methods to perform clustering on the data and solving for the activity probabilities. Greedy methods begin by selecting the time bin that is similar to the most (or least) other bins, and then forming a cluster from it and all other non-discriminable bins. These methods show mediocre performance, as they do not take into account temporal contiguity. Preliminary results indicate that methods that "grow" clusters in time from seed points perform better.

activity analysis↗

Self-organization of cosmic radiation pressure instability. II - One-dimensional simulations

The clustering of statistically uniform discrete absorbing particles moving solely under the influence of radiation pressure from uniformly distributed emitters is studied in a simple one-dimensional model. Radiation pressure tends to amplify statistical clustering in the absorbers; the absorbing material is swept into empty bubbles, the biggest bubbles grow bigger almost as they would in a uniform medium, and the smaller ones get crushed and disappear. Numerical simulations of a one-dimensional system are used to support the conjecture that the system is self-organizing. Simple statistics indicate that a wide range of initial conditions produce structure approaching the same self-similar statistical distribution, whose scaling properties follow those of the attractor solution for an isolated bubble. The importance of the process for large-scale structuring of the interstellar medium is briefly discussed.

Hogan, Craig J.↗

The DESI One-Percent Survey: Modelling the clustering and halo occupation of all four DESI tracers with U CHUU

We present results from a set of mock lightcones for the DESI One-Percent Survey, created from the UCHUU simulation. This 8 h −3 Gpc 3 N-body simulation comprises 2.1 trillion particles and provides high-resolution dark matter (sub)haloes in the framework of the Planck-based ΛCDM cosmology. Employing the subhalo abundance matching (SHAM) technique, we populated the UCHUU (sub)haloes with all four DESI tracers – Bright Galaxy Survey (BGS), luminous red galaxies (LRGs), emission line galaxies (ELGs), and quasars (QSOs) – to z = 2.1. Our method accounts for redshift evolution as well as the clustering dependence on luminosity and stellar mass. The two-point clustering statistics of the DESI One-Percent Survey generally agree with predictions from UCHUU across scales ranging from 0.3 h −1 Mpc to 100 h −1 Mpc for the BGS and across scales ranging from 5 h −1 Mpc to 100 h −1 Mpc for the other tracers. We observed some differences in clustering statistics that can be attributed to incompleteness of the massive end of the stellar mass function of LRGs, our use of a simplified galaxy-halo connection model for ELGs and QSOs, and cosmic variance. We find that at the high precision of UCHUU, the shape of the halo occupation distribution (HOD) of the BGS and LRG samples is smaller bias values, likely due to cosmic variance. The bias dependence on absolute magnitude, stellar mass, and redshift aligns with that of previous surveys. These results provide DESI with tools to generate high-fidelity lightcones for the remainder of the survey and enhance our understanding of the galaxy-halo connection.

cosmology↗

First Constraints on Growth Rate from Redshift-space Ellipticity Correlations of SDSS Galaxies at 0.16 < z < 0.70

We report the first constraints on the growth rate of the universe, f(z)σ 8 (z), with intrinsic alignments (IAs) of galaxies. We measure the galaxy density-intrinsic ellipticity cross-correlation and intrinsic ellipticity autocorrelation functions over 0.16 < z < 0.7 from luminous red galaxies (LRGs) and LOWZ and CMASS galaxy samples in the Sloan Digital Sky Survey (SDSS) and SDSS-III BOSS survey. We detect clear anisotropic signals of IA due to redshift-space distortions. By combining measured IA statistics with the conventional galaxy clustering statistics, we obtain tighter constraints on the growth rate. The improvement is particularly prominent for the LRG, which is the brightest galaxy sample and known to be strongly aligned with underlying dark matter distribution; using the measurements on scales above 10 h -1 Mpc, we obtain $f{\sigma }_{8}={0.5196}_{-0.0354}^{+0.0352}$ (68% confidence level) from the clustering-only analysis and $f{\sigma }_{8}={0.5322}_{-0.0291}^{+0.0293}$ with clustering and IA, meaning 19% improvement. The constraint is in good agreement with the prediction of general relativity, f σ 8 = 0.4937 at z = 0.34. For LOWZ and CMASS samples, the improvement of constraints on f σ 8 is found to be 10% and 3.5%, respectively. Our results indicate that the contribution from IA statistics for cosmological constraints can be further enhanced by carefully selecting galaxies for a shape sample.

79 ASTRONOMY AND ASTROPHYSICS↗

High-precision Galaxy Clustering Predictions from Small-volume Hydrodynamical Simulations via Control Variates

Abstract Cosmological simulations of galaxy formation are an invaluable tool for understanding galaxy formation and its impact on cosmological parameter inference from large-scale structures. However, their high computational cost is a significant obstacle for running simulations that probe cosmological volumes comparable to those analyzed by contemporary large-scale structure experiments. In this work, we explore the possibility of obtaining high-precision galaxy clustering predictions from small-volume hydrodynamical simulations such as MillenniumTNG and FLAMINGO via control variates. In this approach, the hydrodynamical full-physics simulation is paired with a matched low-resolution gravity-only simulation. By learning the galaxy–halo connection from the hydrodynamical simulation and applying it to the gravity-only counterpart, one obtains a galaxy population that closely mimics the one in the more expensive simulation. One can then construct an estimator of galaxy clustering that combines the clustering amplitudes in the small-volume hydrodynamical and gravity-only simulations with clustering amplitudes in a large-volume gravity-only simulation. Depending on the galaxy sample, clustering statistic, and scale, this galaxy clustering estimator can have an effective volume of up to around 100 times the volume of the original hydrodynamical simulation in the nonlinear regime. With this approach, we can construct galaxy clustering predictions from existing simulations that are precise enough for mock analyses of next-generation large-scale structure surveys such as the Dark Energy Spectroscopic Instrument and the Legacy Survey of Space and Time.

Doytcheva, Alexandra (ORCID:0009000111254888)↗

An unsupervised classification technique for multispectral remote sensing data.

Description of a two-part clustering technique consisting of (a) a sequential statistical clustering, which is essentially a sequential variance analysis, and (b) a generalized K-means clustering. In this composite clustering technique, the output of (a) is a set of initial clusters which are input to (b) for further improvement by an iterative scheme. This unsupervised composite technique was employed for automatic classification of two sets of remote multispectral earth resource observations. The classification accuracy by the unsupervised technique is found to be comparable to that by traditional supervised maximum-likelihood classification techniques.

Su, M. Y.↗

Using Clustering to Establish Climate Regimes from PCM Output

A multivariate statistical clustering technique--based on the k-means algorithm of Hartigan has been used to extract patterns of climatological significance from 200 years of general circulation model (GCM) output. Originally developed and implemented on a Beowulf-style parallel computer constructed by Hoffman and Hargrove from surplus commodity desktop PCs, the high performance parallel clustering algorithm was previously applied to the derivation of ecoregions from map stacks of 9 and 25 geophysical conditions or variables for the conterminous U.S. at a resolution of 1 sq km. Now applied both across space and through time, the clustering technique yields temporally-varying climate regimes predicted by transient runs of the Parallel Climate Model (PCM). Using a business-as-usual (BAU) scenario and clustering four fields of significance to the global water cycle (surface temperature, precipitation, soil moisture, and snow depth) from 1871 through 2098, the authors' analysis shows an increase in spatial area occupied by the cluster or climate regime which typifies desert regions (i.e., an increase in desertification) and a decrease in the spatial area occupied by the climate regime typifying winter-time high latitude perma-frost regions. The patterns of cluster changes have been analyzed to understand the predicted variability in the water cycle on global and continental scales. In addition, representative climate regimes were determined by taking three 10-year averages of the fields 100 years apart for northern hemisphere winter (December, January, and February) and summer (June, July, and August). The result is global maps of typical seasonal climate regimes for 100 years in the past, for the present, and for 100 years into the future. Using three-dimensional data or phase space representations of these climate regimes (i.e., the cluster centroids), the authors demonstrate the portion of this phase space occupied by the land surface at all points in space and time. Any single spot on the globe will exist in one of these climate regimes at any single point in time. By incrementing time, that same spot will trace out a trajectory or orbit between and among these climate regimes (or atmospheric states) in phase (or state) space. When a geographic region enters a state it never previously visited, a climatic change is said to have occurred. Tracing out the entire trajectory of a single spot on the globe yields a 'manifold' in state space representing the shape of its predicted climate occupancy. This sort of analysis enables a researcher to more easily grasp the multivariate behavior of the climate system.

Oglesby, Robert↗

Mimas: Preliminary Evidence For Amorphous Water Ice from VIMS

We have conducted a statistical clustering analysis (1,2) on a mosaic of VIMS data cubes obtained on February 13, 2010, for Saturn s satellite Mimas. Seven VIMS cubes were geometrically projected and re-sampled to a common spatial resolution. The clustering technique consists of a partitioning algorithm coupled to a criterion that prevents sub-optimal solutions and tests for the influence of random noise in the measurements. The clustering technique is agnostic about the meaning of the clusters, and scientific interpretation requires their a posteriori evaluation. The preliminary results yielded five clusters, demonstrating that spectral variability across Mimas surface is statistically significant. The ratios of the means calculated for each of the clusters show structure within the 1.6- micron water ice band, as well as the shape and the central wavelength of the strong ice band at 2 micron, that map spatially in patterns apparently related to the topography of Mimas, in particular certain regions in and around Herschel crater. The mean spectra of the five clusters, show similarities with laboratory spectra of amorphous and crystalline H2O ice (3) that are suggestive of the presence of an amorphous ice component in certain regions of Mimas, notably on the central peak of Herschel, on the crater floor, and in faults surrounding the crater. This may represent a mixture of both ice phases, or perhaps a layer of amorphous ice on a base of crystalline ice. Another possible occurrence of amorphous ice appears southwest of Herschel, close to the south pole.

Cruikshank, Dale P.↗