Search NASA⌕ Search

SEARCH · Search NASA

Results for “Bayesian clustering”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

SPT clusters with DES and HST weak lensing. I. Cluster lensing and Bayesian population modeling of multiwavelength cluster datasets

We present a Bayesian population modeling method to analyze the abundance of galaxy clusters identified by the South Pole Telescope (SPT) with a simultaneous mass calibration using weak gravitational lensing data from the Dark Energy Survey (DES) and the Hubble Space Telescope (HST). We discuss and validate the modeling choices with a particular focus on a robust, weak-lensing-based mass calibration using DES data. For the DES Year 3 data, we report a systematic uncertainty in weak-lensing mass calibration that increases from 1% at z = 0.25 to 10% at z = 0.95 , to which we add 2% in quadrature to account for uncertainties in the impact of baryonic effects. We implement an analysis pipeline that joins the cluster abundance likelihood with a multiobservable likelihood for the Sunyaev-Zel’dovich effect, optical richness, and weak-lensing measurements for each individual cluster. We validate that our analysis pipeline can recover unbiased cosmological constraints by analyzing mocks that closely resemble the cluster sample extracted from the SPT-SZ, SPTpol ECS, and SPTpol 500d surveys and the DES Year 3 and HST-39 weak-lensing datasets. This work represents a crucial prerequisite for the subsequent cosmological analysis of the real dataset.

79 ASTRONOMY AND ASTROPHYSICS↗

Evaluation of the procedure 1A component of the 1980 US/Canada wheat and barley exploratory experiment

Several techniques which use clusters generated by a new clustering algorithm, CLASSY, are proposed as alternatives to random sampling to obtain greater precision in crop proportion estimation: (1) Proportional Allocation/relative count estimator (PA/RCE) uses proportional allocation of dots to clusters on the basis of cluster size and a relative count cluster level estimate; (2) Proportional Allocation/Bayes Estimator (PA/BE) uses proportional allocation of dots to clusters and a Bayesian cluster-level estimate; and (3) Bayes Sequential Allocation/Bayesian Estimator (BSA/BE) uses sequential allocation of dots to clusters and a Bayesian cluster level estimate. Clustering in an effective method in making proportion estimates. It is estimated that, to obtain the same precision with random sampling as obtained by the proportional sampling of 50 dots with an unbiased estimator, samples of 85 or 166 would need to be taken if dot sets with AI labels (integrated procedure) or ground truth labels, respectively were input. Dot reallocation provides dot sets that are unbiased. It is recommended that these proportion estimation techniques are maintained, particularly the PA/BE because it provides the greatest precision.

Chapman, G. M.↗

Tool Support for Parametric Analysis of Large Software Simulation Systems

The analysis of large and complex parameterized software systems, e.g., systems simulation in aerospace, is very complicated and time-consuming due to the large parameter space, and the complex, highly coupled nonlinear nature of the different system components. Thus, such systems are generally validated only in regions local to anticipated operating points rather than through characterization of the entire feasible operational envelope of the system. We have addressed the factors deterring such an analysis with a tool to support envelope assessment: we utilize a combination of advanced Monte Carlo generation with n-factor combinatorial parameter variations to limit the number of cases, but still explore important interactions in the parameter space in a systematic fashion. Additional test-cases, automatically generated from models (e.g., UML, Simulink, Stateflow) improve the coverage. The distributed test runs of the software system produce vast amounts of data, making manual analysis impossible. Our tool automatically analyzes the generated data through a combination of unsupervised Bayesian clustering techniques (AutoBayes) and supervised learning of critical parameter ranges using the treatment learner TAR3. The tool has been developed around the Trick simulation environment, which is widely used within NASA. We will present this tool with a GN&C (Guidance, Navigation and Control) simulation of a small satellite system.

Schumann, Johann↗

Copacabana: a probabilistic membership assignment method for galaxy clusters

Cosmological analyses using galaxy clusters in optical/near-infrared photometric surveys require robust characterization of their galaxy content. Precisely determining which galaxies belong to a cluster is crucial. In this paper, we present the COlor Probabilistic Assignment of Clusters And BAyesiaN Analysis (Copacabana) algorithm. Copacabana computes membership probabilities for all galaxies within an aperture centred on the cluster using photometric redshifts, colours, and projected radial probability density functions. We use simulations to validate Copacabana and we show that it achieves up to 89 per cent membership accuracy with a mild dependence on photometric redshift uncertainties and choice of aperture size. We find that the precision of the photometric redshifts has the largest impact on the determination of the membership probabilities followed by the choice of the cluster aperture size. We also quantify how much these uncertainties in the membership probabilities affect the stellar mass–cluster mass scaling relation, a relation that directly impacts cosmology. Using the sum of the stellar masses weighted by membership probabilities (⁠μ * ⁠) as the observable, we find that Copacabana can reach an accuracy of 0.06 dex in the measurement of the scaling relation at low redshift for a Legacy Survey of Space and Time type survey. These results indicate the potential of Copacabana and μ * to be used in cosmological analyses of optically selected clusters in the future.

79 ASTRONOMY AND ASTROPHYSICS↗

A Multi-Armed Bayesian Ordinal Outcome Utility-Based Sequential Trial with a Pairwise Null Clustering Prior

A multi-armed trial based on ordinal outcomes is proposed that leverages a flexible non-proportional odds cumulative logit model and numerical utility scores for each outcome to determine treatment optimality. This trial design uses a Bayesian clustering prior on the treatment effects that encourages the pairwise null hypothesis of no differences between treatments. A group sequential design is proposed to determine which treatments are clinically different with an adaptive decision boundary that becomes more aggressive as the sample size or clinical significance grows, or the number of active treatments decreases. A simulation study is conducted for 3 and 5 treatment arms, which shows that the design has superior operating characteristics (family wise error rate, generalized power, average sample size) compared to utility designs that do not allow clustering, a frequentist proportional odds model, or a permutation test based on empirical mean utilities.

97 MATHEMATICS AND COMPUTING↗

Precision calibration of calorimeter signals in the ATLAS experiment using an uncertainty-aware neural network

The ATLAS experiment at the Large Hadron Collider explores the use of modern neural networks for a multi-dimensional calibration of its calorimeter signal defined by clusters of topologically connected cells (topo-clusters). The Bayesian neural network (BNN) approach not only yields a continuous and smooth calibration function that improves performance relative to the standard calibration but also provides uncertainties on the calibrated energies for each topo-cluster. The results obtained by using a trained BNN are compared to the standard local hadronic calibration and to a calibration provided by training a deep neural network. The uncertainties predicted by the BNN are interpreted in the context of a fractional contribution to the systematic uncertainties of the trained calibration. They are also compared to uncertainty predictions obtained from an alternative estimator employing repulsive ensembles.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

Parametric Analysis of a Hover Test Vehicle using Advanced Test Generation and Data Analysis

Large complex aerospace systems are generally validated in regions local to anticipated operating points rather than through characterization of the entire feasible operational envelope of the system. This is due to the large parameter space, and complex, highly coupled nonlinear nature of the different systems that contribute to the performance of the aerospace system. We have addressed the factors deterring such an analysis by applying a combination of technologies to the area of flight envelop assessment. We utilize n-factor (2,3) combinatorial parameter variations to limit the number of cases, but still explore important interactions in the parameter space in a systematic fashion. The data generated is automatically analyzed through a combination of unsupervised learning using a Bayesian multivariate clustering technique (AutoBayes) and supervised learning of critical parameter ranges using the machine-learning tool TAR3, a treatment learner. Covariance analysis with scatter plots and likelihood contours are used to visualize correlations between simulation parameters and simulation results, a task that requires tool support, especially for large and complex models. We present results of simulation experiments for a cold-gas-powered hover test vehicle.

Gundy-Burlet, Karen↗

Alleviating prior dependencies for DESI DR1 clustering fits through reparameterization

Bayesian analyses of the full-shape clustering of Dark Energy Spectroscopic Instrument (DESI) Data Release 1 (DR1) exhibit prior-volume projection effects, whereby weakly constrained nuisance parameters of the Effective Field Theory of Large Scale Structure (EFTofLSS) shift marginalized cosmological posteriors away from the posterior maximum. We reanalyze DESI DR1 power spectrum multipoles using two complementary mitigation strategies: (i) nonlinear orthogonalization to decorrelate nuisance and cosmological parameter priors, and (ii) a fully reparameterization-invariant Jeffreys prior over all EFTofLSS coefficients, evaluated on-the-fly via closed-form Jacobians. Including data from DESI, Big-Bang Nuclesynthesis and a constraint on $n_{\mathrm{s}}$, baseline priors lead to multi-$σ$ projection in the Hubble parameter $H_{0}$ and dark energy equation of state parameters $w_{0}$ and $w_{a}$; the Jeffreys prior successfully recenters these posteriors to enclose the maximum a posteriori estimate within the 68% credible regions, demonstrating clear mitigation of projection effects for these late-time expansion parameters. A hybrid Jeffreys+baseline-Gaussian configuration controls residual over-broad tails in the physical cold dark matter density $ω_{\mathrm{c}}$ while preserving the volume correction, and is our favoured approach. We compare the credible intervals derived using our methodology to those obtained using Halo Occupation Distribution (HOD)-informed priors and to confidence intervals derived using frequentist profile likelihood analyses, finding agreement in both central values and degeneracy directions in the $w_{0}$--$w_{a}$ plane. This demonstrates that, once projection effects are properly controlled, we can make robust inferences about the late-time cosmological expansion independent of the statistical framework adopted.

Bonici, M. [Waterloo U.; Perimeter Inst. Theor. Ph↗

Bayesian prior construction for uncertainty quantification in first-principles statistical mechanics

First-principles statistical mechanics enables the prediction of thermodynamic and kinetic properties of materials, but is computationally expensive. Many approaches require surrogate models to calculate energies within Monte Carlo or molecular dynamics simulations. Inexpensive surrogates such as cluster expansions enable otherwise intractable calculations by interpolating data from higher accuracy methods, such as Density Functional Theory (DFT). Surrogate models introduce uncertainty into downstream calculations, in addition to any uncertainty inherent to DFT calculations. Bayesian frameworks address this by quantifying uncertainty and incorporating expert knowledge through priors. However, constructing effective priors remains challenging. This work introduces and describes practical strategies for building Bayesian cluster expansions, focusing on basis truncation, hyperparameter selection, and ground state replication. We analyze multiple basis truncation schemes, compare cross-validation to the evidence-approximation for hyperparameter optimization, and provide methods to find and enforce ground-state-preserving models through priors. Additionally, we compare the uncertainties between different approximations to DFT (LDA, PBE, SCAN) against the uncertainty introduced with the use of cluster expansion surrogate models. These approaches are demonstrated on the BCC Li x Mg 1-x and Li x Al 1-x alloys, which are both of interest for solid-state Li batteries. Our results provide guidelines for constructing and utilizing Bayesian cluster expansions, thereby improving the transparency of materials modeling. Furthermore, the approaches and insights developed in this work can be transferred to a wide range of cluster expansion surrogate models, including the atomic cluster expansion and related machine-learned interatomic potential architectures.

Alloy theory↗

Population Subdivision in the Gopher Frog (Rana capito) across the Fragmented Longleaf Pine-Wiregrass Savanna of the Southeastern USA

Delineating genetically distinct population segments of threatened species and quantifying population connectivity are important steps in developing effective conservation and management strategies aimed at preventing extinction. The gopher frog (Rana capito) is a xeric-adapted, pond-breeding species endemic to the Gulf and Atlantic coastal plains of the southeastern United States. This species has experienced extensive habitat loss and fragmentation in the formerly widespread longleaf pine-wiregrass savanna where it lives, resulting in individual abundance declines and population extinctions throughout its range. We used individual-based clustering methods along with Bayesian inference of historical migration based on almost 1500 multilocus microsatellite genotypes to examine genetic structure in this taxon. Clustering analyses identified panhandle and peninsular populations in Florida as distinct genetic clusters separated by the Aucilla River, consistent with the division between the Coastal Plain and peninsular mitochondrial lineages, respectively. Analysis of historical migration indicated an east–west population divergence event followed by immigration to the east. Together, our results indicate that the genetically distinct Coastal Plain and peninsular Florida lineages should be considered separately for conservation and management purposes.

59 BASIC BIOLOGICAL SCIENCES↗

Atacama Cosmology Telescope measurements of a large sample of candidates from the Massive and Distant Clusters of WISE Survey: Sunyaev-Zeldovich effect confirmation of MaDCoWS candidates using ACT

Context. Galaxy clusters are an important tool for cosmology, and their detection and characterization are key goals for current and future surveys. Using data from the Wide-field Infrared Survey Explorer (WISE), the Massive and Distant Clusters of WISE Survey (MaDCoWS) located 2839 significant galaxy overdensities at redshifts 0.7 . z . 1.5, which included extensive follow-up imaging from the Spitzer Space Telescope to determine cluster richnesses. Concurrently, the Atacama Cosmology Telescope (ACT) has produced large area millimeter-wave maps in three frequency bands along with a large catalog of Sunyaev-Zeldovich (SZ)-selected clusters as part of its Data Release 5 (DR5). Aims. We aim to verify and characterize MaDCoWS clusters using measurements of, or limits on, their thermal SZ effect signatures. We also use these detections to establish the scaling relation between SZ mass and the MaDCoWS-defined richness. Methods. Using the maps and cluster catalog from DR5, we explore the scaling between SZ mass and cluster richness. We do this by comparing cataloged detections and extracting individual and stacked SZ signals from the MaDCoWS cluster locations. We use complementary radio survey data from the Very Large Array, submillimeter data from Herschel, and ACT 224 GHz data to assess the impact of contaminating sources on the SZ signals from both ACT and MaDCoWS clusters. We use a hierarchical Bayesian model to fit the mass-richness scaling relation, allowing for clusters to be drawn from two populations: one, a Gaussian centered on the mass-richness relation, and the other, a Gaussian centered on zero SZ signal. Results. We find that MaDCoWS clusters have submillimeter contamination that is consistent with a gray-body spectrum, while the ACT clusters are consistent with no submillimeter emission on average. Additionally, the intrinsic radio intensities of ACT clusters are lower than those of MaDCoWS clusters, even when the ACT clusters are restricted to the same redshift range as the MaDCoWS clusters. We find the best-fit ACT SZ mass versus MaDCoWS richness scaling relation has a slope of p1 = 1.84+0.15 −0.14, where the slope is defined as M ∝ λ p1 15 and λ15 is the richness. We also find that the ACT SZ signals for a significant fraction (∼57%) of the MaDCoWS sample can statistically be described as being drawn from a noise-like distribution, indicating that the candidates are possibly dominated by low-mass and unvirialized systems that are below the mass limit of the ACT sample. Further, we note that a large portion of the optically confirmed ACT clusters located in the same volume of the sky as MaDCoWS are not selected by MaDCoWS, indicating that the MaDCoWS sample is not complete with respect to SZ selection. Finally, we find that the radio loud fraction of MaDCoWS clusters increases with richness, while we find no evidence that the submillimeter emission of the MaDCoWS clusters evolves with richness. Conclusions. We conclude that the original MaDCoWS selection function is not well defined and, as such, reiterate the MaDCoWS collaboration’s recommendation that the sample is suited for probing cluster and galaxy evolution, but not cosmological analyses. We find a best-fit mass-richness relation slope that agrees with the published MaDCoWS preliminary results. Additionally, we find that while the approximate level of infill of the ACT and MaDCoWS cluster SZ signals (1–2%) is subdominant to other sources of uncertainty for current generation experiments, characterizing and removing this bias will be critical for next-generation experiments hoping to constrain cluster masses at the sub-percent level.

large↗

Understanding the Scalability of Bayesian Network Inference Using Clique Tree Growth Curves

One of the main approaches to performing computation in Bayesian networks (BNs) is clique tree clustering and propagation. The clique tree approach consists of propagation in a clique tree compiled from a Bayesian network, and while it was introduced in the 1980s, there is still a lack of understanding of how clique tree computation time depends on variations in BN size and structure. In this article, we improve this understanding by developing an approach to characterizing clique tree growth as a function of parameters that can be computed in polynomial time from BNs, specifically: (i) the ratio of the number of a BN s non-root nodes to the number of root nodes, and (ii) the expected number of moral edges in their moral graphs. Analytically, we partition the set of cliques in a clique tree into different sets, and introduce a growth curve for the total size of each set. For the special case of bipartite BNs, there are two sets and two growth curves, a mixed clique growth curve and a root clique growth curve. In experiments, where random bipartite BNs generated using the BPART algorithm are studied, we systematically increase the out-degree of the root nodes in bipartite Bayesian networks, by increasing the number of leaf nodes. Surprisingly, root clique growth is well-approximated by Gompertz growth curves, an S-shaped family of curves that has previously been used to describe growth processes in biology, medicine, and neuroscience. We believe that this research improves the understanding of the scaling behavior of clique tree clustering for a certain class of Bayesian networks; presents an aid for trade-off studies of clique tree clustering using growth curves; and ultimately provides a foundation for benchmarking and developing improved BN inference and machine learning algorithms.

Mengshoel, Ole J.↗

Characterizing Oscillatory Bursts in Single-Trial EEG Data

Oscillatory bursts in numerous bands ranging from low (theta) to high frequencies (e.g., gamma) undoubtedly play an important role in cortical dynamics. Largely because of the inadequacy of existing analytic techniques. however, oscillatory bursts and their role in cortical processing remains poorly understood. To study oscillatory bursts effectively one must be able to isolate them and characterize them in the single trial. We describe a series of straightforward analysis techniques that produce useful indices of burst characteristics. First, stimulus-evoked responses are estimated using Differentially Variable Component Analysis (dVCA), and are subtracted from the single-trial. The single-trial characteristics of the evoked responses are stored to identify possible correlations with burst activity. Time-frequency (T-F), or wavelet, analyses are then applied to the single trial residuals. While T-F plots have been used in recent studies to identify and isolate bursts, we go further by fitting each burst in the T-F plot with a two-dimensional Gaussian. This provides a set of burst characteristics, such as, center time. burst duration, center frequency. frequency dispersion. and amplitude, all of which contribute to the accurate characterization of the individual burst. The burst phase can also be estimated. Burst characteristics can be quantified with several standard techniques (e.g.. histogramming and clustering), as well as Bayesian techniques (e.g., blocking) to allow a more parametric description analysis of the characteristics of oscillatory bursts, and the relationships of specific parameters to cortical excitability and stimulus integration.

Knuth, K. H.↗

Q-Cluster: Quantum Error Mitigation Through Noise-Aware Unsupervised Learning

Quantum error mitigation (QEM) is critical in reducing the impact of noise in the pre-fault-tolerant era, and is expected to complement error correction in fault-tolerant quantum computing (FTQC). In this work, we propose a novel QEM approach, Q-Cluster, that uses unsupervised learning (clustering) to reshape the measured bit-string distribution. Our approach starts with a simplified bit-flip noise model. It first performs clustering on noisy measurement results, i.e., bit-strings, based on the Hamming distance. The centroid of each cluster is calculated using a qubit-wise majority vote. Next, the noisy distribution is adjusted with the clustering outcomes and the bitflip error rates using Bayesian inference. Our simulation results show that Q-Cluster can mitigate high noise rates (up to 40% per qubit) with the simple bit-flip noise model. However, real quantum computers do not fit such a simple noise model. To address the problem, we (a) apply Pauli twirling to tailor the complex noise channels to Pauli errors, and (b) employ a machine learning model, ExtraTrees regressor, to estimate an effective bit-flip error rate using a feature vector consisting of machine calibration data (gate & measurement error rates), circuit features (number of qubits, numbers of different types of gates, etc.) and the shape of the noisy distribution (entropy). Our experimental results show that our proposed Q-Cluster scheme improves the fidelity by a factor of 1.46x, on average, compared to the unmitigated output distribution, for a set of low-entropy benchmarks on five different IBM quantum machines. Our approach outperforms the state-of-art QEM approaches RZNE [28], M3 [24], Hammer [35], and QBEEP [33] by 1.26x,1.29x,1.47x, and 2.65 x, respectively.

42 ENGINEERING↗

Frequentist cosmological constraints from full-shape clustering measurements in DESI DR1

We present a frequentist analysis of clustering measurements from Data Release 1 of the Dark Energy Spectroscopic Instrument (DESI) using the standard profile likelihood method. While Bayesian inferences for effective field theory models of galaxy clustering can be highly sensitive to prior choices for extended cosmological models, frequentist inferences are not susceptible to such effects. We compare frequentist and Bayesian constraints for the parameter set {σ 8 , H 0 , Ω m , w 0 , w a } using the full-shape power spectrum multipoles, post-reconstruction baryon acoustic oscillation (BAO) measurements, and external datasets from the CMB and type Ia supernovae measurements. The frequentist confidence intervals are significantly shifted relative to the Bayesian credible intervals for the w 0 w a CDM model, unless supernovae data are included. When DESI full-shape and BAO data are fit jointly, we obtain the following 1σ frequentist confidence intervals for ΛCDM (w 0 w a CDM): σ 8 = 0.863 +0.048 -0.040 , H 0 = 68.96 +0.81 -0.80 km s -1 Mpc -1 , Ω m = 0.3034 ± 0.0110 (σ 8 = 0.782 +0.060 -0.036 , H 0 = 63.7 +4.2 -2.0 km s -1 Mpc -1 , Ω m = 0.378 +0.024 -0.047 , w 0 = -0.16 +0.10 -0.50 , w a = -3.0 +1.7 ), corresponding to 0.8σ, 0.3σ, 0.7σ (2.1σ, 4.1σ, 6.5σ, 6.3σ, 6.6σ) shifts between the maximum likelihood estimate and the Bayesian posterior mean for ΛCDM (w 0 w a CDM) respectively.

Bayesian reasoning↗

Weak Lensing Mass Calibration of the ACT DR5 Galaxy Clusters with the DES Year 3 Weak Lensing Data

We use weak gravitational lensing measurements from Year 3 Dark Energy Survey data to calibrate the masses of 443 galaxy clusters selected via the Sunyaev-Zel'dovich effect from Atacama Cosmology Telescope Data Release 5 maps of the cosmic microwave background. We incorporate redshift and SZ measurements for individual clusters into a hierarchical model for the stacked lensing signals and perform Bayesian analyses to constrain the hydrostatic mass bias of the clusters. Our treatment of systematic uncertainties includes a prescription for measuring and accounting for the weak lensing boost factor, consideration of a miscentering effect, as well as marginalization over uncertainties in the source galaxy photometric redshift distributions and shear calibration. The resultant constraints on the normalization of the mass-observable relation have a precision of approximately 7%, with the mean WL halo mass of M $_{500c}$ = 5.4 × 10$^{14}$ M $_{⊙}$. We measure the bias between the true cluster mass and the mass estimated from the SZ signal based on an X-ray-calibrated scaling relation assuming hydrostatic equilibrium, to be 1 - b = 0.74$^{+0.06}$ $_{-0.05}$ over the full sample. When splitting the clusters into high (z = 0.43-0.70) and low (z = 0.15-0.43) redshift bins, we measure 1 - b = 0.58$^{+0.06}$ $_{-0.05}$ and 0.81$^{+0.08}$ $_{-0.06}$, respectively. When introducing additional freedom in redshift and mass to the hydrostatic bias model, we find that 1 - b decreases with redshift (with the power law of -1.8$^{+0.5}$ $_{-0.6}$, 99.95% confidence), consistent with findings from other recent studies, while we do not find any significant trend in mass. We also demonstrate that our result is robust against various systematics such as a scale cut, priors on baryonic and miscentering parameters, and degree of scatter in mass-observable relation. The weak-lensing mass calibration presented in this study will be a useful tool for using the ACT clusters as probes of astrophysics, and as a step towards using their abundance as a cosmological probe.

Shin, T. [Carnegie Mellon U.] (ORCID:0000000263895↗