Search NASA⌕ Search

SEARCH · Search NASA

Results for “ensemble clustering”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

Efficient Agent-Based Cluster Ensembles

Numerous domains ranging from distributed data acquisition to knowledge reuse need to solve the cluster ensemble problem of combining multiple clusterings into a single unified clustering. Unfortunately current non-agent-based cluster combining methods do not work in a distributed environment, are not robust to corrupted clusterings and require centralized access to all original clusterings. Overcoming these issues will allow cluster ensembles to be used in fundamentally distributed and failure-prone domains such as data acquisition from satellite constellations, in addition to domains demanding confidentiality such as combining clusterings of user profiles. This paper proposes an efficient, distributed, agent-based clustering ensemble method that addresses these issues. In this approach each agent is assigned a small subset of the data and votes on which final cluster its data points should belong to. The final clustering is then evaluated by a global utility, computed in a distributed way. This clustering is also evaluated using an agent-specific utility that is shown to be easier for the agents to maximize. Results show that agents using the agent-specific utility can achieve better performance than traditional non-agent based methods and are effective even when up to 50% of the agents fail.

Agogino, Adrian↗

Hydrogen Evolution on Electrode‐Supported Pt n Clusters: Ensemble of Hydride States Governs the Size Dependent Reactivity

Abstract We report the size‐dependent activity and stability of supported Pt 1,4,7,8 for electrocatalytic hydrogen evolution reaction, and show that clusters outperform polycrystalline Pt in activity, with size‐dependent stability. To understand the size effects, we use DFT calculations to study the structural fluxionality under varying potentials. We show that the clusters can reshape under H coverage and populate an ensemble of states with diverse stoichiometry, structure, and thus reactivity. Both experiment and theory suggest that electrocatalytic species are hydridic states of the clusters (≈2 H/Pt). An ensemble‐based kinetic model reproduces the experimental activity trend and reveals the role of metastable states. The stability trend is rationalized by chemical bonding analysis. Our joint study demonstrates the potential‐ and adsorbate‐coverage‐dependent fluxionality of subnano clusters of different sizes and offers a systematic modeling strategy to tackle the complexities.

Zhang, Zisheng↗

Hydrogen Evolution on Electrode–Supported Ptn Clusters: Ensemble of Hydride States Governs the Size Dependent Reactivity

We report the size-dependent activity and stability of supported Pt1,4,7,8 for electrocatalytic hydrogen evolution reaction, and show that clusters outperform polycrystalline Pt in activity, with sizedependent stability. To understand the size effects, we use DFT calculations to study the structural fluxionality under varying potentials. We show that the clusters can reshape under H coverage and populate an ensemble of states with diverse stoichiometry, structure, and thus reactivity. Both experiment and theory suggest that electrocatalytic species are hydridic states of the clusters (~2 H/Pt). An ensemble-based kinetic model reproduces the experimental activity trend and reveals the role of metastable states. Furthermore, the stability trend is rationalized by chemical bonding analysis. Our joint study demonstrates the potential- and adsorbate-coverage-dependent fluxionality of subnano clusters of different sizes and offers a systematic modeling strategy to tackle the complexities.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Cryptate binding energies towards high throughput chelator design: metadynamics ensembles with cluster–continuum solvation

A tiered forcefield/semiempirical/meta-GGA pipeline together with a thermodynamic scheme designed with error cancellation in mind was developed to calculate binding energies of [2.2.2] cryptate complexes of mono- and divalent cations. Stable complexes of Na, K, Rb, Ca, Zn and Pb were generated, revealing consistent cation–N lengths but highly variable cation–O lengths and an amine stacking mechanism potentially augmenting the cation size selectivity. Metadynamics, used for searching the high-dimensional potential energy surface, together with a cluster–continuum model for affordable – yet accurate – solvation modeling, enabled the discovery of more stable geometries than those previously reported. Similar solvation energy curve shapes for lone vs. coordinated ions enabled rapid solvation convergence via the cancellation of errors stemming from finite cluster sizes. In conclusion, an R 2 of 0.850 vs. experimental aqueous binding energies was obtained, validating this scheme as the backbone of a high-throughput workflow for chelator design.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Unravelling the orbits of cluster galaxy populations according to their dominant gas ionization source

ABSTRACT We investigate the kinematical and dynamical properties of cluster galaxy populations classified according to their dominant source of gas ionization, namely: star-forming (SF) galaxies, optical active galactic nuclei (AGNs), mixed SF plus AGN ionization (transition objects, T), and quiescent (Q) galaxies. We stack 8892 member galaxies from 336 relaxed galaxy clusters to build an ensemble cluster and estimate the observed projected profiles of numerical density and velocity dispersion, $\sigma _P(R)$, of each galaxy population. The MAMPOSSt code and the Jeans equations inversion technique are used to constrain the velocity anisotropy profiles of the galaxy populations in both parametric and non-parametric ways. We find that Q (SF) galaxies display the lowest (highest) typical cluster-centric distances and velocity dispersion values. Transition galaxies are more concentrated and tend to exhibit lower velocity dispersion values than SF galaxies. Galaxies that host an optical AGN are as concentrated as Q galaxies but display velocity dispersion values similar to those of the SF population. MAMPOSSt is able to find equilibrium solutions that successfully recover the observed $\sigma _P(R)$ profile only for the Q, T, and AGN populations. We find that the orbits of all populations are consistent with isotropy in the inner regions, becoming increasingly radial with the distance from the cluster centre. These results suggest that Q galaxies are in equilibrium within their clusters, while SF galaxies have more recently arrived in the cluster environment. Finally, the T and AGN populations appear to be in an intermediate dynamical state between those of the SF and Q populations.

Valk, Greique A. (ORCID:0009000827731299)↗

A Novel Data Segmentation Method for Data-driven Phase Identification

This paper presents a smart meter phase identification algorithm for two cases: meter-phase-label-known and meter-phase-label-unknown. To improve the identification accuracy, a data segmentation method is proposed to exclude data segments that are collected when the voltage correlation between smart meters on the same phase is weakened. Then, using the selected data segments, a hierarchical clustering method is used to calculate the correlation distances and cluster the smart meters. If the phase labels are unknown, a Connected-Triple-based Similarity (CTS) method is adapted to further improve the phase identification accuracy of the ensemble clustering method. The methods are developed and tested on both synthetic and real feeder data sets. Here, simulation results show that the proposed phase identification algorithm outperforms the state-of-the-art methods in both accuracy and robustness.

24 POWER TRANSMISSION AND DISTRIBUTION↗

Distilling Knowledge from Ensembles of Cluster-Constrained-Attention Multiple-Instance Learners for Whole Slide Image Classification

The peculiar nature of whole slide imaging (WSI), digitizing conventional glass slides to obtain multiple high resolution images which capture microscopic details of a patient’s histopathological features, has garnered increased interest from the computer vision research community over the last two decades. Given the unique computational space and time complexity inherent to gigapixel-size whole slide image data, researchers have proposed novel machine learning algorithms to aid in the performance of diagnostic tasks in clinical pathology. One effective algorithm represents a Whole slide image as a bag of smaller image patches, which can be represented as low-dimension image patch embeddings. Weakly supervised deep-learning methods, such as cluster-constrained-attention multiple instance learning (CLAM), have shown promising results when combined with image patch embeddings. While traditional ensemble classifiers yield improved task performance, such methods come with a steep cost in model complexity. Through knowledge distillation, it is possible to retain some performance improvements from an ensemble, while minimizing costs to model complexity. In this work, we implement a weakly supervised ensemble using clustering-constrained-attention multiple-instance learners (CLAM), which uses attention and instance-level clustering to identify task salient regions and feature extraction in whole slides. By applying logit-based and attention-based knowledge distillation, we show it is possible to retain some performance improvements resulting from the ensemble at zero cost to model complexity.

Alamudun, Folami↗

Robust design of semi-automated clustering models for 4D-STEM datasets

Materials discovery and design require characterizing material structures at the nanometer and sub-nanometer scale. Four-Dimensional Scanning Transmission Electron Microscopy (4D-STEM) resolves the crystal structure of materials, but many 4D-STEM data analysis pipelines are not suited for the identification of anomalous and unexpected structures. This work introduces improvements to the iterative Non-Negative Matrix Factorization (NMF) method by implementing consensus clustering for ensemble learning. We evaluate the performance of models during parameter tuning and find that consensus clustering improves performance in all cases and is able to recover specific grains missed by the best performing model in the ensemble. The methods introduced in this work can be applied broadly to materials characterization datasets to aid in the design of new materials.

Bruefach, Alexandra (ORCID:0000000209323477)↗

Analytical and EZmock covariance validation for the DESI 2024 results

The estimation of uncertainties in cosmological parameters is an important challenge in Large-Scale-Structure (LSS) analyses. For standard analyses such as Baryon Acoustic Oscillations (BAO) and Full-Shape two approaches are usually considered. First: analytical estimates of the covariance matrix use Gaussian approximations and (nonlinear) clustering measurements to estimate the matrix, which allows a relatively fast and computationally cheap way to generate matrices that adapt to an arbitrary clustering measurement. On the other hand, sample covariances are an empirical estimate of the matrix based on an ensemble of clustering measurements from fast and approximate simulations. While more computationally expensive due to the large amount of simulations and volume required, these allow us to take into account systematics that are impossible to model analytically. In this work we compare these two approaches in order to enable DESI's key analyses. We find that the configuration space analytical estimate performs satisfactorily in BAO analyses and its flexibility in terms of input clustering makes it the fiducial choice for DESI's 2024 BAO analysis. On the contrary, the analytical computation of the covariance matrix in Fourier space does not reproduce the expected measurements in terms of Full-Shape analyses, which motivates the use of a corrected mock covariance for DESI's 2024 Full Shape analysis.

79 ASTRONOMY AND ASTROPHYSICS↗

Cosmological Implications of the Effects of X-Ray Clusters on the Cosmic Microwave Background

We have been carrying forward a program to confront X-ray observations of clusters and their evolution as derived from X-ray observatories with observations of the cosmic microwave background radiation (CMBR). In addition to the material covered in our previous reports (including three published papers), most recently we have explored the effects of a cosmological constant on the predicted Sunyaev-Zel'dovich effect from the ensemble of clusters. In this report we summarize that work from which a paper will be prepared.

Forman, William R.↗

Running Ensemble Workflows at Extreme Scale: Lessons Learned and Path Forward

The ever-increasing volumes of scientific data combined with sophisticated techniques for extracting information from them have led to the increasing popularity of ensemble workflows which are a collection of runs of individual workflows. A traditional approach followed by scientists to run ensembles is to rely on simple scripts to execute different runs and manage resources. This approach is not scalable and is error-prone, thereby motivating the development of workflow management systems that specialize in executing ensembles on HPC clusters. However, when the size of both the ensemble and the target system reach extreme scales, existing workflow management systems face new challenges that hamper their efficient execution. In this paper, we describe our experience scaling an ensemble workflow from the computational biology domain from the early design stages to the execution at extreme scale on Summit, a leadership class supercomputer at the Oak Ridge National Laboratory. We discuss challenges that arise when scaling ensembles to several million runs on thousands of HPC nodes. We identify challenges with composition of the ensemble itself, its execution at large scale, post-processing of the generated data, and scalability of the file system. Based on the experience acquired, we develop a generic vision of the capabilities and abstractions to add to existing workflow management systems to enable the execution of ensemble workflows at extreme scales. We believe that the understanding of these fundamental challenges will help application teams along with workflow system developers with designing the next generation of infrastructure for composing and executing extreme-scale ensemble workflows.

Mehta, Kshitij↗

An X-ray method for detecting substructure in galaxy clusters - Application to Perseus, A2256, Centaurus, Coma, and Sersic 40/6

We use the moments of the X-ray surface brightness distribution to constrain the dynamical state of a galaxy cluster. Using X-ray observations from the Einstein Observatory IPC, we measure the first moment FM, the ellipsoidal orientation angle, and the axial ratio at a sequence of radii in the cluster. We argue that a significant variation in the image centroid FM as a function of radius is evidence for a nonequilibrium feature in the intracluster medium (ICM) density distribution. In simple terms, centroid shifts indicate that the center of mass of the ICM varies with radius. This variation is a tracer of continuing dynamical evolution. For each cluster, we evaluate the significance of variations in the centroid of the IPC image by computing the same statistics on an ensemble of simulated cluster images. In producing these simulated images we include X-ray point source emission, telescope vignetting, Poisson noise, and characteristics of the IPC. Application of this new method to five Abell clusters reveals that the core of each one has significant substructure. In addition, we find significant variations in the orientation angle and the axial ratio for several of the clusters.

Mohr, Joseph J.↗

Using x ray images to detect substructure in a sample of 40 Abell clusters

Using a method for constraining the dynamical state of a galaxy cluster by examining the moments of its x-ray surface brightness distribution, we determine the statistics of cluster substructure for a sample of 40 Abell clusters. Using x-ray observations from the Einstein Observatory Imaging Proportional Counter (IPC), we measure the first moment M1(r), the ellipsoidal orientation angle theta2(r), and the axial ratio eta(r) at several different radii in the cluster. We determine the effects of systematics such as x-ray point source emission, telescope vignetting, Poisson noise, and characteristics of the IPC by measuring the same parameters on an ensemble of simulated cluster images. Due to the small band-pass of the IPC, the ICM emissivity is nearly independent of temperature so the intensity at each point in the IPC images is simply proportional to the emission measure calculated along the line of sight through the cluster (e.g. Fabricant et al. 1980). Therefore, barring a change superposition of two x-ray emitting clusters, a significant variation in the image centroid M1(r) as a function of radius indicates that the center of mass of the intra-cluster medium (ICM) varies with radius. We argue that such a configuration (essentially an m = 1 component in the ICM density distribution) is a non-equilibrium component; it results from an off-center subclump or a recent merger in the ICM.

Mohr, J. J.↗

Observing flow of He II with unsupervised machine learning

Abstract Time dependent observations of point-to-point correlations of the velocity vector field (structure functions) are necessary to model and understand fluid flow around complex objects. Using thermal gradients, we observed fluid flow by recording fluorescence of $${\text{He}}_{2}^{*}$$ He 2 ∗ excimers produced by neutron capture throughout a ~ cm 3 volume. Because the photon emitted by an excited excimer is unlikely to be recorded by the camera, the techniques of particle tracking (PTV) and particle imaging (PIV) velocimetry cannot be applied to extract information from the fluorescence of individual excimers. Therefore, we applied an unsupervised machine learning algorithm to identify light from ensembles of excimers (clusters) and then tracked the centroids of the clusters using a particle displacement determination algorithm developed for PTV.

75 CONDENSED MATTER PHYSICS, SUPERCONDUCTIVITY AND↗