Search NASA⌕ Search

SEARCH · Search NASA

Results for “ensemble clustering”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

Hydrogen Evolution on Electrode‐Supported Pt n Clusters: Ensemble of Hydride States Governs the Size Dependent Reactivity

Abstract We report the size‐dependent activity and stability of supported Pt 1,4,7,8 for electrocatalytic hydrogen evolution reaction, and show that clusters outperform polycrystalline Pt in activity, with size‐dependent stability. To understand the size effects, we use DFT calculations to study the structural fluxionality under varying potentials. We show that the clusters can reshape under H coverage and populate an ensemble of states with diverse stoichiometry, structure, and thus reactivity. Both experiment and theory suggest that electrocatalytic species are hydridic states of the clusters (≈2 H/Pt). An ensemble‐based kinetic model reproduces the experimental activity trend and reveals the role of metastable states. The stability trend is rationalized by chemical bonding analysis. Our joint study demonstrates the potential‐ and adsorbate‐coverage‐dependent fluxionality of subnano clusters of different sizes and offers a systematic modeling strategy to tackle the complexities.

Zhang, Zisheng↗

Hydrogen Evolution on Electrode–Supported Ptn Clusters: Ensemble of Hydride States Governs the Size Dependent Reactivity

We report the size-dependent activity and stability of supported Pt1,4,7,8 for electrocatalytic hydrogen evolution reaction, and show that clusters outperform polycrystalline Pt in activity, with sizedependent stability. To understand the size effects, we use DFT calculations to study the structural fluxionality under varying potentials. We show that the clusters can reshape under H coverage and populate an ensemble of states with diverse stoichiometry, structure, and thus reactivity. Both experiment and theory suggest that electrocatalytic species are hydridic states of the clusters (~2 H/Pt). An ensemble-based kinetic model reproduces the experimental activity trend and reveals the role of metastable states. Furthermore, the stability trend is rationalized by chemical bonding analysis. Our joint study demonstrates the potential- and adsorbate-coverage-dependent fluxionality of subnano clusters of different sizes and offers a systematic modeling strategy to tackle the complexities.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Cryptate binding energies towards high throughput chelator design: metadynamics ensembles with cluster–continuum solvation

A tiered forcefield/semiempirical/meta-GGA pipeline together with a thermodynamic scheme designed with error cancellation in mind was developed to calculate binding energies of [2.2.2] cryptate complexes of mono- and divalent cations. Stable complexes of Na, K, Rb, Ca, Zn and Pb were generated, revealing consistent cation–N lengths but highly variable cation–O lengths and an amine stacking mechanism potentially augmenting the cation size selectivity. Metadynamics, used for searching the high-dimensional potential energy surface, together with a cluster–continuum model for affordable – yet accurate – solvation modeling, enabled the discovery of more stable geometries than those previously reported. Similar solvation energy curve shapes for lone vs. coordinated ions enabled rapid solvation convergence via the cancellation of errors stemming from finite cluster sizes. In conclusion, an R 2 of 0.850 vs. experimental aqueous binding energies was obtained, validating this scheme as the backbone of a high-throughput workflow for chelator design.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Unravelling the orbits of cluster galaxy populations according to their dominant gas ionization source

ABSTRACT We investigate the kinematical and dynamical properties of cluster galaxy populations classified according to their dominant source of gas ionization, namely: star-forming (SF) galaxies, optical active galactic nuclei (AGNs), mixed SF plus AGN ionization (transition objects, T), and quiescent (Q) galaxies. We stack 8892 member galaxies from 336 relaxed galaxy clusters to build an ensemble cluster and estimate the observed projected profiles of numerical density and velocity dispersion, $\sigma _P(R)$, of each galaxy population. The MAMPOSSt code and the Jeans equations inversion technique are used to constrain the velocity anisotropy profiles of the galaxy populations in both parametric and non-parametric ways. We find that Q (SF) galaxies display the lowest (highest) typical cluster-centric distances and velocity dispersion values. Transition galaxies are more concentrated and tend to exhibit lower velocity dispersion values than SF galaxies. Galaxies that host an optical AGN are as concentrated as Q galaxies but display velocity dispersion values similar to those of the SF population. MAMPOSSt is able to find equilibrium solutions that successfully recover the observed $\sigma _P(R)$ profile only for the Q, T, and AGN populations. We find that the orbits of all populations are consistent with isotropy in the inner regions, becoming increasingly radial with the distance from the cluster centre. These results suggest that Q galaxies are in equilibrium within their clusters, while SF galaxies have more recently arrived in the cluster environment. Finally, the T and AGN populations appear to be in an intermediate dynamical state between those of the SF and Q populations.

Valk, Greique A. (ORCID:0009000827731299)↗

A Novel Data Segmentation Method for Data-driven Phase Identification

This paper presents a smart meter phase identification algorithm for two cases: meter-phase-label-known and meter-phase-label-unknown. To improve the identification accuracy, a data segmentation method is proposed to exclude data segments that are collected when the voltage correlation between smart meters on the same phase is weakened. Then, using the selected data segments, a hierarchical clustering method is used to calculate the correlation distances and cluster the smart meters. If the phase labels are unknown, a Connected-Triple-based Similarity (CTS) method is adapted to further improve the phase identification accuracy of the ensemble clustering method. The methods are developed and tested on both synthetic and real feeder data sets. Here, simulation results show that the proposed phase identification algorithm outperforms the state-of-the-art methods in both accuracy and robustness.

24 POWER TRANSMISSION AND DISTRIBUTION↗

Distilling Knowledge from Ensembles of Cluster-Constrained-Attention Multiple-Instance Learners for Whole Slide Image Classification

The peculiar nature of whole slide imaging (WSI), digitizing conventional glass slides to obtain multiple high resolution images which capture microscopic details of a patient’s histopathological features, has garnered increased interest from the computer vision research community over the last two decades. Given the unique computational space and time complexity inherent to gigapixel-size whole slide image data, researchers have proposed novel machine learning algorithms to aid in the performance of diagnostic tasks in clinical pathology. One effective algorithm represents a Whole slide image as a bag of smaller image patches, which can be represented as low-dimension image patch embeddings. Weakly supervised deep-learning methods, such as cluster-constrained-attention multiple instance learning (CLAM), have shown promising results when combined with image patch embeddings. While traditional ensemble classifiers yield improved task performance, such methods come with a steep cost in model complexity. Through knowledge distillation, it is possible to retain some performance improvements from an ensemble, while minimizing costs to model complexity. In this work, we implement a weakly supervised ensemble using clustering-constrained-attention multiple-instance learners (CLAM), which uses attention and instance-level clustering to identify task salient regions and feature extraction in whole slides. By applying logit-based and attention-based knowledge distillation, we show it is possible to retain some performance improvements resulting from the ensemble at zero cost to model complexity.

Alamudun, Folami↗

Robust design of semi-automated clustering models for 4D-STEM datasets

Materials discovery and design require characterizing material structures at the nanometer and sub-nanometer scale. Four-Dimensional Scanning Transmission Electron Microscopy (4D-STEM) resolves the crystal structure of materials, but many 4D-STEM data analysis pipelines are not suited for the identification of anomalous and unexpected structures. This work introduces improvements to the iterative Non-Negative Matrix Factorization (NMF) method by implementing consensus clustering for ensemble learning. We evaluate the performance of models during parameter tuning and find that consensus clustering improves performance in all cases and is able to recover specific grains missed by the best performing model in the ensemble. The methods introduced in this work can be applied broadly to materials characterization datasets to aid in the design of new materials.

Bruefach, Alexandra (ORCID:0000000209323477)↗

Analytical and EZmock covariance validation for the DESI 2024 results

The estimation of uncertainties in cosmological parameters is an important challenge in Large-Scale-Structure (LSS) analyses. For standard analyses such as Baryon Acoustic Oscillations (BAO) and Full-Shape two approaches are usually considered. First: analytical estimates of the covariance matrix use Gaussian approximations and (nonlinear) clustering measurements to estimate the matrix, which allows a relatively fast and computationally cheap way to generate matrices that adapt to an arbitrary clustering measurement. On the other hand, sample covariances are an empirical estimate of the matrix based on an ensemble of clustering measurements from fast and approximate simulations. While more computationally expensive due to the large amount of simulations and volume required, these allow us to take into account systematics that are impossible to model analytically. In this work we compare these two approaches in order to enable DESI's key analyses. We find that the configuration space analytical estimate performs satisfactorily in BAO analyses and its flexibility in terms of input clustering makes it the fiducial choice for DESI's 2024 BAO analysis. On the contrary, the analytical computation of the covariance matrix in Fourier space does not reproduce the expected measurements in terms of Full-Shape analyses, which motivates the use of a corrected mock covariance for DESI's 2024 Full Shape analysis.

79 ASTRONOMY AND ASTROPHYSICS↗

Running Ensemble Workflows at Extreme Scale: Lessons Learned and Path Forward

The ever-increasing volumes of scientific data combined with sophisticated techniques for extracting information from them have led to the increasing popularity of ensemble workflows which are a collection of runs of individual workflows. A traditional approach followed by scientists to run ensembles is to rely on simple scripts to execute different runs and manage resources. This approach is not scalable and is error-prone, thereby motivating the development of workflow management systems that specialize in executing ensembles on HPC clusters. However, when the size of both the ensemble and the target system reach extreme scales, existing workflow management systems face new challenges that hamper their efficient execution. In this paper, we describe our experience scaling an ensemble workflow from the computational biology domain from the early design stages to the execution at extreme scale on Summit, a leadership class supercomputer at the Oak Ridge National Laboratory. We discuss challenges that arise when scaling ensembles to several million runs on thousands of HPC nodes. We identify challenges with composition of the ensemble itself, its execution at large scale, post-processing of the generated data, and scalability of the file system. Based on the experience acquired, we develop a generic vision of the capabilities and abstractions to add to existing workflow management systems to enable the execution of ensemble workflows at extreme scales. We believe that the understanding of these fundamental challenges will help application teams along with workflow system developers with designing the next generation of infrastructure for composing and executing extreme-scale ensemble workflows.

Mehta, Kshitij↗

Observing flow of He II with unsupervised machine learning

Abstract Time dependent observations of point-to-point correlations of the velocity vector field (structure functions) are necessary to model and understand fluid flow around complex objects. Using thermal gradients, we observed fluid flow by recording fluorescence of $${\text{He}}_{2}^{*}$$ He 2 ∗ excimers produced by neutron capture throughout a ~ cm 3 volume. Because the photon emitted by an excited excimer is unlikely to be recorded by the camera, the techniques of particle tracking (PTV) and particle imaging (PIV) velocimetry cannot be applied to extract information from the fluorescence of individual excimers. Therefore, we applied an unsupervised machine learning algorithm to identify light from ensembles of excimers (clusters) and then tracked the centroids of the clusters using a particle displacement determination algorithm developed for PTV.

75 CONDENSED MATTER PHYSICS, SUPERCONDUCTIVITY AND↗

Ensemble Effects on Hydroxide Bond Dissociation Free Energies in Polyoxovanadate Clusters

Understanding structure-property relationships is foundational to numerous modern chemistries, such as proton-coupled electron transfer (PCET). However, an experimentally measured property is the result of the behavior from an ensemble of molecules. Neglecting ensemble effects, especially under complex chemical environments, may obfuscate these relationships and lead to discrepancies between theory and experiment. In this work, we demonstrate the impact of configurational entropy and local chemical environments on hydroxide bond dissociation free energies [BDFE- (O−H)] for a set of polyoxovanadate nanoclusters, at ambient conditions. The O−H bond strengths are investigated via density functional theory (DFT) coupled with statistical thermodynamic analysis and bilinear modeling, and compared with previous experimental results on the same systems, namely electrochemical solutions of: [V 6 O 13−x (OH) x (TRIOL R ) 2 ] −2 (x = 2, 4, 6; R = NO 2 , Me) and [V 6 O 11−x (OMe) 2 (OH) x (TRIOL NO 2 ) 2 ] −2 (x = 2, 4). Interestingly, we find that ensemble effects, even at room temperature, can account for a significant portion of the BDFE(O−H) trend with the degree of reduction via H atom binding, which cannot be fully captured by single-structure, static DFT calculations. Moreover, we find that the ensemble effects may be replicated statistically, requiring only enumeration of energetically accessible H-binding sites. With the ensemble effects resolved, we present a simple bilinear model to reconcile remaining biases between experiment and ensemble-informed theory, which corelate with clusterspecific electronic environment differences. The bilinear model achieves outstanding accuracy vs experiments with a root-mean squared error of 0.4 kcal/mol. Finally, based on the physicochemical characteristics of hydrogen interaction with polyoxometalates, we present a simple methodology that captures the BDFE(O−H) trend while dramatically reducing required DFT calculations by 98% and achieving accuracy within 1 kcal/mol. Overall, this work elucidates the roles and structural origins of configurational entropy and chemical effects on polyoxometalate hydroxide bond energies, with potential applicability to various atomically precise metal oxide systems. Importantly, it introduces models for rapid and highly accurate property calculations in connection with experiments.

Cluster chemistry↗

Catalytic Activity of an Ensemble of Sites for CO 2 Hydrogenation to Methanol on a ZrO 2 -on-Cu Inverse Catalyst

The significant increase in CO 2 emissions from heavy fossil fuel utilization has raised serious concerns, highlighting the need for effective methods to convert CO 2 into value-added chemicals. Here, in this work, we report a computational investigation on the catalytic activity of ZrO 2 -on-Cu inverse catalysts for CO 2 hydrogenation to methanol, considering highly dispersed ZrO 2 trimers on Cu (111). Such clusters present a large ensemble of formate-containing configurations, Zr 3 O n (OH) m (OCHO) l , making the evaluation of the catalytic activity very challenging. We found that the sites on the various catalyst configurations exhibit markedly different activities for formate hydrogenation, despite their similar free energy and composition. To understand these differences in reactivity, we examined the structural and electronic nature of the low free-energy catalyst configurations and identified that the energy of the lowest unoccupied orbital of the reacting formate, modified by its binding with the catalytic site, is a descriptor for the reaction energy of the formate hydrogenation step. From there, we screened an ensemble of catalyst structures using this descriptor to predict highly active metastable catalyst configurations and computed the reaction pathways and transition states for formate hydrogenation. From this investigation, we distinguished reactive from nonreactive sites and formate species on the ZrO 2 /Cu inverse catalyst based on structural and electronic features. We showed that rare metastable configurations control the activity. Additionally, an efficient method for examining the reactivity of a large number of coexisting catalyst structures was developed.

catalysts↗

Deducing subnanometer cluster size and shape distributions of heterogeneous supported catalysts

Abstract Infrared (IR) spectra of adsorbate vibrational modes are sensitive to adsorbate/metal interactions, accurate, and easily obtainable in-situ or operando. While they are the gold standards for characterizing single-crystals and large nanoparticles, analogous spectra for highly dispersed heterogeneous catalysts consisting of single-atoms and ultra-small clusters are lacking. Here, we combine data-based approaches with physics-driven surrogate models to generate synthetic IR spectra from first-principles. We bypass the vast combinatorial space of clusters by determining viable, low-energy structures using machine-learned Hamiltonians, genetic algorithm optimization, and grand canonical Monte Carlo calculations. We obtain first-principles vibrations on this tractable ensemble and generate single-cluster primary spectra analogous to pure component gas-phase IR spectra. With such spectra as standards, we predict cluster size distributions from computational and experimental data, demonstrated in the case of CO adsorption on Pd/CeO 2 (111) catalysts, and quantify uncertainty using Bayesian Inference. We discuss extensions for characterizing complex materials towards closing the materials gap.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Structural characterization of an intrinsically disordered protein complex using integrated small-angle neutron scattering and computing

Characterizing structural ensembles of intrinsically disordered proteins (IDPs) and intrinsically disordered regions (IDRs) of proteins is essential for studying structure–function relationships. Due to the different neutron scattering lengths of hydrogen and deuterium, selective labeling and contrast matching in small-angle neutron scattering (SANS) becomes an effective tool to study dynamic structures of disordered systems. However, experimental timescales typically capture measurements averaged over multiple conformations, leaving complex SANS data for disentanglement. We hereby demonstrate an integrated method to elucidate the structural ensemble of a complex formed by two IDRs. We use data from both full contrast and contrast matching with residue-specific deuterium labeling SANS experiments, microsecond all-atom molecular dynamics (MD) simulations with four molecular mechanics force fields, and an autoencoder-based deep learning (DL) algorithm. From our combined approach, we show that selective deuteration provides additional information that helps characterize structural ensembles. We find that among the four force fields, a99SB-disp and CHARMM36m show the strongest agreement with SANS and NMR experiments. In addition, our DL algorithm not only complements conventional structural analysis methods but also successfully differentiates NMR and MD structures which are indistinguishable on the free energy surface. Finally, we present an ensemble that describes experimental SANS and NMR data better than MD ensembles generated by one single force field and reveal three clusters of distinct conformations. Our results demonstrate a new integrated approach for characterizing structural ensembles of IDPs.

59 BASIC BIOLOGICAL SCIENCES↗

Dark Energy Survey Year 6 results: Clustering redshifts and importance sampling of self-organized-maps 𝑛⁡(𝑧) realizations for 3 × 2 ⁢pt samples

This work is part of a series establishing the redshift framework for the 3 × 2 ⁢pt analysis of the Dark Energy Survey Year 6 (DES Y6). For DES Y6, photometric redshift distributions are estimated using self-organizing maps (SOMs), calibrated with spectroscopic and many-band photometric data. To overcome limitations from color-redshift degeneracies and incomplete spectroscopic coverage, we enhance this approach by incorporating clustering-based redshift constraints (clustering-z, or WZ) from angular cross-correlations with BOSS and eBOSS galaxies and eBOSS quasar samples. We define a WZ likelihood and apply importance sampling to a large ensemble of SOM-derived 𝑛⁡(𝑧) realizations, selecting those consistent with the clustering measurements to produce a posterior sample for each lens and source bin. The analysis uses angular scales corresponding to 1.5–5 Mpc to optimize signal-to-noise ratio while mitigating modeling uncertainties and marginalizes over redshift-dependent galaxy bias and other systematics informed by the N-body simulation CARDINAL . While a sparser spectroscopic reference sample limits WZ constraining power at 𝑧 >1.1, particularly for source bins, we demonstrate that combining SOM with WZ improves redshift accuracy and enhances the overall cosmological constraining power of DES Y6. As a result, we estimate an improvement in 𝑆 8 of approximately 10% for cosmic shear and 3 ×2⁢pt analysis, primarily due to the WZ calibration of the source samples.

Cosmological parameters↗