Search NASA⌕ Search

SEARCH · Search NASA

Results for “distributed clustering methods”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

Designing a parallel Feel-the-Way clustering algorithm on HPC systems

This paper introduces a new parallel clustering algorithm, named Feel-the-Way clustering algorithm, that provides better or equivalent convergence rate than the traditional clustering methods by optimizing the synchronization and communication costs. Our algorithm design centers on how to optimize three factors simultaneously: reduced synchronizations, improved convergence rate, and retained same or comparable optimization cost. To compare the optimization cost, we use the Sum of Square Error (SSE) cost as the metric, which is the sum of the square distance between each data point and its assigned clusters. Compared with the traditional MPI k-means algorithm, the new Feel-the-Way algorithm requires less communications among participating processes. As for the convergence rate, the new algorithm requires fewer number of iterations to converge. As for the optimization cost, it obtains the SSE costs that are close to the k-means algorithm. In the paper, we first design the full-step Feel-the-Way k-means clustering algorithm that can significantly reduce the number of iterations that are required by the original k-means clustering method. Next, we improve the performance of the full-step algorithm by adopting an optimized sampling-based approach, named reassignment-history-aware sampling. Our experimental results show that the optimized sampling-based Feel-the-Way method is significantly faster than the widely used k-means clustering method, and can provide comparable optimization costs. More extensive experiments with several synthetic datasets and real-world datasets (e.g., MNIST, CIFAR-10, ENRON, and PLACES-2) show that the new parallel algorithm can outperform the open source MPI k-means library by up to 110% on a high-performance computing system using 4,096 CPU cores. In addition, the new algorithm can take up to 51% fewer iterations to converge than the k-means clustering algorithm.

97 MATHEMATICS AND COMPUTING↗

Efficient Agent-Based Cluster Ensembles

Numerous domains ranging from distributed data acquisition to knowledge reuse need to solve the cluster ensemble problem of combining multiple clusterings into a single unified clustering. Unfortunately current non-agent-based cluster combining methods do not work in a distributed environment, are not robust to corrupted clusterings and require centralized access to all original clusterings. Overcoming these issues will allow cluster ensembles to be used in fundamentally distributed and failure-prone domains such as data acquisition from satellite constellations, in addition to domains demanding confidentiality such as combining clusterings of user profiles. This paper proposes an efficient, distributed, agent-based clustering ensemble method that addresses these issues. In this approach each agent is assigned a small subset of the data and votes on which final cluster its data points should belong to. The final clustering is then evaluated by a global utility, computed in a distributed way. This clustering is also evaluated using an agent-specific utility that is shown to be easier for the agents to maximize. Results show that agents using the agent-specific utility can achieve better performance than traditional non-agent based methods and are effective even when up to 50% of the agents fail.

Agogino, Adrian↗

Source Analysis of Ozone Pollution in Liaoyuan City’s Atmosphere Based on Machine Learning Models and HYSPLIT Clustering Method

Firstly, this study investigates the spatiotemporal distribution characteristics of the ozone (O 3 ) pollution in Liaoyuan City using monitoring data from 2015 to 2024. Then, three machine learning models (ML)—random forest (RF), support vector machine (SVM), and artificial neural network (ANN)—are employed to quantify the influence of meteorological and non-meteorological factors on O 3 concentrations. Finally, the HYSPLIT clustering method and CMAQ model are utilized to analyze inter-regional transport characteristics, identifying the causes of O 3 pollution. The results indicate that O 3 pollution in Liaoyuan exhibits a distinct seasonal pattern, with the highest concentrations found in spring and summer, peaking in the afternoon. Among the three ML models, the random forest model demonstrates the best predictive performance (R 2 = 0.9043). Feature importance identifies NO 2 as the primary driving factor, followed by meteorological conditions in the second quarter and land surface characteristics. Furthermore, regional transport significantly contributes to O 3 pollution, with approximately 80% of air mass trajectories in heavily polluted episodes originating from adjacent industrial areas and the sea. The combined effects of transboundary precursors and O 3 transport with local emissions and meteorological conditions further increase the O 3 pollution level. This study highlights the need to strengthen coordinated NO X and VOCs emission reductions and enhance regional joint prevention and control strategies in China.

HYSPLIT clustering↗

Enhancing Active Distribution Systems Resilience by Fully Distributed Self-Healing Strategy

Distributed restoration can exploit smart grid technologies to enhance the resilience of active distribution networks toward a self-healing smart grid. However, the large number of decision variables, especially the binary ones for reconfiguration, bring challenges to developing scalable distributed distribution service restoration (DDSR) strategies. This paper proposes a fully distributed solution procedure based on the alternating direction method of multipliers (ADMM) for mixed-integer programming problems and applies to develop the DDSR framework. The method consists of relax-drive-polish phases, 1) relaxing binary variables, and applying the convex ADMM as a warm start; 2) driving the solutions toward Boolean values through a proximal operator; 3) fixing the obtained binding binary variables and solving the rest of the problem to polish results and achieve a high-quality suboptimal solution. Then, an autonomous clustering strategy and consensus ADMM are integrated with the proposed method to realize the fully distributed cluster-based framework of DDSR. This framework can first determine DER scheduling and switch status for reconfiguration to energize the out-of-service areas from local faults, and then provide the load restoration solution in a distributed manner for total blackouts in large-scale distribution networks. Furthermore, the effectiveness and scalability of the proposed DDSR framework are demonstrated through testing on the IEEE 123-node, IEEE 8500-node, and synthetic 100k-node test feeders.

24 POWER TRANSMISSION AND DISTRIBUTION↗

X-ray halos in galaxies and clusters of galaxies - Theory

X-ray measurements provide an excellent method to determine the amount and distribution of the dark matter in clusters. Unfortunately, accurate temperature profiles, necessary to this method, are currently not available. However, if the intracluster gas is assumed to have a monotonically decreasing temperature, one finds that the dark matter is strongly concentrated to the cluster center, and has a mass which only exceeds the known baryonic mass by a factor of about three. On a second topic, cooling flows are shown to be a very common feature of cluster central and normal elliptical galaxies. The cooling gas is probably ultimately converted into low mass stars.

Sarazin, Craig L.↗

Dark Energy Survey Year 3 results: calibration of lens sample redshift distributions using clustering redshifts with BOSS/eBOSS

ABSTRACT We present clustering redshift measurements for Dark Energy Survey (DES) lens sample galaxies used in weak gravitational lensing and galaxy clustering studies. To perform these measurements, we cross-correlate with spectroscopic galaxies from the Baryon Acoustic Oscillation Survey (BOSS) and its extension, eBOSS. We validate our methodology in simulations, including a new technique to calibrate systematic errors that result from the galaxy clustering bias, and we find that our method is generally unbiased in calibrating the mean redshift. We apply our method to the data, and estimate the redshift distribution for 11 different photometrically selected bins. We find general agreement between clustering redshift and photometric redshift estimates, with differences on the inferred mean redshift found to be below |Δz| = 0.01 in most of the bins. We also test a method to calibrate a width parameter for redshift distributions, which we found necessary to use for some of our samples. Our typical uncertainties on the mean redshift ranged from 0.003 to 0.008, while our uncertainties on the width ranged from 4 to 9 per cent. We discuss how these results calibrate the photometric redshift distributions used in companion papers for DES Year 3 results.

79 ASTRONOMY AND ASTROPHYSICS↗

Watershed zonation through hillslope clustering for tractably quantifying above- and below-ground watershed heterogeneity and functions

Abstract. In this study, we develop a watershed zonation approach for characterizing watershed organization and functions in a tractable manner by integrating multiple spatial data layers. We hypothesize that (1) a hillslope is an appropriate unit for capturing the watershed-scale heterogeneity of key bedrock-through-canopy properties and for quantifying the co-variability of these properties representing coupled ecohydrological and biogeochemical interactions, (2) remote sensing data layers and clustering methods can be used to identify watershed hillslope zones having the unique distributions of these properties relative to neighboring parcels, and (3) property suites associated with the identified zones can be used to understand zone-based functions, such as response to early snowmelt or drought and solute exports to the river. We demonstrate this concept using unsupervised clustering methods that synthesize airborne remote sensing data (lidar, hyperspectral, and electromagnetic surveys) along with satellite and streamflow data collected in the East River Watershed, Crested Butte, Colorado, USA. Results show that (1) we can define the scale of hillslopes at which the hillslope-averaged metrics can capture the majority of the overall variability in key properties (such as elevation, net potential annual radiation, and peak snow-water equivalent – SWE), (2) elevation and aspect are independent controls on plant and snow signatures, (3) near-surface bedrock electrical resistivity (top 20 m) and geological structures are significantly correlated with surface topography and plan species distribution, and (4) K-means, hierarchical clustering, and Gaussian mixture clustering methods generate similar zonation patterns across the watershed. Using independently collected data, we show that the identified zones provide information about zone-based watershed functions, including foresummer drought sensitivity and river nitrogen exports. The approach is expected to be applicable to other sites and generally useful for guiding the selection of hillslope-experiment locations and informing model parameterization.

58 GEOSCIENCES↗

A new step forward in realistic cluster lens mass modelling: analysis of Hubble Frontier Field Cluster Abell S1063 from joint lensing, X-ray, and galaxy kinematics data

We present a new method to simultaneously and self-consistently model the mass distribution of galaxy clusters that combines constraints from strong lensing features, X-ray emission, and galaxy kinematics measurements. We are able to successfully decompose clusters into their collisionless and collisional mass components thanks to the X-ray surface brightness, as well as use the dynamics of cluster members, to obtain more accurate masses exploiting the fundamental plane of elliptical galaxies. Knowledge from all observables is included through a consistent Bayesian approach in the likelihood or in physically motivated priors. We apply this method to the galaxy cluster Abell S1063 and produce a mass model that we publicly release with this paper. The resulting mass distribution presents different ellipticities for the intra-cluster gas and the other large-scale mass components as well as deviation from elliptical symmetry in the main halo. We assess the ability of our method to recover the masses of the different elements of the cluster using a mock cluster based on a simplified version of our Abell S1063 model. Thanks to the wealth of mutliwavelength information provided by the mass model and the detected X-ray emission, we also found evidence for an ongoing merger event with gas sloshing from a smaller infalling structure into the main cluster. In agreement with previous findings, the total mass, gas profile, and gas mass fraction are all consistent with small deviations from the hydrostatic equilibrium. This new mass model for Abell S1063 is publicly available, as the lenstool extension used to construct it.

79 ASTRONOMY AND ASTROPHYSICS↗

Intelligent Energy Optimizer for Residential Buildings

Demand-side management in the buildings is essential for meeting grid flexibility needs in a highly renewable energy scenario. Appliance load monitoring helps decision making for demand-side management by providing the information on operation status/power consumption from different appliances in the buildings. Nonintrusive load monitoring (NILM) is an attractive option for appliance load monitoring using because it has lower cost for sensors and helps mitigate privacy concerns. In this study, the team used an event detection technique followed by two different methods for event classification. The results from k-means clustering showed that the events from a single appliance are often distributed in multiple clusters. Thus, the unsupervised method of NILM using k-means clustering used in this study was not very suitable for load disaggregation. The results from NILM showed that the F1 score for event classification was 0.77 for a heat pump water heater and very low for other appliances using the rule-based classification.

24 POWER TRANSMISSION AND DISTRIBUTION↗

Optimizing the shape of photometric redshift distributions with clustering cross-correlations

We present an optimization method for the assignment of photometric galaxies to a chosen set of redshift bins. This is achieved by combining simulated annealing, an optimization algorithm inspired by solid-state physics, with an unsupervised machine learning method, a self-organizing map (SOM) of the observed colours of galaxies. Starting with a sample of galaxies that is divided into redshift bins based on a photometric redshift point estimate, the simulated annealing algorithm repeatedly reassigns SOM-selected subsamples of galaxies, which are close in colour, to alternative redshift bins. We optimize the clustering cross-correlation signal between photometric galaxies and a reference sample of galaxies with well-calibrated redshifts. Depending on the effect on the clustering signal, the reassignment is either accepted or rejected. By dynamically increasing the resolution of the SOM, the algorithm eventually converges to a solution that minimizes the number of mismatched galaxies in each tomographic redshift bin and thus improves the compactness of their corresponding redshift distribution. This method is demonstrated on the synthetic Legacy Survey of Space and Time cosmoDC2 catalogue. We find a significant decrease in the fraction of catastrophic outliers in the redshift distribution in all tomographic bins, most notably in the highest redshift bin with a decrease in the outlier fraction from 57 percent to 16 percent.

79 ASTRONOMY AND ASTROPHYSICS↗

Weak Gravitational Lensing around Low Surface Brightness Galaxies in the DES Year 3 Data

We present galaxy-galaxy lensing measurements using a sample of low surface brightness galaxies (LSBGs) drawn from the Dark Energy Survey Year 3 (Y3) data as lenses. LSBGs are diffuse galaxies with a surface brightness dimmer than the ambient night sky. These dark-matter-dominated objects are intriguing due to potentially unusual formation channels that lead to their diffuse stellar component. Given the faintness of LSBGs, using standard observational techniques to characterize their total masses proves challenging. Weak gravitational lensing, which is less sensitive to the stellar component of galaxies, could be a promising avenue to estimate the masses of LSBGs. Our LSBG sample consists of 23,790 galaxies separated into red and blue color types at g - i ≥ 0.60 and g - i < 0.60 , respectively. Combined with the DES Y3 shear catalog, we measure the tangential shear around these LSBGs and find signal-to-noise ratios of 6.67 for the red sample, 2.17 for the blue sample, and 5.30 for the full sample. We use the clustering redshifts method to obtain redshift distributions for the red and blue LSBG samples. Assuming all red LSBGs are satellites, we fit a simple model to the measurements and estimate the host halo mass of these LSBGs to be log(M host /M ⊙ ) = $12.98^{+0.10}_{-0.11}$. We place a 95% upper bound on the subhalo mass at log(M sub /M ⊙ ) < 11.51. By contrast, we assume the blue LSBGs are centrals, and place a 95% upper bound on the halo mass at log(M host /M ⊙ ) < 11.84. We find that the stellar-to-halo mass ratio of the LSBG samples is consistent with that of the general galaxy population. This work illustrates the viability of using weak gravitational lensing to constrain the halo masses of LSBGs.

79 ASTRONOMY AND ASTROPHYSICS↗

A Zonal Volt/VAR Control Mechanism for High PV Penetration Distribution Systems

This paper presents a zonal Volt/VAR control scheme for regulating voltage in unbalanced 3-phase distribution systems using inverter-based resources (IBR). First, the dependency between nodal voltage changes and IBR reactive power injections is derived via voltage sensitivity studies. Then, a fast-incremental clustering method is used to divide the distribution circuit into weakly-coupled zones based on correlations between nodal voltage sensitivities. The weak-coupling allows the voltage to be regulated independently within each zone using a rule-based voltage controller to dispatch IBR for voltage corrections. Simulation results on actual distribution feeder show that the proposed zone-based Volt/VAR control method maintains system voltages within their operational limits while reducing the runtime of the Volt/VAR controller from tens of second to a couple of milliseconds compared with centralized, optimization-based Volt/VAR control methods.

Alrushoud, Asmaa↗

Non-local corrections to the typical medium theory of Anderson localization

We use the recently developed finite cluster typical medium approach to study the Anderson localization transition in three dimensions. Applying our method to the box and binary alloy disorder distributions, we find a fast convergence with the cluster size. We demonstrate the importance of the typical medium environment and the non-local spatial correlations for the proper characterization of the localization transition. As the cluster size increases, our typical medium cluster method recovers the correct critical disorder strength for the transition. Our findings highlight the importance of the non-local cluster corrections for capturing the localization behavior of the mobility edge trajectories. Our results demonstrate that the typical medium cluster approach developed here provides a consistent and systematic description of the Anderson localization transition in the framework of the effective medium embedding schemes.

71 CLASSICAL AND QUANTUM MECHANICS, GENERAL PHYSIC↗

Mesh-based multiphysics coupling acceleration for fusion neutronics through clustering for fusion blanket applications

Accurate modeling of particle transport within fusion blankets is essential for predicting performance metrics such as heat deposition and the tritium breeding ratio (TBR). However, high-fidelity coupling of thermal fluids from computational fluid dynamics (CFD) to neutronics simulations often incurs significant computational costs due to the complexity of surface intersection calculations in Monte Carlo codes. This paper presents an accelerated multiphysics coupling method for neutronics that utilizes hierarchical agglomerative clustering to map complex material property distributions to a neutronics model. Implemented within the fusion reactor design and assessment (FREDA) framework, the method leverages existing Python packages to automate the creation of clustered geometries for OpenMC. The approach is demonstrated on a sector model of an ARC-class tokamak with an immersion molten salt blanket, and an simple geometry with varying isotopic concentrations. Results show that the clustering method significantly reduces computational burden without compromising fidelity, providing a foundation for agile iteration of neutronics simulations involving multiple coupled material properties.

Bae, Jin Whan [ORNL] (ORCID:0000000326548907)↗

Real Space Quantum Cluster Formulation for the Typical Medium Theory of Anderson Localization

We develop a real space cluster extension of the typical medium theory (cluster-TMT) to study Anderson localization. By construction, the cluster-TMT approach is formally equivalent to the real space cluster extension of the dynamical mean field theory. Applying the developed method to the 3D Anderson model with a box disorder distribution, we demonstrate that cluster-TMT successfully captures the localization phenomena in all disorder regimes. As a function of the cluster size, our method obtains the correct critical disorder strength for the Anderson localization in 3D, and systematically recovers the re-entrance behavior of the mobility edge. From a general perspective, our developed methodology offers the potential to study Anderson localization at surfaces within quantum embedding theory. This opens the door to studying the interplay between topology and Anderson localization from first principles.

36 MATERIALS SCIENCE↗

Transient nucleation in condensed systems

Using classical nucleation theory we consider transient nucleation occurring in a one-component, condensed system under isothermal conditions. We obtain an exact closed-form expression for the time dependent cluster populations. In addition, a more versatile approach is developed: a numerical simulation technique which models directly the reactions by which clusters are produced. This simulation demonstrates the evolution of cluster populations and nucleation rate in the transient regime. Results from the simulation are verified by comparison with exact analytical solutions for the steady state. Experimental methods for measuring transient nucleation are assessed, and it is demonstrated that the observed behavior depends on the method used. The effect of preexisting cluster distributions is studied. Previous analytical and numerical treatments of transient nucleation are compared to the solutions obtained from the simulation. The simple expressions of Kashchiev are shown to give good descriptions of the nucleation behavior.

Kelton, K. F.↗

Technical support for creating an artificial intelligence system for feature extraction and experimental design

Techniques for classifying objects into groups or clases go under many different names including, most commonly, cluster analysis. Mathematically, the general problem is to find a best mapping of objects into an index set consisting of class identifiers. When an a priori grouping of objects exists, the process of deriving the classification rules from samples of classified objects is known as discrimination. When such rules are applied to objects of unknown class, the process is denoted classification. The specific problem addressed involves the group classification of a set of objects that are each associated with a series of measurements (ratio, interval, ordinal, or nominal levels of measurement). Each measurement produces one variable in a multidimensional variable space. Cluster analysis techniques are reviewed and methods for incuding geographic location, distance measures, and spatial pattern (distribution) as parameters in clustering are examined. For the case of patterning, measures of spatial autocorrelation are discussed in terms of the kind of data (nominal, ordinal, or interval scaled) to which they may be applied.

Glick, B. J.↗