Search NASA⌕ Search

SEARCH · Search NASA

Results for “Dynamic clustering algorithm”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

Automatic Traffic Queue-End Identification using Location-Based Waze User Reports

Traffic queues, especially queues caused by non-recurrent events such as incidents, are unexpected to high-speed drivers approaching the end of queue (EOQ) and become safety concerns. Though the topic has been extensively studied, the identification of EOQ has been limited by the spatial-temporal resolution of traditional data sources. This study explores the potential of location-based crowdsourced data, specifically Waze user reports. It presents a dynamic clustering algorithm that can group the location-based reports in real time and identify the spatial-temporal extent of congestion as well as the EOQ. The algorithm is a spatial-temporal extension of the density-based spatial clustering of applications with noise (DBSCAN) algorithm for real-time streaming data with an adaptive threshold selection procedure. Here, the proposed method was tested with 34 traffic congestion cases in the Knoxville, Tennessee area of the United States. It is demonstrated that the algorithm can effectively detect spatial-temporal extent of congestion based on Waze report clusters and identify EOQ in real-time. The Waze report-based detection are compared to the detection based on roadside sensor data. The results are promising: The EOQ identification time of Waze is similar to the EOQ detection time of traffic sensor data, with only 1.1 min difference on average. In addition, Waze generates 1.9 EOQ detection points every mile, compared to 1.8 detection points generated by traffic sensor data, suggesting the two data sources are comparable in respect of reporting frequency. The results indicate that Waze is a valuable complementary source for EOQ detection where no traffic sensors are installed.

99 GENERAL AND MISCELLANEOUS↗

Scalable edge clustering of dynamic graphs via weighted line graphs

Timestamped relational datasets consisting of records (or connections) between pairs of entities are ubiquitous in network science. For applications like peer-to-peer communication, email, various social network interactions, and computer network security, it is useful to organize these records into groups based on how and when they are occurring. Weighted line graphs offer a natural way to model how records are related in such datasets but for large real-world graph topologies, building and utilizing the line graph is prohibitively expensive. Here, we present the framework to cluster the edges of a dynamic graph via the associated line graph that contains two major contributions. The first is a method to work with the line graph implicitly and the second is a distributed scale implementation of an agglomerative hierarchical graph clustering algorithm. We outline a novel hierarchical dynamic graph edge clustering approach that efficiently breaks massive relational datasets into small sets of edges containing events at various timescales. This is in stark contrast to traditional graph clustering algorithms that prioritize highly connected (clique-like) community structures. Our approach relies on constructing a sufficient subgraph of a weighted line graph and applying a hierarchical agglomerative clustering. This approach is related to scalable techniques from spatial clustering, nonlinear-dimension reduction, topological data analysis, and draws particular inspiration from HDBSCAN. As an edge clustering, this method yields an overlapping node clustering. Our algorithm is parallelizable and we demonstrate efficient clustering of a billion-scale, real-world dynamic graph into small edge sets that correlate in topology and time. The entire clustering process for a graph with tens of billions of edges takes just a few minutes of run time on 256 nodes of a distributed compute environment. We argue how the output of the edge clustering is useful for a multitude of data visualization and powerful machine learning tasks, both involving the original massive dynamic graph data and metadata associated with the nodes and edges. Finally, we describe how this approach can be extended to dynamic hypergraphs and dynamic graphs/hypergraphs with unstructured data living on vertices and edges.

Data Analysis↗

Dynamic PRA-Based Estimation of PWR Coping Time Using a Surrogate Model for Accident Tolerant Fuel

In this study, we propose an interpolation-based response surface surrogate methodology to manage a large number of scenarios in dynamic probabilistic risk assessment. It adopts the shape Dynamic Time Warping algorithm to cluster the interpolation neighborhood from time series sample data. The interpolation method was adapted from Taylor Kriging to allow a reduced-order model of the Taylor series. In order to demonstrate its applicability to complex issues in risk assessment for nuclear engineering, an example risk response surface to estimate emergency core cooling system (ECCS) criteria for triplex silicon carbide (SiC) accident-tolerant fuel was constructed. The response surface was exploited to estimate the cumulative failure probability of the fuel cladding structure due to the uncertainties in operator actions and safety systems. The functional failures were assessed based on a combination of individual layer failures computed by coupling Risk Analysis Virtual Environment software with a pressurized water reactor 1000-MW(electric) RELAP5 model and the in-house fuel performance assessment module. Results showed that SiC cladding failure probability spiked less than 1 min after a large-break loss-of- coolant accident whenever the current ECCS criteria for Zircaloy-4 (Zr-4) cladding was used. However, it still provides an increased safety margin of three orders of magnitude compared to Zr-4. This positive margin could be utilized to relax active ECCS requirements by allowing deviations of up to 450 s in its actuation time. The proposed surrogate methodology generated a response surface of SiC cladding failure probability reasonably well, with a significant savings of computation time. This methodology is expected to be useful in the analysis of system response with complex uncertainty sources.

22 GENERAL STUDIES OF NUCLEAR REACTORS↗

Jacobian-scaled K-means clustering for physics-informed segmentation of reacting flows

This work introduces Jacobian-scaled K-means (JSK-means) clustering, which is a physicsinformed clustering strategy centered on the K-means framework. The method allows for the injection of underlying physical knowledge into the clustering procedure through a distance function modification: instead of leveraging conventional Euclidean distance vectors, the JSKmeans procedure operates on distance vectors scaled by matrices obtained from dynamical system Jacobians evaluated at the cluster centroids. The goal of this work is to show how the JSKmeans algorithm - without modifying the input dataset - produces clusters that capture regions of dynamical similarity, in that the clusters are redistributed towards high-sensitivity regions in phase space and are described by similarity in the source terms of samples instead of the samples themselves. The algorithm is demonstrated on a complex reacting flow simulation dataset (a channel detonation configuration), where the dynamics in the thermochemical composition space are known through the highly nonlinear and stiff Arrhenius-based chemical source terms. Interpretations of cluster partitions in both physical space and composition space reveal how JSK-means shifts clusters produced by standard K-means towards regions of high chemical sensitivity (e.g., towards regions of peak heat release rate near the detonation reaction zone). Furthermore, the findings presented here illustrate the benefits of utilizing Jacobian-scaled distances in clustering techniques, and the JSK-means method in particular displays promising potential for improving former partition-based modeling strategies in reacting flow (and other multi-physics) applications.

Clustering↗

DONKEY: A Flexible and Accurate Algorithm for Clustering

We propose an accurate clustering algorithm suitable for the varied and multidimensional data sets that correspond to temporal snapshots from on-the-fly nonadiabatic trajectory-based simulations of photoexcited dynamics. The algorithm approximates the underlying probability density function using variable kernel density estimation, with local maxima corresponding to cluster centers. Each data point is then assigned to one of the maxima by employing a maximization procedure. Finally, clusters artificially separated by minor fluctuations in the probability density are merged. The algorithm does not require parameter tuning, which ensures flexibility and reduces the risk of bias. It is tested on several synthetic data sets, where it consistently outperforms conventional clustering algorithms. As a final example, the algorithm is applied to the excited dynamics of the norbornadiene ⇌ quadricyclane (C 7 H 8 ) molecular photoswitch, demonstrating how distinct reaction pathways can be identified.

algorithms↗

Investigation of plasmon relaxation mechanisms using nonadiabatic molecular dynamics

Hot carriers generated from the decay of plasmon excitation can be harvested to drive a wide range of physical or chemical processes. However, their generation efficiency is limited by the concomitant phonon-induced relaxation processes by which the energy in excited carriers is transformed into heat. However, simulations of dynamics of nanoscale clusters are challenging due to the computational complexity involved. Here, in this paper, we adopt our newly developed Trajectory Surface Hopping (TSH) nonadiabatic molecular dynamics algorithm to simulate plasmon relaxation in Au 20 clusters, taking the atomistic details into account. The electronic properties are treated within the Linear Response Time-Dependent Tight-binding Density Functional Theory (LR-TDDFTB) framework. The relaxation of plasmon due to coupling to phonon modes in Au 20 beyond the Born–Oppenheimer approximation is described by the TSH algorithm. The numerically efficient LR-TDDFTB method allows us to address a dense manifold of excited states to ensure the inclusion of plasmon excitation. Starting from the photoexcited plasmon states in Au 20 cluster, we find that the time constant for relaxation from plasmon excited states to the lowest excited states is about 2.7 ps, mainly resulting from a stepwise decay process caused by low-frequency phonons of the Au 20 cluster. Furthermore, our simulations show that the lifetime of the phonon-induced plasmon dephasing process is ~10.4 fs and that such a swift process can be attributed to the strong nonadiabatic effect in small clusters. Our simulations demonstrate a detailed description of the dynamic processes in nanoclusters, including plasmon excitation, hot carrier generation from plasmon excitation dephasing, and the subsequent phonon-induced relaxation process.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Automated characterization of spatial and dynamical heterogeneity in supercooled liquids via implementation of machine learning

Abstract A computational approach by an implementation of the principle component analysis (PCA) with K -means and Gaussian mixture (GM) clustering methods from machine learning algorithms to identify structural and dynamical heterogeneities of supercooled liquids is developed. In this method, a collection of the average weighted coordination numbers ( W C N s ‾ ) of particles calculated from particles’ positions are used as an order parameter to build a low-dimensional representation of feature (structural) space for K -means clustering to sort the particles in the system into few meso-states using PCA. Nano-domains or aggregated clusters are also formed in configurational (real) space from a direct mapping using associated meso-states’ particle identities with some misclassified interfacial particles. These classification uncertainties can be improved by a co-learning strategy which utilizes the probabilistic GM clustering and the information transfer between the structural space and configurational space iteratively until convergence. A final classification of meso-states in structural space and domains in configurational space are stable over long times and measured to have dynamical heterogeneities. Armed with such a classification protocol, various studies over the thermodynamic and dynamical properties of these domains indicate that the observed heterogeneity is the result of liquid–liquid phase separation after quenching to a supercooled state.

75 CONDENSED MATTER PHYSICS, SUPERCONDUCTIVITY AND↗

Continuous momentum dependence in the dynamical cluster approximation

The dynamical cluster approximation (DCA) is a quantum cluster extension to the single-site dynamical mean-field theory that incorporates spatially nonlocal dynamic correlations systematically and nonperturbatively. The DCA + algorithm addresses the cluster shape dependence of the DCA and improves the convergence with cluster size by introducing a lattice self-energy with continuous momentum dependence. However, we show that the DCA + algorithm is plagued by a fundamental problem when its self-consistency equations are formulated using the bare Green's function of the cluster. This problem is most severe in the strongly correlated regime at low doping, where the DCA + self-energy becomes overly metallic and local, and persists to cluster sizes where the standard DCA has long converged. In view of the failure of the DCA + algorithm, we propose to complement DCA simulations with a post-interpolation procedure for single-particle and two-particle correlation functions to preserve continuous momentum dependence and the associated benefits in the DCA. We demonstrate the effectiveness of this practical approach with results for the half-filled and hole-doped two-dimensional Hubbard model.

75 CONDENSED MATTER PHYSICS, SUPERCONDUCTIVITY AND↗

Predicting images for the dynamics of stellar clusters ( π-DOC ): a deep learning framework to predict mass, distance, and age of globular clusters

ABSTRACT Dynamical mass estimates of simple systems such as globular clusters (GCs) still suffer from up to a factor of 2 uncertainty. This is primarily due to the oversimplifications of standard dynamical models that often neglect the effects of the long-term evolution of GCs. Here, we introduce a new approach to measure the dynamical properties of GCs, based on the combination of a deep-learning framework and the large amount of data from direct N-body simulations. Our algorithm, π-DOC (Predicting Images for the Dynamics Of stellar Clusters) is composed of two convolutional networks, trained to learn the non-trivial transformation between an observed GC luminosity map and its associated mass distribution, age, and distance. The training set is made of V-band luminosity and mass maps constructed as mock observations from N-body simulations. The tests on π-DOC demonstrate that we can predict the mass distribution with a mean error per pixel of 27 per cent, and the age and distance with an accuracy of 1.5 Gyr and 6 kpc, respectively. In turn, we recover the shape of the mass-to-light profile and its global value with a mean error of 12 per cent, which implies that we efficiently trace mass segregation. A preliminary comparison with observations indicates that our algorithm is able to predict the dynamical properties of GCs within the limits of the training set. These encouraging results demonstrate that our deep-learning framework and its forward modelling approach can offer a rapid and adaptable tool competitive with standard dynamical models.

Chardin, Jonathan↗

Improving the Accuracy of Clustering Electric Utility Net Load Data using Dynamic Time Warping

Identifying patterns in electric utility net load data in a time-series format is very useful in preparing the operation for next day. Machine learning algorithms have been used in other domains and those concepts are applied in this paper on real-world net load measurement data. Clustering is the practice of grouping data with similar characteristics as determined by the distance measure. The K-means clustering algorithm is utilized here with actual electric utility data. The paper uses the standard distance measure, Euclidean distance (ED), and compares its performance against the dynamic time warping (DTW) measure. An actual case study with real data is presented, and DTW distance measure-based method observed to result better accuracy compared to the ED based method for substation net load measurements predominantly with residential customers.

clustering↗

Distribution Cutoff for Clusters near the Gel Point

The mechanical and dynamic properties of developing networks near the gel point are susceptible to the distribution of clusters coexisting with percolating networks. The distribution of cluster numbers follows a broad power law, wrapped by a cutoff function that rapidly decays at a characteristic size. The form of the cutoff function has been speculated based on known results from lattice percolation and, in certain cases, solved. We obtained this cutoff function from simulated dynamic clusters of polymeric precursor chains using a hybrid Monte Carlo algorithm. The results obtained from three different precursor chain lengths are consistent with each other and are consistent with the expectation from lattice percolation.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Sub-system quantum dynamics using coupled cluster downfolding techniques

In this paper, we discuss extending the sub-system embedding sub-algebra coupled cluster (SES-CC) formalism and the double unitary coupled cluster (DUCC) ansatz to the time domain. As we demonstrated in earlier studies, it is possible, using these formalisms, to calculate the energy of the entire system as an eigenvalue of downfolded/effective Hamiltonian in the active space, that is identifiable with the sub-system of the composite system. In these studies, we demonstrated that downfolded Hamiltonians integrate out Fermionic degrees of freedom that do not correspond to the physics encapsulated by the active space. We extend these results to the time-dependent Schrödinger equation, showing that a similar construct is possible to partition a system into a sub-system that varies slowly in time and a remaining subsystem that corresponds to fast oscillations. This time dependent formalism allows coupled cluster quantum dynamics to be extended to larger systems and for the formulation of novel quantum algorithms based on the quantum Lanczos approach, which have recently been considered in the literature.

coupled cluster, Electron correlation, quantum dyn↗

Electromagnetic Transient (EMT) Simulation Algorithm for Evaluation of Photovoltaic (PV) Generation Systems

Use of inverter-based resources facilitating renewable energy resources such as photovoltaic (PV) generation is increasing rapidly with decreasing costs and reduced emissions associated. To accommodate such rapid growth of inverter-based resources like PV systems, electromagnetic transient (EMT) simulation models of both PV systems and grids are required to analyze the interaction of PVs in the grid (like the post-event analysis). In addition, the EMT simulation would help with the planning of future power grid with a large number of PVs as well as other inverter-based distributed generation systems. In this paper, the EMT simulation models of PV systems and grids are developed based on the differential algebraic equations (DAEs) representing their EMT dynamics. Furthermore, advanced simulation algorithms including numerical stiffness-based hybrid discretization, DAE clustering and aggregation, multi-order integration, and matrix splitting approaches are applied to accelerate the EMT simulation. The proposed algorithm was applied to 125 PV inverters within 52-bus medium-voltage (MV) distribution grid.

Choi, Jongchan↗

Optimizing the shape of photometric redshift distributions with clustering cross-correlations

We present an optimization method for the assignment of photometric galaxies to a chosen set of redshift bins. This is achieved by combining simulated annealing, an optimization algorithm inspired by solid-state physics, with an unsupervised machine learning method, a self-organizing map (SOM) of the observed colours of galaxies. Starting with a sample of galaxies that is divided into redshift bins based on a photometric redshift point estimate, the simulated annealing algorithm repeatedly reassigns SOM-selected subsamples of galaxies, which are close in colour, to alternative redshift bins. We optimize the clustering cross-correlation signal between photometric galaxies and a reference sample of galaxies with well-calibrated redshifts. Depending on the effect on the clustering signal, the reassignment is either accepted or rejected. By dynamically increasing the resolution of the SOM, the algorithm eventually converges to a solution that minimizes the number of mismatched galaxies in each tomographic redshift bin and thus improves the compactness of their corresponding redshift distribution. This method is demonstrated on the synthetic Legacy Survey of Space and Time cosmoDC2 catalogue. We find a significant decrease in the fraction of catastrophic outliers in the redshift distribution in all tomographic bins, most notably in the highest redshift bin with a decrease in the outlier fraction from 57 percent to 16 percent.

79 ASTRONOMY AND ASTROPHYSICS↗

Graphic contrastive learning analyses of discontinuous molecular dynamics simulations: Study of protein folding upon adsorption

A comprehensive understanding of the interfacial behaviors of biomolecules holds great significance in the development of biomaterials and biosensing technologies. In this work, we used discontinuous molecular dynamics (DMD) simulations and graphic contrastive learning analysis to study the adsorption of ubiquitin protein on a graphene surface. Our high-throughput DMD simulations can explore the whole protein adsorption process including the protein structural evolution with sufficient accuracy. Contrastive learning was employed to train a protein contact map feature extractor aiming at generating contact map feature vectors. Subsequently, these features were grouped using the k-means clustering algorithm to identify the protein structural transition stages throughout the adsorption process. The machine learning analysis can illustrate the dynamics of protein structural changes, including the pathway and the rate-limiting step. Our study indicated that the protein–graphene surface hydrophobic interactions and the π–π stacking were crucial to the seven-stage adsorption process. Upon adsorption, the secondary structure and tertiary structure of ubiquitin disintegrated. The unfolding stages obtained by contrastive learning-based algorithm were not only consistent with the detailed analyses of protein structures but also provided more hidden information about the transition states and pathway of protein adsorption process and structural dynamics. Our combination of efficient DMD simulations and machine learning analysis could be a valuable approach to studying the interfacial behaviors of biomolecules.

97 MATHEMATICS AND COMPUTING↗

Possibilities and Limitations of Kinematically Identifying Stars from Accreted Ultra-faint Dwarf Galaxies

Abstract The Milky Way has accreted many ultra-faint dwarf galaxies (UFDs), and stars from these galaxies can be found throughout our Galaxy today. Studying these stars provides insight into galaxy formation and early chemical enrichment, but identifying them is difficult. Clustering stellar dynamics in 4D phase space ( E , L z , J r , J z ) is one method of identifying accreted structure that is currently being utilized in the search for accreted UFDs. We produce 32 simulated stellar halos using particle tagging with the Caterpillar simulation suite and thoroughly test the abilities of different clustering algorithms to recover tidally disrupted UFD remnants. We perform over 10,000 clustering runs, testing seven clustering algorithms, roughly twenty hyperparameter choices per algorithm, and six different types of data sets each with up to 32 simulated samples. Of the seven algorithms, HDBSCAN most consistently balances UFD recovery rates and cluster realness rates. We find that, even in highly idealized cases, the vast majority of clusters found by clustering algorithms do not correspond to real accreted UFD remnants and we can generally only recover 6% of UFDs remnants at best. These results focus exclusively on groups of stars from UFDs, which have weak dynamic signatures compared to the background of other stars. The recoverable UFD remnants are those that accreted recently, z accretion ≲ 0.5. Based on these results, we make recommendations to help guide the search for dynamically linked clusters of UFD stars in observational data. We find that real clusters generally have higher median energy and J r , providing a way to help identify real versus fake clusters. We also recommend incorporating chemical tagging as a way to improve clustering results.

79 ASTRONOMY AND ASTROPHYSICS↗

Ab Initio Study of the Beryllium Isotopes 7 Be to 12 Be

We present a systematic ab initio study of the low-lying states in beryllium isotopes from 7 Be to 12 Be using nuclear lattice effective field theory with the N 3 ⁢LO interaction. Our calculations achieve good agreement with experimental data for energies, radii, and electromagnetic properties. We introduce a novel, model-independent method to quantify nuclear shapes, uncovering a distinct pattern in the interplay between positive and negative parity states across the isotopic chain. By combining Monte Carlo sampling of the many-body density operator with a novel nucleon-grouping algorithm, the prominent two-center cluster structures, the emergence of one-neutron halo, complex nuclear molecular dynamics such as 𝜋 orbital and 𝜎 orbital, emerge naturally.

binding energy & masses↗