Search NASA⌕ Search

SEARCH · Search NASA

Results for “clustering analysis (CA)”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

Structuring Nutrient Yields throughout Mississippi/Atchafalaya River Basin Using Machine Learning Approaches

To minimize the eutrophication pressure along the Gulf of Mexico or reduce the size of the hypoxic zone in the Gulf of Mexico, it is important to understand the underlying temporal and spatial variations and correlations in excess nutrient loads, which are strongly associated with the formation of hypoxia. This study’s objective was to reveal and visualize structures in high-dimensional datasets of nutrient yield distributions throughout the Mississippi/Atchafalaya River Basin (MARB). For this purpose, the annual mean nutrient concentrations were collected from thirty-three US Geological Survey (USGS) water stations scattered in the upper and lower MARB from 1996 to 2020. Eight surface water quality indicators were selected to make comparisons among water stations along the MARB over the past two decades. Principal component analysis (PCA) was used to comprehensively evaluate the nutrient yields across thirty-three USGS monitoring stations and identify the major contributing nutrient loads. The results showed that all samples could be analyzed using two main components, which accounted for 81.6% of the total variance. The PCA results showed that yields of orthophosphate (OP), silica (SI), nitrate–nitrites (NO 3 -NO 2 ), and total suspended sediment (TSS) are major contributors to nutrient yields. It also showed that land-planted crops, density of population, domestic and industrial discharges, and precipitation are fundamental causes of excess nutrient loads in MARB. These factors are of great significance for the excess nutrient load management and pollution control of the Mississippi River. It was found that the average nutrient yields were stable within the sub-MARB area, but the large nitrogen yields in the upper MARB and the large phosphorus yields in the lower MARB were of great concern. t-distributed stochastic neighbor embedding (t-SNE) revealed interesting nonlinear and local structures in nutrient yield distributions. Clustering analysis (CA) showed the detailed development of similarities in the nutrient yield distribution. Moreover, PCA, t-SNE, and CA showed consistent clustering results. This study demonstrated that the integration of dimension reduction techniques, PCA, and t-SNE with CA techniques in machine learning are effective tools for the visualization of the structures of the correlations in high-dimensional datasets of nutrient yields and provide a comprehensive understanding of the correlations in the distributions of nutrient loads across the MARB.

54 ENVIRONMENTAL SCIENCES↗

Streamlining heterologous expression of top carbonic anhydrases in Escherichia coli : bioinformatic and experimental approaches

Carbonic anhydrase (CA) enzymes facilitate the reversible hydration of CO 2 to bicarbonate ions and protons. Identifying efficient and robust CAs and expressing them in model host cells, such as Escherichia coli, enables more efficient engineering of these enzymes for industrial CO 2 capture. However, expression of CAs in E. coli is challenging due to the possible formation of insoluble protein aggregates, or inclusion bodies. This makes the production of soluble and active CA protein a prerequisite for downstream applications. In this study, we streamlined the process of CA expression by selecting seven top CA candidates and used two bioinformatic tools to predict their solubility for expression in E. coli. The prediction results place these enzymes in two categories: low and high solubility. Our expression of high solubility score CAs (namely CA5-SspCA, CA6-SazCAtrunc, CA7-PabCA and CA8-PhoCA) led to significantly higher protein yields (5 to 75 mg purified protein per liter) in flask cultures, indicating a strong correlation between the solubility prediction score and protein expression yields. Furthermore, phylogenetic tree analysis demonstrated CA class-specific clustering patterns for protein solubility and production yields. Unexpectedly, we also found that the unique N-terminal, 11-amino acid segment found after the signal sequence (not present in its homologs), was essential for CA6-SazCA activity. Overall, this work demonstrated that protein solubility prediction, phylogenetic tree analysis, and experimental validation are potent tools for identifying top CA candidates and then producing soluble, active forms of these enzymes in E. coli. The comprehensive approaches we report here should be extendable to the expression of other heterogeneous proteins in E. coli.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Scalable edge clustering of dynamic graphs via weighted line graphs

Timestamped relational datasets consisting of records (or connections) between pairs of entities are ubiquitous in network science. For applications like peer-to-peer communication, email, various social network interactions, and computer network security, it is useful to organize these records into groups based on how and when they are occurring. Weighted line graphs offer a natural way to model how records are related in such datasets but for large real-world graph topologies, building and utilizing the line graph is prohibitively expensive. Here, we present the framework to cluster the edges of a dynamic graph via the associated line graph that contains two major contributions. The first is a method to work with the line graph implicitly and the second is a distributed scale implementation of an agglomerative hierarchical graph clustering algorithm. We outline a novel hierarchical dynamic graph edge clustering approach that efficiently breaks massive relational datasets into small sets of edges containing events at various timescales. This is in stark contrast to traditional graph clustering algorithms that prioritize highly connected (clique-like) community structures. Our approach relies on constructing a sufficient subgraph of a weighted line graph and applying a hierarchical agglomerative clustering. This approach is related to scalable techniques from spatial clustering, nonlinear-dimension reduction, topological data analysis, and draws particular inspiration from HDBSCAN. As an edge clustering, this method yields an overlapping node clustering. Our algorithm is parallelizable and we demonstrate efficient clustering of a billion-scale, real-world dynamic graph into small edge sets that correlate in topology and time. The entire clustering process for a graph with tens of billions of edges takes just a few minutes of run time on 256 nodes of a distributed compute environment. We argue how the output of the edge clustering is useful for a multitude of data visualization and powerful machine learning tasks, both involving the original massive dynamic graph data and metadata associated with the nodes and edges. Finally, we describe how this approach can be extended to dynamic hypergraphs and dynamic graphs/hypergraphs with unstructured data living on vertices and edges.

Data Analysis↗

Unsupervised atomic data mining via multi-kernel graph autoencoders for machine learning force fields

Constructing a chemically diverse dataset while avoiding sampling bias is critical to training efficient and generalizable force fields. However, in computational chemistry and materials science, many common dataset generation techniques are prone to oversampling regions of the potential energy surface. Furthermore, these regions can be difficult to identify and isolate from each other or may not align well with human intuition, making it challenging to systematically remove bias in the dataset. While traditional clustering and pruning (down-sampling) approaches can be useful for this, they can often lead to information loss or a failure to properly identify distinct regions of the potential energy surface due to difficulties associated with the high dimensionality of atomic descriptors. In this work, we introduce the Multi-kernel Edge Attention-based Graph Autoencoder (MEAGraph) model, an unsupervised approach for analyzing atomic datasets. MEAGraph combines multiple linear kernel transformations with attention-based message passing to capture geometric sensitivity and enable effective dataset pruning without relying on labels or extensive training. Demonstrated applications on niobium, tantalum, and iron datasets show that MEAGraph efficiently groups similar atomic environments, allowing for the use of basic pruning techniques for removing sampling bias. This approach provides an effective method for representation learning and clustering that can be used for data analysis, outlier detection, and dataset optimization.

Materials science↗

Critical Simulation Pipeline for COG Suites [Poster]

The CRItical Simulation Pipeline (CRISP) is a Python package for automating validation of reactor criticality benchmarks. CRISP supplies COG—a multi-particle radiation transport code maintained by the Nuclear Criticality Safety Division—with a pipeline to calculate k eff performance for 400+ benchmark experiments with 3,400+ configurations from the International Criticality Safety Benchmark Evaluation Project (ICSBEP). The pipeline includes four stages: materials configuration, input card templating, cluster submission, and results analysis. CRISP includes a command-line interface to facilitate user interaction.

73 NUCLEAR PHYSICS AND RADIATION PHYSICS↗

Density functional modeling of the binding energies between aluminosilicate oligomers and different metal cations

Interactions between negatively charged aluminosilicate species and positively charged metal cations are critical to many important engineering processes and applications, including sustainable cements and aluminosilicate glasses. In an effort to probe these interactions, here we have calculated the pair-wise interaction energies (i.e., binding energies) between aluminosilicate dimer/trimer and 17 different metal cations M n+ (M n+ = Li + , Na + , K + , Cu + , Cu 2+ , Co 2+ , Zn 2+ , Ni 2+ , Mg 2+ , Ca 2+ , Ti 2+ , Fe 2+ , Fe 3+ , Co 3+ , Cr 3+ , Ti 4+ and Cr 6+ ) using a density functional theory (DFT) approach. Analysis of the DFT-optimized structural representations for the clusters (dimer/trimer + M n+ ) shows that their structural attributes (e.g., interatomic distances) are generally consistent with literature observations on aluminosilicate glasses. The DFT-derived binding energies are seen to vary considerably depending on the type of cations (i.e., charge and ionic radii) and aluminosilicate species (i.e., dimer or trimer). A survey of the literature reveals that the difference in the calculated binding energies between different M n+ can be used to explain many literature observations associated with the impact of metal cations on materials properties (e.g., glass corrosion, mineral dissolution, and ionic transport). Analysis of all the DFT-derived binding energies reveals that the correlation between these energy values and the ionic potential and field strength of the metal cations are well captured by 2nd order polynomial functions ( R 2 values of 0.99–1.00 are achieved for regressions). Given that the ionic potential and field strength of a given metal cation can be readily estimated using well-tabulated ionic radii available in the literature, these simple polynomial functions would enable rapid estimation of the binding energies of a much wider range of cations with the aluminosilicate dimer/trimer, providing guidance on the design and optimization of sustainable cements and aluminosilicate glasses and their associated applications. Finally, the limitations associated with using these simple model systems to model complex interactions are also discussed.

Gong, Kai↗

Geometric Interpretation of the Cluster Location Problem Part I: Theory

We present a new framing of the seismic location problem using principles drawn from differential geometry. Our interpretation relies upon the common assumption that travel times observed across a network are continuous, differentiable functions of source location. In consequence, travel‐time functions constitute a differentiable map between the source region and a Riemannian manifold. The manifold is said to be the image of the source region embedded in a generally high‐dimension travel‐time vector space. A cluster of events in the source region has an image of discrete points on the manifold, that, except in the simplest cases, cannot be viewed directly. However, it is possible to project the image of a cluster into a tangent space of the manifold for direct visualization. The projection operator can be computed directly from the data without a velocity model, but produces a distorted rendering of the cluster geometry. With a model we can predict the distortions and correct them to estimate cluster geometry. We develop these points with the simplest possible example, one for which direct visualization of the manifold is possible, using the example as an introduction to the relevant concepts from differential geometry in a familiar setting. The tangent space, a local linearization of the manifold, plays a key role. We develop a metric to estimate the limits of linearization, that is, to determine when the curvature of the manifold invalidates the linear assumption. We also examine the interplay of model error, inadequate network geometry, and pick error. We then generalize our results from the simple case to the general case of 3D source regions observed by general networks. Although we do suggest a new “project and correct” method for location, we do not develop it into a practical algorithm. In conclusion, our intention rather is to highlight new analytical methods grounded in differential geometry.

East Pacific Ocean Islands↗

Accelerating the Structure Exploration of Diverse Bi–Pt Nanoclusters via Physics‐Informed Machine Learning Potential and Particle Swarm Optimization

Bimetallic Bi–Pt nanoclusters exhibit diverse structural motifs, including core-shell, Janus, and mixed alloy configurations, due to the unique bonding characteristics between Bi and Pt atoms. Using density functional theory refinements from ChIMES physically machine-learned potential and CALYPSO particle swarm optimization global searches, 34 Bi20-Pt20 nanoclusters are systematically classified. The results reveal that Bi atoms predominantly occupy surface sites, driven by charge transfer effects. Cohesive energy trends alone prove insufficient for structure differentiation, necessitating a data-driven approach employing principal component analysis and K-means clustering. Furthermore, vibrational, electronic, and infrared spectral analyses provide additional insights into structure-property relationships. The findings offer an original framework for the automated classification and analysis of bimetallic nanoclusters, enhancing the understanding of their stability and functional properties.

bimetallic nanoparticles↗

CASM Monte Carlo: Calculations of the thermodynamic and kinetic properties of complex multicomponent crystals

Monte Carlo techniques play a central role in statistical mechanics approaches that connect macroscopic thermodynamic and kinetic properties to the electronic structure of a material. This paper describes the implementation of Monte Carlo techniques for the study of multicomponent crystalline materials within the Clusters Approach to Statistical Mechanics (CASM) software suite, and demonstrates their use in model systems to calculate free energies and kinetic coefficients, study phase transitions, and construct phase diagrams from first principles. Many crystal structures are complex, with multiple sublattices occupied by differing sets of chemical species, along with the presence of vacancies or interstitial species. This imposes constraints on concentration variables, the form of thermodynamic potentials, and the values of kinetic transport coefficients. The framework used by CASM to formulate thermodynamic potentials and kinetic transport coefficients accounting for arbitrarily complex crystal structures is presented and demonstrated with examples of increasing complexity. Additionally, an overview of the capabilities of the CASM software specific to Monte Carlo methods is given, and a new CASM software package is introduced, casm-flow, which helps automate the setup, submission, management, and analysis of Monte Carlo simulations.

Cluster expansion↗

State-level suicide mortality insights: a comparative study of VHA veterans and the whole US population

Background: Suicide is a leading cause of death in the US Comparative State-level spatial analysis between Veterans Health Administration (VHA veterans) and the whole US population can reveal differences in conditions for targeted interventions and intricate geographical patterns. Methods: The study population contains 2018 and 2019 suicide deaths of VHA veterans and the whole US population. They were used to calculate state-level rates. States were classified by whether their VHA veteran and whole US population rates were above or below respective mean rates. Local Moran’s I was leveraged to examine spatial autocorrelation. Results: State-level suicide mortality rates and disparities among states were generally higher for VHA veterans (2018: 37.3 ± 7.2; 2019: 46.8 ± 8.3) than for the whole US population (2018: 16.6 ± 4.3; 2019: 16.4 ± 4.4). For both populations, there were statistically significant clusters with high suicide rates. Over one-fourth of states demonstrated inverse relationships, with rates above mean for one group but below for other. VHA veterans are at higher risk with over one-third of states had greater than average veteran suicide risk ratio. Conclusions: VHA veterans are at higher risk than the whole population across all states. Mortality disparities among states and clusters of states with high and low rates suggest targeted interventions and cooperative health strategies may help address these differences.

60 APPLIED LIFE SCIENCES↗

Memoirs of Mass Accretion: Probing the Edges of Intracluster Light in Simulated Galaxy Clusters

The diffuse starlight extending throughout massive galaxy clusters, known as intracluster light (ICL), has the potential to be read as a memoir of mass accretion: informative, individual, and yet imperfect. Here, we combine dark-matter-only zoom-in simulations from the Symphony suite with the Nimbus “star-tagging” model of the stellar halo to assess how much information about the mass assembly of an individual galaxy cluster can be gleaned from idealized measurements of ICL outskirts. We show that the edges of a cluster’s stellar profile—the primary (R sp⋆,1 ) and secondary (R sp⋆,2 ) stellar “splashback” radii—are sensitive to both continuous mass accretion histories (MAHs) and discrete merger events, making them potentially powerful probes of a cluster’s past. We find that R sp⋆,1 strongly correlates with the cluster’s mass ∼1 dynamical time ago, while R sp⋆,2 traces more recent MAH to a slightly lesser degree. In combination, these features can further distinguish between clusters that have and have not undergone a major merger within the past dynamical time. We use both to predict realistic cluster MAHs with the MultiCAM framework. These outer ICL features are significantly more sensitive to mass accretion and merger histories than the stellar mass gap and halo concentration, and perform comparably to the commonly used X-ray-based tracer of relaxedness, x off . While our analysis is idealized, the relevant ICL features are potentially detectable in next-generation deep imaging of nearby clusters. This work highlights the promise of ICL measurements and lays the groundwork for more detailed forecasts of their power.

79 ASTRONOMY AND ASTROPHYSICS↗

Nucleation and growth of polar clusters with in-phase tilts into a long-range ferroelectric matrix in a sodium niobate based complex relaxor

In this study, we have investigated the temperature dependence of atomic ordering at multiple length scales in a lead-free sodium niobate-based relaxor, i.e., 0.75 NaNbO 3 -0.25 Ba 0.9⁢ Ca 0.1⁢ TiO 3 (NN-25BCT) via synchrotron x-ray diffraction, Raman spectroscopy, and pair distribution function analysis. High-resolution synchrotron x-ray powder diffraction (SXRD) measurements reveal a ferroelectric phase transition in the relaxor ferroelectric NN-25BCT below the Vogel-Fulcher freezing temperature (𝑇 VF ≈ 270 K). In addition, SXRD analysis demonstrates the competition between in-phase octahedral tilting and ferroelectric order at the long-range scale using mode crystallography. On the other hand, Raman spectroscopic analysis provides evidence of polar ordering for 𝑇 > 𝑇 VF (with tetragonal symmetry) persisting up to the Burns temperature (𝑇 B ). Furthermore, pair distribution function (PDF) analysis reveals the presence of a polar antiferrodistortive tetragonal phase with 𝑃⁢4⁢𝑏𝑚 space group at short ranges throughout the studied temperatures (i.e., 110 K ≤ 𝑇 ≤500 K), irrespective of nonpolar long-range ordering above 𝑇 VF . Therefore, our measurements provide direct evidence for the presence of polar ordering at short ranges and their gradual transformation into long-range polar ordering using an integrated multiscale structural analysis. In conclusion, as a result of a transition from relaxor to a ferroelectric phase in the vicinity of room temperature, NN-25BCT can be exploited for applications in pyroelectric detectors, electrocaloric devices, and multilayered ceramic capacitors.

36 MATERIALS SCIENCE↗

Similarities in Meteorological Composites Among Different Atmospheric River Detection Tools During Landfall Over Western Coastal North America

Many atmospheric river detection tools (ARDTs) have been developed over the past few decades to identify atmospheric rivers (ARs). Different ARDTs have been observed to capture a variety of frequencies, shapes, and sizes of ARs. Due to this, questions have arisen about the underlying phenomena associated with the detected ARs: do all ARDTs detect the same meteorological phenomena? In this paper, we assess eight ARDTs and investigate the underlying synoptic scale phenomena during landfalling ARs along the west coast of North America. We find that during landfalling AR events, prevalent low-pressure and high-pressure systems converge and enhance moisture influx toward the landfalling site. We identify that all eight ARDTs identify AR conditions associated with baroclinic waves, with the region of intense integrated vapor transport (IVT) located downstream of the upper level (500 hPa) trough. The magnitude of IVT is enhanced by the strength of the pressure gradients in the confluence region. Although the ARDTs assessed agree on the general phenomena, there are however subtle differences in each ARDT per the clustering analysis we performed. We conclude that the eight ARDTs identify similar underpinning synoptic scale meteorological phenomena.

54 ENVIRONMENTAL SCIENCES↗

Electronic Spin Relaxation and Clustering in High-Pressure High-Temperature Synthesized Microcrystalline Diamond Particles with Reduced Nitrogen Content

The negatively charged nitrogen-vacancy (NV – ) color center in diamonds is widely studied because of numerous applications of this unique quantum system in sensing and quantum information sciences. While substitutional nitrogen is required to form the NV – centers in diamond, it also yields other paramagnetic defects─primarily the neutrally charged substitutional nitrogen centers (P1)─that decrease NV – spin coherence, which in turn degrades performance in applications. Herein, we investigate high-pressure high-temperature synthesized diamond microparticles (ca. 140–185 μm) having lower─ranging from 3 to 38 ppm─than the typical nitrogen content of type 1b diamond ( ca. 100 ppm and higher) typically used for the production of fluorescent diamond particles with NV – centers. A suite of electron paramagnetic resonance, optically detected magnetic resonance, and nuclear magnetic resonance methods are used to characterize spin properties of P1 and NV – centers in the particles. Upon decreasing the nitrogen content from 29 to 3 ppm, the ensemble NV – T 2 relaxation time increased by about 3-fold as measured directly in the Hahn Echo experiment at magnetic field of 1.2 T. Analysis of electronic relaxation of P1 centers revealed the existence of at least two distinct populations of P1 centers, consisting of fast and slower relaxing spins and allowed for an estimation of local concentrations. Even with <10 ppm nitrogen contents, the analysis indicated a highly heterogeneous distribution of P1 centers, suggesting the possibility of P1 spin clustering even at low nitrogen concentrations. The combined data demonstrate that the particles prepared from HPHT diamond with a low nitrogen content offer improved spin properties that are beneficial for NV – sensing applications.

Carbon↗

The Many-Body Expansion for Metals: II. Nonadditive Terms in Clusters Composed of Metals with n s 1 , n s 2 , and n s 2 p 1 Configurations

The many-body expansion (MBE) was applied to homometallic and heterometallic trimers of metals with ns 1 , ns 2 , and ns 2 p 1 configurations to investigate its convergence, the magnitude and nature (stabilizing/destabilizing) of the individual terms and seek an understanding of their variation across the different families of clusters. In particular, we examined the series of alkali metals (Li 3 , Na 3 , K 3 , Li 2 Na, LiNa 2 ), alkali metal borides (Li 2 B and LiB 2 ), and alkaline earth metals (Be 3 , Be 2 Mg, BeMg 3 , and Mg 3 ) trimers, as well as sodium clusters Na n , n = 2–5. Here, we found that there is no uniform contribution (stabilizing or destabilizing) across the series in the different families of trimers. For instance, the 2-B term stabilizes the ground states of the Na 3 (doublet), Na 4 (singlet), and Na 5 (doublet) clusters, and the 3-B term destabilizes them; however, the opposite holds for the quartet state of the Li 3 , Li 2 Na, LiNa 2 , and Na 3 clusters (destabilizing 2-B, stabilizing 3-B). Substituting Li with B in the quartet state of Li 3 results in a significant reduction of the 3-B term amounting to 16% (Li 2 B) and 5% (LiB 3 ) of the binding energy. On the contrary, the ground states of the alkaline earth metal clusters (Be 3 , Be 2 Mg, BeMg 3 , and Mg 3 ) are stabilized by the 3-B term, while the 2-B term destabilizes them. Overall, we find that the 3-B terms significantly stabilize the high-spin multiplicity states of the ns 1 configurations and the low-spin states of the ns 2 configurations. Finally, as the size of the metal increases, the contribution of the 3-B term to the binding energy decreases due to the longer metal–metal bond distances.

Alkali metals↗

Osprey Framework v0.2.2

The Alpha Berkeley Framework is a software architecture for building agentic AI systems that coordinate multi-step workflows in scientific and industrial environments. It is based on a plan-first orchestration model, where natural language requests are translated into execution plans with explicit dependencies and optional human approval. The framework includes capability classification, which selects relevant tools on a per-task basis to keep orchestration efficient as the number of available tools grows. It incorporates task extraction methods that compress conversational context and integrate external resources such as databases, APIs, and knowledge bases into structured, machine-readable tasks. Execution is supported by modular services with checkpointing, artifact management, and error handling, allowing workflows to be paused, inspected, and resumed. The system is designed for deployment in production environments, supporting both local and containerized execution as well as integration with HPC clusters. Interfaces include command-line tools, browser-based workflows, and containerized services. The framework has been demonstrated in tutorial examples and deployed at the Advanced Light Source, where it coordinates accelerator control and analysis workflows.

Hellert, Thorsten [Lawrence Berkeley National Labo↗

Direct observation of ultrafast cluster dynamics in supercritical carbon dioxide using X-ray Photon Correlation Spectroscopy

Supercritical fluids exhibit distinct thermodynamic and transport properties, making them of particular interest for a wide range of scientific and engineering applications. These anomalous properties emerge from structural heterogeneities due to the formation of molecular clusters at conditions above the critical point. While the static behavior of these clusters and their effects on the thermodynamic response functions have been recognized, the relation between the ultrafast cluster dynamics and transport properties remains elusive. By measuring the intermediate scattering function in carbon dioxide at conditions near the critical point with X-ray photon correlation spectroscopy, we directly capture the cross-over dynamics between 4 and 13 picoseconds, revealing the transition between ballistic and diffusive motion. Complementary analysis using large-scale molecular dynamics simulations reveals that this behavior arises from collisions between unbound molecules and clusters. This study provides direct evidence of the ultrafast momentum exchange between clusters, which has significant impact on transport properties, solvation processes, and reaction kinetics in supercritical fluids.

carbon capture and storage↗

Unsupervised Segmentation and Clustering Workflow for Efficient Processing of 4D-STEM and 5D-STEM Data

Four-dimensional scanning transmission electron microscopy (4D-STEM) enables mapping of diffraction information with nanometer-scale spatial resolution, offering detailed insight into local structure, orientation, and strain. However, as data dimensionality and sampling density increase, particularly for in situ scanning diffraction experiments (5D-STEM), robust segmentation of structurally consistent behavior across sequential measurements becomes essential for efficient and physically meaningful analysis. Here, we introduce a clustering framework that identifies crystallographically distinct domains from 4D-STEM datasets. By using local diffraction-pattern similarity as a metric, the method extracts closed contours delineating spatially contiguous regions. This approach produces cluster-averaged diffraction patterns that improve signal quality while reducing data volume by orders of magnitude, enabling rapid and accurate orientation, phase, and strain mapping. We demonstrate the applicability of this approach to in situ liquid-cell 4D-STEM data of gold nanoparticle growth. Our method provides a scalable and generalizable route for spatially coherent segmentation, data compression, and quantitative structure–strain mapping across diverse 4D-STEM modalities. The full analysis code and example workflows are publicly available to support reproducibility and reuse.

4D-STEM↗