Search NASA⌕ Search

SEARCH · Search NASA

Results for “clustering analysis (CA)”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

Structuring Nutrient Yields throughout Mississippi/Atchafalaya River Basin Using Machine Learning Approaches

To minimize the eutrophication pressure along the Gulf of Mexico or reduce the size of the hypoxic zone in the Gulf of Mexico, it is important to understand the underlying temporal and spatial variations and correlations in excess nutrient loads, which are strongly associated with the formation of hypoxia. This study’s objective was to reveal and visualize structures in high-dimensional datasets of nutrient yield distributions throughout the Mississippi/Atchafalaya River Basin (MARB). For this purpose, the annual mean nutrient concentrations were collected from thirty-three US Geological Survey (USGS) water stations scattered in the upper and lower MARB from 1996 to 2020. Eight surface water quality indicators were selected to make comparisons among water stations along the MARB over the past two decades. Principal component analysis (PCA) was used to comprehensively evaluate the nutrient yields across thirty-three USGS monitoring stations and identify the major contributing nutrient loads. The results showed that all samples could be analyzed using two main components, which accounted for 81.6% of the total variance. The PCA results showed that yields of orthophosphate (OP), silica (SI), nitrate–nitrites (NO 3 -NO 2 ), and total suspended sediment (TSS) are major contributors to nutrient yields. It also showed that land-planted crops, density of population, domestic and industrial discharges, and precipitation are fundamental causes of excess nutrient loads in MARB. These factors are of great significance for the excess nutrient load management and pollution control of the Mississippi River. It was found that the average nutrient yields were stable within the sub-MARB area, but the large nitrogen yields in the upper MARB and the large phosphorus yields in the lower MARB were of great concern. t-distributed stochastic neighbor embedding (t-SNE) revealed interesting nonlinear and local structures in nutrient yield distributions. Clustering analysis (CA) showed the detailed development of similarities in the nutrient yield distribution. Moreover, PCA, t-SNE, and CA showed consistent clustering results. This study demonstrated that the integration of dimension reduction techniques, PCA, and t-SNE with CA techniques in machine learning are effective tools for the visualization of the structures of the correlations in high-dimensional datasets of nutrient yields and provide a comprehensive understanding of the correlations in the distributions of nutrient loads across the MARB.

54 ENVIRONMENTAL SCIENCES↗

Multi-viewpoint clustering analysis

In this paper, we address the feasibility of partitioning rule-based systems into a number of meaningful units to enhance the comprehensibility, maintainability and reliability of expert systems software. Preliminary results have shown that no single structuring principle or abstraction hierarchy is sufficient to understand complex knowledge bases. We therefore propose the Multi View Point - Clustering Analysis (MVP-CA) methodology to provide multiple views of the same expert system. We present the results of using this approach to partition a deployed knowledge-based system that navigates the Space Shuttle's entry. We also discuss the impact of this approach on verification and validation of knowledge-based systems.

Mehrotra, Mala↗

The methodology of multi-viewpoint clustering analysis

One of the greatest challenges facing the software engineering community is the ability to produce large and complex computer systems, such as ground support systems for unmanned scientific missions, that are reliable and cost effective. In order to build and maintain these systems, it is important that the knowledge in the system be suitably abstracted, structured, and otherwise clustered in a manner which facilitates its understanding, manipulation, testing, and utilization. Development of complex mission-critical systems will require the ability to abstract overall concepts in the system at various levels of detail and to consider the system from different points of view. Multi-ViewPoint - Clustering Analysis MVP-CA methodology has been developed to provide multiple views of large, complicated systems. MVP-CA provides an ability to discover significant structures by providing an automated mechanism to structure both hierarchically (from detail to abstract) and orthogonally (from different perspectives). We propose to integrate MVP/CA into an overall software engineering life cycle to support the development and evolution of complex mission critical systems.

Mehrotra, Mala↗

Streamlining heterologous expression of top carbonic anhydrases in Escherichia coli : bioinformatic and experimental approaches

Carbonic anhydrase (CA) enzymes facilitate the reversible hydration of CO 2 to bicarbonate ions and protons. Identifying efficient and robust CAs and expressing them in model host cells, such as Escherichia coli, enables more efficient engineering of these enzymes for industrial CO 2 capture. However, expression of CAs in E. coli is challenging due to the possible formation of insoluble protein aggregates, or inclusion bodies. This makes the production of soluble and active CA protein a prerequisite for downstream applications. In this study, we streamlined the process of CA expression by selecting seven top CA candidates and used two bioinformatic tools to predict their solubility for expression in E. coli. The prediction results place these enzymes in two categories: low and high solubility. Our expression of high solubility score CAs (namely CA5-SspCA, CA6-SazCAtrunc, CA7-PabCA and CA8-PhoCA) led to significantly higher protein yields (5 to 75 mg purified protein per liter) in flask cultures, indicating a strong correlation between the solubility prediction score and protein expression yields. Furthermore, phylogenetic tree analysis demonstrated CA class-specific clustering patterns for protein solubility and production yields. Unexpectedly, we also found that the unique N-terminal, 11-amino acid segment found after the signal sequence (not present in its homologs), was essential for CA6-SazCA activity. Overall, this work demonstrated that protein solubility prediction, phylogenetic tree analysis, and experimental validation are potent tools for identifying top CA candidates and then producing soluble, active forms of these enzymes in E. coli. The comprehensive approaches we report here should be extendable to the expression of other heterogeneous proteins in E. coli.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Determination of Cluster Distances from Chandra Imaging Spectroscopy and Sunyaev-Zeldovich Effect Measurements: Analysis Methods and Initial Results - I

X-ray and Sunyaev-Zeldovich Effect data ca,n be combined to determine the distance to galaxy clusters. High-resolution X-ray data are now available from the Chandra Observatory, which provides both spatial and spectral information, and interferometric radio measurements of the Sunyam-Zeldovich Effect are available from the BIMA and 0VR.O arrays. We introduce a Monte Carlo Markov chain procedure for the joint analysis of X-ray and Sunyaev-Zeldovich Effect data. The advantages of this method are the high computational efficiency and the ability to measure the full probability distribution of all parameters of interest, such as the spatial and spectral properties of the cluster gas and the cluster distance. We apply this technique to the Chandra X-ray data and the OVRO radio data for the galaxy cluster Abell 611. Comparisons with traditional likelihood-ratio methods reveal the robustness of the method. This method will be used in a follow-up paper to determine the distance of a large sample of galaxy clusters for which high-resolution Chandra X-ray and BIMA/OVRO radio data are available.

Bonamente, Massimiliano↗

Scalable edge clustering of dynamic graphs via weighted line graphs

Timestamped relational datasets consisting of records (or connections) between pairs of entities are ubiquitous in network science. For applications like peer-to-peer communication, email, various social network interactions, and computer network security, it is useful to organize these records into groups based on how and when they are occurring. Weighted line graphs offer a natural way to model how records are related in such datasets but for large real-world graph topologies, building and utilizing the line graph is prohibitively expensive. Here, we present the framework to cluster the edges of a dynamic graph via the associated line graph that contains two major contributions. The first is a method to work with the line graph implicitly and the second is a distributed scale implementation of an agglomerative hierarchical graph clustering algorithm. We outline a novel hierarchical dynamic graph edge clustering approach that efficiently breaks massive relational datasets into small sets of edges containing events at various timescales. This is in stark contrast to traditional graph clustering algorithms that prioritize highly connected (clique-like) community structures. Our approach relies on constructing a sufficient subgraph of a weighted line graph and applying a hierarchical agglomerative clustering. This approach is related to scalable techniques from spatial clustering, nonlinear-dimension reduction, topological data analysis, and draws particular inspiration from HDBSCAN. As an edge clustering, this method yields an overlapping node clustering. Our algorithm is parallelizable and we demonstrate efficient clustering of a billion-scale, real-world dynamic graph into small edge sets that correlate in topology and time. The entire clustering process for a graph with tens of billions of edges takes just a few minutes of run time on 256 nodes of a distributed compute environment. We argue how the output of the edge clustering is useful for a multitude of data visualization and powerful machine learning tasks, both involving the original massive dynamic graph data and metadata associated with the nodes and edges. Finally, we describe how this approach can be extended to dynamic hypergraphs and dynamic graphs/hypergraphs with unstructured data living on vertices and edges.

Data Analysis↗

Mass and spatial distribution of carbonaceous component in Comet Halley

Cometary grains containing large amounts of carbon and/or organic matter were discovered by in situ measurements of cometary dust composition during VEGA and GIOTTO fly-by missions. In accordance with the classification for the data of PUMA-1 and PUMA-2 mass-spectrometers on board the VEGA spacecraft, particles with a ratio of C to any rock-forming element (Mg, Si, Fe, Ca etc.) greater than 10, were categorized as CHON. There are 464 such particles in PUMA-1 data and 51 in PUMA-2 data. Application of cluster analysis to these grains revealed several distinct compositional classes, namely: (H,C,N,O), (H,C,N), (H,C), (H,C,O), (C,N), (C,O), (C,N,O), and (C). Similar classes were identified among particles analyzed by PIA. Also, about a third of all particles fell into groups (H) and (O) characterized by abundances of these elements beyond chemically reasonable limits.

Fomenkova, M.↗

Unsupervised atomic data mining via multi-kernel graph autoencoders for machine learning force fields

Constructing a chemically diverse dataset while avoiding sampling bias is critical to training efficient and generalizable force fields. However, in computational chemistry and materials science, many common dataset generation techniques are prone to oversampling regions of the potential energy surface. Furthermore, these regions can be difficult to identify and isolate from each other or may not align well with human intuition, making it challenging to systematically remove bias in the dataset. While traditional clustering and pruning (down-sampling) approaches can be useful for this, they can often lead to information loss or a failure to properly identify distinct regions of the potential energy surface due to difficulties associated with the high dimensionality of atomic descriptors. In this work, we introduce the Multi-kernel Edge Attention-based Graph Autoencoder (MEAGraph) model, an unsupervised approach for analyzing atomic datasets. MEAGraph combines multiple linear kernel transformations with attention-based message passing to capture geometric sensitivity and enable effective dataset pruning without relying on labels or extensive training. Demonstrated applications on niobium, tantalum, and iron datasets show that MEAGraph efficiently groups similar atomic environments, allowing for the use of basic pruning techniques for removing sampling bias. This approach provides an effective method for representation learning and clustering that can be used for data analysis, outlier detection, and dataset optimization.

Materials science↗

Critical Simulation Pipeline for COG Suites [Poster]

The CRItical Simulation Pipeline (CRISP) is a Python package for automating validation of reactor criticality benchmarks. CRISP supplies COG—a multi-particle radiation transport code maintained by the Nuclear Criticality Safety Division—with a pipeline to calculate k eff performance for 400+ benchmark experiments with 3,400+ configurations from the International Criticality Safety Benchmark Evaluation Project (ICSBEP). The pipeline includes four stages: materials configuration, input card templating, cluster submission, and results analysis. CRISP includes a command-line interface to facilitate user interaction.

73 NUCLEAR PHYSICS AND RADIATION PHYSICS↗

Geometric Interpretation of the Cluster Location Problem Part I: Theory

We present a new framing of the seismic location problem using principles drawn from differential geometry. Our interpretation relies upon the common assumption that travel times observed across a network are continuous, differentiable functions of source location. In consequence, travel‐time functions constitute a differentiable map between the source region and a Riemannian manifold. The manifold is said to be the image of the source region embedded in a generally high‐dimension travel‐time vector space. A cluster of events in the source region has an image of discrete points on the manifold, that, except in the simplest cases, cannot be viewed directly. However, it is possible to project the image of a cluster into a tangent space of the manifold for direct visualization. The projection operator can be computed directly from the data without a velocity model, but produces a distorted rendering of the cluster geometry. With a model we can predict the distortions and correct them to estimate cluster geometry. We develop these points with the simplest possible example, one for which direct visualization of the manifold is possible, using the example as an introduction to the relevant concepts from differential geometry in a familiar setting. The tangent space, a local linearization of the manifold, plays a key role. We develop a metric to estimate the limits of linearization, that is, to determine when the curvature of the manifold invalidates the linear assumption. We also examine the interplay of model error, inadequate network geometry, and pick error. We then generalize our results from the simple case to the general case of 3D source regions observed by general networks. Although we do suggest a new “project and correct” method for location, we do not develop it into a practical algorithm. In conclusion, our intention rather is to highlight new analytical methods grounded in differential geometry.

East Pacific Ocean Islands↗

Where was the Iron Synthesized in Cassiopeia A?

We investigate the properties of Fe-rich knots on the east limb of the Cassiopeia A supernova remnant observed with Chandra/AXAF CCD Imaging Spectrometer (ACIS). Using analysis methods developed in a companion paper, we constrain the ejecta density profile and the Lagrangian mass coordinates of the knots from their fitted ionization age and electron temperature. Fe-rich knots which also have strong emission from Si, S, Ar, and Ca are clustered around mass coordinates q approx. equal to 0.35 - 0.4 in the shocked ejecta of 2 solar masses; this places them 0.7 - 0.8 solar masses out from the center (or 2 - 2.1 solar masses, allowing for the mass of a compact object). We also find an Fe clump that is evidently devoid of line emission from lower mass elements, as would be expected for a region that had undergone alpha-rich freeze out. This clump has a similar mass coordinate to the other Fe knots.

Hwang, Una↗

Intermediate Element Abundances In Galaxy Clusters

We present the average abundances of the intermediate elements obtained by performing a stacked analysis of all the galaxy clusters in the archive of the X-ray telescope AKA. We determine the abundances of Fe, Si, S, and Ni as a function of cluster temperature (mass) from 1 - 10 keV, and place strong upper limits on the abundances of Ca and Ar. In general, Si and Ni are overabundant with respect to Fe, while Ar and Ca are very underabundant. The discrepancy between the abundances of Si, S , Ar, and Ca indicate that the alpha-elements do not behave homogeneously as a single group. We show that the abundances of the most well-determined elements Fe, Si, and S in conjunction with recent theoretical supernovae yields do not give a consistent solution for the fraction of material produced by Type Ia and Type II supernovae at any temperature or mass. The general trend is for higher temperature clusters to have more of their metals produced in Type II supernovae than in Type Ias. The inconsistency of our results with abundances in the Milky Way indicate that spiral galaxies are not the dominant metal contributors to the intracluster medium (ICM). The pattern of elemental abundances requires an additional source of metals beyond standard SNIa and SNII enrichment. The properties of this new source are well matched to those of Type II supernovae with very massive, metal-poor progenitor stars. These results are consistent with a significant fraction of the ICM metals produced by an early generation of population III stars.

White, Nicholas E.↗

Accelerating the Structure Exploration of Diverse Bi–Pt Nanoclusters via Physics‐Informed Machine Learning Potential and Particle Swarm Optimization

Bimetallic Bi–Pt nanoclusters exhibit diverse structural motifs, including core-shell, Janus, and mixed alloy configurations, due to the unique bonding characteristics between Bi and Pt atoms. Using density functional theory refinements from ChIMES physically machine-learned potential and CALYPSO particle swarm optimization global searches, 34 Bi20-Pt20 nanoclusters are systematically classified. The results reveal that Bi atoms predominantly occupy surface sites, driven by charge transfer effects. Cohesive energy trends alone prove insufficient for structure differentiation, necessitating a data-driven approach employing principal component analysis and K-means clustering. Furthermore, vibrational, electronic, and infrared spectral analyses provide additional insights into structure-property relationships. The findings offer an original framework for the automated classification and analysis of bimetallic nanoclusters, enhancing the understanding of their stability and functional properties.

bimetallic nanoparticles↗

Orthopyroxenes as recorders of diogenite petrogenesis: Major and minor element systematics

As a part of our research to better understand magmatic processes in the Eucrite Parent Body, we have initiated an ambitious program of study of major, minor and trace elements in orthopyroxene from diogenites. This paper reports preliminary results for major and minor elements in orthopyroxenes for a suite of 13 diogenites: Aioun El Atrouss, ALH 84001, ALH A 77256, EET 87530, Ellemeet, Garland, Ibbenburen, Johnstown, Manegoan, peckelsheim, Roda, Shalka, and Tatahouine. A companion paper by Shearer et al. reports new trace element data for ALH 84001, ALH A 77256, Ibbenburen, and Tatahouine. We have presently collected over 800 high quality pyroxene microprobe analyses for Si, Al, Ca, Na, Mn, Fe, Mg, Cr, and Ti. The chemical systematics observed for these orthopyroxenes reflect original magmatic mineral/melt partitioning plus later trapped liquid/mineral equilibration, subsolids, exsolution, and mineral/mineral metamorphic reactions. We have therefore avoided, at this point, any attempt to use statistical analysis to group (e.g. factor or cluster analysis) these orthopyroxenes chemically.

Papike, J. J.↗

CASM Monte Carlo: Calculations of the thermodynamic and kinetic properties of complex multicomponent crystals

Monte Carlo techniques play a central role in statistical mechanics approaches that connect macroscopic thermodynamic and kinetic properties to the electronic structure of a material. This paper describes the implementation of Monte Carlo techniques for the study of multicomponent crystalline materials within the Clusters Approach to Statistical Mechanics (CASM) software suite, and demonstrates their use in model systems to calculate free energies and kinetic coefficients, study phase transitions, and construct phase diagrams from first principles. Many crystal structures are complex, with multiple sublattices occupied by differing sets of chemical species, along with the presence of vacancies or interstitial species. This imposes constraints on concentration variables, the form of thermodynamic potentials, and the values of kinetic transport coefficients. The framework used by CASM to formulate thermodynamic potentials and kinetic transport coefficients accounting for arbitrarily complex crystal structures is presented and demonstrated with examples of increasing complexity. Additionally, an overview of the capabilities of the CASM software specific to Monte Carlo methods is given, and a new CASM software package is introduced, casm-flow, which helps automate the setup, submission, management, and analysis of Monte Carlo simulations.

Cluster expansion↗

State-level suicide mortality insights: a comparative study of VHA veterans and the whole US population

Background: Suicide is a leading cause of death in the US Comparative State-level spatial analysis between Veterans Health Administration (VHA veterans) and the whole US population can reveal differences in conditions for targeted interventions and intricate geographical patterns. Methods: The study population contains 2018 and 2019 suicide deaths of VHA veterans and the whole US population. They were used to calculate state-level rates. States were classified by whether their VHA veteran and whole US population rates were above or below respective mean rates. Local Moran’s I was leveraged to examine spatial autocorrelation. Results: State-level suicide mortality rates and disparities among states were generally higher for VHA veterans (2018: 37.3 ± 7.2; 2019: 46.8 ± 8.3) than for the whole US population (2018: 16.6 ± 4.3; 2019: 16.4 ± 4.4). For both populations, there were statistically significant clusters with high suicide rates. Over one-fourth of states demonstrated inverse relationships, with rates above mean for one group but below for other. VHA veterans are at higher risk with over one-third of states had greater than average veteran suicide risk ratio. Conclusions: VHA veterans are at higher risk than the whole population across all states. Mortality disparities among states and clusters of states with high and low rates suggest targeted interventions and cooperative health strategies may help address these differences.

60 APPLIED LIFE SCIENCES↗

Memoirs of Mass Accretion: Probing the Edges of Intracluster Light in Simulated Galaxy Clusters

The diffuse starlight extending throughout massive galaxy clusters, known as intracluster light (ICL), has the potential to be read as a memoir of mass accretion: informative, individual, and yet imperfect. Here, we combine dark-matter-only zoom-in simulations from the Symphony suite with the Nimbus “star-tagging” model of the stellar halo to assess how much information about the mass assembly of an individual galaxy cluster can be gleaned from idealized measurements of ICL outskirts. We show that the edges of a cluster’s stellar profile—the primary (R sp⋆,1 ) and secondary (R sp⋆,2 ) stellar “splashback” radii—are sensitive to both continuous mass accretion histories (MAHs) and discrete merger events, making them potentially powerful probes of a cluster’s past. We find that R sp⋆,1 strongly correlates with the cluster’s mass ∼1 dynamical time ago, while R sp⋆,2 traces more recent MAH to a slightly lesser degree. In combination, these features can further distinguish between clusters that have and have not undergone a major merger within the past dynamical time. We use both to predict realistic cluster MAHs with the MultiCAM framework. These outer ICL features are significantly more sensitive to mass accretion and merger histories than the stellar mass gap and halo concentration, and perform comparably to the commonly used X-ray-based tracer of relaxedness, x off . While our analysis is idealized, the relevant ICL features are potentially detectable in next-generation deep imaging of nearby clusters. This work highlights the promise of ICL measurements and lays the groundwork for more detailed forecasts of their power.

79 ASTRONOMY AND ASTROPHYSICS↗

Nucleation and growth of polar clusters with in-phase tilts into a long-range ferroelectric matrix in a sodium niobate based complex relaxor

In this study, we have investigated the temperature dependence of atomic ordering at multiple length scales in a lead-free sodium niobate-based relaxor, i.e., 0.75 NaNbO 3 -0.25 Ba 0.9⁢ Ca 0.1⁢ TiO 3 (NN-25BCT) via synchrotron x-ray diffraction, Raman spectroscopy, and pair distribution function analysis. High-resolution synchrotron x-ray powder diffraction (SXRD) measurements reveal a ferroelectric phase transition in the relaxor ferroelectric NN-25BCT below the Vogel-Fulcher freezing temperature (𝑇 VF ≈ 270 K). In addition, SXRD analysis demonstrates the competition between in-phase octahedral tilting and ferroelectric order at the long-range scale using mode crystallography. On the other hand, Raman spectroscopic analysis provides evidence of polar ordering for 𝑇 > 𝑇 VF (with tetragonal symmetry) persisting up to the Burns temperature (𝑇 B ). Furthermore, pair distribution function (PDF) analysis reveals the presence of a polar antiferrodistortive tetragonal phase with 𝑃⁢4⁢𝑏𝑚 space group at short ranges throughout the studied temperatures (i.e., 110 K ≤ 𝑇 ≤500 K), irrespective of nonpolar long-range ordering above 𝑇 VF . Therefore, our measurements provide direct evidence for the presence of polar ordering at short ranges and their gradual transformation into long-range polar ordering using an integrated multiscale structural analysis. In conclusion, as a result of a transition from relaxor to a ferroelectric phase in the vicinity of room temperature, NN-25BCT can be exploited for applications in pyroelectric detectors, electrocaloric devices, and multilayered ceramic capacitors.

36 MATERIALS SCIENCE↗