Search NASA⌕ Search

SEARCH · Search NASA

Results for “community clustering”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

VC3: Virtual Clusters for Community Computation

A traditional HPC computing facility provides a large amount of computing power but has a fixed environment designed to satisfy local needs. This makes it very challenging for users to deploy complex applications that span multiple sites and require specific application software, scheduling middleware, or sharing policies. This project addressed many of these challenges by making it possible for researchers to easily aggregate and share resources, install custom software environments, and deploy clustering frameworks across multiple HPC facilities through the concept of “virtual clusters”. We designed and implemented a prototype virtual cluster facility that enabled unprivileged users to create dynamic aggregations of computing power across multiple sites, deployed with custom middleware and complex software dependencies. This service is hosted at the University of Chicago and available through the site virtualclusters.org.

97 MATHEMATICS AND COMPUTING↗

VC3: Virtual Clusters for Community Computation (Final Technical Report)

A traditional HPC computing facility provides a large amount of computing power but has a fixed environment designed to satisfy local needs. This makes it very challenging for users to deploy complex applications that span multiple sites and require specific application software, scheduling middleware, or sharing policies. This project addressed many of these challenges by making it possible for researchers to easily aggregate and share resources, install custom software environments, and deploy clustering frameworks across multiple HPC facilities through the concept of “virtual clusters”. We designed and implemented a prototype virtual cluster facility that enabled unprivileged users to create dynamic aggregations of computing power across multiple sites, deployed with custom middleware and complex software dependencies.

97 MATHEMATICS AND COMPUTING↗

Disruption-Robust Community Detection Using Consensus Clustering in Complex Networks

Topological (graph-theoretic) analysis of critical infrastructure networks provides insight on several aspects of resilience. Graph clustering or community detection, which identifies densely connected components in a graph, has been employed for analysis. In this paper, we propose employing consensus clustering, which is a technique to determine consensus from a collection of different clusters on an input, such that the resulting clustering is robust to disruptions, where a disruption is represented as loss of one or more vertices or edges in the graph. Using two critical infrastructure networks as case studies, we empirically demonstrate the need to compute consensus clustering in order to address the drastic changes in the topology due to disruptions in the network.

Hussain, Md Taufique↗

Single- and two-particle finite size effects in interacting lattice systems

Simulations of extended quantum systems are typically performed by extrapolating results of a sequence of finite-system-size simulations to the thermodynamic limit. In the quantum Monte Carlo community, twist-averaging was pioneered as an efficient strategy to eliminate one-body finite size effects. In the dynamical mean field community, cluster generalizations of the dynamical mean field theory were formulated to study systems with nonlocal correlations. In this work, we put the twist-averaging and the dynamical cluster approximation variant of the dynamical mean field theory onto equal footing, discuss commonalities and differences, and compare results from both techniques to the standard periodic boundary technique. At the example of Hubbard-type models with local, short-range and Yukawa-like longer range interactions we show that all methods converge to the same limit, but that the convergence speed differs in practice. We show that embedding theories are an effective tool for managing both one-body and two-body finite size effects, in particular if interactions are averaged over twist angles.

36 MATERIALS SCIENCE↗

Missing microbial eukaryotes and misleading meta-omic conclusions

Meta-omics is commonly used for large-scale analyses of microbial eukaryotes, including species or taxonomic group distribution mapping, gene catalog construction, and inference on the functional roles and activities of microbial eukaryotes in situ. Here, we explore the potential pitfalls of common approaches to taxonomic annotation of protistan meta-omic datasets. We re-analyze three environmental datasets at three levels of taxonomic hierarchy in order to illustrate the crucial importance of database completeness and curation in enabling accurate environmental interpretation. We show that taxonomic membership of sequence clusters estimates community composition more accurately than returning exact sequence labels, and overlap between clusters can address database shortcomings. Clustering approaches can be applied to diverse environments while continuing to exploit the wealth of annotation data collated in databases, and selecting and evaluating these databases is a critical part of correctly annotating protistan taxonomy in environmental datasets. We argue that ongoing curation of genetic resources is crucial in accurately annotating protists in in situ meta-omic datasets. Moreover, we propose that precise taxonomic annotation of meta-omic data is a clustering problem rather than a feasible alignment problem.

59 BASIC BIOLOGICAL SCIENCES↗

Topological data analysis of task-based fMRI data from experiments on schizophrenia

We use methods from computational algebraic topology to study functional brain networks, in which nodes represent brain regions and weighted edges represent similarity of fMRI time series from each region. With these tools, which allow one to characterize topological invariants such as loops in high-dimensional data, we are able to gain understanding into low-dimensional structures in networks in a way that complements traditional approaches based on pairwise interactions. In the present paper, we analyze networks constructed from task-based fMRI data from schizophrenia patients, healthy controls, and healthy siblings of schizophrenia patients using persistent homology, which allows us to explore the persistence of topological structures such as loops at different scales in the networks. We use persistence landscapes, persistence images, and Betti curves to create output summaries from our persistent-homology calculations, and we study the persistence landscapes and images using k-means clustering and community detection. Based on our analysis of persistence landscapes, we find that the members of the sibling cohort have topological features (specifically, their 1-dimensional loops) that are distinct from the other two cohorts. From the persistence images, we are able to distinguish all three subject groups and to determine the brain regions in the loops (with four or more edges) that allow us to make these distinctions.

60 APPLIED LIFE SCIENCES↗

Graph theory and nighttime imagery based microgrid design

Reducing the duration and frequency of blackouts in remote communities poses an engineering challenge for grid operators. Outage effects can also be mitigated locally through microgrids. This paper develops a systematic procedure to account for these challenges by creating microgrids prioritizing high value assets within vulnerable communities. Nighttime satellite imagery is used to identify vulnerable communities. Using an asset classification and rating system, multi-asset clusters within these communities are prioritized. Infrastructure data, geographic information systems, satellite imagery, and spectral clustering are used to form and rank microgrid candidates. A microgrid sizing algorithm is included to guide through the microgrid design process. Finally, an application of the methodology is presented using real event, location, and asset data.

24 POWER TRANSMISSION AND DISTRIBUTION↗

Choice of 16S ribosomal RNA primers affects the microbiome analysis in chicken ceca

We evaluated the effect of applying different sets of 16S rRNA primers on bacterial composition, diversity, and predicted function in chicken ceca. Cecal contents from Ross 708 birds at 1, 3, and 5 weeks of age were collected for DNA isolation. Eight different primer pairs targeting different variable regions of the 16S rRNA gene were employed. DNA sequences were analyzed using open-source platform QIIME2 and the Greengenes database. PICRUSt2 was used to determine the predicted function of bacterial communities. Changes in bacterial relative abundance due to 16S primers were determined by GLMs. The average PCR amplicon size ranged from 315 bp (V3) to 769 bp (V4–V6). Alpha- and beta-diversity, taxonomic composition, and predicted functions were significantly affected by the primer choice. Beta diversity analysis based on Unweighted UniFrac distance matrix showed separation of microbiota with four different clusters of bacterial communities. Based on the alpha- and beta-diversity and taxonomic composition, variable regions V1–V3(1) and (2), and V3–V4 and V3–V5 were in most consensus. Our data strongly suggest that selection of particular sets of the 16S rRNA primers can impact microbiota analysis and interpretation of results in chicken as was shown previously for humans and other animal species.

59 BASIC BIOLOGICAL SCIENCES↗

Search for brown dwarfs in IC 1396 with Subaru HSC: interpreting the impact of environmental factors on substellar population

ABSTRACT Young stellar clusters are predominantly the hub of star formation and hence, ideal to perform comprehensive studies over the least explored substellar regime. Various unanswered questions like the mass distribution in brown dwarf regime and the effect of diverse cluster environment on brown dwarf formation efficiency still plague the scientific community. The nearby young cluster, IC 1396 with its feedback-driven environment, is ideal to conduct such study. In this paper, we adopt a multiwavelength approach, using deep Subaru HSC along with other data sets and machine learning techniques to identify the cluster members complete down to ∼ 0.03 M⊙ in the central 22 arcmin area of IC 1396. We identify 458 cluster members including 62 brown dwarfs which are used to determine mass distribution in the region. We obtain a star-to-brown dwarf ratio of ∼ 6 for a stellar mass range 0.03–1 M⊙ in the studied cluster. The brown dwarf fraction is observed to increase across the cluster as radial distance from the central OB-stars increases. This study also compiles 15 young stellar clusters to check the variation of star-to-brown dwarf ratio relative to stellar density and ultraviolet (UV) flux ranging within 4–2500 stars pc−2 and 0.7–7.3 G0, respectively. The brown dwarf fraction is observed to increase with stellar density but the results about the influence of incident UV flux are inconclusive within this range. This is the deepest study of IC 1396 as of yet and it will pave the way to understand various aspects of brown dwarfs using spectroscopic observations in future.

Gupta, Saumya (ORCID:0000000161843958)↗

Clusters of Regional Precipitation Seasonality Change in the Community Earth System Model, Version 2

Abstract The likely changes to precipitation seasonality with warming are both impactful and not well understood. This work aims to describe areas that experience similar changes to seasonal precipitation irrespective of the original underlying precipitation seasonality. We train a self-organizing map on the difference between the seasonal cycle of precipitation in the past and in a high-warming future climate as represented by the Community Earth System Model, version 2, to create regions with similar changes in precipitation seasonality. This method is applied separately over land and ocean surfaces because of the differing processes leading to precipitation over each. This method indicates that future changes in seasonal precipitation are most varied in the tropics because of a southward shift in the intertropical convergence zone. The seasonal shifts found over midlatitude oceans indicate a poleward shift in atmospheric river activity. We find a correspondence between certain land-based precipitation changes and Köppen climate classification. The seasonality of large-scale and convective precipitation is examined for each region. The relationship between the seasonal changes to precipitation and associated atmospheric processes is discussed. These processes include atmospheric rivers, the intertropical convergence zone, tropical cyclones, and monsoons.

54 ENVIRONMENTAL SCIENCES↗

MIBiG 3.0: a community-driven effort to annotate experimentally validated biosynthetic gene clusters

Abstract With an ever-increasing amount of (meta)genomic data being deposited in sequence databases, (meta)genome mining for natural product biosynthetic pathways occupies a critical role in the discovery of novel pharmaceutical drugs, crop protection agents and biomaterials. The genes that encode these pathways are often organised into biosynthetic gene clusters (BGCs). In 2015, we defined the Minimum Information about a Biosynthetic Gene cluster (MIBiG): a standardised data format that describes the minimally required information to uniquely characterise a BGC. We simultaneously constructed an accompanying online database of BGCs, which has since been widely used by the community as a reference dataset for BGCs and was expanded to 2021 entries in 2019 (MIBiG 2.0). Here, we describe MIBiG 3.0, a database update comprising large-scale validation and re-annotation of existing entries and 661 new entries. Particular attention was paid to the annotation of compound structures and biological activities, as well as protein domain selectivities. Together, these new features keep the database up-to-date, and will provide new opportunities for the scientific community to use its freely available data, e.g. for the training of new machine learning models to predict sequence-structure-function relationships for diverse natural products. MIBiG 3.0 is accessible online at https://mibig.secondarymetabolites.org/.

59 BASIC BIOLOGICAL SCIENCES↗

Subsurface microbial community structure shifts along the geological features of the Central American Volcanic Arc

Subduction of the Cocos and Nazca oceanic plates beneath the Caribbean plate drives the upward movement of deep fluids enriched in carbon, nitrogen, sulfur, and iron along the Central American Volcanic Arc (CAVA). These compounds fuel diverse subsurface microbial communities that in turn alter the distribution, redox state, and isotopic composition of these compounds. Microbial community structure and functions vary according to deep fluid delivery across the arc, but less is known about how microbial communities differ along the axis of a convergent margin as geological features (e.g., extent of volcanism and subduction geometry) shift. Here, we investigate changes in bacterial 16S rRNA gene amplicons and geochemical analysis of deeply-sourced seeps along the southern CAVA, where subduction of the Cocos Ridge alters the geological setting. We find shifts in community composition along the convergent margin, with communities in similar geological settings clustering together independently of the proximity of sample sites. Microbial community composition correlates with geological variables such as host rock type, maturity of hydrothermal fluid and slab depth along different segments of the CAVA. This reveals tight coupling between deep Earth processes and subsurface microbial activity, controlling community distribution, structure and composition along a convergent margin.

Science & Technology - Other Topics↗

Metagenomic insights into the taxonomy, function, and dysbiosis of prokaryotic communities in octocorals

Background. In octocorals (Cnidaria Octocorallia), the functional relationship between host health and its symbiotic consortium has yet to be determined. Here, we employed comparative metagenomics to uncover the distinct functional and phylogenetic features of the microbiomes of healthy Eunicella gazella, Eunicella verrucosa, and Leptogorgia sarmentosa tissues, in contrast with the microbiomes found in seawater and sediments. We further explored how the octocoral microbiome shifts to a pathobiome state in E. gazella. Results. Multivariate analyses based on 16S rRNA genes, Clusters of Orthologous Groups of proteins (COGs), Protein families (Pfams), and secondary metabolite-biosynthetic gene clusters annotated from 20 Illumina-sequenced metagenomes each revealed separate clustering of the prokaryotic communities of healthy tissue samples of the three octocoral species from those of necrotic E. gazella tissue and surrounding environments. While the healthy octocoral microbiome was distinguished by so-far uncultivated Endozoicomonadaceae, Oceanospirillales, and Alteromonadales phylotypes in all host species, a pronounced increase of Flavobacteriaceae and Alphaproteobacteria, originating from seawater, was observed in necrotic E. gazella tissue. Increased abundances of eukaryotic-like proteins, exonucleases, restriction endonucleases, CRISPR/Cas proteins, and genes encoding for heat-shock proteins, inorganic ion transport, and iron storage distinguished the prokaryotic communities of healthy octocoral tissue regardless of the host species. An increase of arginase and nitric oxide reductase genes, observed in necrotic E. gazella tissues, suggests the existence of a mechanism for suppression of nitrite oxide production by which octocoral pathogens may overcome the host’s immune system. Conclusions. This is the first study to employ primer-less, shotgun metagenome sequencing to unveil the taxonomic, functional, and secondary metabolism features of prokaryotic communities in octocorals. Our analyses reveal that the octocoral microbiome is distinct from those of the environmental surroundings, is host genus (but not species) specific, and undergoes large, complex structural changes in the transition to the dysbiotic state. Host-symbiont recognition, abiotic-stress response, micronutrient acquisition, and an antiviral defense arsenal comprising multiple restriction endonucleases, CRISPR/Cas systems, and phage lysogenization regulators are signatures of prokaryotic communities in octocorals. We argue that these features collectively contribute to the stabilization of symbiosis in the octocoral holobiont and constitute beneficial traits that can guide future studies on coral reef conservation and microbiome therapy.

59 BASIC BIOLOGICAL SCIENCES↗

Enhancing Cluster Identification in Atom Probe Tomography Data Using Transfer Learning

Atom probe tomography (APT) has enabled the direct visualization of solute clusters, providing valuable insights into material structures. This clustering is crucial for understanding the nanoscale composition and behavior of materials, which can significantly influence their mechanical and physical properties. However, the widely used clustering methods in the APT community face challenges such as subjective parametric selection and limited applicability, particularly in dealing with overlapping clusters, nested clusters, and artifacts across different scales, such as precipitates and dislocations. To address these challenges, we present a framework based on density-based cluster analysis that aims to be less dependent on user input, reproducible, and robust.

Density-based clustering↗

Fungal diversity and function in metagenomes sequenced from extreme environments

Fungi are increasingly recognized as key players in various extreme environments. Here we present an analysis of publicly-sourced metagenomes from global extreme environments, focusing on fungal taxonomy and function. The majority of 855 selected metagenomes contained scaffolds assigned to fungi. Relative abundance of fungi was as high as 10% of protein-coding genes with taxonomic annotation, with up to 289 fungal genera per sample. Despite taxonomic clustering by environment, fungal communities were more dissimilar than archaeal and bacterial communities, both for within- and between-environment comparisons. Relatively abundant fungal classes in extreme environments included Dothideomycetes, Eurotiomycetes, Leotiomycetes, Pezizomycetes, Saccharomycetes, and Sordariomycetes. Broad generalists and prolific aerial spore formers were the most relatively abundant fungal genera detected in most of the extreme environments, bringing up the question of whether they are actively growing in those environments or just surviving as spores. More specialized fungi were common in some environments, such as zoosporic taxa in cryosphere water and hot springs. Relative abundances of genes involved in adaptation to general, thermal, oxidative, and osmotic stress were greatest in soda lake, acid mine drainage, and cryosphere water samples.

60 APPLIED LIFE SCIENCES↗

Discovery of correlated electron molecular orbital materials using graph representations

Correlated electron molecular orbital (CEMO) materials host emergent electronic states built from molecular orbitals localized over clusters of transition metal ions yet have historically been discovered sporadically and generally been treated as isolated case studies. Here we establish CEMO materials as a systematically discoverable class and introduce a graph-based framework to identify, classify, and organize transition-metal cluster motifs in inorganic solids. Starting from crystal structures in the Materials Project, we construct transition metal connectivity graphs, extract cluster motifs using a bond-cutting algorithm, and determine cluster point groups, effective cluster sublattice dimensionality, and translational symmetry. Applying this approach in a high-throughput screen of 34,548 compounds yields 5,306 cluster-containing materials, including 2,627 stable or metastable compounds with isolated clusters and 984 materials featuring mixed-metal clusters. The resulting dataset reveals symmetry and element dependent trends in cluster formation. By integrating cluster classification with flat band lattice topology and battery-relevant information, we provide further relevant information to multiple scientific communities. The accompanying open dataset, Cluster Finder software, and interactive web platform enable systematic exploration of cluster driven electronic phenomena and establish a general pathway for discovering correlated quantum materials and functional materials with cluster-based or extended metal-metal bonding in inorganic solids.

Akhond, Md. Rajbanul [Department of Chemistry, 800↗

Exploring the Landscape of Distributed Graph Clustering on Leadership Supercomputers

The rapid growth of large-scale datasets in fields like biology and social networks has driven the need for advanced graph analytics techniques. Community detection, a fundamental task in graph analytics, identifies closely connected groups of nodes within a network, providing valuable insights across various disciplines. This study focuses on two classic community detection methods, the Louvain algorithm and Markov Clustering (MCL), and evaluates the performance of two prominent distributed community detection algorithms: HiPDPL-GPU, our prior implementation, and HipMCL. We conduct experiments on GPU-accelerated heterogeneous HPC systems, Summit and Frontier, to assess their performance under varying conditions. Our objective is to identify the strengths and weaknesses of these algorithms in terms of scalability, and quality of solutions. We evaluate these algorithms on a diverse set of 70+ networks spanning 13 domains, with sizes ranging up to 4.2 billion edges. Our results demonstrate that HiPDPL-GPU consistently outperforms HipMCL, especially for large-scale networks. HiPDPL-GPU achieves significantly faster runtimes (47x to 1439x), higher modularity scores, and improved scalability. These findings highlight HiPDPL-GPU as a promising solution for efficient and effective large-scale graph analytics in diverse application domains, and provide insights into the feasibility of using MCL-based approaches for certain application domains.

Community detection, graph algorithms↗