Search NASA⌕ Search

SEARCH · Search NASA

Results for “Hierarchical clustering”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

Statistical relationships across epigenomes using large-scale hierarchical clustering

Recent advances in genomics and sequencing platforms have revolutionized our ability to create immense data sets, particularly for studying epigenetic regulation of gene expression. However, the avalanche of epigenomic data is difficult to parse for biological interpretation given nonlinear complex patterns and relationships. This attractive challenge in epigenomic data lends itself to machine learning for discerning infectivity and susceptibility. In this study, we explore over 3000 epigenomes of uninfected individuals and provide a framework to characterize the relationships among epigenetic modifiers, their modifiers, genetic loci, and specific immune cell types across all chromosomes using hierarchical clustering. Hierarchical clustering of epigenomic data revealed consistent epigenetic patterns across chromosomes, demonstrating that variation due to epigenetic modifiers is greater than variation between cell types. Gene Ontology and KEGG pathway analyses indicated significant enrichment of genes involved in chromatin remodeling, mRNA splicing, immune responses, and the regulation of microRNAs and snoRNAs. Epigenetic modifiers frequently formed biologically relevant clusters, including the cohesin complex, RNA Polymerase II transcription factors, and PRC2 complex members. These clustering behaviors remained consistent across all chromosomes, supported by entropy analysis and high Adjusted Rand Index scores, indicating robust cross-chromosomal similarity. Co-occurrence analysis further revealed specific sets of modifiers that consistently appeared together within clusters, reflecting shared biological functions and interactions. Validation using another dataset confirmed the reproducibility of these clustering patterns and modifier co-occurrence relationships, underscoring the reliability and generalizability of the methodology.

97 MATHEMATICS AND COMPUTING↗

Galaxy formation through hierarchical clustering

Analytic methods for studying the formation of galaxies by gas condensation within massive dark halos are presented. The present scheme applies to cosmogonies where structure grows through hierarchical clustering of a mixture of gas and dissipationless dark matter. The simplest models consistent with the current understanding of N-body work on dissipationless clustering, and that of numerical and analytic work on gas evolution and cooling are adopted. Standard models for the evolution of the stellar population are also employed, and new models for the way star formation heats and enriches the surrounding gas are constructed. Detailed results are presented for a cold dark matter universe with Omega = 1 and H(0) = 50 km/s/Mpc, but the present methods are applicable to other models. The present luminosity functions contain significantly more faint galaxies than are observed.

White, Simon D. M.↗

Exact hierarchical clustering in one dimension

The present adhesion model-based one-dimensional simulations of gravitational clustering have yielded bound-object catalogs applicable in tests of analytical approaches to cosmological structure formation. Attention is given to Press-Schechter (1974) type functions, as well as to their density peak-theory modifications and the two-point correlation function estimated from peak theory. The extent to which individual collapsed-object locations can be predicted by linear theory is significant only for objects of near-characteristic nonlinear mass.

Williams, B. G.↗

Dynamics of the baryonic component in hierarchical clustering universes

I present self-consistent 3-D simulations of the formation of virialized systems containing both gas and dark matter in a flat universe. A fully Lagrangian code based on the Smoothed Particle Hydrodynamics technique and a tree data structure has been used to evolve regions of comoving radius 2-3 Mpc. Tidal effects are included by coarse-sampling the density of the outer regions up to a radius approx. 20 Mpc. Initial conditions are set at high redshift (z greater than 7) using a standard Cold Dark Matter perturbation spectrum and a baryon mass fraction of 10 percent (omega(sub b) = 0.1). Simulations in which the gas evolves either adiabatically or radiates energy at a rate determined locally by its cooling function were performed. This allows us to investigate with the same set of simulations the importance of radiative losses in the formation of galaxies and the equilibrium structure of virialized systems where cooling is very inefficient. In the absence of radiative losses, the simulations can be rescaled to the density and radius typical of galaxy clusters. A summary of the main results is presented.

Navarro, Julio↗

ASCot, the NASA Analogy Software Cost Tool Suite: expanding our estimation horizons

The NASA Analogy Software Costing Tool Suite (ASCoT) consists of a cluster-based analogy estimator for estimating software development effort, a K-Nearest Neighbors (KNN) analogy estimator for estimating effort and delivered lines of code, a simple regression-based cost estimating relationship (CER) model that estimates cost in dollars, and a probabilistic version of COCOMO II. In this paper we document the analogy algorithms as well as summarize the results of the performance of the KNN and the principle components (PCA) cluster analogy models. KNN performance is assessed by varying the number of inputs and number of neighbors. Four different clustering methods: K-means, Spectral Clustering, Hierarchical Clustering, and Principle Components Analysis (PCA), and their respective evaluation criterion are described in detail. The comparative performance of all four estimation models is assessed using magnitude of relative error (MRE) measurements.

Menzies, Tim↗

Multi–Step Nucleation of a Crystalline Silicate Framework via a Structurally Precise Prenucleation Cluster

Hierarchical nucleation pathways are ubiquitous in the synthesis of minerals and materials. In the case of zeolites and metal–organic frameworks, pre-organized multi-ion “secondary building units” (SBUs) have been proposed as fundamental building blocks. However, detailing the progress of multi-step reaction mechanisms from monomeric species to stable crystals and defining the structures of the SBUs remains an unmet challenge. Furthermore, combining in situ nuclear magnetic resonance, small-angle X-ray scattering, and atomic force microscopy, we show that crystallization of the framework silicate, cyclosilicate hydrate, occurs through an assembly of cubic octameric Q 3 8 polyanions formed through cross-linking and polymerization of smaller silicate monomers and other oligomers. These Q 3 8 are stabilized by hydrogen bonds with surrounding H 2 O and tetramethylammonium ions (TMA + ). When Q 3 8 levels reach a threshold of ≈32 % of the total silicate species, nucleation occurs. Further growth proceeds through the incorporation of [(TMA) x (Q 3 8 )•n H 2 O] (x–8) clathrate complexes into step edges on the crystals.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

A systematic analysis and data mining of opioid-related adverse events submitted to the FAERS database

The opioid epidemic has become a serious national crisis in the United States. An indepth systematic analysis of opioid-related adverse events (AEs) can clarify the risks presented by opioid exposure, as well as the individual risk profiles of specific opioid drugs and the potential relationships among the opioids. In this study, 92 opioids were identified from the list of all Food and Drug Administration (FDA)-approved drugs, annotated by RxNorm and were classified into 13 opioid groups: buprenorphine, codeine, dihydrocodeine, fentanyl, hydrocodone, hydromorphone, meperidine, methadone, morphine, oxycodone, oxymorphone, tapentadol, and tramadol. A total of 14,970,399 AE reports were retrieved and downloaded from the FDA Adverse Events Reporting System (FAERS) from 2004, Quarter 1 to 2020, Quarter 3. After data processing, Empirical Bayes Geometric Mean (EBGM) was then applied which identified 3317 pairs of potential risk signals within the 13 opioid groups. Based on these potential safety signals, a comparative analysis was pursued to provide a global overview of opioid-related AEs for all 13 groups of FDA-approved prescription opioids. The top 10 most reported AEs for each opioid class were then presented. Both network analysis and hierarchical clustering analysis were conducted to further explore the relationship between opioids. Results from the network analysis revealed a close association among fentanyl, oxycodone, hydrocodone, and hydromorphone, which shared more than 22 AEs. In addition, much less commonly reported AEs were shared among dihydrocodeine, meperidine, oxymorphone, and tapentadol. On the contrary, the hierarchical clustering analysis further categorized the 13 opioid classes into two groups by comparing the full profiles of presence/absence of AEs. The results of network analysis and hierarchical clustering analysis were not only consistent and cross-validated each other but also provided a better and deeper understanding of the associations and relationships between the 13 opioid groups with respect to their adverse effect profiles.

Research & Experimental Medicine↗

PANDORA: A Parallel Dendrogram Construction Algorithm for Single Linkage Clustering on GPU

This paper introduces Pandora, a parallel algorithm for computing dendrograms, the hierarchical cluster trees for single linkage clustering (SLC). Current parallel approaches construct dendrograms by partitioning a minimum spanning tree and removing edges. However, they struggle with skewed, hard-to-parallelize real-world dendrograms. Consequently, computing dendrograms is the sequential bottleneck in HDBSCAN*[21], a popular SLC variant. Pandora uses recursive tree contraction to address this limitation. Pandora contracts nodes to construct progressively smaller trees. It computes the smallest contracted dendrogram and expands it by inserting contracted edges. This recursive strategy is highly parallel, skew-independent, work-optimal, and well-suited for GPUs and multicores. We develop a performance portable implementation of Pandora in Kokkos[31] and evaluate its performance on multicore CPUs and multi-vendor GPUs (e.g., Nvidia, AMD) for dendrogram construction in HDBSCAN*. Multithreaded Pandora is 2.2x faster than the current best-multithreaded implementation. Our GPU version achieves 6-20x speedup on AMD GPUs and 10-37x on NVIDIA GPUs over multithreaded Pandora. Pandora removes HDBSCAN*’s sequential bottleneck, greatly boosting efficiency, particularly with GPUs.

Sao, Piyush↗

Correlations in cosmic density fields

A method is proposed to place constraints on the functional form of the high-order correlation functions zeta(sub n) that arise in cosmic density fields at large scales. This technique is based on a mass-in-cell statistic and a difference of mass in partitions of a cell. The relationship between these measures is sensitive to the formal structure of the zeta(sub n) as well as their amplitudes. This relationship is quantified in several theoretical models of structure, based on the hierarchical clustering paradigm. The results lead to a test for specific types of hierarchical clustering that is sensitive to correlations of all orders. The method is applied to examples of simulated large-scaled structure dominated by cold dark matter. In the preliminary study, the hierarchical paradigm appears to be a realistic approximation over a broad range of the scales. Furthermore, there is evidence that graphs of low-order vertices are dominant. On the basis of simulated data a phenomological model is specified that gives a good representation of clustering from linear scales to the strongly clustered regime (zeta(sub 2) approximately 500).

Bromley, B. C.↗

Transport in the Subtropical Lowermost Stratosphere during CRYSTAL-FACE

We use in situ measurements of water vapor (H2O), ozone (O3), carbon dioxide (CO2), carbon monoxide (CO), nitric oxide (NO), and total reactive nitrogen (NO(y)) obtained during the CRYSTAL-FACE campaign in July 2002 to study summertime transport in the subtropical lowermost stratosphere. We use an objective methodology to distinguish the latitudinal origin of the sampled air masses despite the influence of convection, and we calculate backward trajectories to elucidate their recent geographical history. The methodology consists of exploring the statistical behavior of the data by performing multivariate clustering and agglomerative hierarchical clustering calculations, and projecting cluster groups onto principal component space to identify air masses of like composition and hence presumed origin. The statistically derived cluster groups are then examined in physical space using tracer-tracer correlation plots. Interpretation of the principal component analysis suggests that the variability in the data is accounted for primarily by the mean age of air in the stratosphere, followed by the age of the convective influence, and lastly by the extent of convective influence, potentially related to the latitude of convective injection [Dessler and Sherwuud, 2004]. We find that high-latitude stratospheric air is the dominant source region during the beginning of the campaign while tropical air is the dominant source region during the rest of the campaign. Influence of convection from both local and non-local events is frequently observed. The identification of air mass origin is confirmed with backward trajectories, and the behavior of the trajectories is associated with the North American monsoon circulation.

Pittman, Jasna V.↗

Transport in the Subtropical Lowermost Stratosphere during the Cirrus Regional Study of Tropical Anvils and Cirrus Layers-Florida Area Cirrus Experiment

We use in situ measurements of water vapor (H2O), ozone (O3), carbon dioxide (CO2), carbon monoxide (CO), nitric oxide (NO), and total reactive nitrogen (NOy) obtained during the CRYSTAL-FACE campaign in July 2002 to study summertime transport in the subtropical lowermost stratosphere. We use an objective methodology to distinguish the latitudinal origin of the sampled air masses despite the influence of convection, and we calculate backward trajectories to elucidate their recent geographical history. The methodology consists of exploring the statistical behavior of the data by performing multivariate clustering and agglomerative hierarchical clustering calculations and projecting cluster groups onto principal component space to identify air masses of like composition and hence presumed origin. The statistically derived cluster groups are then examined in physical space using tracer-tracer correlation plots. Interpretation of the principal component analysis suggests that the variability in the data is accounted for primarily by the mean age of air in the stratosphere, followed by the age of the convective influence, and last by the extent of convective influence, potentially related to the latitude of convective injection (Dessler and Sherwood, 2004). We find that high-latitude stratospheric air is the dominant source region during the beginning of the campaign while tropical air is the dominant source region during the rest of the campaign. Influence of convection from both local and nonlocal events is frequently observed. The identification of air mass origin is confirmed with backward trajectories, and the behavior of the trajectories is associated with the North American monsoon circulation.

Pittman, Jasna V.↗

Detecting CAN Masquerade Attacks with Signal Clustering Similarity

Vehicular Controller Area Networks (CANs) are susceptible to cyber attacks of different levels of sophistication. Fabrication attacks are the easiest to administer—an adversary simply sends (extra) frames on a CAN—but also the easiest to detect because they disrupt frame frequency. To overcome time-based detection methods, adversaries must administer masquerade attacks by sending frames in lieu of (and therefore at the expected time of) benign frames but with malicious payloads. Research efforts have proven that CAN attacks, and masquerade attacks in particular, can affect vehicle functionality. Examples include causing unintended acceleration, deactivation of vehicle’s brakes, as well as steering the vehicle. We hypothesize that masquerade attacks modify the nuanced correlations of CAN signal time series and how they cluster together. Therefore, changes in cluster assignments should indicate anomalous behavior. We confirm this hypothesis by leveraging our previously developed capability for reverse engineering CAN signals (i.e., CAN-D [Controller Area Network Decoder]) and focus on advancing the state of the art for detecting masquerade attacks by analyzing time series extracted from raw CAN frames. Specifically, we demonstrate that masquerade attacks can be detected by computing time series clustering similarity using hierarchical clustering on the vehicle’s CAN signals (time series) and comparing the clustering similarity across CAN captures with and without attacks. We test our approach in a previously collected CAN dataset with masquerade attacks (i.e., the ROAD dataset) and develop a forensic tool as a proof of concept to demonstrate the potential of the proposed approach for detecting CAN masquerade attacks.

Moriano Salazar, Pablo↗

Statistical framework to assess long-term spatio-temporal climate changes: East River mountainous watershed case study

Abstract Evaluation of long-term temporal and spatial climatic change in mountainous regions is a critical challenge because of the interactive effects of multiple land and climatic factors and processes. Here we present the application of the statistical framework to the assessment of changes of climatic conditions, using data from 17 meteorological stations across the East River watershed near Crested Butte, Colorado, USA, and spanning the period from 1966 to 2021. The framework is developed based on (1) a time-series analysis of daily, monthly, and yearly averaged meteorological parameters (temperature, relative humidity, precipitation, wind speed, etc.), (2) evaluation and time series analysis of potential evapotranspiration (ET o ), actual evapotranspiration (ET), aridity index (AI), standard precipitation index (SPI) and standard precipitation-evapotranspiration index (SPEI), and (3) a temporal-spatial climatic zonation of the studied area based on the hierarchical clustering and PCA analysis of the SPEI, because the SPEI can be considered an integrative characteristic of the changes of climatic conditions. The Budyko model, with the application of the Penman–Monteith equation for the estimation of ET o , was used to determine the ET. The time series analysis of the AI is used to identify the periods with energy limited and water limited conditions. Hierarchical clustering of site locations for the three temporal segments of the SPEI showed a significant temporal-spatial shifts, indicating that dynamic climatic processes drive zonation patterns. Therefore, the watershed climatic zonation requires periodic re-evaluation based on the structural time series analysis of meteorological and water balance data.

54 ENVIRONMENTAL SCIENCES↗

Testing higher-order Lagrangian perturbation theory against numerical simulation. 1: Pancake models

We present results showing an improvement of the accuracy of perturbation theory as applied to cosmological structure formation for a useful range of quasi-linear scales. The Lagrangian theory of gravitational instability of an Einstein-de Sitter dust cosmogony investigated and solved up to the third order is compared with numerical simulations. In this paper we study the dynamics of pancake models as a first step. In previous work the accuracy of several analytical approximations for the modeling of large-scale structure in the mildly non-linear regime was analyzed in the same way, allowing for direct comparison of the accuracy of various approximations. In particular, the Zel'dovich approximation (hereafter ZA) as a subclass of the first-order Lagrangian perturbation solutions was found to provide an excellent approximation to the density field in the mildly non-linear regime (i.e. up to a linear r.m.s. density contrast of sigma is approximately 2). The performance of ZA in hierarchical clustering models can be greatly improved by truncating the initial power spectrum (smoothing the initial data). We here explore whether this approximation can be further improved with higher-order corrections in the displacement mapping from homogeneity. We study a single pancake model (truncated power-spectrum with power-spectrum with power-index n = -1) using cross-correlation statistics employed in previous work. We found that for all statistical methods used the higher-order corrections improve the results obtained for the first-order solution up to the stage when sigma (linear theory) is approximately 1. While this improvement can be seen for all spatial scales, later stages retain this feature only above a certain scale which is increasing with time. However, third-order is not much improvement over second-order at any stage. The total breakdown of the perturbation approach is observed at the stage, where sigma (linear theory) is approximately 2, which corresponds to the onset of hierarchical clustering. This success is found at a considerable higher non-linearity than is usual for perturbation theory. Whether a truncation of the initial power-spectrum in hierarchical models retains this improvement will be analyzed in a forthcoming work.

Buchert, T.↗

Probabilisitc Geobiological Classification Using Elemental Abundance Distributions and Lossless Image Compression in Recent and Modern Organisms

Last year we presented techniques for the detection of fossils during robotic missions to Mars using both structural and chemical signatures[Storrie-Lombardi and Hoover, 2004]. Analyses included lossless compression of photographic images to estimate the relative complexity of a putative fossil compared to the rock matrix [Corsetti and Storrie-Lombardi, 2003] and elemental abundance distributions to provide mineralogical classification of the rock matrix [Storrie-Lombardi and Fisk, 2004]. We presented a classification strategy employing two exploratory classification algorithms (Principal Component Analysis and Hierarchical Cluster Analysis) and non-linear stochastic neural network to produce a Bayesian estimate of classification accuracy. We now present an extension of our previous experiments exploring putative fossil forms morphologically resembling cyanobacteria discovered in the Orgueil meteorite. Elemental abundances (C6, N7, O8, Na11, Mg12, Ai13, Si14, P15, S16, Cl17, K19, Ca20, Fe26) obtained for both extant cyanobacteria and fossil trilobites produce signatures readily distinguishing them from meteorite targets. When compared to elemental abundance signatures for extant cyanobacteria Orgueil structures exhibit decreased abundances for C6, N7, Na11, All3, P15, Cl17, K19, Ca20 and increases in Mg12, S16, Fe26. Diatoms and silicified portions of cyanobacterial sheaths exhibiting high levels of silicon and correspondingly low levels of carbon cluster more closely with terrestrial fossils than with extant cyanobacteria. Compression indices verify that variations in random and redundant textural patterns between perceived forms and the background matrix contribute significantly to morphological visual identification. The results provide a quantitative probabilistic methodology for discriminating putatitive fossils from the surrounding rock matrix and &om extant organisms using both structural and chemical information. The techniques described appear applicable to the geobiological analysis of meteoritic samples or in situ exploration of the Mars regolith. Keywords: cyanobacteria, microfossils, Mars, elemental abundances, complexity analysis, multifactor analysis, principal component analysis, hierarchical cluster analysis, artificial neural networks, paleo-biosignatures

Storrie-Lombardi, Michael C.↗

Overall Performance Losses and Activated Mechanisms in Double Glass and Glass-backsheet Photovoltaic Modules with Monofacial and Bifacial PERC Cells, under Accelerated Exposures

Commercial PV modules have various packaging choices nowadays, which influence their long-term reliability. This study compared the degradation behaviors of sixteen module variants from two brands with varying encapsulant materials (EVA or POE), encapsulant types, module architectures (GB or DG), and cell types (monofacial or bifacial) using null hypothesis testing to determine statistical significant findings. The modules were exposed for 2,520 hours under two accelerated exposures: modified damp heat (mDH) and modified damp heat with full-spectrum light (mDH+FSL). For both brands, two DG module variants with UV-Cutoff rear encapsulant are found to have significantly lower average power loss than the module variants of EVA+GB with opaque rear encapsulant after each accelerated exposure. Metallization interconnect corrosion is identified as the primary degradation mechanism. Furthermore, unsupervised hierarchical clustering finds that the degradation behaviors of modules from one brand with a more strict manufacturing quality control depends on module architectures only.

14 SOLAR ENERGY↗

Sulfur in Cometary Dust

The computer-intensive project consisted of the analysis and synthesis of existing data on composition of comet Halley dust particles. The main objective was to obtain a complete inventory of sulfur containing compounds in the comet Halley dust by building upon the existing classification of organic and inorganic compounds and applying a variety of statistical techniques for cluster and cross-correlational analyses. A student hired for this project wrote and tested the software to perform cluster analysis. The following tasks were carried out: (1) selecting the data from existing database for the proposed project; (2) finding access to a standard library of statistical routines for cluster analysis; (3) reformatting the data as necessary for input into the library routines; (4) performing cluster analysis and constructing hierarchical cluster trees using three methods to define the proximity of clusters; (5) presenting the output results in different formats to facilitate the interpretation of the obtained cluster trees; (6) selecting groups of data points common for all three trees as stable clusters. We have also considered the chemistry of sulfur in inorganic compounds.

Fomenkova, M. N.↗