Search NASA⌕ Search

SEARCH · Search NASA

Results for “Network-based clustering”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

A Network-Based Algorithm for Clustering Multivariate Repeated Measures Data

The National Aeronautics and Space Administration (NASA) Astronaut Corps is a unique occupational cohort for which vast amounts of measures data have been collected repeatedly in research or operational studies pre-, in-, and post-flight, as well as during multiple clinical care visits. In exploratory analyses aimed at generating hypotheses regarding physiological changes associated with spaceflight exposure, such as impaired vision, it is of interest to identify anomalies and trends across these expansive datasets. Multivariate clustering algorithms for repeated measures data may help parse the data to identify homogeneous groups of astronauts that have higher risks for a particular physiological change. However, available clustering methods may not be able to accommodate the complex data structures found in NASA data, since the methods often rely on strict model assumptions, require equally-spaced and balanced assessment times, cannot accommodate missing data or differing time scales across variables, and cannot process continuous and discrete data simultaneously. To fill this gap, we propose a network-based, multivariate clustering algorithm for repeated measures data that can be tailored to fit various research settings. Using simulated data, we demonstrate how our method can be used to identify patterns in complex data structures found in practice.

Koslovsky, Matthew↗

A Neural Network Approach for Identifying Particle Pitch Angle Distributions in Van Allen Probes Data

Analysis of particle pitch angle distributions (PADs) has been used as a means to comprehend a multitude of different physical mechanisms that lead to flux variations in the Van Allen belts and also to particle precipitation into the upper atmosphere. In this work we developed a neural network-based data clustering methodology that automatically identifies distinct PAD types in an unsupervised way using particle flux data. One can promptly identify and locate three well-known PAD types in both time and radial distance, namely, 90deg peaked, butterfly, and flattop distributions. In order to illustrate the applicability of our methodology, we used relativistic electron flux data from the whole month of November 2014, acquired from the Relativistic Electron-Proton Telescope instrument on board the Van Allen Probes, but it is emphasized that our approach can also be used with multiplatform spacecraft data. Our PAD classification results are in reasonably good agreement with those obtained by standard statistical fitting algorithms. The proposed methodology has a potential use for Van Allen belt's monitoring.

Souza, V. M.↗

MEDUSA - An overset grid flow solver for network-based parallel computer systems

Continuing improvement in processing speed has made it feasible to solve the Reynolds-Averaged Navier-Stokes equations for simple three-dimensional flows on advanced workstations. Combining multiple workstations into a network-based heterogeneous parallel computer allows the application of programming principles learned on MIMD (Multiple Instruction Multiple Data) distributed memory parallel computers to the solution of larger problems. An overset-grid flow solution code has been developed which uses a cluster of workstations as a network-based parallel computer. Inter-process communication is provided by the Parallel Virtual Machine (PVM) software. Solution speed equivalent to one-third of a Cray-YMP processor has been achieved from a cluster of nine commonly used engineering workstation processors. Load imbalance and communication overhead are the principal impediments to parallel efficiency in this application.

Smith, Merritt H.↗

High performance FPGA embedded system for machine learning based tracking and trigger in sPhenix and EIC

We present a comprehensive end-to-end pipeline to classify triggers versus background events in this paper. This pipeline makes online decisions to select signal data and enables the intelligent trigger system for efficient data collection in the Data Acquisition System (DAQ) of the upcoming sPHENIX and future EIC (Electron-Ion Collider) experiments. Starting from the coordinates of pixel hits that are lightened by passing particles in the detector, the pipeline applies three-stage of event processing (hits clustering, track reconstruction, and trigger detection) and labels all processed events with the binary tag of trigger versus background events. The pipeline consists of deterministic algorithms such as clustering pixels to reduce event size, tracking reconstruction to predict candidate edges, and advanced graph neural network-based models for recognizing the entire jet pattern. In particular, we apply the message-passing graph neural network to predict links between hits and reconstruct tracks and a hierarchical pooling algorithm (DiffPool) to make the graph-level trigger detection. We obtain an impressive performance (≥70% accuracy) for trigger detection with only 3200 neuron weights in the end-to-end pipeline. We deploy the end-to-end pipeline into a field-programmable gate array (FPGA) and accelerate the three stages with speedup factors of 1152, 280, and 21, respectively.

Instruments & Instrumentation↗

Programmable photonic integrated meshes for modular generation of optical entanglement links

Abstract Large-scale generation of quantum entanglement between individually controllable qubits is at the core of quantum computing, communications, and sensing. Modular architectures of remotely-connected quantum technologies have been proposed for a variety of physical qubits, with demonstrations reported in atomic and all-photonic systems. However, an open challenge in these architectures lies in constructing high-speed and high-fidelity reconfigurable photonic networks for optically-heralded entanglement among target qubits. Here we introduce a programmable photonic integrated circuit (PIC), realized in a piezo-actuated silicon nitride (SiN)-in-oxide CMOS-compatible process, that implements an N × N Mach–Zehnder mesh (MZM) capable of high-speed execution of linear optical transformations. The visible-spectrum photonic integrated mesh is programmed to generate optical connectivity on up to N = 8 inputs for a range of optically-heralded entanglement protocols. In particular, we experimentally demonstrated optical connections between 16 independent pairwise mode couplings through the MZM, with optical transformation fidelities averaging 0.991 ± 0.0063. The PIC’s reconfigurable optical connectivity suffices for the production of 8-qubit resource states as building blocks of larger topological cluster states for quantum computing. Our programmable PIC platform enables the fast and scalable optical switching technology necessary for network-based quantum information processors.

47 OTHER INSTRUMENTATION↗

Neural network-based model of galaxy power spectrum: fast full-shape galaxy power spectrum analysis

ABSTRACT We present a neural network-based emulator for the galaxy redshift-space power spectrum that enables several orders of magnitude acceleration in the galaxy clustering parameter inference, while preserving 3$\sigma$ accuracy better than 0.5 per cent up to $k_{\mathrm{max}}$ = 0.25 $\, h\text{Mpc}^{-1}$ within Lambda-cold dark matter ($\Lambda$CDM) and around 0.5 per cent $w_0$–$w_a$CDM. Our surrogate model only emulates the galaxy bias-invariant terms of one-loop perturbation theory predictions, these terms are then combined analytically with galaxy bias terms, counter-terms, and stochastic terms in order to obtain the non-linear redshift-space galaxy power spectrum. This allows us to avoid any galaxy bias prescription in the training of the emulator, which makes it more flexible. Moreover, we include the redshift $z \in [0,1.4]$ in the training which further avoids the need for re-training the emulator. We showcase the performance of the emulator in recovering the cosmological parameters of $\Lambda$CDM by analysing the suite of 25 AbacusSummit simulations that mimic the Dark Energy Spectroscopic Instrument luminous red galaxies at $z=0.5$ and 0.8, together as the emission line galaxies at $z=0.8$. We obtain similar performance in all cases, demonstrating the reliability of the emulator for any galaxy sample at any redshift in $0 \lt z \lt 1.4$. We will make our emulator public at github repository.

Trusov, Svyatoslav (ORCID:0000000224146720)↗

Understanding Proton Movement in [Fe-Fe] Hydrogenases

Nature uses specialized metalloenzymes to carry out small molecule activation reactions, including CO 2 fixation, O 2 activation, and proton reduction, with unparalleled efficiency, rates, and selectivity. The latter reactions are performed by hydrogenases, protein metallo-complexes that interconverts H 2 to protons and electrons (H 2 oxidation) and the reverse reaction (H 2 production) with incredibly low energy input and amazingly fast kinetics. Reproducing both the activity and efficiency of metalloenzymes in sustainable anthropogenic systems remains one of the “holy grails” of inorganic chemistry. However, identifying the precise molecular components responsible for these desirable properties has been challenging in the natural metalloenzymes, hindering efforts to develop analogous processes in synthetic compounds. Considering the inherent complexity of a metalloenzyme and the many interactions, both strong and weak, that contribute to the function of an enzyme, we have elected to model natural metalloenzymes on a biochemical platform. Towards this end, we have a developed structural, functional, and mechanistic mimic of the [Ni-Fe] hydrogenases within a robust protein scaffold, rubredoxin, to understand the influence of the secondary coordination environment on the metal center. This involved preparing a series of three rubredoxin constructs containing a single point mutation at Val positions and making the NMR chemical shift assignments for the paramagnetic (nickel-substituted) and non-paramagnetic (zinc-substituted) form. These physical studies were complemented with in silico molecular dynamic studies on proteins with [Fe-S] clusters to engineer new proton channels to test in vitro . To assist our search for new natural metalloprotein scaffolds within the vast number of sequenced genomes, we created a neural network-based program to identify proteins with specific metal-binding sites. These computational and physical studies with metalloenzymes provide direct insight into the fundamental chemical principles driving the natural systems and offer design principles for developing catalysts that utilize analogous principles.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Performance Evaluation in Network-Based Parallel Computing

Network-based parallel computing is emerging as a cost-effective alternative for solving many problems which require use of supercomputers or massively parallel computers. The primary objective of this project has been to conduct experimental research on performance evaluation for clustered parallel computing. First, a testbed was established by augmenting our existing SUNSPARCs' network with PVM (Parallel Virtual Machine) which is a software system for linking clusters of machines. Second, a set of three basic applications were selected. The applications consist of a parallel search, a parallel sort, a parallel matrix multiplication. These application programs were implemented in C programming language under PVM. Third, we conducted performance evaluation under various configurations and problem sizes. Alternative parallel computing models and workload allocations for application programs were explored. The performance metric was limited to elapsed time or response time which in the context of parallel computing can be expressed in terms of speedup. The results reveal that the overhead of communication latency between processes in many cases is the restricting factor to performance. That is, coarse-grain parallelism which requires less frequent communication between processes will result in higher performance in network-based computing. Finally, we are in the final stages of installing an Asynchronous Transfer Mode (ATM) switch and four ATM interfaces (each 155 Mbps) which will allow us to extend our study to newer applications, performance metrics, and configurations.

Dezhgosha, Kamyar↗

Neural network acceleration of large-scale structure theory calculations

Here, we make use of neural networks to accelerate the calculation of power spectra required for the analysis of galaxy clustering and weak gravitational lensing data. For modern perturbation theory codes, evaluation time for a single cosmology and redshift can take on the order of two seconds. In combination with the comparable time required to compute linear predictions using a Boltzmann solver, these calculations are the bottleneck for many contemporary large-scale structure analyses. Here, we construct neural network-based surrogate models for Lagrangian perturbation theory (LPT) predictions of matter power spectra, real and redshift space galaxy power spectra, and galaxy-matter cross power spectra that attain ~ 0.1% (at one sigma) accuracy over a broad range of scales in a ωCDM parameter space. The neural network surrogates can be evaluated in approximately one millisecond, a factor of 1000 times faster than the full Boltzmann code and LPT computations. In a simulated full-shape redshift space galaxy power spectrum analysis, we demonstrate that the posteriors obtained using our surrogates are accurate compared to those obtained using the full LPT model. We make our surrogate models public at https://github.com/sfschen/EmulateLSS, so that others may take advantage of the speed gains they provide to enable rapid iteration on analysis settings, something that is essential in complex contemporary large-scale structure analyses.

79 ASTRONOMY AND ASTROPHYSICS↗

Mitigating imaging systematics for DESI 2024 emission Line Galaxies and beyond

Emission Line Galaxies (ELGs) are one of the main tracers that the Dark Energy Spectroscopic Instrument (DESI) uses to probe the universe. However, they are afflicted by strong spurious correlations between target density and observing conditions known as imaging systematics. In this paper, we present the imaging systematics mitigation applied to the DESI Data Release 1 (DR1) large-scale structure catalogs used in the DESI 2024 cosmological analyses. We also explore extensions of the fiducial treatment. This includes a combined approach, through forward image simulations (Obiwan) in conjunction with neural network-based regression, to obtain an angular selection function that mitigates the imaging systematics observed in the DESI DR1 ELGs target density. We further derive a line of sight selection function from the forward model that removes the strong redshift dependence between imaging systematics and low redshift ELGs. Combining both angular and redshift-dependent systematics, we construct a three-dimensional selection function and assess the impact of all selection functions on clustering statistics. We quantify differences between these extended treatments and the fiducial treatment in terms of the measured 2-point statistics. We find that the results are generally consistent with the fiducial treatment and conclude that the differences are far less than the imaging systematics uncertainty included in DESI 2024 full-shape measurements. We extend our investigation to the ELGs at 0.6 < z < 0.8, i.e., beyond the redshift range (0.8 < z < 1.6) adopted for the DESI clustering catalog, and demonstrate that determining the full three-dimensional selection function is necessary in this redshift range. Our tests showed that all changes are consistent with statistical noise for BAO analyses indicating they are robust to even severe imaging systematics. Specific tests for the full-shape analysis will be presented in a companion paper.

79 ASTRONOMY AND ASTROPHYSICS↗

Machine Learning of All Mycobacterium tuberculosis H37Rv RNA-seq Data Reveals a Structured Interplay between Metabolism, Stress Response, and Infection

Mycobacterium tuberculosis is one of the most consequential human bacterial pathogens, posing a serious challenge to 21st century medicine. A key feature of its pathogenicity is its ability to adapt its transcriptional response to environmental stresses through its transcriptional regulatory network (TRN). While many studies have sought to characterize specific portions of the M. tuberculosis TRN, and some studies have performed system-level analysis, few have been able to provide a network-based model of the TRN that also provides the relative shifts in transcriptional regulator activity triggered by changing environments. Here, we compiled a compendium of nearly 650 publicly available, high quality M. tuberculosis RNA-sequencing data sets and applied an unsupervised machine learning method to obtain a quantitative, top-down TRN. It consists of 80 independently modulated gene sets known as “iModulons,” 41 of which correspond to known regulons. These iModulons explain 61% of the variance in the organism’s transcriptional response. We show that iModulons (i) reveal the function of poorly characterized regulons, (ii) describe the transcriptional shifts that occur during environmental changes such as shifting carbon sources, oxidative stress, and infection events, and (iii) identify intrinsic clusters of regulons that link several important metabolic systems, including lipid, cholesterol, and sulfur metabolism. This transcriptome-wide analysis of the M. tuberculosis TRN informs future research on effective ways to study and manipulate its transcriptional regulation and presents a knowledge-enhanced database of all published high-quality RNA-seq data for this organism to date.

59 BASIC BIOLOGICAL SCIENCES↗

OctoFAS: A Two-Level Fair Scheduler That Increases Fairness in Network-Based Key-Value Storage

We identified a fairness problem in a network-based key-value storage system using Intel Storage Performance Development Kit (SPDK) in a multitenant environment. In such an environment, each tenant’s I/O service rate is not fairly guaranteed compared to that of other tenants. To address the fairness problem, we propose OctoFAS, a two-level fair scheduler designed to improve overall throughput and fairness among tenants. The two-level scheduler of OctoFAS consists of (i) inter-core scheduling and (ii) intra-core scheduling. Through inter-core scheduling, OctoFAS addresses the load imbalance problem that is inherent in SPDK on the storage server by dynamically migrating I/O requests from overloaded cores to underloaded cores, thereby increasing overall throughput. Intra-core scheduling prioritizes handling requests from starving tenants over well-fed tenants within core-specific event queues to ensure fair I/O services among multiple tenants. OctoFAS is deployed on a Linux cluster with SPDK. Through extensive evaluations, we found that OctoFAS ensures that the total system throughput remains high and balanced, while enhancing fairness by approximately 10% compared to the baseline, when both scheduling levels operate in a hybrid fashion.

97 MATHEMATICS AND COMPUTING↗

Acceleration of Graph Neural Network-Based Prediction Models in Chemistry via Co-Design Optimization on Intelligence Processing Units

Atomic structure prediction and associated property calculations are the bedrock of chemical physics. Since high-fidelity ab initio modeling techniques for computing the structure and properties can be prohibitively expensive, this motivates the development of machine-learning (ML) models that make these predictions more efficiently. Training graph neural networks over large atomistic databases introduces unique computational challenges such as the need to process millions of small graphs with variable size and support communication patterns that are distinct from learning over large graphs such as social networks. We demonstrate a novel hardware-software co-design approach to scale up the training of atomistic graph neural networks (GNN) for structure and property prediction. First, to eliminate redundant computation and memory associated with alternative padding techniques and to improve throughput via minimizing communication, we formulate the effective coalescing of the batches of variable-size atomistic graphs as the bin packing problem and introduce a hardware-agnostic algorithm to pack these batches. In addition, we propose hardware-specific optimizations including a planner and vectorization for the gather-scatter operations targeted for Graphcore’s Intelligence Processing Unit (IPU), as well as model-specific optimizations such as merged communication collectives and optimized softplus. Putting these all together, we demonstrate the effectiveness of the proposed co-design approach by providing an implementation of a well-established atomistic GNN on the Graphcore IPUs. We evaluate the training performance on multiple atomistic graph databases with varying degrees of graph counts, sizes and sparsity. Here, we demonstrate that such a co-design approach can reduce the training time of atomistic GNNs and can improve the performance by up to 1.5× compared to the baseline implementation of the model on the IPUs. Additionally, we compare our IPU implementation with a Nvidia GPU-based implementation and show that our atomistic GNN implementation on the IPUs can run 1.8× faster on average compared to the execution time on the GPUs.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Visualizing novel connections and genetic similarities across diseases using a network-medicine based approach

Understanding the genetic relationships between human disorders could lead to better treatment and prevention strategies, especially for individuals with multiple comorbidities. A common resource for studying genetic-disease relationships is the GWAS Catalog, a large and well curated repository of SNP-trait associations from various studies and populations. Some of these populations are contained within mega-biobanks such as the Million Veteran Program (MVP), which has enabled the genetic classification of several diseases in a large well-characterized and heterogeneous population. Here we aim to provide a network of the genetic relationships among diseases and to demonstrate the utility of quantifying the extent to which a given resource such as MVP has contributed to the discovery of such relations. We use a network-based approach to evaluate shared variants among thousands of traits in the GWAS Catalog repository. Our results indicate many more novel disease relationships that did not exist in early studies and demonstrate that the network can reveal clusters of diseases mechanistically related. Finally, we show novel disease connections that emerge when MVP data is included, highlighting methodology that can be used to indicate the contributions of a given biobank.

59 BASIC BIOLOGICAL SCIENCES↗

Leveraging BERT and Network-Based Attention Analysis for Identifying Treatment Milestones in EHRs

This study introduces a sophisticated data-driven framework for analyzing Electronic Health Records (EHRs) using transformer-based models to identify and disentangle overlapping treatment contexts. The framework leverages a preprocessing pipeline that transforms structured procedural codes into semantically enriched descriptive text, enabling the use of attention mechanisms to cluster medical events into treatment milestones—cohesive and distinct components of care processes. The methodology is rigorously validated using synthetic datasets derived from the MIMIC-III database, designed to simulate the heterogeneity and overlapping procedural contexts characteristic of real-world EHR scenarios. Quantitative evaluation highlights the framework’s robustness in disentangling concurrent care pathways, with attention metrics and unsupervised clustering approaches demonstrating the ability to preserve intra-context relationships while distinguishing inter-context dependencies. By addressing challenges inherent in data heterogeneity, this approach provides a foundation for uncovering complex treatment patterns, advancing clinical decision-making, and optimizing resource allocation in diverse healthcare environments.

Kim, Minsu [ORNL] (ORCID:0000000224185535)↗

A network approach for multiscale catchment classification using traits

Abstract. The classification of river catchments into groups with similar biophysical characteristics is useful to understand and predict their hydrological behavior. The increasing availability of remote sensing and other large-scale geospatial datasets has enabled the use of advanced data-driven approaches to classify catchments using traits such as topography, geology, climate, land cover, land use, and human influence. Unsupervised clustering algorithms based on the Euclidean distance are commonly used for trait-based classification but are not suitable for highly dimensional data. In this study we present a new network-based method for multi-scale catchment classification, which can be applied to large datasets and used to determine the traits associated with different catchment groups. In this framework, two networks are analyzed in parallel: the first being where the nodes are traits and the second being where the nodes are catchments. In both cases, edges represent pairwise similarity, and a network cluster detection algorithm is used for the classification. The trait network is used to investigate redundancy in the trait data and to condense this information into a small number of interpretable categories. The catchments network is used to classify the catchments into clusters and to identify representative catchments for the different groups using the degree centrality metric. We apply this method to classify 9067 river catchments across the contiguous United States at both regional and continental scales using 274 non-categorical traits. At the continental scale, we identify 25 interpretable trait categories and 34 catchment clusters of sizes greater than 50. We find that catchments with similar trait categories are typically located in the same region, with different spatial patterns emerging among clusters dominated by natural and anthropogenic traits. We also find that the catchment clusters exhibit distinct hydrological behavior based on an analysis of streamflow indices. This network approach provides several advantages over traditional means of classification, including better separation of clusters, the use of alternate similarity metrics that are more suitable for highly dimensional data, and reducing redundancy in the trait information. The paired catchment–trait networks enable analysis of hydrological behavior using the dominant trait categories for each catchment cluster. The approach can be used at multiple spatial scales since the network topologies adjust automatically to reflect the trait patterns at the scale of investigation. Finally, the representative catchments identified as hub nodes in the network can be used to guide transferable observational and modeling strategies. The method is broadly applicable beyond hydrology for classification of other complex systems that utilize different types of trait datasets.

54 ENVIRONMENTAL SCIENCES↗

OCTOKV: An Agile Network-Based Key-Value Storage System with Robust Load Orchestration

In this paper, we propose OctoKV, an innovative network-based key-value storage system. OctoKV addresses the repetitive address translation overhead associated with traditional key-value stores running on file systems on the client side. To mitigate this overhead, we implemented the key-value store on the server side using NVMe-oF and a user-level NVMe driver. In particular, we employed fine-grained resource monitoring and load balancing based on heuristics to optimize I/O performance. OctoKV is deployed on a Linux cluster with Intel SPDK. The extensive evaluation shows that OctoKV achieves lower I/O response times in comparison to traditional approaches where key-value stores run on the client side. Also, the proposed load balancing strategies efficiently enhance I/O response times by equally distributing the workload from overloaded cores to other cores.

Khan, Awais↗