Search NASA⌕ Search

SEARCH · Search NASA

Results for “DBSCAN”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

Metric DBSCAN

SAND2025-11725O Metric DBSCAN is an implementation of the popular DBSCAN clustering algorithm that works in general metric spaces. DBSCAN is a clustering algorithm, a fundamental building block in machine learning. It takes a set of objects and, given some notion of distance, identifies coherent groups of objects. With Metric DBSCAN, users can provide an arbitrary function to compute distance. Nearly all existing implementations of DBSCAN restrict distance to one of a few formulations. Metric DBScan accomplishes this cleanly and efficiently. The Python source code is on Github. Sandia National Laboratories is a multimission laboratory managed and operated by National Technology & Engineering Solutions of Sandia, LLC, a wholly owned subsidiary of Honeywell International Inc., for the U.S. Department of Energy’s National Nuclear Security Administration under contract DE-NA0003525.

Dalbey, Keith↗

Fast tree-based algorithms for DBSCAN for low-dimensional data on GPUs

DBSCAN is a well-known density-based clustering algorithm to discover arbitrary shape clusters. While conceptually simple in serial, the algorithm is challenging to efficiently parallelize on manycore GPU architectures. Common pitfalls, such as asynchronous range query calls, result in high thread execution divergence in many implementations. In this paper, we propose a new framework for GPU-accelerated DBSCAN, and describe two tree-based algorithms within that framework. Both algorithms fuse the search for neighbors with updating cluster information, but differ in their treatment of dense regions of the data. We show that the time taken to compute clusters is at most twice that of determination of the neighbors. We compare the proposed algorithms with existing CPU and GPU implementations, and demonstrate their competitiveness and performance using a fast traversal structure (bounding volume hierarchy) for low dimensional data. We also show that the memory usage can be reduced by processing object neighbors dynamically without storing them.

Prokopenko, Andrey↗

Geographic_Distribution_of_Populus_trichocarpa_Genotypes_by_DBSCAN_Cluster

Aninteractive mapshowingPopulus trichocarpaGWAS sub-population structure identified by DBSCAN clustering, which were derived from a UMAP projection of the top 8 PCs of LD-pruned pangenome SNP data. Geographic origins are searchable by genotype or river system using the search bar.

09 BIOMASS FUELS↗

Machine Learning–Based Condition Monitoring of a Circulating Water System of a Canadian Nuclear Plant

With the need to maintain long-term reliable energy using nuclear power plants, there is an underlying demand to ensure that the maintenance of plant components and systems is also done in an efficient and cost-effective manner. One way to achieve this is by moving from time-based maintenance to condition-based maintenance. The research presented in this paper focuses on applying statistical and machine-learning-based methods to capture anomalies within data for fault detection to further develop into condition monitoring. This paper focuses on system data for a circulating water system (CWS) of a pressurized heavy-water reactor for detecting anomalies. The different methodologies used for detecting and capturing anomalies in the CWS data are matrix profile, density-based spatial clustering of applications with noise (DBSCAN), and support vector machines (SVMs). Matrix profile and DBSCAN are used to distinguish between normal data and anomalous data. This paper presents a hybrid method using DBSCAN and SVM when a portion of the data is used for DBSCAN to generate clusters. This portion of data is then used to train the SVM along with the clusters generated by DBSCAN as output. SVM is then tested on unseen data as a predictive tool, which can work in real time to categorize data points as either normal or anomalous. This paper presents results that show the high accuracies of DBSCAN and SVM in capturing anomalies within the data for a CWS for fault detection. Thus, the maintenance plan would be focused on component condition rather than a time-based schedule by switching to an automated system to identify and predict faults within a CWS.

21 SPECIFIC NUCLEAR REACTORS AND ASSOCIATED PLANTS↗

Advances in ArborX to support exascale applications

ArborX is a performance portable geometric search library developed as part of the Exascale Computing Project (ECP). In this paper, we explore a collaboration between ArborX and a cosmological simulation code HACC. Large cosmological simulations on exascale platforms encounter a bottleneck due to the in-situ analysis requirements of halo finding, a problem of identifying dense clusters of dark matter (halos). This problem is solved by using a density-based DBSCAN clustering algorithm. With each MPI rank handling hundreds of millions of particles, it is imperative for the DBSCAN implementation to be efficient. In addition, the requirement to support exascale supercomputers from different vendors necessitates performance portability of the algorithm. We describe how this challenge problem guided ArborX development, and enhanced the performance and the scope of the library. We explore the improvements in the basic algorithms for the underlying search index to improve the performance, and describe several implementations of DBSCAN in ArborX. Further, we report the history of the changes in ArborX and their effect on the time to solve a representative benchmark problem, as well as demonstrate the real world impact on production end-to-end cosmology simulations.

97 MATHEMATICS AND COMPUTING↗

Visualizing Corridors in Terminal Airspace using Trajectory Clustering

Context: Advances in battery and automation technology have made routine air taxi and cargo transport in urban areas a business model that can be attained by emerging aviation innovators. The community vision and work to enable these novel operations is discussed using the term ‘Urban Air Mobility’ or UAM. Small, piloted, airspace vehicles that fly with a few passengers do operate in urban areas today, and these vehicles can be studied as an early proxy for this future UAM traffic. Aim: We seek to identify corridors already in daily operation and their properties. Method: We applied DBSCAN and HDBSCAN to Dallas Forth-Worth TRACON flight data to identify corridors in use, their density, and devised a method to annotate landing sites used in these corridors with site metadata. Results: While DBSCAN was unable to group similar trajectories, we we were able to successfully identify corridors using HDBSCAN, measure their density and annotate them. Conclusion: The applied method can successfully identify corridors in daily operation with additional metadata to help domain expert understand the intent of UAM corridors.

UAM, Trajectory, TRACON, Clustering, DBSCAN, HDBSC↗

Visualizing Corridors in Terminal Airspace Using Trajectory Clustering

Context: Advances in battery and automation technology have made routine air taxi and cargo transport in urban areas a business model that can be attained by emerging aviation innovators. The community vision and work to enable these novel operations is discussed using the term ‘Urban Air Mobility’ or UAM. Small, piloted, airspace vehicles that fly with a few passengers do operate in urban areas today, and these vehicles can be studied as an early proxy for this future UAM traffic. Aim: We seek to identify corridors already in daily operation and their properties. Method: We applied DBSCAN and HDBSCAN to Dallas Forth-Worth TRACON flight data to identify corridors in use, their density, and devised a method to annotate landing sites used in these corridors with site metadata. Results: While DBSCAN was unable to group similar trajectories, we we were able to successfully identify corridors using HDBSCAN, measure their density and annotate them. Conclusion: The applied method can successfully identify corridors in daily operation with additional metadata to help domain expert understand the intent of UAM corridors.

UAM Trajectory, TRACON, Clustering, DBSCAN, HDBSCA↗

Machine Learning-Driven Quantification of CO2 Plume Dynamics at Illinois Basin Decatur Project Sites Using Microseismic Data

This study utilizes machine learning to quantify CO2 plume extents by analyzing microseismic data from the Illinois Basin Decatur Project (IBDP). Leveraging a unique dataset of well logs, microseismic records, and CO2 injection metrics, this work aims to predict the temporal evolution of subsurface CO2 saturation plumes. The findings illustrate that machine learning can predict plume dynamics, revealing vertical clustering of microseismic events over distinct time periods within certain proximities to the injection well, consistent with an invasion percolation model. The buoyant CO2 plume partially trapped within sandstone intervals periodically breaches localized barriers or baffles, which act as leaky seals and impede vertical migration until buoyancy overcomes gravity and capillary forces, leading to breakthroughs along vertical zones of weakness. Between different unsupervised clustering techniques, K-Means and DBSCAN were applied and analyzed in detail, where K-means outperformed DBSCAN in this specific study by indicating the combination of the highest Silhouette Score and the lowest Davies–Bouldin Index. The predictive capability of machine learning models in quantifying CO2 saturation plume extension is significant for real-time monitoring and management of CO2 sequestration sites. The models exhibit high accuracy, validated against physical models and injection data from the IBDP, reinforcing the viability of CO2 geological sequestration as a climate change mitigation strategy and enhancing advanced tools for safe management of these operations.

Iyegbekedo, Ikponmwosa↗

Explainable AI for Multivariate Time Series Pattern Exploration: Latent Space Visual Analytics With Temporal Fusion Transformer and Variational Autoencoders in Power Grid Event Diagnosis

Detecting and analyzing complex patterns in multivariate time-series data is crucial for decision-making in urban and environmental system operations. However, challenges arise from the high dimensionality, intricate complexity, and interconnected nature of complex patterns, which hinder the understanding of their underlying physical processes. Existing AI methods often face limitations in interpretability, computational efficiency, and scalability, reducing their applicability in real-world scenarios. This paper proposes a novel visual analytics framework that integrates two generative AI models, Temporal Fusion Transformer (TFT) and Variational Autoencoders (VAEs), to reduce complex patterns into lower-dimensional latent spaces and visualize them in 2D using dimensionality reduction techniques such as PCA, t-SNE, and UMAP with DBSCAN. These visualizations, presented through coordinated and interactive views and tailored glyphs, enable intuitive exploration of complex multivariate temporal patterns, identifying patterns’ similarities and uncover their potential correlations for a better interpretability of the AI outputs. The framework is demonstrated through a case study on power grid signal data, where it identifies multi-label grid event signatures, including faults and anomalies with diverse root causes. Additionally, novel metrics and visualizations are introduced to validate the models and assess the performance, efficiency, and consistency of latent maps generated by VAE, which have been utilized in prior studies for latent space cartography and used as a benchmark in this study, and the emerging TFT architecture under various configurations. These analyses provide actionable insights for model parameter tuning and reliability improvements. Comparative results highlight that TFT achieves shorter run times and superior scalability to diverse time-series data shapes compared to VAE. This work advances fault diagnosis in multivariate time series, fostering explainable AI to support critical system operations.

Explainable AI↗

ArborX 2.0

ArborX library tackles a problem of efficiently finding geometric objects that are close in space. Variations of this problem, such as finding the nearest neighbors of a point, or finding all objects within a certain distance, are inherent components of applications in many fields. The data may be large so that solving the problem efficiently may require significant computational resources, such as multiple processors or accelerators such as general purpose GPUs. ArborX' main advantage in its ability to solve large problems efficiently utilizing a combination of distributed and on-node parallelism. ArborX can be run efficiently on a wide variety of hardware, including GPUs from different vendors, which distinguishes it from other available libraries which typically choose only few of these. The other advantage is that it supports both types of user problems: spatial problems (useful for intersections and finding objects within certain distance), and nearest neighbor problems. ArborX also supports flexible interface in its interaction with a user. Particularly, it allows a user to call user's own function on a positive match, a functionality not rarely available in other libraries. ArborX implements construction and traversal algorithms using efficient tree structures, such as bounding volume hierarchy (BVH). At its core, ArborX uses linear BVH for its low construction cost and sufficient quality. ArborX implements both spatial and nearest-neighbor traversal algorithms. ArborX also provides several clustering algorithms (minimum spanning tree, DBSCAN, HDBSCAN*), interpolation using minimum least squares and ray tracing. ArborX is written using C++, and is parallelized using the message passing interface (MPI) for the distributed communication, and the Kokkos library for on-node parallelism. This approach allows ArborX to be run on a wide variety of hardware, from common laptops and desktops to supercomputers while using the same codebase.

Prokopenko, Andrey [Oak Ridge National Laboratory ↗

Frosted Tracks

SAND2025-01893O Frosted Tracks is a software tool to group trajectories according to sequences of their behavior. The goal is to start with a very large number of trajectories and identify groups that exhibit similar behavior patterns. The application combines TICC and Metric DBSCAN clustering algorithms for behavioral segmentation and labeling of air/sea trajectory data. Sandia National Laboratories is a multimission laboratory managed and operated by National Technology & Engineering Solutions of Sandia, LLC, a wholly owned subsidiary of Honeywell International Inc., for the U.S. Department of Energy’s National Nuclear Security Administration under contract DE-NA0003525.

Dalbey, Keith↗

Poplar

SAND2025-00683O Poplar is a software tool that generates a phylogenetic tree from input gene and genome sequences. It integrates established tools to identify genes within genomes, group sequences, construct gene trees, and infer a species tree. Poplar processes nucleotide sequences, identifies similar sequences using Nucleotide BLAST, groups them with DBSCAN, aligns sequences with MAFFT, constructs gene trees with RAxML-NG, and infers a species tree using ASTRAL-Pro3. This pipeline provides a structured approach to phylogenetic analysis, facilitating the study of evolutionary relationships among species. Sandia National Laboratories is a multimission laboratory managed and operated by National Technology & Engineering Solutions of Sandia, LLC, a wholly owned subsidiary of Honeywell International Inc., for the U.S. Department of Energy’s National Nuclear Security Administration under contract DE-NA0003525.

Krishnakumar, Raga↗

SCULPT (Supervised Clustering and Uncovering Latent Patterns with Training) v1

SCULPT (Supervised Clustering and Uncovering Latent Patterns with Training) is a comprehensive data visualization and analysis application focused on working with COLTRIMS (COLd Target Recoil Ion Momentum Spectroscopy) data, which is used in atomic and molecular physics experiments. The application offers several powerful features: - Data uploading and processing capabilities for COLTRIMS files - Multiple visualization methods using UMAP (Uniform Manifold Approximation and Projection) for dimensionality reduction - Interactive selection of data points across multiple views - Feature engineering through various methods: - Manual feature selection from calculated physics parameters - Deep autoencoder for dimension reduction - Genetic programming for discovering meaningful features - Mutual information-based feature selection - Multiple clustering approaches (DBSCAN, KMeans, Agglomerative) - Quality metrics for evaluating clustering results - Export capabilities for selections and generated features

Daoud, Hazem [Lawrence Berkeley National Laborator↗

Anomaly Detection On Network Traffic Using A Sequence Model Approach

The code ingests Zeek logs derived from network packet captures, applies a DBSCAN model, Gower distance to build a feature set as input into a transformer model that will output anomaly scores and provides the option to set a threshold as to what’s considered an anomaly.

Quach, Anna [Idaho National Laboratory (INL), Idah↗

Analysis of Superconducting Magnet Quench Antenna Data

Quenching poses a serious problem for superconducting magnets operating at high currents. It occurs when the material transitions from the superconducting to the normal state, which leads to heating and potential damage to the magnet. To understand and mitigate quenching, the Magnet Department at Fermilab is developing and testing superconducting magnet quench antenna arrays. This study delves into the anomalous events preceding the quench during magnet training by analyzing the collected data. With the moving average and Fast Fourier Transform techniques, we investigate the trends and frequency patterns of the data. Moreover, we introduce an unsupervised anomaly detection algorithm based on Principal Component Analysis and DBSCAN clustering. It can autonomously identify events within background noise, without relying on any predefined event features. Our analysis reveals that the spatio-temporal distribution of these anomalous events has little connection to the quench location, indicating that a majority of them bear no relation to the quenching process.

43 PARTICLE ACCELERATORS↗

Clustering Acoustic Background Noise in the Stratosphere Using Machine Learning

Infrasound, characterized by low-frequency sound inaudible to humans (<20 Hz), emanates from natural and anthropogenic sources. Its efficacy for monitoring phenomena necessitates robust sensing networks. Traditional ground-based infrasound sensors have limitations due to atmospheric dynamics and noise interference. Balloon-bore sensors have emerged as an alternative, offering reduced noise and improved capabilities. This study bridges clustering algorithms with balloon borne infrasound data, a domain yet to be explored. Employing K-Means, DBSCAN, and GMM algorithms on normalized and reshaped data and only normalized data from a New Zealand-based NASA balloon flight, insights into background noise at stratospheric altitudes were revealed. Despite challenges arising from distinguishing signals amid unique background noise, this research provides vital reference material for noise analysis and calibration. Beyond infrasound event capture, the dataset enriches comprehension of background noise characteristics in the southern hemisphere.

47 OTHER INSTRUMENTATION↗

FY 2026 Midyear Report: Seismic Monitoring of Underground Vibration Sources Using Distributed Acoustic Sensing and Seismometers

Safeguards-relevant temporal changes in underground facilities can be observed using geophysical monitoring techniques. Seismic waves, in particular, provide valuable insights into subsurface activities and can serve as an important tool for detecting anomalous events that may indicate containment breaches at geological repositories. This midyear report summarizes ongoing efforts to automatically and rapidly detect and locate anomalous vibration signals that could be indicative of potential containment breaches. Previous work during FY25 focused on compiling continuous seismic datasets from two underground sites and developing a database of continuous waveforms and ground-truth event data derived from multiple sensing modalities. Building on this foundation, we are adapting anomaly detection and geolocation algorithms to explore methods for monitoring underground activities using two relatively low-maintenance sensing technologies: a dense surface geophone array deployed at the Pleasant Gap mine in Pennsylvania, and a three-dimensional fiber-optic cable array for distributed acoustic sensing (DAS) installed in the subsurface at the Sanford Underground Research Facility (SURF) in South Dakota. This report summarizes work conducted during the first two quarters of FY26, during which we refined a dynamic power spectral density (PSD)-based detector, applied it independently to each geophone station, and then combined the per‑station detections with density-based spatial clustering of applications with noise (DBSCAN) to cluster events and produce spatial maps over a nine‑day interval. In addition, we outline plans for a field trial at the Waste Isolation Pilot Plant (WIPP) in New Mexico to compare traditional seismic monitoring approaches with DAS techniques and to evaluate the benefits of combined data analysis. Activities during the past two quarters have included the preparation and submission of a Field Test Plan to WIPP for approval, as well as submission to headquarters for review and feedback.

58 GEOSCIENCES↗