Search NASA⌕ Search

SEARCH · Search NASA

Results for “data volume”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 37 records · Page 2

SCILLA Secondary Aerosol Volume Concentration Airborne Data

This data set contains secondary aerosol volume concentration in cubic micrometers per cubic centimeter (um^3/cm^3). The aerosol was generated in an oxidation flow reactor (OFR) and its volume concentration measured with a scanning mobility particle sizer (SMPS). Both components were designed and built at the University of California Riverside. Ambient particles were removed with a Teflon filter upstream of the OFR and then pure ammonium sulfate particles were added to create a stable aerosol surface area on which low-volatility oxidation products condense. High concentrations of hydroxyl radical (OH) formed inside the OFR accelerate the oxidative chemistry that would typically occur over a period of several days in the atmosphere, resulting in the production of secondary aerosol from precursor gases present in the ambient air. Concentrations of added ozone and water vapor were controlled to produce the desired and approximately constant level of photochemical aging. The reactor temperature was controlled, while the pressure was not, and was always slightly lower than that of the sampled outside air. Sampled air was pulled from above the Naval Postgraduate School’s (NPS) Twin Otter aircraft through a rear-facing, ¼” OD PFA Teflon tube. The sample stream was split between the OFR and several gas analyzers (NOx, CO, O3, and H2O). Though not contained in these files, data from an aerosol mass spectrometer intermittently operated downstream of the OFR are also available.

{"secondary aerosol concentration",aerosol_concent↗

Validating the galaxy and quasar catalog-level blinding scheme for the DESI 2024 analysis

In the era of precision cosmology, ensuring the integrity of data analysis through blinding techniques is paramount — a challenge particularly relevant for the Dark Energy Spectroscopic Instrument (DESI). DESI represents a monumental effort to map the cosmic web, with the goal to measure the redshifts of tens of millions of galaxies and quasars. Given the data volume and the impact of the findings, the potential for confirmation bias poses a significant challenge. To address this, we implement and validate a comprehensive blind analysis strategy for DESI Data Release 1 (DR1), tailored to the specific observables DESI is most sensitive to: Baryonic Acoustic Oscillations (BAO), Redshift-Space Distortion (RSD) and primordial non-Gaussianities (PNG). We carry out the blinding at the catalog level, implementing shifts in the redshifts of the observed galaxies to blind for BAO and RSD signals and weights to blind for PNG through a scale-dependent bias. We validate the blinding technique on mocks as well as on data by applying a second blinding layer to perform a series of sanity checks; the latter allows probing complexities in real data not captured in mocks. We find that the blinding strategy alters the data vector in a controlled way, and the BAO and RSD analysis choices are robust to blinding. The successful validation of the blinding strategy paves the way for the unblinded DESI DR1 analysis, alongside future blind analyses with DESI and other surveys.

79 ASTRONOMY AND ASTROPHYSICS↗

Unsupervised Segmentation and Clustering Workflow for Efficient Processing of 4D-STEM and 5D-STEM Data

Four-dimensional scanning transmission electron microscopy (4D-STEM) enables mapping of diffraction information with nanometer-scale spatial resolution, offering detailed insight into local structure, orientation, and strain. However, as data dimensionality and sampling density increase, particularly for in situ scanning diffraction experiments (5D-STEM), robust segmentation of structurally consistent behavior across sequential measurements becomes essential for efficient and physically meaningful analysis. Here, we introduce a clustering framework that identifies crystallographically distinct domains from 4D-STEM datasets. By using local diffraction-pattern similarity as a metric, the method extracts closed contours delineating spatially contiguous regions. This approach produces cluster-averaged diffraction patterns that improve signal quality while reducing data volume by orders of magnitude, enabling rapid and accurate orientation, phase, and strain mapping. We demonstrate the applicability of this approach to in situ liquid-cell 4D-STEM data of gold nanoparticle growth. Our method provides a scalable and generalizable route for spatially coherent segmentation, data compression, and quantitative structure–strain mapping across diverse 4D-STEM modalities. The full analysis code and example workflows are publicly available to support reproducibility and reuse.

4D-STEM↗

A Versatile Simulated Data Transport Layer for in Situ Workflows Performance Evaluation

In situ processing does not only allow scientific applications to face the explosion in data volume and velocity but also to address the time constraints of many simulation-analysis workflows by providing scientists with early insights about their applications at runtime. Multiple frameworks implement the concept of a data transport layer (DTL) to enable such in situ workflows. These tools are very versatile, directly or indirectly access the data generated on the same node, another node of the same compute cluster, or a completely distinct node, and allow data publishers and subscribers to run on the same computing resources or not. This versatility puts on researchers the onus of taking key decisions related to resource allocation and how to transport data to ensure the most efficient execution of their in situ workflows. However, domain scientists and workflow practitioners lack the appropriate tools to assess the respective performance of particular design and deployment options. In this paper we introduce a versatile simulated DTL designed to provide researchers with insights on the respective performance of different execution scenarios of in situ workflows. This open-source, standalone library builds on the SimGrid toolkit and can be linked to any SimGrid-based simulator. It facilitates the evaluation of the performance behavior, at scale, of different data transport configurations and the study of the effects of resource allocation strategies. We demonstrate the scalability, versatility, and accuracy of this simulated DTL by reproducing the execution of two synthetic benchmarks and of a real-world in situ workflow composed of an MPI application and a parallel data analysis. Results of simulations run on a single core show that the proposed library can simulate the interactions of tens of thousands of simulated processes deployed on two interconnected commodity clusters in a few seconds, and the execution by a thousand simulated processes of an in situ workflow in less than three minutes.

Suter, Fred [ORNL] (ORCID:0000000319021955)↗

Latency Analysis of the Nexus Digital Twin Framework

Real-time digital catalogs are increasingly relied upon to track metadata and connect disparate data sources for cloud-based data integration efforts. One such tool, Deeplynx Nexus is supporting real-time digital twin efforts through event-driven data integration and time-series queries. Nexus’s usefulness for these applications depends critically on how quickly individual records can be uploaded and downloaded, since delays directly affect the responsiveness of any system built on top of it. However, the actual latency a user should expect from Nexus has not been systematically measured before, particularly for the small, frequent transactions typical of live sensor feeds. Here we show that single-record round-trip latency is 61.1 ms on a local Nexus instance and 391.7 ms on the hosted production infrastructure, a roughly 6.4x difference driven primarily by fixed per-request overhead rather than data volume. This overhead dominates at small scale: comparing single-record and ten-record trials suggests approximately 56 ms of each single-record request is fixed connection and authentication cost rather than data-transfer time, meaning batching even a handful of records is substantially more efficient than transmitting them individually. At large batch sizes, this pattern reverses for uploads, which converge to near parity between local and hosted environments by 25,000-50,000 records, while download latency remains persistently 5.7-6.4x slower on hosted infrastructure even at scale. These results suggest that Nexus deployments intended for real-time digital twin applications should prioritize record batching over single-record transactions, and that download-path optimization on hosted infrastructure offers the largest remaining opportunity to reduce latency at scale. We anticipate these baseline measurements will serve as a reference point for future digital twin projects evaluating whether Nexus’s latency profile meets their real-time requirements, and as a benchmark for tracking the effect of future infrastructure or API changes.

99 - GENERAL AND MISCELLANEOUS↗

Efficient Anomaly Detection Driven By Different Machine Learning Architectures And Models

The rapid growth and ubiquitous adoption of the internet and cyber-physical systems (CPS) have fundamentally transformed modern communication, work, and human-system interactions. While networks now form the backbone of critical digital ecosystems, enabling seamless data transmission across diverse, interconnected systems, this increased connectivity also expands the attack surface, making real-time detection of network intrusions and anomalies a pressing challenge. Detecting unusual activities within network infrastructure requires advanced data traffic analysis to differentiate between legitimate and malicious interactions. Traditional approaches to network anomaly detectionâ??such as rule-based and signature-based systemsâ??often depend on predefined patterns to identify known anomalies, limiting their effectiveness against emerging, stealthy, or previously unseen threats. These conventional methods suffer from high false alarm rates and fail to adapt to the ever-evolving nature of network traffic, particularly in large-scale, decentralized environments where data volume, velocity, and variety are constantly increasing. This dissertation presents artificial intelligence (AI)-driven approaches to anomaly detection that leverage graphics processing unit (GPU)-enabled high-performance computing (HPC) platforms for processing massive network traffic data and monitoring the components of cyber-physical systems (CPS) for potentially hazardous conditions. The research advances several key contributions: (1) Designing efficient machine learning techniques for CPS condition monitoring and anomaly detection; (2) enabling federated learning (FL) frameworks that enable distributed detection while preserving data privacy and system resilience; (3) exploring graph-based methodologies combining graph neural networks (GNN) and graph machine learning (ML) approaches for the Internet of Things (IoT) and automotive network security, and (4) performing distributed edge computing optimizations that integrate FL with scalable technologies for reduced communication overhead. Through extensive experiments, these methodologies demonstrate that complex anomaly detection and condition monitoring tasks can be achieved while balancing computational efficiency and detection accuracy through fine-grained network information processing. The frameworks developed in this research establish a robust foundation for network anomaly detection, providing scalable, adaptive, and privacy-preserving solutions for safeguarding CPS and IoT networks in an increasingly interconnected digital landscape. The practical implications of these research findings are significant, as they can inform the development of next-generation network security systems and contribute to the protection of critical infrastructure against sophisticated cyber attacks.

Marfo, William↗

Even Higher-Level Synthesis: An Exploration of AI Hardware Accelerators using HLS4ML

With the rise of artificial intelligence, the popularization of deep learning, and a constantly evolving industry, the demand for flexible and efficient tools has never been greater. As algorithms grow more complex, their runtime and energy consumption increase exponentially. Customized hardware accelerators, long used for specific mathematical operations, remain essential for managing modern applications' computational and power demands. Hardware accelerators can speed up complex computations by orders of magnitude, but their manual design and verification processes are often challenging and time-consuming. High-Level Synthesis (HLS) provides a solution by transforming high-level algorithm descriptions, typically written in C++ or SystemC, into synthesizable RTL suitable for hardware implementation. This approach reduces development time for RTL engineers while offering flexibility beyond what traditional handwritten RTL can provide. We extended this capability to the machine-learning domain with the open-source framework hls4ml, which allows neural networks trained in Python frameworks like Tensorflow or PyTorch to be synthesized into efficient hardware representations for the traditional FPGA and ASIC flows. This breakthrough addresses the growing need for reduced design turnaround and easy verification of ML hardware accelerators with low latency and power efficiency constraints. During this tutorial, we will demonstrate how Python complements HLS by simplifying the ML design process, bridging the gap between software and hardware development. Attendees will explore how we translate neural networks modeled in Python into fixed-point C++ models suitable for HLS workflows. We will dive into strategies like Value-Range Analysis and Quantization-Aware Training, which optimize these designs for deployment and evaluate their accuracy, power consumption, and energy efficiency. To exemplify these concepts, experts from Fermilab will share their experiences applying this technology to high-energy physics experiments, where real-time, low-latency processing is critical. Over the years, Fermilab engineers have demonstrated how deep neural networks, optimized for hardware using hls4ml, can meet the stringent requirements of trigger systems at the CERN Large Hadron Collider. These systems rely on rapid decision-making to process immense data volumes while retaining only the most relevant events for further analysis. The application of hls4ml has also been extended to innovative technologies like smart pixel arrays. These smart pixels integrate ML inference capabilities directly into sensor devices, enabling localized data processing at the pixel level. This approach drastically reduces the need to transmit raw data to external processing units, significantly decreasing power consumption and latency. By embedding neural networks within the pixel architecture, the smart pixels can identify and prioritize relevant data in real time, providing a highly efficient solution for edge computing in scenarios such as particle detectors and imaging systems. Fermilab's work highlights the potential of hardware-accelerated ML in scenarios where both speed and power efficiency are mission-critical. Through this tutorial, attendees will gain valuable insights into the challenges and solutions of deploying ML in hardware. Understanding how HLS and hls4ml streamline the development of neural network-based hardware accelerators is fundamental for the industry's future. Participants will learn how these technologies are shaping the future of AI and scientific computing.

Di Guglielmo, Giuseppe [Fermilab]↗

ARMing the Edge: Demonstration of Edge Computing Field Campaign Report

Edge computing enables “next-to-instrument” control and intelligent data volume reduction and the potential for autonomous, adaptive measurement strategies such as for automated control of scan strategies for Doppler lidar (DL). Instruments with narrow bandwidth connections (e.g., ship and remote sites) can do scene determination and save phenomenon-appropriate data. For example, Doppler spectrum can be saved when clouds are detected by the instrument or automatic moment detection can take place in camera images and only preserve spectrum when non-monomodal spectra are detected. Automated control at the edge involves changing the sampling (temporal or scanning strategy) of an instrument to suit the phenomena both present and being studied (Jackson et al. 2020). Both data processing and instrument control introduces the possibility of a software-defined instrument.

54 ENVIRONMENTAL SCIENCES↗

NEWTS Economic Data Dashboard: Critical Materials for Energy

The NEWTS Economic Data Dashboard: Critical Materials for Energy is an economic screening tool for assessing the concentration and potential value of the 18 critical materials for energy in fossil energy-related wastewater across the United States. Datasets used to develop the dashboard and complete economic calculations are available as supplementary downloads. These resources were developed primarily using geochemical composition and volume data from the NEWTS Integrated dataset (version 1.0). Energy-related wastewater types presented in the dashboard include produced water (PW), brackish groundwater (BW), acid mine drainage (AMD), coal combustion residual leachate (CCRL), power plant flue gas desulfurization wastewater (FGD), and geothermal fluids. The concentrations of the following critical minerals were included in the analysis, when available: Al, Co, Cu, Dy, F, Ga, Ge, C, Ir, Li, Mg, Mn, Nd, Ni, Pt, Pr, Si, Tb. This economic screening tool was built to support identification of promising critical mineral feedstocks and economic research targets.

Critical Minerals; Critical Minerals and Materials↗

Smart pixel sensors: towards on-sensor filtering of pixel clusters with deep learning

Highly granular pixel detectors allow for increasingly precise measurements of charged particle tracks. Next-generation detectors require that pixel sizes will be further reduced, leading to unprecedented data rates exceeding those foreseen at the High- Luminosity Large Hadron Collider. Signal processing that handles data incoming at a rate of $\mathcal{O}$(40 MHz) and intelligently reduces the data within the pixelated region of the detector at rate will enhance physics performance at high luminosity and enable physics analyses that are not currently possible. Using the shape of charge clusters deposited in an array of small pixels, the physical properties of the traversing particle can be extracted with locally customized neural networks. In this first demonstration, we present a neural network that can be embedded into the on-sensor readout and filter out hits from low momentum tracks, reducing the detector's data volume by 57.1%–75.7%. The network is designed and simulated as a custom readout integrated circuit with 28 nm CMOS technology and is expected to operate at less than 300 μW with an area of less than 0.2 mm 2 . The temporal development of charge clusters is investigated to demonstrate possible future performance gains, and there is also a discussion of future algorithmic and technological improvements that could enhance efficiency, data reduction, and power per area.

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND ↗

A Science Gateway for the Repeatable Analysis of Machine Learning Predicted Gravity Anomalies

In recent years, deep learning has become an increasingly popular alternative for modeling in geoscience applications due to its scalability and efficiency. However, the interpretability, compute, data volume, and hyperparameter tuning requirements of deep learning models make development and monitoring difficult. Furthermore, model explainability and communicating results obtained by these models to users or domain experts is a challenge, as domain experts in geoscience also need to have a deep understanding of how those models function in order to support their scientific works. Here, we describe a science gateway and machine learning pipeline for predicting gravity anomalies from geophysical data. The gateway, built on open-source technologies, provides a holistic view of the pipeline through interactive visualizations aimed at enabling efficient exploratory data analysis. The repeatability, reproducibility, and monitoring capabilities of this overall system allow us to iterate and analyze at scale. Using this pipeline and gateway, we can repeatedly produce accurate high-resolution gravity anomaly datasets. By describing the underlying technologies, implementation, and results, here we provide a foundation for the broader adoption of science gateways into cross-cutting geoscience and machine learning research projects as a means to improve the scientific discovery and collaboration in the geophysics and computational sciences community.

58 GEOSCIENCES↗

Final technical report for DE-SC0022255: Discovering Physically Meaningful Structures from Climate Extreme Data

The past two decades have witnessed natural disasters and extreme weather events that affect millions of people. At the same time, the data volume from high-resolution climate models, satellite, in-situ and ground-based measurements have substantially increased to petabyte scales. These new and readily accessible datasets create the previously missing pipeline required for scientific machine learning (ML) and therefore new opportunities for improved understanding and prediction capability of climate extreme events. This project developed a deep latent variable model framework to discover physically meaningful hidden structures from high-dimensional, spatiotemporal climate extreme data.

97 MATHEMATICS AND COMPUTING↗

Profiles of Radiative Fluxes at ENA

Profiles of radiative fluxes observed at the Atmospheric Radiation Measurement (ARM)’s Eastern North Atlantic (ENA) observatory along with the ancillary measurements are reported. The below-cloud drizzle properties were derived by combining the data from the ceilometer and Ka-band ARM Zenith Radar (KAZR) following the technique explained by Ghate et al. (2021 JAMC). The cloud and drizzle water path values were derived from the brightness temperatures reported by the microwave radiometer following the technique of Cadeddu et al. (2020 AMT). The cloud water path was then scaled to the KAZR-reported radar reflectivity to calculate profiles of liquid water content (LWC). Following the analysis from Ghate et al. (2023 JGR), cloud droplet effective radius was calculated using the number concentration value of 100 cm-3. The cloud properties, along with the thermodynamic properties, served as an input to the Rapid Radiative Transfer Model (RRTM) to yield profiles of radiative fluxes at a 1-minute temporal and 50-m vertical resolution. The fluxes were then averaged to hourly temporal resolution for analysis. In Mitra et al. (2025 JClim), the calculated profiles were compared against those derived from the satellite measurements (SYN1deg). Flux profiles from the SYN1deg and the thermodynamic and cloud properties used for deriving them are also reported here. Both all-sky and clear-sky radiative flux profiles were calculated. Due to the large data volume, the surface and top-of-atmosphere (TOA) radiative fluxes for the six-year period, and the hourly profiles of the radiative fluxes for January 2018, are submitted here. Full profiles of radiative fluxes calculated from the thermodynamic and cloud properties measured at the ENA site at 1-minute temporal and 50-m vertical resolution for a six-year period are available from the authors. Six files here correspond to the following data: 1_ENARAD_CERES_with_cld_amount_timeseries.nc: Time-series of hourly values of RRTM-simulated values of upwelling and downwelling fluxes at the surface and TOA, observed boundary-layer cloud fractions, and upwelling and downwelling fluxes from the SYN1deg from July 2015 to January 2022. 2_CERES_2018_at_CERES_levels.nc: SYN1deg radiative fluxes at six levels for the year 2018. 3_ENARad_2018_at_CERES_levels.nc: RRTM calculated fluxes at the SYN1deg vertical levels for the year 2018. 4_ENARad_rrtminputs_hourly_201801.nc: Thermodynamic and cloud properties used as an input to the RRTM for January 2018. 5_CERES_inputs_hourly_201801.nc: Thermodynamic and cloud properties utilized by SYN1deg algorithm for January 2018. 6_ENARAD_hourly_201801.nc: Full profiles of hourly averaged radiative fluxes from the RRTM simulations for January 2018.

Atmosphere↗

W-Band ARM Scanning Cloud Radar (WSACR) 2nd Generation CF-Radial Spectral Data, Vertically-Pointing Scan, Cross Mode (a1)

ARM's scanning cloud radars are fully coherent dual-frequency, dual-polarization Doppler radars mounted on a common scanning pedestal. Each pedestal includes a Ka-band radar (2kW peak power) and the deployment location determines whether the second radar is a W-band (WSACR; 1.7 kW peak power) or X-band (XSACR; 20 kW peak power). Beamwidths at Ka and W bands are roughly matched at 0.3 degrees. Due to the narrow antenna beamwidth, ARM’s scanning cloud radars use scanning strategies that are unlike typical weather radars. Rather than focusing on plan position indicator, or PPI, scans, the Ka-SACR uses range height indicator, or RHI, scans at numerous azimuths to obtain cloud volume data. Measurements collected with the W-SACR are copolar and cross-polar radar reflectivity, Doppler velocity, spectra width and spectra when not scanning, and linear depolarization ration.

54 ENVIRONMENTAL SCIENCES↗

W-Band ARM Scanning Cloud Radar (WSACR) 2nd Generation CF-Radial Spectral Data, Vertically-Pointing Scan, Co-Polarization Mode (a1)

ARM's scanning cloud radars are fully coherent dual-frequency, dual-polarization Doppler radars mounted on a common scanning pedestal. Each pedestal includes a Ka-band radar (2kW peak power) and the deployment location determines whether the second radar is a W-band (WSACR; 1.7 kW peak power) or X-band (XSACR; 20 kW peak power). Beamwidths at Ka and W bands are roughly matched at 0.3 degrees. Due to the narrow antenna beamwidth, ARM’s scanning cloud radars use scanning strategies that are unlike typical weather radars. Rather than focusing on plan position indicator, or PPI, scans, the Ka-SACR uses range height indicator, or RHI, scans at numerous azimuths to obtain cloud volume data. Measurements collected with the W-SACR are copolar and cross-polar radar reflectivity, Doppler velocity, spectra width and spectra when not scanning, and linear depolarization ration.

54 ENVIRONMENTAL SCIENCES↗

W-Band ARM Scanning Cloud Radar (WSACR) 2nd Generation CF-Radial Spectral Data, Vertically-Pointing Scan, Cross-Polarization Model (a1)

ARM's scanning cloud radars are fully coherent dual-frequency, dual-polarization Doppler radars mounted on a common scanning pedestal. Each pedestal includes a Ka-band radar (2kW peak power) and the deployment location determines whether the second radar is a W-band (WSACR; 1.7 kW peak power) or X-band (XSACR; 20 kW peak power). Beamwidths at Ka and W bands are roughly matched at 0.3 degrees. Due to the narrow antenna beamwidth, ARM’s scanning cloud radars use scanning strategies that are unlike typical weather radars. Rather than focusing on plan position indicator, or PPI, scans, the Ka-SACR uses range height indicator, or RHI, scans at numerous azimuths to obtain cloud volume data. Measurements collected with the W-SACR are copolar and cross-polar radar reflectivity, Doppler velocity, spectra width and spectra when not scanning, and linear depolarization ration.

54 ENVIRONMENTAL SCIENCES↗

Anomaly Detection and Approximate Similarity Searches of Transients in Real-time Data Streams

Abstract We present Lightcurve Anomaly Identification and Similarity Search ( LAISS ), an automated pipeline to detect anomalous astrophysical transients in real-time data streams. We deploy our anomaly detection model on the nightly Zwicky Transient Facility (ZTF) Alert Stream via the ANTARES broker, identifying a manageable ∼1–5 candidates per night for expert vetting and coordinating follow-up observations. Our method leverages statistical light-curve and contextual host galaxy features within a random forest classifier, tagging transients of rare classes ( spectroscopic anomalies), of uncommon host galaxy environments ( contextual anomalies), and of peculiar or interaction-powered phenomena ( behavioral anomalies). Moreover, we demonstrate the power of a low-latency (∼ms) approximate similarity search method to find transient analogs with similar light-curve evolution and host galaxy environments. We use analogs for data-driven discovery, characterization, (re)classification, and imputation in retrospective and real-time searches. To date, we have identified ∼50 previously known and previously missed rare transients from real-time and retrospective searches, including but not limited to superluminous supernovae (SLSNe), tidal disruption events, SNe IIn, SNe IIb, SNe I-CSM, SNe Ia-91bg-like, SNe Ib, SNe Ic, SNe Ic-BL, and M31 novae. Lastly, we report the discovery of 325 total transients, all observed between 2018 and 2021 and absent from public catalogs (∼1% of all ZTF Astronomical Transient reports to the Transient Name Server through 2021). These methods enable a systematic approach to finding the “needle in the haystack” in large-volume data streams. Because of its integration with the ANTARES broker, LAISS is built to detect exciting transients in Rubin data.

79 ASTRONOMY AND ASTROPHYSICS↗

A Semi-Supervised Learning Method for the Identification of Bad Exposures in Large Imaging Surveys

As the data volume of astronomical imaging surveys rapidly increases, traditional methods for image anomaly detection, such as visual inspection by human experts, are becoming impractical. We introduce a machine-learning-based approach to detect poor-quality exposures in large imaging surveys, with a focus on the DECam Legacy Survey (DECaLS) in regions of low extinction (i.e., E ( B − V ) < 0.04 ). Our semi-supervised pipeline integrates a vision transformer (ViT), trained via self-supervised learning (SSL), with a k-Nearest Neighbor (kNN) classifier. We train and validate our pipeline using a small set of labeled exposures observed by surveys with the Dark Energy Camera (DECam). A clustering-space analysis of where our pipeline places images labeled in good and bad categories suggests that our approach can efficiently and accurately determine the quality of exposures. Applied to new imaging being reduced for DECaLS Data Release 11, our pipeline identifies 780 problematic exposures, which we subsequently verify through visual inspection. Being highly efficient and adaptable, our method offers a scalable solution for quality control in other large imaging surveys.

Luo, Yufeng (ORCID:0000000246230683)↗