Search NASASearch

SEARCH · Search NASA

Results for “Streaming data”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 127 records · Page 7

The Sensor Dilemma in Intelligent Transportation Systems: Evaluating Radar, Lidar and Camera: Preprint

Intelligent transportation systems (ITS) are at the forefront in advancing the way we interact with and perceive the transportation network. This revolution is fueled by the significant advancement in sensor perception technologies such as radar, lidar, and video imaging, which are the most popular modalities for ITS. Real-time perception data from these sensors allow intelligent infrastructure-side decision-making to improve the energy, efficiency, and safety at traffic intersections. As traffic departments across the United States transition from traditional loop detectors and emulators and embrace newer technologies, they are often left with a dilemma in choosing a sensor technology for infrastructure-based perception that is reliable, inexpensive, and easy to set up and that has robust performance in varying weather conditions. However, choosing a sensor that checks all these boxes is not straightforward, as every sensor type has unique benefits and drawbacks. Radar is excellent at detecting long-range vehicles and weather resistance but lacks high resolution. Lidar is expensive and weather-sensitive, while cameras provide rich visual data at a low cost but are constrained by lighting and visibility. This study examines radar, lidar, and camera sensor capabilities to ascertain whether any of these qualifies as the "best" sensor for ITS perception. While no single sensor can meet all the demands of ITS, a hybrid approach combining multiple sensor modalities like radar, lidar, and cameras offers the most robust solution for enhancing the safety and efficiency of ITS. Through this evaluation, we hope to draw attention to the necessity of the National Renewable Energy Laboratory's infrastructure perception and control framework, which presents a multisensor track data fusion engine to assimilate multiple data streams in order to provide robust and reliable perception.

33 ADVANCED PROPULSION SYSTEMS

Neutrino beam bunch structure reconstruction with precision timing in the ICARUS liquid argon time projection chamber

The ICARUS detector has been operating smoothly since 2021 as the far detector in the Short Baseline Neutrino (SBN) program at Fermilab, collecting neutrino interactions from both the Booster Neutrino Beam (BNB) and off-axis from the Neutrinos at the Main Injector (NuMI) beam. Analysis of neutrino interactions in ICARUS requires mitigation of substantial cosmogenic backgrounds. This is achieved by using an external Cosmic Ray Tagger (CRT) and a Photomultiplier Tube (PMT) system embedded in the liquid argon. The intrinsic neutrino beam bunch structure, inherited time structure from the Radio Frequency (RF) system used to accelerate the protons, can be resolved at the ICARUS detector using precise timing information. Located at shallow depth, ICARUS is exposed to a high flux of cosmic rays that can be mistaken for neutrino interactions. To mitigate this background, the CRT and a 3-meter-thick concrete overburden were installed. To better model backgrounds, ICARUS makes use of an overlay technique where simulated neutrino events are superimposed on detector beam-off data. PMTs installed within a Time Projection Chamber detect argon scintillation light emitted by high energy charged particles passing through the chamber and provide the event timing of neutrino interactions. The nanosecond-level beam bunch structure is reconstructed with the PMT system and can be used to further understand backgrounds and enhance neutrino physics capabilities. In this thesis, I will discuss background mitigation techniques using precision timing and present a novel technique to select neutrino events from our unbiased data stream using the beam bunch structure.

Heggestuen, Anna [Colorado State U.] (ORCID:000000

EXCLUSIVE NEUTRAL PION ELECTROPRODUCTION CROSS SECTION MEASUREMENTSWITHANEUTRALPARTICLE SPECTROMETER

Deep Virtual Compton Scattering (DVCS), the exclusive electron-proton scattering process ep ¿e'p'¿, provides access to generalized parton distributions (GPDs), which correlate information about the longitudinal momentum and transverse spatial structure of quarks inside the nucleon. Experiment E12-13-010 in Hall C at Jefferson Lab was designed to take high-precision measurements of the DVCS cross section over an extended kinematic range using the newly commissioned Neutral Particle Spectrometer (NPS). The NPS features a high-resolution electromagnetic calorimeter and a streaming data acquisition system optimized for operation at high luminosities. This thesis presents the detector and analysis work carried out to support the NPS DVCS program. In particular, it focuses on the hardware design, calibration, and performance of the calorimeter. A development of a waveform reconstruction analysis of the calorimeter signals enabled improved extraction of pulse amplitudes and times. The waveform analysis was also extended to operate in a multithreaded environment, substantially reducing processing time for large datasets. Analysis of exclusive neutral pion electroproduction events in the calorimeter gives a strong validation of the calorimeter’s performance and resolution. Together these developments establish a foundation for future analyses and extraction of the DVCS cross section and its use in constraining the GPDs.

Kerver, Mitchell [Old Dominion Univ., Norfolk, VA

Kernelized approaches to streaming compression of scientific data

In this paper three algorithms are developed for the streaming compression of scientific data. The algorithms presented are reliant on the theory of vector-valued reproducing kernel Hilbert spaces and operator valued kernel. Further, the scientific data is modeled as a snapshot of time dependent vector field F(x, t) over a manifold M and the recovery of the data is framed as a learning problem. These processes are then appropriately modified and ana lyzed for the streaming scenario in which data is generated without the ability to revisit past entries.

97 MATHEMATICS AND COMPUTING

Signatures of a Tidally Induced Spiral Arm at the Anticenter of the Milky Way and a Kinematically Extended Anticenter Stream Using DESI Data Release 2

Using the Dark Energy Spectroscopic Instrument (DESI) Milky Way Survey, we examine the six-dimensional space of the anticenter region of the Milky Way stellar disk (150° < Galactic longitude < 220°) using 61,883 main-sequence turnoff stars. We focus on two well-known stellar overdensities in the anticenter: the Monoceros Ring (MRi) and Anticenter Stream (ACS). We find that the MRi overdensity has kinematic signatures consistent with a tidally induced spiral arm, a type of dynamic spiral arm created by an interaction with a satellite galaxy, most likely the Sagittarius dwarf spheroidal galaxy (Sgr). We use the kinematics of the MRi to calculate the two most recent passage times of Sgr, finding 0.25 ± 0.09 Gyr and 1.10 ± 0.23 Gyr from the present day. We validate that the ACS is kinematically decoupled from the MRi because they are moving in opposite radial and vertical directions. We find that the kinematics associated with the ACS extends beyond our defined overdensity. The features we see in the ACS region are likely part of a broader distribution of stars with the same kinematic signature as detected in other places, like the vertical wave in the outer disk and phase spiral.

Lambert, Mika [University of California, Santa Cru

Streaming Large-Scale Microscopy Data to a Supercomputing Facility

Data management is a critical component of modern experimental workflows. As data generation rates increase, transferring data from acquisition servers to processing servers via conventional file-based methods is becoming increasingly impractical. The 4D Camera at the National Center for Electron Microscopy generates data at a nominal rate of 480 Gbit s -1 (87,000 frames s -1 ⁠), producing a 700 GB dataset in 15 s. To address the challenges associated with storing and processing such quantities of data, we developed a streaming workflow that utilizes a high-speed network to connect the 4D Camera’s data acquisition system to supercomputing nodes at the National Energy Research Scientific Computing Center, bypassing intermediate file storage entirely. In this work, we demonstrate the effectiveness of our streaming pipeline in a production setting through an hour-long experiment that generated over 10 TB of raw data, yielding high-quality datasets suitable for advanced analyses. Additionally, we compare the efficacy of this streaming workflow against the conventional file-transfer workflow by conducting a postmortem analysis on historical data from experiments performed by real users. Our findings show that the streaming workflow significantly improves data turnaround time, enables real-time decision-making, and minimizes the potential for human error by eliminating manual user interactions.

4D-STEM

Stream discharge and temperature data collected within the East and Taylor Watershed, Colorado for the Lawrence Berkeley National Laboratory Watershed Function Science Focus Area (water years 2019 to 2025)

This dataset contains stream discharge and temperature data for water years 2019 to 2025 from the East and Taylor Watersheds in Colorado, United States. This data was collected to understand hydrological processes occurring in the East River and Taylor River Watersheds, Colorado, which is part of the Lawrence Berkeley National Laboratory Watershed Function Scientific Focus Area. Data includes instantaneous observed discharge using salt dilution and acoustic doppler velocimeter techniques, raw pressure transducer downloaded data, sub-hourly temperature as well as corrected water level and associated stream discharge and mean daily values. Notes on water level corrections, rating curve development and metadata provided. A rating curve is the translation of depth to streamflow. The rating curve can be used as a quantitative measure of the “quality of the data.” Data within this dataset is formatted using ESS-DIVE’s Hydrological Monitoring Reporting Format. This data package contains (1) a zip file (Stream_Discharge_Data_WY19-WY25.zip) containing stream discharge and temperature data organized by location; (2) an InstallationMethods file (InstallationMethods.csv) describing metadata about the installation; (3) a file-level metadata (flmd.csv) file that lists each file contained in the dataset with associated metadata; (4) a data dictionary (dd.csv) file that contains terms/column_headers used throughout the files along with a definition, units, and data type; (5) a locations metadata file (locations.csv); (6) and a sensor metadata file (sensors.csv). All data files are in non-proprietary formats (csv, png, or pdf formats). Please contact Rosemary Carroll, Curtis Beutler, or Austin Shirley for any support in accessing the files. Update on 2023-05-12: Additional data from WYs 2021 and 2022 were added. Additionally, the dataset was converted using ESS-DIVE’s Hydrological Monitoring Reporting Format. Data files were reformatted to match reporting format guidance, new metadata files were added, and files were converted from excel to CSV. Update on 2025-05-16: Additional data from WYs 2022 (for locations not previously included), 2023, and 2024 were added. An additional descriptive PDF (WFSFA_Streamflow_Hydrograph_Disclaimer.pdf) was added. Metadata files were updated to reflect the addition of new data and locations. Update on 2026-05-18: Additional data from WY 2025 were added, including a new location Upper Trail Creek (TR-TCG2). Metadata files were updated to reflect the addition of new data.

54 ENVIRONMENTAL SCIENCES

Levoglucosan data from five coastal streams impacted by the 2020 CZU Lightning Complex Fires, California, United States

This dataset includes levoglucosan data for five coastal California (United States) streams impacted by the 2020 CZU Lightning Complex Fires which burned from August 16th through September 22nd. Levoglucosan is a highly soluble and biolabile fraction of pyrogenic carbon. The five watersheds (San Lorenzo River, Pescadero Creek, Majors Creek, Laguna Creek, and Scott Creek) were impacted by the fires with watersheds experiencing a range of burn severity and extents. Grab samples were collected from each stream between October 2020 and May 2021, targeting both baseflow and event flow hydrologic conditions. Additional biogeochemistry data (i.e., organic and black carbon concentrations) can be found in a separate data package (https://doi.org/10.4211/hs.421c0226bb38460c8393d67fe0c4f802). This data package consists of one main data folder that contains (1) readme; (2) file-level metadata; (3) data dictionary; (4) field metadata with international generic sample numbers (IGSN); (5) methods codes; and (6) levoglucosan data. All files are .csv or .pdf.

2020 CZU Lightning Complex Fires

An Active Learning-Based Streaming Pipeline for Reduced Data Training of Structure Finding Models in Neutron Diffractometry

Structure determination workloads in neutron diffractometry are computationally expensive and routinely require several hours to many days to determine the structure of a material from its neutron diffraction patterns. The potential for machine learning models trained on simulated neutron scattering patterns to significantly speed up these tasks have been reported recently. However, the amount of simulated data needed to train these models grows exponentially with the number of structural parameters to be predicted and poses a significant computational challenge. To overcome this challenge, we introduce a novel batch-mode active learning (AL) policy that uses uncertainty sampling to simulate training data drawn from a probability distribution that prefers labelled examples about which the model is least certain. We confirm its efficacy in training the same models with ∼ 75% less training data while improving the accuracy. We then discuss the design of an efficient stream-based training workflow that uses this AL policy and present a performance study on two heterogeneous platforms to demonstrate that, compared with a conventional training workflow, the streaming workflow delivers ∼ 20% shorter training time without any loss of accuracy.

Wang, Tianle [Brookhaven National Laboratory (BNL)

Surface Water Quality Data from Beaver-Impacted Streams; Trail Creek and East River, Colorado 2025

This data package contains surface water chemistry measurements collected in 2025 to evaluate how beaver damming and low-tech process-based stream restoration influence water quality and metal mobility in mountainous headwater systems of the Upper Colorado River Basin. Sampling was conducted at Trail Creek (Taylor Park watershed, Colorado), a tributary undergoing restoration through installation of low-tech process-based structures (i.e., beaver dam analogs), and at off-channel beaver ponds within the East River floodplain (East River watershed, Colorado). Samples were collected along longitudinal transects spanning upstream control reaches, beaver-influenced ponded reaches, and downstream segments. Additional samples were collected from near-surface pore waters within a beaver dam seepage face. The dataset includes concentrations of major and trace elements measured by inductively coupled plasma–mass spectrometry (ICP-MS) and inductively coupled plasma–optical emission spectrometry (ICP-OES), major anions measured by ion chromatography (IC), and dissolved organic carbon (DOC; reported as non-purgeable organic carbon, NPOC). Samples were size-fractionated at 0.45 micrometers (µm), 0.22 µm, and 0.02 µm to distinguish particulate (>0.45 µm), colloidal (0.22–0.02 µm), and dissolved (<0.02 µm) fractions. The data package consists of comma-separated value (.csv) files containing tabulated chemical concentration data, sample metadata (site identifiers, geographic coordinates, sampling dates, fraction type), and quality control flags. All files are provided in open, non-proprietary formats that can be accessed using standard data analysis software such as Microsoft Excel, R, Python, MATLAB, or other programs capable of reading .csv files. Units, detection limits, and analytical methods are documented in accompanying metadata files. The dataset is designed to support analyses of (1) how beaver impoundment and restoration structures alter elemental partitioning and transport, (2) the role of iron and organic carbon in mediating trace metal mobility, and (3) reach-scale changes in water quality across restoration gradients. This work was supported by the Watershed Function Science Focus Area at Lawrence Berkeley National Laboratory funded by the US Department of Energy, Office of Science, Biological and Environmental Research under Contract No. DE-AC02-05CH11231. Part of this work was performed at SLAC Accelerator Laboratory funded by the US Department of Energy, Office of Science, Biological and Environmental Research under Contract No. DE-AC02-76SF00515.

Anions

PvaPy streaming framework for real-time data processing

User facility upgrades, new measurement techniques, advances in data analysis algorithms as well as advances in detector capabilities result in an increasing amount of data collected at X-ray beamlines. Some of these data must be analyzed and reconstructed on demand to help execute experiments dynamically and modify them in real time. In turn, this requires a computing framework for real-time processing capable of moving data quickly from the detector to local or remote computing resources, processing data, and returning results to users. In this paper, we discuss the streaming framework built on top of PvaPy, a Python API for the EPICS pvAccess protocol. We describe the framework architecture and capabilities, and discuss scientific use cases and applications that benefit from streaming workflows implemented on top of this framework. We also illustrate the framework's performance in terms of achievable data-processing rates for various detector image sizes.

EPICS pvAccess

Integration Development and Testing of Rear Transition Monitor for Beam Current Monitoring System

Addressing baseline effects in accelerator environments is crucial for accurate data acquisition and analysis, since baseline effects can obscure signal clarity and impact the reliability of beam current monitoring systems. There are many potential contributors to baseline noise, such as variations in beam dynamics, electromagnetic interference from nearby equipment, or RF interference. Previous applications of noise reduction systems don t sufficiently filter sources of asynchronous noise, so a new algorithm was implemented. A simulation dataset was created to replicate beam conditions and a Red Pitaya FPGA was used to collect data through the streaming application. A Python script was developed to implement noise reduction algorithms and efforts were made to integrate real-time data streaming with the Redis platform and Acnet Front End infrastructure.

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND

Data Assimilation for Robust UQ Within Agent-Based Simulation on HPC Systems

Agent-based simulation provides a powerful tool for in silico system modeling. However, these simulations do not provide built-in methods for uncertainty quantification (UQ). Within these types of models a typical approach to UQ is to run multiple realizations of the model then compute aggregate statistics. This approach is limited due to the compute time required for a solution. When faced with an emerging biothreat, public health decisions need to be made quickly and solutions for integrating near real-time data with analytic tools are needed. We propose an integrated Bayesian UQ framework for agent-based models based on sequential Monte Carlo sampling. Given streaming or static data about the evolution of an emerging pathogen this Bayesian framework provides a distribution over the parameters governing the spread of a disease through a population. These estimates of the spread of a disease may be provided to public health agencies seeking to abate the spread. By coupling agent-based simulations with Bayesian modeling in a data assimilation, our proposed framework provides a powerful tool for modeling dynamical systems in silico. We propose a method which reduces model error and provides a range of realistic possible outcomes. Moreover, our method addresses two primary limitations of ABMs: the lack of UQ and an inability to assimilate data. Our proposed framework combines the flexibility of an agent-based model with UQ provided by the Bayesian paradigm in a workflow which scales well to HPC systems. We provide algorithmic details and results on a simulated outbreak with both static and streaming data.

Spannaus, Adam [ORNL] (ORCID:0000000225213657)

Bayesian estimation of HIV acquisition dates for prevention trials

Accurate timing estimates of when participants acquire HIV in HIV prevention trials are necessary for determining antibody levels at acquisition. The Antibody-Mediated Prevention (AMP) Studies showed that a passively administered broadly neutralizing antibody can prevent the acquisition of HIV from a neutralization-sensitive virus. We developed a pipeline for estimating the date of detectable HIV acquisition (DDA) in AMP Study participants using diagnostic and viral sequence data. Using a Bayesian strategy that combines three streams of data (REN [rev/vpu/env/Δnef] sequence, GP [gag/Δpol] sequence, and diagnostic) where their 95% credible intervals overlap based on pre-specified criteria and decision rules. We evaluated the performance of our AMP pipeline using PacBio viral sequence data from 41 participants across two prospective acute HIV acquisition cohort studies, FRESH and RV217, with twice-weekly sampling. These cohort studies enrolled young women in South Africa and men and women in Kenya and Thailand, respectively, with a high likelihood of HIV acquisition. In evaluating performance, “true DDA” was the center of bounds between last-negative and first-positive RNA diagnostic tests (median time 4 days, range 2–7 days); bias was the mean difference between estimated and true DDA. Using diagnostic data alone yielded timing estimates with a bias of 2.4 days and root mean square error (RMSE) of 7.9 days. These results were improved using sequence + diagnostic data (bias 1.5 days, RMSE 6.9 days), as well as by restricting sequence-based estimation to samples from ≤5 weeks post-DDA (bias 0.2 days, RMSE 7.8 days).

59 BASIC BIOLOGICAL SCIENCES

Real Time implementation of Artificial Intelligence compression algorithm for High-Speed Streaming Readout signals

The new generation of high-energy physics experiments plans to acquire data in streaming mode. With this approach, it is possible to access the information of the whole detector (organized in time slices) for optimal and lossless triggering of data acquisitions. With this approach, data rates, especially in large detectors, are often very high, and the network is likely to be the bottleneck for the entire Streaming Read Out system. The aim of this work is to study the implementation of a lossy compression algorithm based on Artificial Intelligence: an Autoencoder. With Machine Learning it is possible to achieve a high compression ratio and fast inference time with only a small degradation of the signals, almost negligible for the specific application. This work explores different configurations of the Autoencoder and the implementation on different hardware. Different Autoencoder configurations are explored to find the best trade-off between compression ratio and reconstruction loss, both for signals and energy spectrum. Different hardware implementations are also explored to find the best platform to achieve real-time performance for the specific application.

Rossi, Fabio (ORCID:0009000385713885)

Joint Modeling of GD-1 and C-19 as Old Streams

DESI observational data for the GD-1 and C-19 streams are compared to stream simulations in a common evolving multi-halo potential of a Milky Way-like galaxy based on a cosmological simulation. The goal is to find the best match of the stream velocity spread and the density power spectrum stream density to simulations having either CDM or WDM subhalos. The cocoon velocity width integrated over the length of the stream is independent of orbital blurring along the stream and the power spectrum integrates over the width of the stream, sidestepping the geometric details of the streams. Streams develop from star clusters inserted at $\simeq$1 Gyr after the Big Bang and evolved for 13 Gyr to their current orbital positions. Streams in a CDM subhalo population provide the best match to the velocity width, with streams younger than 10 Gyr ruled out as insufficiently hot. The progenitor star cluster masses, which determine the fraction of stars released at late times which comprise the stream core, are found to be $\simeq 8\times 10^4 M_\odot$ for GD-1 and $\simeq 4\times 10^4 M_\odot$ for C-19, although the mass depends on the star cluster half mass radius. Stream heating leads to stream lumpiness which is measurable for the relatively large and clean GD-1 dataset. The stream density power spectrum measured along the length of the DESI GD-1 sample is in good agreement with CDM simulations, with 1.7 to 1.9 times more power than WDM 7 keV and 5.5 keV simulations.

Carlberg, Raymond G. [Toronto U.] (ORCID:000000027

Remote sensing images, DEM, and point clouds associated with “Accuracy evaluation of cost-effective 3D reconstruction approaches for hydrobiogeochemical processes in non-perennial stream riverbeds”

This data package is associated with the publication “Accuracy evaluation of cost-effective 3D reconstruction approaches for hydrobiogeochemical processes in non-perennial stream riverbeds” published in Frontiers in Environmental Science, Environmental Informatics and Remote Sensing (Bao et al., 2026; doi: 10.3389/fenvs.2026.1725258). This data package includes the drone photos for a section of Umtanum Creek in Washington, Unted States. The photos were used to reconstruct the 3-dimensional (3D) digital elevation model (DEM) of the riverbed for the investigated stream section. The reconstruction results from four approaches are provided: (1) unoccupied aerial vehicle (UAV, colloquially known as drone) imagery-based Structure-from-Motion (SfM), (2) a machine learning-based 3D reconstruction model, Visual Geometry Grounded Deep Structure from Motion (VGGSfM), (3) Visual Geometry Grounded Transformer for long sequence of images (VGGT-Long), and (4) handheld smartphone LiDAR scanning. The ground truth measurements by tripod-mounted optical level kit and ground control points GPS locations for evaluating the accuracy of the four reconstruction approaches are also provided in this data package. A preliminary version of this data package was published in October 2025 at the time of manuscript submission. It was updated in March 2026, at the time of manuscript acceptance, to include additional metadata (this readme, data dictionary, and file level metadata). The data did not change. For details on how to navigate data packages generated by this project, see https://data.ess-dive.lbl.gov/portals/PNNLRiverCorridorSFA/About. In addition to a readme, this data package also includes a file-level metadata (FLMD) file that describes each file and a data dictionary (DD) that describes all column/row headers and variable definitions. This dataset is comprised of (1) 8 folders; (2) the detailed flight configuration html files; (3) field metadata; (4) a readme; (5) a data dictionary; and (6) file-level metadata. The folders “2024_10_18_d01” and “2024_10_18_d02” contain the original drone photos for the two drone flights (d01 and d02) on October 18, 2024. The reconstruction results from each of the approaches are in the folders called “ODM_SfM”, “VGGSfM”, “VGGTLong”, and “LiDAR”. The ground truth measurements are in the folder called “optical_level_kit”. Lastly, results comparing the different approaches are in the folder called “comparisons”. All files are .csv, .html, .jpg, .obj, .txt, and .npy. For information on using the .obj and .npy files, see the readme files within the same folder as the files.

54 ENVIRONMENTAL SCIENCES