Search NASASearch

SEARCH · Search NASA

Results for “Streaming data”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

A Brief Survey of Data Streaming Technologies

Streaming data is data that is emitted at variable volumes in a continuous, incremental manner with the goal of low-latency processing often at a different physical location. Network infrastructure is used to facilitate the connection between data sources and sinks, and must be robust to handle the requirements of the workflow. The U.S. Department of Energy Office of Science (DOE SC) a federal agency supporting fundamental scientific research for energy and the Nation’s largest supporter of basic research in the physical sciences. DOE SC has the responsibility for operating $\mathbf{1 0}$ National Laboratories, and 28 scientific user facilities supporting advanced supercomputers, particle accelerators, large x-ray light sources, neutron scattering sources, and other specialized facilities for nanoscience and genomics. This paper investigates the state of streaming data workfows, and details some of the approaches to this challenging problem.

Kissel, Ezra

From Edge to HPC: Investigating Cross-Facility Data Streaming Architectures

In this paper, we investigate three cross-facility data streaming architectures, Direct Streaming (DTS), Proxied Streaming (PRS), and Managed Service Streaming (MSS). We examine their architectural variations in data flow paths and deployment feasibility, and detail their implementation using the Data Streaming to HPC (DS2HPC) architectural framework and the SciStream memory-to-memory streaming toolkit on the production-grade Advanced Computing Ecosystem (ACE) infrastructure at Oak Ridge Leadership Computing Facility (OLCF). We present a workflow-specific evaluation of these architectures using three synthetic workloads derived from the streaming characteristics of scientific workflows. Through simulated experiments, we measure streaming throughput, round-trip time, and overhead under work sharing, work sharing with feedback, and broadcast and gather messaging patterns commonly found in AI-HPC communication motifs. Our study shows that DTS offers a minimal-hop path, resulting in higher throughput and lower latency, whereas MSS provides greater deployment feasibility and scalability across multiple users but incurs significant overhead. PRS lies in between, offering a scalable architecture whose performance matches DTS in most cases.

George, Anjus [ORNL] (ORCID:0000000179737061)

Uncertainty guided online ensemble for non-stationary data streams in fusion science

Machine Learning (ML) is poised to play a pivotal role in the development and operation of next-generation fusion devices. Fusion data shows non-stationary behavior with distribution drifts, resulted by both experimental evolution and machine wear-and-tear. ML models assume stationary distribution and fail to maintain performance when encountered with such non-stationary data streams. Online learning techniques have been leveraged in other domains, however it has been largely unexplored for fusion applications. In this paper, we investigate online learning for continuous adaptation to drifting data streams in the prediction of Toroidal Field (TF) coils deflection at the DIII-D fusion facility. We further address the short-term performance degradation inherent to standard online learning, which arises because ground truth is unavailable at prediction time. To mitigate this issue, we propose an uncertainty-guided online ensemble framework. The method leverages the Deep Gaussian Process Approximation (DGPA) for calibrated uncertainty estimation and uses these uncertainty measures to guide a meta-algorithm that aggregates predictions from learners trained over different historical horizons. Our results show that online learning reduces prediction error by 80% compared to a static model. The online ensemble and the proposed uncertainty-guided ensemble further reduce error by approximately 6%, and 10% respectively, relative to standard single-model online learning, while also providing calibrated uncertainty estimates to support operational decision-making.

AI

Online learning of quadratic manifolds from streaming data for nonlinear dimensionality reduction and nonlinear model reduction

Here, this work introduces an online greedy method for constructing quadratic manifolds from streaming data, designed to enable in situ analysis of numerical simulation data on the Petabyte scale. Unlike traditional batch methods, which require all data to be available upfront and take multiple passes over the data, the proposed online greedy method incrementally updates quadratic manifolds in one pass as data points are received, eliminating the need for expensive disk input/output operations as well as storing and loading data points once they have been processed. A range of numerical examples demonstrate that the online greedy method learns accurate quadratic manifold embeddings while being capable of processing data that far exceed common disk input/output capabilities and volumes as well as main-memory sizes.

97 MATHEMATICS AND COMPUTING

Streaming Data in HPC Workflows Using ADIOS

The “IO Wall” problem, in which the gap between computation rate and data access rate grows continuously, poses significant problems to scientific workflows which have traditionally relied upon using the filesystem for intermediate storage between workflow stages. One way to avoid this problem in scientific workflows is to stream data directly from producers to consumers and avoiding storage entirely. However, the manner in which this is accomplished is key to both performance and usability. This paper presents the Sustainable Staging Transport, an approach which allows direct streaming between traditional file writers and readers with few application changes. SST is an ADIOS “engine”, accessible via standard ADIOS APIs, and because ADIOS allows engines to be chosen at run-time, many existing file-oriented ADIOS workflows can utilize SST for direct application-to-application communication without any source code changes. This paper describes the design of SST and presents performance results from various applications that use SST, for feeding model training with simulation data with substantially higher bandwidth than the theoretical limits of Frontier’s file system, for strong coupling of separately developed applications for multiphysics multiscale simulation, or for in situ analysis and visualization of data to complete all data processing shortly after the simulation finishes.

Podhorszki, Norbert [ORNL] (ORCID:000000019647542X

A tale of two towers: comparing NEON and AmeriFlux data streams at Bartlett Experimental Forest

Long-term ecological data are essential for detecting impacts of climate change and other global change factors, and for making informed predictions about future change. However, long-term measurements are rarely replicated at the site level, which raises questions about their representativeness. We used a multiscale approach to evaluate the agreement of parallel observations from AmeriFlux and NEON (National Ecological Observatory Network) towers at Bartlett Experimental Forest, New Hampshire, USA. The two towers are separated by a horizontal distance of 93 m. Here, we focused our analysis on standard meteorological variables; fluxes of CO 2 , sensible heat, and latent heat measured by eddy covariance; and phenology derived from PhenoCam imagery. Results suggest excellent agreement between AmeriFlux and NEON in meteorology and phenology, and good agreement in fluxes at the half-hourly scale. However, large disagreements in CO 2 and latent heat fluxes occurred at the annual scale, with implications especially for the forest carbon balance. The AmeriFlux tower measurements indicate a site that is close to carbon-neutral (-8 ± 65 g C m -2 y -1 , mean ± 1 SD), whereas the NEON tower measurements indicate a forest that is a carbon sink (-137 ± 10 g C m -2 y -1 ). Causes of this disagreement may include measurement height (26 m vs. 35 m), which resulted in different flux footprints being measured by the two towers, and differences in the flux measurement systems. Our results suggest the need for caution when attempting to merge long-term flux data from two different measurement platforms, and when using measurements from any one measurement platform to inform decision-making on issues related to carbon accounting or natural climate solutions.

Carbon cycle

A Digital Twin for an Inverter-Based Resource Power Plant: Real-time data streaming unlocks situation awareness

Here, this study presents the development and successful implementation of a digital twin specifically designed for a grid-connected IBR power plant. By integrating a reduced-order model of the IBR system and dynamically updating the grid impedance with real-time data, the digital twin effectively captures and replicates the behavior of the physical system. Its accuracy and reliability are validated through critical test scenarios, including a three-phase fault and a line-tripping event. The results confirm that the digital twin closely emulates its physical counterpart, demonstrating its strong potential for real-time analysis, system monitoring, and predictive decision making in modern power systems.

Digital twins

Characterizing the GD-1 Stream with DESI DR2 Data: Thin Stream and Hot Cocoon

GD-1 is among the longest, coldest stellar streams in the Milky Way, making it an ideal target for probing dark matter substructure through dynamical heating. We present a catalog of 608 spectroscopically confirmed GD-1 members from the first three years of Dark Energy Spectroscopic Instrument (DESI) observations. This constitutes the largest homogeneous spectroscopic sample of GD-1, doubling the number of members previously available only through heterogeneous compilations combining multiple surveys with different systematics. Using these data, we derive updated stream tracks in sky position, proper motion, and radial velocity that extend over $100^\circ$ of the stream. We apply a Gaussian mixture model to decompose the stream into a dynamically cold thin component ($σ_V = 2.49\pm 0.28$ km s$^{-1}$, width $= 0.23\pm0.01^\circ$) and a kinematically hot cocoon ($σ_V = 6.13\pm0.75$ km s$^{-1}$, width $= 2.18\pm0.17^\circ$). The cocoon contains $\sim30\%$ of members and its velocity dispersion is consistent with $\sim11$ Gyr of heating by cold dark matter subhalos. We also detect a large proper motion dispersion ($41.36\pm4.98$ km s$^{-1}$) along the stream direction in the cocoon component. This feature indicates a significant line-of-sight distance spread in the cocoon, and its origin will be further explored in a forthcoming paper. These measurements demonstrate the power of DESI spectroscopy for characterizing the multi-component phase-space structure of stellar streams and constraining small-scale dark matter substructure.

Jarvis, Emma [Toronto U.] (ORCID:0009000656127336)

Real Time Phasor Analytics (RTPA) and RTPA-SCR System Strength Online Tool

This presentation showcases the Real-Time Phasor Analytics (RTPA) framework for monitoring inertia and assessing system strength in power grids. RTPA is an open-source tool designed to standardize access to data from Power Management Units (PMUs) and Phasor Data Concentrators (PDCs). It facilitates real-time connectivity to multiple PDCs in accordance with the IEEE C37.118-2 standard and supports asynchronous data stream integration. Additionally, RTPA can simulate a PDC server streaming C37.118-2 data and provides Python bindings for seamless interaction with the framework, eliminating the need for direct Rust programming.

24 POWER TRANSMISSION AND DISTRIBUTION

A Data-Agnostic, Continuous Machine Learning Framework for Application in High Energy Physics and Beyond: Phase 1 Final Scientific/Technical Report

This Phase 1 effort has focused on the development of continual learning frameworks for use in machine learning, specifically in the applied context of High Energy Physics (HEP). Machine learning (ML) is a transformative technology by which computers, typically through the use of neural networks, are able to perform tasks with proficiency that rivals or surpasses that of human users. Model Degradation & Catastrophic Forgetting are two undesired phenomena which can occur in ML where the performance of a model degrades when either deployed on novel data streams, or trained on novel data which are sufficiently different than the data the models were initially trained on. A natural example where these sorts of effects can be observed is in the performance of detectors in harsh environments, where the detector signature may change over the lifetime of the detector as it ages and deteriorates — precisely what occurs in the experiments conducted in HEP. Real world HEP data is therefore an excellent test-ground and use-case for Continual Learning paradigms, which are techniques used in ML to counteract these problems. Ensemble learning is one such technique, where multiple smaller models are trained on subsets of the overall data and are ensembled together during inference. The intuition behind this technique is that, although there are shifts in the distributions which govern the incoming data streams, these shifts are not expected to be homogeneous or global. If a sufficient diversity in solutions within the various sub-models has been achieved, then at least one sub-model is expected to retain its performance within the overall ensemble. One further strength of this approach is that the architectures of the various models do not need to be identical, and in fact even different modalities of data can naturally be combined in this way. This work focused on applying ensemble learning techniques to derive results using two main datasets, anomaly detection in HEP data & time-series forecasting in semiconductor manufacturing data. Semiconductor manufacturing involves data with surprising similarity to that of HEP (e.g. wafer maps look very similar to digi-occupancy maps) and Cerium Lab’s prominence within the semiconductor industry makes semiconductor manufacturing a natural opportunity for commercialization of this work. Our efforts have led to two strong results. The first is that we evaluated the proposed ensembling techniques using previously proposed machine learning architectures for use in anomaly detection, namely AutoEncoder based models and their derivatives. We also developed new architectures which have not been evaluated in this context before. In fact, this work marks the first use of Vision Transformers for anomaly detection in HEP. Second, we demonstrated that ensemble learning significantly improves model performance in scenarios prone to degradation, validating its effectiveness across both HEP and semiconductor datasets. These results further support ensemble learning as a powerful strategy for mitigating catastrophic forgetting and maintaining robust performance in evolving data environments.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS

StOKeDMD: Streaming Occupation kernel dynamic mode decomposition

Dynamic mode decomposition (DMD) has become a common technique for constructing surrogate models for dynamical systems from observed system states. The Occupation Kernel DMD (OKDMD) method proposed in (Rosenfeld et al., 2022) and (Rosenfeld et al., 2024) is a Liouville operator based method that builds surrogate models from system state trajectories. Here, this paper proposes an extension of OKDMD to the case when the system states are observed in a streaming fashion, i.e., only a small fraction of the state trajectory is available at a given time. The developed method, Streaming Occupation Kernel DMD (StOKeDMD), accommodates the streaming data input by leveraging properties of specific choices of kernel functions and occupation kernels. We apply the StoKeDMD method as a compression method for streaming data, analyze the memory complexity, and demonstrate the performance of StoKeDMD in the compression of streaming data generated from a Lorenz system and a fluid flow simulation.

97 MATHEMATICS AND COMPUTING

EJFAT Scientific Perspective

Presented new computing model to the test by deploying the EJFAT system alongside a data-stream processing framework running the production-level CLAS12 event reconstruction application. In this experiment, a continuous stream of CLAS12 Level-1 identified events was processed in real-time using the EJFAT load balancer, distributing the workload across 90 computing nodes located across the U.S. This marks the first-ever large-scale, real-time distributed data stream processing experiment, demonstrating that scientific data-streaming pipelines can efficiently scale across four dimensions, thanks to EJFAT’s advanced hardware and software capabilities.

Gyurjyan, Vardan [Thomas Jefferson National Accele

Continual learning in the presence of repetition

Continual learning (CL) provides a framework for training models in ever-evolving environments. Although re-occurrence of previously seen objects or tasks is common in real-world problems, the concept of repetition in the data stream is not often considered in standard benchmarks for CL. Unlike with the rehearsal mechanism in buffer-based strategies, where sample repetition is controlled by the strategy, repetition in the data stream naturally stems from the environment. This report provides a summary of the CLVision challenge at CVPR 2023, which focused on the topic of repetition in class-incremental learning. The report initially outlines the challenge objective and then describes three solutions proposed by finalist teams that aim to effectively exploit the repetition in the stream to learn continually. The experimental results from the challenge highlight the effectiveness of ensemble-based solutions that employ multiple versions of similar modules, each trained on different but overlapping subsets of classes. This report underscores the transformative potential of taking a different perspective in CL by employing repetition in the data stream to foster innovative strategy design.

Class-incremental learning

WHONDRS Surface Water Geochemistry and Organic Matter Characterization Data from Streams Distributed across Latin America

This dataset supports a broader study examining global transferability of stream biogeochemistry and was generated in collaboration with the MicroSudAqua (µSudAqua) network (https://microsudaqua.netlify.app/en/). The dataset provides surface water geochemistry (dissolved organic carbon, total dissolved nitrogen, cations) and organic matter characterization (FTICR-MS) from streams in Argentina, Brazil, Chile, and Colombia. Samples were collected across stream orders (1st to 6th order) within five basins. Related data were collected and will be published separately in collaboration with the µSudAqua network. For details on how to navigate data packages generated by this project, see https://data.ess-dive.lbl.gov/portals/PNNLRiverCorridorSFA/About. In addition to this readme, this data package also includes a file-level metadata (FLMD) file that describes each file and a data dictionary (DD) that describes all column/row headers and variable definitions. This dataset is comprised of (1) a folder of field photos; (2) a folder of surface water sample data, (3) a folder of raw Fourier transform ion cyclotron resonance mass spectrometry (FTICR-MS) data; (4) file-level metadata; (5) data dictionary; (6) field metadata; (7) readme; (8) international generic sample number (IGSN) mapping file; and (9) field protocol. The sample data subfolder contains (1) dissolved organic carbon (DOC, measured as non-purgeable organic carbon, NPOC) data and averages; (2) total dissolved nitrogen data and averages; (3) anions and averages; (4) methods codes; (5) FTICR-MS methods; and (15) a subfolder of 9.4 Tesla (9.4T) FTICR-MS data. This folder contains the processed data and three subfolders, one containing the .xml files, one containing the water CoreMS output files, and the other containing instructions and scripts for processing the files in CoreMS (https://github.com/EMSL-Computing/CoreMS). All files are .csv, .pdf, .R, .xml, .d, .html, .Rmd, .py, .cal, .json, .jpg, .jpeg, .png, .mov, or .mp4.

Anions

Machine Learning Digital Twin for Lithium Ion Battery State of Health Predictions

A digital twin system has been established to model the long term degradation of the state of health of lithium ion batteries. Two data streams result from the computational model of the system and from the physical experiment measurements. A machine learning pipeline has been developed at the nexus of these data streams. Leveraging the unique data sources in multiple transfer learning approaches has lead to the development of multiple cell specific machine learning digital twins. Our discussion on this first of its kind technology will cover challenges in data ingestion, scalability, architecture and orchestration, and research findings.

25 - ENERGY STORAGE

DEVELOPMENT AND DEMONSTRATION TESTBED FOR THE REMOTE OPERATIONS AND MONITORING OF MICROREACTORS

The nuclear industry is rapidly developing many advanced-reactor concepts for near-term deployment in both traditional and non-traditional nuclear-powered applications. One such category of advanced reactor is the microreactor, a class of reactor with less than 20MWth power output, intended for applications where the economics or logistics of traditional power sources are difficult. This includes applications such as remote communities, mining sites, defense installations, or humanitarian and disaster-relief missions. One key enabling feature for the successful deployment of microreactors is a remote operations capability. Remote operations provide monitoring and control capabilities which can significantly reduce staffing costs by eliminating the need for licensed operators at each reactor facility and improve the economic viability for microreactor deployment. A remote concept of operations is not currently an established capability in the nuclear industry. In addition, no demonstration microreactor is expected to complete construction or go critical until at least 2026. This leaves two major capability gaps: the successful demonstration of a remote concept of operations for microreactors and a test bed suitable for said demonstration. Both gaps must be addressed in order to advance the remote concepts of nuclear operation and, more broadly, microreactors themselves from paper to reality. This paper aims to fill these gaps and describes a test bed that would support development and deployment of a remote concept of nuclear operations, initial experimental results from that test bed, and the application of the test bed and experimental results for a digital-twin-based remote concept of operations underdevelopment at Idaho National Laboratory (INL). The platform chosen as a remote concept of nuclear operations test bed is the Single Primary Heat Extraction and Removal Emulator, known as SPHERE, located at INL. SPHERE is a small-scale non-nuclear test bed that emulates thermal behavior of a microreactor. The small-scale and non-nuclear nature of SHPERE limit safety concerns associated with remote operations while still providing the physical response representative of a microreactor. A network connection was added to SPHERE that enables remote-monitoring capability. This allows for real-time data streaming to networked workstations, data historians, and human-machine interfaces (HMIs). These are all critical components in a remote concept of operations, thus providing a robust development and demonstration platform. An initial experiment was performed using the SPHERE remote operations testbed. This included running a comprehensively instrumented SPHERE through a series of steady-state and transient operating scenarios in both normal and abnormal operating conditions, all while streaming live test data to a remote HMI and data warehouse. This initial experiment served three purposes: (1) characterizing the response of SPHERE, (2) demonstrating the remote connection to SPHERE, and (3) providing a baseline data set for development of a digital-twin-based remote concept of operations that is under development at INL.

22 GENERAL STUDIES OF NUCLEAR REACTORS

DEVELOPMENT AND DEMONSTRATION TESTBED FOR THE REMOTE OPERATIONS AND MONITORING OF MICROREACTORS

The nuclear industry is rapidly developing many advanced-reactor concepts for near-term deployment in both traditional and non-traditional nuclear-powered applications. One such category of advanced reactor is the microreactor, a class of reactor with less than 20MWth power output, intended for applications where the economics or logistics of traditional power sources are difficult. This includes applications such as remote communities, mining sites, defense installations, or humanitarian and disaster-relief missions. One key enabling feature for the successful deployment of microreactors is a remote operations capability. Remote operations provide monitoring and control capabilities which can significantly reduce staffing costs by eliminating the need for licensed operators at each reactor facility and improve the economic viability for microreactor deployment. A remote concept of operations is not currently an established capability in the nuclear industry. In addition, no demonstration microreactor is expected to complete construction or go critical until at least 2026. This leaves two major capability gaps: the successful demonstration of a remote concept of operations for microreactors and a test bed suitable for said demonstration. Both gaps must be addressed in order to advance the remote concepts of nuclear operation and, more broadly, microreactors themselves from paper to reality. This paper aims to fill these gaps and describes a test bed that would support development and deployment of a remote concept of nuclear operations, initial experimental results from that test bed, and the application of the test bed and experimental results for a digital-twin-based remote concept of operations underdevelopment at Idaho National Laboratory (INL). The platform chosen as a remote concept of nuclear operations test bed is the Single Primary Heat Extraction and Removal Emulator, known as SPHERE, located at INL. SPHERE is a small-scale non-nuclear test bed that emulates thermal behavior of a microreactor. The small-scale and non-nuclear nature of SHPERE limit safety concerns associated with remote operations while still providing the physical response representative of a microreactor. A network connection was added to SPHERE that enables remote-monitoring capability. This allows for real-time data streaming to networked workstations, data historians, and human-machine interfaces (HMIs). These are all critical components in a remote concept of operations, thus providing a robust development and demonstration platform. An initial experiment was performed using the SPHERE remote operations testbed. This included running a comprehensively instrumented SPHERE through a series of steady-state and transient operating scenarios in both normal and abnormal operating conditions, all while streaming live test data to a remote HMI and data warehouse. This initial experiment served three purposes: (1) characterizing the response of SPHERE, (2) demonstrating the remote connection to SPHERE, and (3) providing a baseline data set for development of a digital-twin-based remote concept of operations that is under development at INL.

22 - GENERAL STUDIES OF NUCLEAR REACTORS