Search NASASearch

SEARCH · Search NASA

Results for “data pipeline”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

U.S. Hydropower Development Pipeline Data, 2026

The U.S. Hydropower Development Pipeline dataset provides a comprehensive, regularly updated view of proposed and potential hydropower projects across the United States. This resource compiles information from federal agencies and other public sources to track non-powered dams considered for electrification, proposed hydropower facilities at stream reaches with no existing dams, conduit exemptions, and emerging pumped storage hydropower proposals. The dataset includes project characteristics such as location, development status, technology type, ownership category, and other attributes that support analysis of future hydropower trends. It is designed to help researchers, planners, policymakers, and stakeholders assess national‑scale development patterns, understand the evolving hydropower landscape, and explore opportunities and challenges associated with new hydropower deployment. The dataset is updated annually to reflect changes in project status, new proposals entering the pipeline, and projects that are cancelled, completed, or otherwise removed from active consideration. Note: Capacity additions to existing hydropower plants are not included in this database due to reliance on a proprietary data source.

Johnson, Megan [ORNL] (ORCID:0000000290141741)

RectifHydPlus Data Pipeline

The RectifHydPlus Data Pipeline is an open source and fully reproducible data processing pipeline for creating RectifHydPlus—a dataset of historical monthly net electricity generation for all US hydropower plants (>10MW). The pipeline is coded in R, applying tidyverse libraries and code principles, and using the targets data pipeline framework. All data inputs to the RectifHydPlus Data Pipeline are available from public sources. References to all data inputs, as well as instructions for running the RectifHydPlus Data Pipeline, are available on the GitLab code repository: https://code.ornl.gov/turnersw/rectifhydplus

Turner, SeanWilliam Donald [Oak Ridge National Lab

RectifHydPlus Data Pipeline v1.1.0

The RectifHydPlus Data Pipeline is an open source and fully reproducible data processing pipeline for creating RectifHydPlus—a dataset of historical monthly net electricity generation for all US hydropower plants (>10MW). The pipeline is coded in R, applying tidyverse libraries and code principles, and using the targets data pipeline framework. All data inputs to the RectifHydPlus Data Pipeline are available from public sources. References to all data inputs, as well as instructions for running the RectifHydPlus Data Pipeline, are available on the GitLab code repository: https://code.ornl.gov/turnersw/rectifhydplus

Turner, SeanWilliam Donald [Oak Ridge National Lab

AAM National Campaign Tech Talk: Data Pipeline Familiarization

NASA's AWS-based Data Pipeline allows real-time data submission and ingestion, with immediate monitoring of data rate, coverage and ingestion quality. Problems are immediately discovered and can be corrected with agility by both Partners and NASA during the a simulation or flight event​.

Data Pipeline

Step-oriented pipeline data processing system

Architecture for step-oriented pipeline data processing is disclosed utilizing a plurality of cascaded modules, each module including a programmable general purpose processor and a read/write random access memory. The memory of each module, shared with the next module in cascade serves as an output memory for the processor and the input memory for the next processor. An additional memory is provided to serve as the input memory of the first module, and each module is provided with a memory, which may be a read-out memory, to store a program for the processor. Each module is further provided with a logic network for resolving a potential memory sharing conflict by awarding priority to the processor of the module.

Castleman, Kenneth R.

POWER DATA PIPELINE

SF-25-081 Utility software for creating high-performance data pipelines to extract, load, and transform raw electric power systems measurements. For use with anomaly detection models training workflows. The software supports the project: Adaptive Cybersecurity for DER: A Game-Theoretic and Machine Learning approach for Real-Time Threat Detection and Mitigation

Plathottam, Silby Jose [Argonne National Laborator

The Kepler End-to-End Data Pipeline: From Photons to Far Away Worlds

The Kepler mission is described in overview and the Kepler technique for discovering exoplanets is discussed. The design and implementation of the Kepler spacecraft, tracing the data path from photons entering the telescope aperture through raw observation data transmitted to the ground operations team is described. The technical challenges of operating a large aperture photometer with an unprecedented 95 million pixel detector are addressed as well as the onboard technique for processing and reducing the large volume of data produced by the Kepler photometer. The technique and challenge of day-to-day mission operations that result in a very high percentage of time on target is discussed. This includes the day to day process for monitoring and managing the health of the spacecraft, the annual process for maintaining sun on the solar arrays while still keeping the telescope pointed at the fixed science target, the process for safely but rapidly returning to science operations after a spacecraft initiated safing event and the long term anomaly resolution process.The ground data processing pipeline, from the point that science data is received on the ground to the presentation of preliminary planetary candidates and supporting data to the science team for further evaluation is discussed. Ground management, control, exchange and storage of Kepler's large and growing data set is discussed as well as the process and techniques for removing noise sources and applying calibrations to intermediate data products.

data archiving

Dynamic Black-Level Correction and Artifact Flagging in the Kepler Data Pipeline

Instrument-induced artifacts in the raw Kepler pixel data include time-varying crosstalk from the fine guidance sensor (FGS) clock signals, manifestations of drifting moiré pattern as locally correlated nonstationary noise and rolling bands in the images which find their way into the calibrated pixel time series and ultimately into the calibrated target flux time series. Using a combination of raw science pixel data, full frame images, reverse-clocked pixel data and ancillary temperature data the Keplerpipeline models and removes the FGS crosstalk artifacts by dynamically adjusting the black level correction. By examining the residuals to the model fits, the pipeline detects and flags spatial regions and time intervals of strong time-varying blacklevel (rolling bands ) on a per row per cadence basis. These flags are made available to downstream users of the data since the uncorrected rolling band artifacts could complicate processing or lead to misinterpretation of instrument behavior as stellar. This model fitting and artifact flagging is performed within the new stand-alone pipeline model called Dynablack. We discuss the implementation of Dynablack in the Kepler data pipeline and present results regarding the improvement in calibrated pixels and the expected improvement in cotrending performances as a result of including FGS corrections in the calibration. We also discuss the effectiveness of the rolling band flagging for downstream users and illustrate with some affected light curves.

Clarke, B. D.

Integrase-On-Demand-Pipeline Data Set

Files needed to run the Integrase-On-Demand-Pipeline, a program designed to provide users with a list of putative attachment site and integrase pairs for a prokaryotic genome of interest. isles.pkl: Serialized python-object file, containing a dictionary of attachment site sequences and reference genomic island information extracted from the Genomic island database ints.gff: Gene format file containing annotations for all integrases referenced in isles.pkl. The source genome, gene coordinates, integrase name, protein IDs and amino acid sequence included. reps.msh: Binary file containing 1000 128-bit MurmurHash3 hashes for >80,000 genomes

McClain, Hannah Marie [Sandia National Laboratorie

L-PBF High-Throughput Data Pipeline Approach for Multi-modal Integration

Abstract Metal-based additive manufacturing requires active monitoring solutions for assessing part quality. Multiple sensors and data streams, however, generate large heterogeneous data sets that are impractical for manual assessment and characterization. In this work, an automated pipeline is developed that enables feature extraction from high-speed camera video and multi-modal data analysis. The framework removes the need for manual assessment through the utilization of deep learning techniques and training models in a weakly supervised paradigm. We demonstrate this pipeline’s capability over 700,000 high-speed camera frames. The pipeline successfully extracts melt pool and spatter geometries and links them to corresponding pyrometry, radiography, and processparameter information. 715 individual prints are examined to reveal melt pool areas that exceeds 0.07 mm 2 and pyrometry signal over a threshold (375 pyrometry units) were more likely to have defects. These automated processes enable massive throughput of characterization techniques.

36 MATERIALS SCIENCE

Design Choices in Anomaly Detection for Industrial Control Systems: Insights from Gas Pipeline Data

Industrial control systems (ICS) remain vulnerable to increasingly sophisticated cyberattacks, yet evaluating anomaly detection models in these environments is challenging due to temporal dependencies, missing-not-at-random patterns, and extremely imbalanced datasets. These factors make common practices—especially random data splits and naïve imputation—prone to severe temporal leakage, which can inflate reported performance and obscure real-world limitations. In this work, we systematically examine classical machine learning models, temporal deep learning architecture, and tensor-decomposition–based methods on a gas-pipeline dataset using a fully temporally separated evaluation pipeline designed to mimic realistic deployment conditions. Our findings show that proper temporal handling and MNAR-aware preprocessing significantly alter the relative performance of popular anomaly-detection methods, providing practical guidance for designing reliable, leakage-resistant ICS intrusion-detection systems.

97 MATHEMATICS AND COMPUTING

The Kepler End-to-End Data Pipeline: From Photons to Far Away Worlds

Launched by NASA on 6 March 2009, the Kepler Mission has been observing more than 100,000 targets in a single patch of sky between the constellations Cygnus and Lyra almost continuously for the last two years looking for planetary systems using the transit method. As of October 2011, the Kepler spacecraft has collected and returned to Earth just over 290 GB of data, identifying 1235 planet candidates with 25 of these candidates confirmed as planets via ground observation. Extracting the telltale signature of a planetary system from stellar photometry where valid signal transients can be small as a 40 ppm is a difficult and exacting task. The end-to end processing of determining planetary candidates from noisy, raw photometric measurements is discussed.

Cygnus constallation