Search NASA⌕ Search

SEARCH · Search NASA

Results for “pipelines”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 181 records · Page 10

The Kepler End-to-End Data Pipeline: From Photons to Far Away Worlds

The Kepler mission is described in overview and the Kepler technique for discovering exoplanets is discussed. The design and implementation of the Kepler spacecraft, tracing the data path from photons entering the telescope aperture through raw observation data transmitted to the ground operations team is described. The technical challenges of operating a large aperture photometer with an unprecedented 95 million pixel detector are addressed as well as the onboard technique for processing and reducing the large volume of data produced by the Kepler photometer. The technique and challenge of day-to-day mission operations that result in a very high percentage of time on target is discussed. This includes the day to day process for monitoring and managing the health of the spacecraft, the annual process for maintaining sun on the solar arrays while still keeping the telescope pointed at the fixed science target, the process for safely but rapidly returning to science operations after a spacecraft initiated safing event and the long term anomaly resolution process.The ground data processing pipeline, from the point that science data is received on the ground to the presentation of preliminary planetary candidates and supporting data to the science team for further evaluation is discussed. Ground management, control, exchange and storage of Kepler's large and growing data set is discussed as well as the process and techniques for removing noise sources and applying calibrations to intermediate data products.

data archiving↗

AAM National Campaign Tech Talk: Data Pipeline Familiarization

NASA's AWS-based Data Pipeline allows real-time data submission and ingestion, with immediate monitoring of data rate, coverage and ingestion quality. Problems are immediately discovered and can be corrected with agility by both Partners and NASA during the a simulation or flight event​.

Data Pipeline↗

HyBlend: Pipeline CRADA Cost and Emissions Analysis

This presentation provides an overview of National Renewable Energy Laboratory and Argonne National Laboratory research as part of the pipeline blending CRADA (a HyBlend project) to the DOE's 2024 Annual Merit Review.

BlendPATH↗

Data readiness pipeline patterns for scientific AI at scale: Insights from climate, fusion, life sciences, and materials

This article examines how data readiness for AI principles apply to large scientific datasets used to train foundation models. We analyze archetypal workflows across four representative domains—climate, nuclear fusion, life sciences, and materials—to identify common preprocessing patterns and domain‐specific constraints. We introduce a two‐dimensional readiness model that combines canonical preprocessing patterns with a five‐level operational readiness scale, both tailored to high‐performance computing (HPC) environments. This construct helps outline key challenges in transforming large‐scale scientific data into formats suitable for scalable AI training. Together, these dimensions form a conceptual maturity matrix that characterizes scientific data readiness and guides infrastructure development toward standardized, cross‐domain support for scalable and reproducible AI for science. Finally, we evaluate this maturity matrix in the context of case studies including ClimaX (climate), AFLOW (materials), OpenFold (proteomics), and DIII‐D fusion disruption‐prediction workflows, from which we distill lessons learned and provide recommendations to guide practitioners in developing robust AI‐readiness pipelines. Finally, we discuss remaining cross‐cutting challenges that persist across scientific domains.

97 MATHEMATICS AND COMPUTING↗

Single‐Cell Nanodroplet Processing Proteomics Pipeline for Analysis of Human‐Derived Microglia

Single-cell omics tools provide unique insights into heterogeneous cell populations and their responses to stimuli. For example, single-cell RNA sequencing has identified several transcriptionally distinct populations of microglia, which are resident immune cells of the central nervous system (CNS) that are responsive to CNS injury, infection, and neurodegeneration. To date, single-cell studies of microglia have focused on RNA-sequencing or cytometry by time of flight (CyTOF), which provide indirect readouts of protein abundance or quantification of a limited number of targets. Herein, we present a workflow based on FACS-assisted isolation, cryopreservation, and nanodroplet-based processing for single-cell mass spectrometry proteomics analysis of the postmortem human brain cortex-derived microglia. From a single microglial cell, 1039 proteins could be identified on average. As a proof-of-principle, we applied single-cell proteomics for exploring the heterogeneity of brain microglia at the cellular level. This pilot proteomics data partially recapitulates the prior microglia subtypes. Specifically, we determined that mitochondrial proteins, in particular members of NADH dehydrogenase (Complex I), cytochrome b-c1 (Complex III), cytochrome c oxidase (Complex IV), F1-ATPase (Complex V), and Na+/K+-ATPase complex, drive variation across microglia. This pipeline offers the potential for identifying functionally and analytically relevant protein targets for microglia in Alzheimer's disease and other neurological disorders.

59 BASIC BIOLOGICAL SCIENCES↗

A network-enabled pipeline for gene discovery and validation in non-model plant species

Identifying key regulators of important genes in non-model crop species is challenging due to limited multi-omics resources. To address this, we introduce the network-enabled gene discovery pipeline NEEDLE, a user-friendly tool that systematically generates coexpression gene network modules, measures gene connectivity, and establishes network hierarchy to pinpoint key transcriptional regulators from dynamic transcriptome datasets. After validating its accuracy with two independent datasets, we applied NEEDLE to identify transcription factors (TFs) regulating the expression of cellulose synthase-like F6 ( CSLF6 ), a crucial cell wall biosynthetic gene, in Brachypodium and sorghum. Our analyses uncover regulators of CSLF6 and also shed light on the evolutionary conservation or divergence of gene regulatory elements among grass species. These results highlight NEEDLE’s capability to provide biologically relevant TF predictions and demonstrate its value for non-model plant species with dynamic transcriptome datasets.

59 BASIC BIOLOGICAL SCIENCES↗

Protocol for applying a network-enabled gene discovery pipeline to non-model plant species

Identifying upstream regulators of key genes is essential for understanding gene regulatory mechanisms and translating these insights into functional targets. Here, we present a protocol for applying the network-enabled gene discovery pipeline (NEEDLE) to non-model plant species. We describe steps for environment setup, data preparation, computational analysis, expected outputs, and parameter considerations. NEEDLE integrates RNA sequencing (RNA-seq) processing, weighted gene co-expression analysis (WGCNA), Gene Network Inference with Ensemble of trees (GENIE3), and promoter conservation analysis to prioritize candidate transcriptional regulators.

Plant Sciences↗

Automated Gold Nanorod Spectral Morphology Analysis Pipeline

The development of a colloidal synthesis procedure to produce nanomaterials with high shape and size purity is often a time-consuming, iterative process. This is often due to quantitative uncertainties in the required reaction conditions and the time, resources, and expertise intensive characterization methods required for quantitative determination of nanomaterial size and shape. Absorption spectroscopy is often the easiest method for colloidal nanomaterial characterization. However, due to the lack of a reliable method to extract nanoparticle shapes from absorption spectroscopy, it is generally treated as a more qualitative measure for metal nanoparticles. This work demonstrates a gold nanorod (AuNR) spectral morphology analysis tool, called AuNR-SMA, which is a fast and accurate method to extract quantitative structural information from colloidal AuNR absorption spectra. To demonstrate the practical utility of this model, we apply it to three distinct applications. First, we demonstrate this model's utility as an automated analysis tool in a high-throughput AuNR synthesis procedure by generating quantitative size information from optical spectra. Second, we use the predictions generated by this model to train a machine learning model to predict the resulting AuNR size distributions under specified reaction conditions. Third, we apply this model to spectra extracted from the literature where no size distributions are reported and impute unreported quantitative information on AuNR synthesis. This approach can potentially be extended to any other nanocrystal system where absorption spectra are size dependent, and accurate numerical simulation of absorption spectra is possible. In addition, this pipeline could be integrated into automated synthesis apparatuses to provide interpretable data from simple measurements, help explore the synthesis science of nanoparticles in a rational manner, or facilitate closed-loop workflows.

36 MATERIALS SCIENCE↗

A Performant, Scalable Processing Pipeline for High‐Quality and FAIR Environmental Sensor Data

High-resolution environmental monitoring is necessary to record, understand, and predict biogeochemical and ecological changes particularly in coastal systems but brings significant challenges in processing and making rapidly available the resulting data. The COMPASS-FME project established a network of coastal observational sites across the Chesapeake Bay and western Lake Erie regions extensively instrumented with soil, vegetation, and weather sensors logging data every 15 min. Our data processing framework, written in R and completely open source, prioritizes rapid model-experiment iteration and makes biogeochemical data rapidly available for quality assurance/quality control, analysis, and model ingestion. This pipeline is distinguished by a standardized and modular approach to data curation, extensive metadata and documentation, and its high performance. These attributes combine to make biogeochemical data rapidly accessible across COMPASS-FME and the broader community. Flexible, powerful, and reproducible approaches to handling high-volume environmental data are crucial for accelerating biogeosciences research.

Pennington, Stephanie C. [Pacific Northwest Nation↗

Enabling high-throughput enzyme discovery and engineering with a low-cost, robot-assisted pipeline

Abstract As genomic databases expand and artificial intelligence tools advance, there is a growing demand for efficient characterization of large numbers of proteins. To this end, here we describe a generalizable pipeline for high-throughput protein purification using small-scale expression in E. coli and an affordable liquid-handling robot. This low-cost platform enables the purification of 96 proteins in parallel with minimal waste and is scalable for processing hundreds of proteins weekly per user. We demonstrate the performance of this method with the expression and purification of the leading poly(ethylene terephthalate) hydrolases reported in the literature. Replicate experiments demonstrated reproducibility and enzyme purity and yields (up to 400 µg) sufficient for comprehensive analyses of both thermostability and activity, generating a standardized benchmark dataset for comparing these plastic-degrading enzymes. The cost-effectiveness and ease of implementation of this platform render it broadly applicable to diverse protein characterization challenges in the biological sciences.

36 MATERIALS SCIENCE↗

Crystal generation using the fully differentiable pipeline and latent space optimization

We present a materials generation framework that couples a symmetry-conditioned variational autoencoder with a differentiable SO(3) power spectrum objective to steer candidates toward a specified local environment under the crystallographic constraints. In particular, we implement a fully differentiable pipeline that performs batch-wise optimization on both direct and latent crystallographic representations. Using the GPU acceleration, the implementation achieves about fivefold speed compared to our previous CPU workflow, while yielding comparable outcomes. In addition, we introduce the optimization strategy that alternatively performs optimization on the direct and latent crystal representations. This dual-level relaxation approach can effectively escape local minima defined by different objective gradients, thus increasing the success rate of generating complex structures satisfying the target local environments. This framework can be extended to systems consisting of multi-components and multi-environments, providing a scalable route to generate material structures with the target local environment.

conditional VAE↗

nf-core/proteinfamilies: a scalable pipeline for the generation of protein families

The growth of metagenomics-derived amino acid sequence data has transformed our understanding of protein function, microbial diversity, and evolutionary relationships. However, the vast majority of these proteins remain functionally uncharacterized. Grouping the millions of such uncharacterized sequences with the few experimentally characterized ones allows the transfer of annotations, while the inspection of conserved residues with multiple sequence alignments can provide clues to function, even in the absence of existing functional information. To address the challenges associated with this data surge and the need to group sequences, we present a scalable, open-source, parametrizable Nextflow pipeline (nf-core/proteinfamilies) that generates nascent protein families or assigns new proteins to existing families. The computational benchmarks demonstrated that resource usage scales approximately linearly with input size, and the biological benchmarks showed that the generated protein families closely resemble manually curated families in widely used databases.

Nextflow↗

Toward an AI-Powered Software Pipeline for Real-Time Tracking and Analysis of Wildfire and Smoke

Real-time tracking of wildfires and smoke is crucial for effective response, minimizing damage, protecting lives, and efficiently managing resources during fire emergencies. We develop a web-based AI-powered pipeline that detects wildfires in aerial video and estimates deployment-relevant behavior metrics, including cumulative burned area, burned-area growth rate, fire spread direction, and smoke dispersion. The system combines a YOLO-based detector with YCbCr-based fire segmentation, HSV-based smoke segmentation, Farneback optical flow, and centroid-based spatiotemporal tracking. Using ground sampling distance (GSD), pixel-level fire masks are converted to physical burned-area measurements by correlating fire pixel counts with camera altitude and tilt angle. We benchmark YOLO variants and non-YOLO baselines (GoogLeNet, CNN, DBN, Autoencoder, U-Net, and AlexNet) on the IEEE FLAME dataset and a newly created aerial frame dataset, Wildfire-DB. Cross-dataset evaluation uses a strict threshold-transfer protocol: decision thresholds are selected on FLAME validation and transferred unchanged to Wildfire-DB to quantify generalization under domain shift. YOLOv6 achieves the strongest cross-dataset frame-level fire detection on Wildfire-DB (ROC-AUC 0.8200, PR-AUC 0.8044, and transferred-threshold F1 0.7596). For tracking-oriented deployment requiring oriented localization, YOLO11-OBB provides the most reliable cross-dataset behavior among OBB-capable models while remaining computationally feasible. To analyze the feasibility of UAV deployment, we further measure inference efficiency using synchronized GPU and CPU power logs on a fixed workload of 1569 frames. YOLO-family models process the video in 5.73–12.47 seconds with net energy of 1247.28–1775.39 J, substantially lower latency and energy than heavier classification and reconstruction baselines. Overall, model optimality depends on operational objectives: YOLOv6 is best for cross-dataset detection robustness, whereas YOL...

Color segmentation↗

From Sim to Real: A Pipeline for Training and Deploying Traffic Smoothing Cruise Controllers

Designing and validating controllers for connected and automated vehicles to enhance traffic flow presents significant challenges, from the complexity of replicating real-world stop-and-go traffic dynamics in simulation, to the intricacies involved in transitioning from simulation to actual deployment. In this work, we present a full pipeline from data collection to controller deployment. Specifically, we collect 772 km of driving data from the I-24 in Tennessee, and use it to build a one-lane simulator, placing simulated vehicles behind real-world trajectories. Using policy-gradient methods with an asymmetric critic, we improve fuel efficiency by over 10% when simulating congested scenarios. Our comprehensive approach includes reinforcement learning for controller training, software verification, hardware validation and setup, and navigating various sim-to-real challenges. Furthermore, we analyze the controller's behavior and wave-smoothing properties, and deploy it on four Toyota Rav4’s in a real-world validation experiment on the I-24. Lastly, we release the driving dataset, the simulator and the trained controller, to enable future benchmarking and controller design.

42 ENGINEERING↗

Investigating In-Situ Fracture Behaviors of Polymer Pipeline Materials in Hydrogen and Hydrogen-Methane Blended Gas Environments

To reduce carbon emissions, the US natural gas infrastructure is seen as a primary solution for efficiently transporting hydrogen gas. Blending hydrogen gas with natural gas and transporting it across a national infrastructure could save significant infrastructure costs. To properly operate the infrastructure under the new gas system, it is critical to understand material compatibility with hydrogen under various conditions. The Blended Gas CRADA, a Hyblend project, is established to determine the material compatibility of existing natural gas pipes with hydrogen gas. In this study, we investigate the in-plane fracture behaviors of MDPEMarlex and HDPEGDB exposed to hydrogen and hydrogen-methane blended gas. Single-edge notch bending geometry is used. All tests are executed in-situ with the gas environment. The experimental results show a significant effect of the gas environment on HDPEGDB specimens, reducing 5% (H2) to 42% (Blended gas) of specific fracture energy compared to non-aged specimens. For the MDPEMarlex, the effects of the gas environment have increased the specific fracture energy by 10% (H2) to 15% (Blended gas). Fracture surfaces of the tested samples are observed using an electronic microscope. The in-plane fracture surface of HDPEGDB shows a pronounced dimple fracture pattern after exposure to hydrogen and blended gas. The expanded fracture pattern contributes to lower the specific fracture energy. These observations provide critical information for validating polymer pipeline materials when interact with hydrogen and hydrogen-blend gas.

Ko, Seunghyun↗

Integrase-On-Demand-Pipeline Data Set

Files needed to run the Integrase-On-Demand-Pipeline, a program designed to provide users with a list of putative attachment site and integrase pairs for a prokaryotic genome of interest. isles.pkl: Serialized python-object file, containing a dictionary of attachment site sequences and reference genomic island information extracted from the Genomic island database ints.gff: Gene format file containing annotations for all integrases referenced in isles.pkl. The source genome, gene coordinates, integrase name, protein IDs and amino acid sequence included. reps.msh: Binary file containing 1000 128-bit MurmurHash3 hashes for >80,000 genomes

McClain, Hannah Marie [Sandia National Laboratorie↗

Machine Learning Automation Pipeline

Machine Learning Automation Pipeline (MLAP) is a package to perform machine learning (ML) analysis in a step by step manner, starting with data extraction until analysis and prediction. The scripts provide the users option to chose an action such as "Extract", "Prep", and "Train" and numerous cases can be launched with just a single command. The inputs for each case are provided using a JSON file. The simulation results of several cases can be assessed using an automated process and analyzed for various metrics pertinent to ML analysis.

Jha, Pankaj↗

Disclosure of the XRD pipeline software

The XRD pipeline software is a program to reduce 2D X-ray diffraction data from area detectors with advanced algorithms on automatic masking and image segmentation, which help characterization of multiple phases recorded in the data and facilitate data analysis.SF-25-114

XU, WENQIAN [Argonne National Laboratory (ANL), Ar↗