Search NASA⌕ Search

SEARCH · Search NASA

Results for “pipeline data processing”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 217 records · Page 12

Processing and Managing the Kepler Mission's Treasure Trove of Stellar and Exoplanet Data

The Kepler telescope launched into orbit in March 2009, initiating NASAs first mission to discover Earth-size planets orbiting Sun-like stars. Kepler simultaneously collected data for 160,000 target stars at a time over its four-year mission, identifying over 4700 planet candidates, 2300 confirmed or validated planets, and over 2100 eclipsing binaries. While Kepler was designed to discover exoplanets, the long term, ultra- high photometric precision measurements it achieved made it a premier observational facility for stellar astrophysics, especially in the field of asteroseismology, and for variable stars, such as RR Lyraes. The Kepler Science Operations Center (SOC) was developed at NASA Ames Research Center to process the data acquired by Kepler from pixel-level calibrations all the way to identifying transiting planet signatures and subjecting them to a suite of diagnostic tests to establish or break confidence in their planetary nature. Detecting small, rocky planets transiting Sun-like stars presents a variety of daunting challenges, from achieving an unprecedented photometric precision of 20 parts per million (ppm) on 6.5-hour timescales, supporting the science operations, management, processing, and repeated reprocessing of the accumulating data stream. This paper describes how the design of the SOC meets these varied challenges, discusses the architecture of the SOC and how the SOC pipeline is operated and is run on the NAS Pleiades supercomputer, and summarizes the most important pipeline features addressing the multiple computational, image and signal processing challenges posed by Kepler.

high performance computing↗

The Zwicky Transient Facility: System Overview, Performance, and First Results

The Zwicky Transient Facility (ZTF) is a new optical time-domain survey that uses the Palomar 48 inch Schmidt telescope. A custom-built wide-field camera provides a 47 deg ^(2) field of view and 8 s readout time, yielding more than an order of magnitude improvement in survey speed relative to its predecessor survey, the Palomar Transient Factory. We describe the design and implementation of the camera and observing system. The ZTF data system at the Infrared Processing and Analysis Center provides near-real-time reduction to identify moving and varying objects. We outline the analysis pipelines, data products, and associated archive. Finally, we present on-sky performance analysis and first scientific results from commissioning and the early survey. ZTF’s public alert stream will serve as a useful precursor for that of the Large Synoptic Survey Telescope.

Eric C. Bellm↗

Scalable Hybrid Learning Techniques for Scientific Data Compression

Data compression is becoming critical for storing scientific data because many scientific applications need to store large amounts of data and post process this data for scientific discovery. Unlike image and video compression algorithms that limit errors to primary data (PD), scientists require compression techniques that accurately preserve derived quantities of interest (QoIs). Here, this article presents a physics-informed compression technique implemented as an end-to-end, scalable, GPU-based pipeline for data compression that addresses this requirement. Our hybrid compression technique combines machine learning techniques and standard compression methods. Specifically, we combine an autoencoder, an error-bounded lossy compressor to provide guarantees on raw data error, and a constraint satisfaction post-processing step to preserve the QoIs within a minimal error (generally less than floating point error). The effectiveness of the data compression pipeline is demonstrated by compressing nuclear fusion simulation data generated by a large-scale fusion code, XGC, which produces hundreds of terabytes of data in a single day. Our approach works within the ADIOS framework and results in compression by a factor of more than 150 while requiring only a few percent of the computational resources necessary for generating the data, making the overall approach highly effective for practical scenarios.

ITER↗

Development of Computational Environmental Microbiome Workflows for the Laboratory and the International Space Station

Identification of microorganisms in the spaceflight environment is critical for crew health risk assessment on the International Space Station (ISS). Since 2017, nanopore sequencing technology has been used to support thein situ identification of microbial species during spaceflight. Beginning in 2018, a culture-independent, swab-to-sequencer method was implemented onboard the ISS to provide a more thorough insight of the ISS microbiome. Eliminating microbial culture enables identification of difficult-to-culture organisms, reduces risks associated with potentially pathogenic cultures, and could significantly reduce the time from sample-to-answer. However, this molecular-based approach generates large metagenomic datasets that require substantial computational resources for analysis. To process nanopore-generated sequencing data, the JSC Microbiology Laboratory established a bioinformatics workflow on Amazon EC2 under the security guidance of the NASA Science Managed Cloud Environment (SMCE).This resource allows for the development, testing, and accessing of computational tools for processing large and complex datasets. The work described here will address the downlinking of data from the ISS, the automated pipeline developed to identify targeted bacterial and fungal organisms, and the time from sampling onboard to microbial identification. The pipelines have been enhanced to address high and low biomass samples using optimization based on sample source (air, water, or surface) and type of collection (filter, colony, or swab).The resulting microbiome data can be assessed beyond microbial identifications to gain understanding toward population changes over time, potential selective environmental pressures, and evaluating correlations with a wide range of additional data sets. Metagenome analysis pipelines in development could allow for simultaneous identification of microbial species, gene function, and gene pathways present in the environment. Beyond the ground processing, the developed analysis pipeline is currently deployed onboard the ISS to allow for near real-time assessments of the ISS microbiome. This study serves as a critical foundation for exploration missions, where rapid microbiome analyses will be required.

G. Marie Sharp↗

Accelerating the identification of novel secondary metabolites in bioenergy plant root exudates using MicroED

Small molecule metabolites drive inter- and intraspecies communication and dependencies in diverse biological systems, yet a large proportion of these important chemical compounds remain uncharacterized in plants and microbes. Approximately 90% of the metabolites in root exudate profiles are unknown compounds, despite the importance of root exudate composition in plant-microbe interactions. We need advanced analytical capabilities that will support rapid discovery and structural elucidation of metabolites from biological samples that may be limited in quantity and high in complexity. To fill this gap, this project aimed to develop an integrated workflow involving metabolite extraction, separation, and crystallization from plant root exudates followed by characterization using nuclear magnetic resonance (NMR) spectroscopy, mass spectrometry, and microcrystal electron diffraction (MicroED). Using crude root exudates from sorghum, this project successfully developed higher throughput exudate fractionation strategies to obtain pure compounds for crystallization and identified crystals in multiple fractions that diffracted. Additional efforts to increase the throughput of high-quality crystal generation for MicroED, such as crystallization screening and crystallization chaperone exploration, will be needed to further advance root exudate metabolite identification. The overall optimized sample preparation process can then be integrated with the existing data collection and data analysis pipelines for MicroED at PNNL to facilitate more rapid natural product discovery.

59 BASIC BIOLOGICAL SCIENCES↗

Chlamydomonas reinhardtii responses to Fe-excess, Fe-deficiency, and Fe-limitation in either photoautotrophic or mixotrophic growth

A systems level analysis of Chlamydomonas reinhardtii grown photoautotrophically or mixotrophically with a reduced carbon source, acetate, under four different defined Fe stages of Fe-replete, Fe-deficient, Fe-limited, or Fe-excess. Samples were digested with trypsin, labeled with TMT 10-Plex, then analyzed by LC-MS/MS. Data was searched with MS-GF+ using PNNL's DMS Processing pipeline. [doi:10.25345/C5707X12X] [dataset license: CC0 1.0 Universal (CC0 1.0)]

59 BASIC BIOLOGICAL SCIENCES↗

Iron-starvation induces photosystem I antenna remodeling in green algae

Dunaliella salina and Dunaliella tertiolecta are extremophile, marine algae that can survive in very low Fe conditions. In this study, we used TMT-proteomics to compare the Fe starvation responses to the Fe replete responses. Samples were digested with trypsin, labeled with TMT 10-Plex, then analyzed by LC-MS/MS. Data was searched with MS-GF+ using PNNL's DMS Processing pipeline.

59 BASIC BIOLOGICAL SCIENCES↗

Simplified microprocessor design for VLSI control applications

A design technique for microprocessors combining the simplicity of reduced instruction set computers (RISC's) with the richer instruction sets of complex instruction set computers (CISC's) is presented. They utilize the pipelined instruction decode and datapaths common to RISC's. Instruction invariant data processing sequences which transparently support complex addressing modes permit the formulation of simple control circuitry. Compact implementations are possible since neither complicated controllers nor large register sets are required.

Cameron, K.↗

TPSAS-NF1676L-32493-DND

The Committee on Earth Observation Satellites (CEOS) System Engineering Office (SEO) has supported the Open Data Cube (ODC) initiative to provide a data architecture solution that has value to its global users and increases the impact of EO satellite data. ODC is an open-source platform for processing satellite data. We have developed software products and tools around the core ODC that would help users perform machine learning on EO satellite data. The recent United Nations (UN) Sustainable Development Agenda provides a shared blueprint for peace and prosperity for people and for the planet, considering our current situation and helping to create a plan. The core of this agenda is a set of seventeen Sustainable Development Goals (SDGs), which represent an urgent call for action by all countries - both developed and developing - in a global partnership. The CEOS SEO team has recently developed and released a set of innovative Jupyter notebooks addressing UN SDGs 6.6.1 (spatial extents of water-related ecosystems), 11.3.1 (ratio of land consumption rate to population growth rate), and 15.3.1 (proportion of land that is degraded over total land area). These notebooks empower users by providing features that will assist with streamlining analysis ready data retrieval, processing, and visualization. We have recently incorporated several machine learning techniques in these notebooks. In this paper, we present the lessons learned from our experience on classifying land using supervised and unsupervised machine learning techniques using ODC framework for UN SDGs. We identify the current limitations of ODC to seamlessly support machine learning techniques. We propose features that would help machine learning, specifically within the ODC framework. We propose a thematic indexing/loading of data for both unsupervised learning as well as data annotation/labeling pipeline. Currently, ODC supports machine learning by separating data-management from the analysis process. It works as a mechanism to load cubes of data. ODC does not natively support features that are vital in machine learning such as validation splits, fair/balanced sampling, establishing load size constraints, etc. We believe that our proposed features will empower users by providing features that bring machine learning techniques closed to ODC. Enhancements to ODC to better accommodate machine learning techniques can assist in fulfilling UN SDGs such as 6.3.2, 6.4.2, 6.6.1, 11.3.1, 14.1.1, 15.1.1, 15.3.1, and 15.4.2.

Syed R Rizvi↗

TESS Science Processing Operations Center Pipeline Status and Updates

The past eighteen months have seen a number of important changes for the TESS Science Processing Operations Center (SPOC) and our archival data products as TESS embarked upon its first extended mission. First, the SPOC developed and deployed a new 20-sec cadence pipeline, promising to unveil exciting new astrophysics at these short timescales for up to 1000 targets per observing sector. We also developed an FFI light curve pipeline that creates light curves and associated data products for up to 160,000 targets in each sector and archive these as High-Level Science Products (HLSP) at the Mikulski Archive for Space Telescopes (MAST). Soon we plan to perform transiting planet searches on these light curves and to release Data Validation reports and associated data products to the MAST. We also present results from the first multi-year transiting planet search of sectors 1 through 36. Finally, we discuss major changes to the SPOC pipeline that motivated the reprocessing of the first year of data, including the application of target-and cadence-specific scattered light flags, and an update to the sky background correction algorithm to mitigate bias in the original algorithm for dim and/or severely crowded stars.

TESS↗

Intelligent alarming

This talk discusses the importance of providing a process operator with concise information about a process fault including a root cause diagnosis of the problem, a suggested best action for correcting the fault, and prioritization of the problem set. A decision tree approach is used to illustrate one type of approach for determining the root cause of a problem. Fault detection in several different types of scenarios is addressed, including pump malfunctions and pipeline leaks. The talk stresses the need for a good data rectification strategy and good process models along with a method for presenting the findings to the process operator in a focused and understandable way. A real time expert system is discussed as an effective tool to help provide operators with this type of information. The use of expert systems in the analysis of actual versus predicted results from neural networks and other types of process models is discussed.

Braden, W. B.↗

Adapt: A Weather Radar Data Analysis and Nowcasting Platform for Informed Adaptive Scanning

SF-26-021 Adapt is a data processing platform for real-time data analysis, short term prediction of targets convective cells and tracking for archived data. It provides tools for downloading, processing, segmenting, projecting, analyzing, and visualizing storm cell data from weather radar. The pipeline includes cell detection, motion estimation using optical flow, cell property extraction, and persistence to NetCDF and SQLite/Parquet for guiding adaptive scanning.

Raut, Bhupendra Ashokrao [Argonne National Laborat↗

Long-term measurements of ice nucleating particles at Atmospheric Radiation Measurement (ARM) sites worldwide

Ice nucleating particles (INPs) play a critical role in cloud microphysics and precipitation formation, yet long-term, spatially extensive observational datasets remain limited. Here, we present one of the most comprehensive publicly available datasets of immersion-mode INP concentrations using a single analytical method, generated through the U.S. Department of Energy's (DOE) Atmospheric Radiation Measurement (ARM) user facility. INP filter samples have been collected across a broad range of environments – including agricultural plains, Arctic coastlines, high-elevation mountain sites, marine regions, and urban areas – via fixed observatories, mobile facility deployments, and vertically-resolved tethered balloon system operations. We describe the standardized processing and quality assurance pipeline, from filter collection and processing using the Ice Nucleation Spectrometer to final data products archived on the ARM Data Discovery portal. The dataset includes both total INP concentrations and selectively treated samples, allowing for classification of biological, organic, and inorganic INP types. It features a continuous 5-year record of INP measurements from a central U.S. site, with data collection still ongoing. Seasonal and site-specific differences in INP concentrations are illustrated through intercomparisons at −10 and −20 °C, revealing distinct regional sources and atmospheric drivers. We also outline mechanisms for researchers to access existing data, request additional sample analyses, and propose future field campaigns involving ARM INP measurements. This dataset supports a wide range of scientific applications, from observational and mechanistic studies to model development, and provides critical constraints on aerosol-cloud interactions across diverse atmospheric regimes (Creamean et al., 2024, 2020b; https://doi.org/10.5439/1770816).

Creamean, Jessie M. [Colorado State Univ., Fort Co↗

Onboard Experiment Data Support Facility

An onboard array structure has been devised for end to end processing of data from multiple spaceborne sensors. The array constitutes sets of programmable pipeline processors whose elements perform each assigned function in 0.25 microseconds. This space shuttle computer system can handle data rates from a few bits to over 100 megabits per second.

Source record↗

Development of the Science Data System for the International Space Station Cold Atom Lab

Cold Atom Laboratory (CAL) is a facility that will enable scientists to study ultra-cold quantum gases in a microgravity environment on the International Space Station (ISS) beginning in 2016. The primary science data for each experiment consists of two images taken in quick succession. The first image is of the trapped cold atoms and the second image is of the background. The two images are subtracted to obtain optical density. These raw Level 0 atom and background images are processed into the Level 1 optical density data product, and then into the Level 2 data products: atom number, Magneto-Optical Trap (MOT) lifetime, magnetic chip-trap atom lifetime, and condensate fraction. These products can also be used as diagnostics of the instrument health. With experiments being conducted for 8 hours every day, the amount of data being generated poses many technical challenges, such as downlinking and managing the required data volume. A parallel processing design is described, implemented, and benchmarked. In addition to optimizing the data pipeline, accuracy and speed in producing the Level 1 and 2 data products is key. Algorithms for feature recognition are explored, facilitating image cropping and accurate atom number calculations.

bose einstein condensate↗

Dynamic Black-Level Correction and Artifact Flagging in the Kepler Data Pipeline

Instrument-induced artifacts in the raw Kepler pixel data include time-varying crosstalk from the fine guidance sensor (FGS) clock signals, manifestations of drifting moiré pattern as locally correlated nonstationary noise and rolling bands in the images which find their way into the calibrated pixel time series and ultimately into the calibrated target flux time series. Using a combination of raw science pixel data, full frame images, reverse-clocked pixel data and ancillary temperature data the Keplerpipeline models and removes the FGS crosstalk artifacts by dynamically adjusting the black level correction. By examining the residuals to the model fits, the pipeline detects and flags spatial regions and time intervals of strong time-varying blacklevel (rolling bands ) on a per row per cadence basis. These flags are made available to downstream users of the data since the uncorrected rolling band artifacts could complicate processing or lead to misinterpretation of instrument behavior as stellar. This model fitting and artifact flagging is performed within the new stand-alone pipeline model called Dynablack. We discuss the implementation of Dynablack in the Kepler data pipeline and present results regarding the improvement in calibrated pixels and the expected improvement in cotrending performances as a result of including FGS corrections in the calibration. We also discuss the effectiveness of the rolling band flagging for downstream users and illustrate with some affected light curves.

Clarke, B. D.↗

Robust Mosaicking of Stereo Digital Elevation Models from the Ames Stereo Pipeline

Robust estimation method is proposed to combine multiple observations and create consistent, accurate, dense Digital Elevation Models (DEMs) from lunar orbital imagery. The NASA Ames Intelligent Robotics Group (IRG) aims to produce higher-quality terrain reconstructions of the Moon from Apollo Metric Camera (AMC) data than is currently possible. In particular, IRG makes use of a stereo vision process, the Ames Stereo Pipeline (ASP), to automatically generate DEMs from consecutive AMC image pairs. However, the DEMs currently produced by the ASP often contain errors and inconsistencies due to image noise, shadows, etc. The proposed method addresses this problem by making use of multiple observations and by considering their goodness of fit to improve both the accuracy and robustness of the estimate. The stepwise regression method is applied to estimate the relaxed weight of each observation.

Kim, Tae Min↗