Search NASASearch

SEARCH · Search NASA

Results for “pipeline data processing”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 55 records · Page 3

Photometer Performance Assessment in Kepler Science Data Processing

This paper describes the algorithms of the Photometer Performance Assessment (PPA) software component in the science data processing pipeline of the Kepler mission. The PPA performs two tasks: One is to analyze the health and performance of the Kepler photometer based on the long cadence science data down-linked via Ka band approximately every 30 days. The second is to determine the attitude of the Kepler spacecraft with high precision at each long cadence. The PPA component is demonstrated to work effectively with the Kepler flight data.

Li, Jie

Dynamic Black-Level Correction and Artifact Flagging for Kepler Pixel Time Series

Methods applied to the calibration stage of Kepler pipeline data processing [1] (CAL) do not currently use all of the information available to identify and correct several instrument-induced artifacts. These include time-varying crosstalk from the fine guidance sensor (FGS) clock signals, and manifestations of drifting moire pattern as locally correlated nonstationary noise, and rolling bands in the images which find their way into the time series [2], [3]. As the Kepler Mission continues to improve the fidelity of its science data products, we are evaluating the benefits of adding pipeline steps to more completely model and dynamically correct the FGS crosstalk, then use the residuals from these model fits to detect and flag spatial regions and time intervals of strong time-varying black-level which may complicate later processing or lead to misinterpretation of instrument behavior as stellar activity.

Kolodziejczak, J. J.

Kepler Data Validation I: Architecture, Diagnostic Tests, and Data Products for Vetting Transiting Planet Candidates

The Kepler Mission was designed to identify and characterize transiting planets in the Kepler Field of View and to determine their occurrence rates. Emphasis was placed on identification of Earth-size planets orbiting in the Habitable Zone of their host stars. Science data were acquired for a period of four years. Long-cadence data with 29.4 min sampling were obtained for approx. 200,000 individual stellar targets in at least one observing quarter in the primary Kepler Mission. Light curves for target stars are extracted in the Kepler Science Data Processing Pipeline, and are searched for transiting planet signatures. A Threshold Crossing Event is generated in the transit search for targets where the transit detection threshold is exceeded and transit consistency checks are satisfied. These targets are subjected to further scrutiny in the Data Validation (DV) component of the Pipeline. Transiting planet candidates are characterized in DV, and light curves are searched for additional planets after transit signatures are modeled and removed. A suite of diagnostic tests is performed on all candidates to aid in discrimination between genuine transiting planets and instrumental or astrophysical false positives. Data products are generated per target and planet candidate to document and display transiting planet model fit and diagnostic test results. These products are exported to the Exoplanet Archive at the NASA Exoplanet Science Institute, and are available to the community. We describe the DV architecture and diagnostic tests, and provide a brief overview of the data products. Transiting planet modeling and the search for multiple planets on individual targets are described in a companion paper. The final revision of the Kepler Pipeline code base is available to the general public through GitHub. The Kepler Pipeline has also been modified to support the Transiting Exoplanet Survey Satellite (TESS) Mission which is expected to commence in 2018.

data analysis

Kepler Planet Detection Metrics: Statistical Bootstrap Test

This document describes the data produced by the Statistical Bootstrap Test over the final three Threshold Crossing Event (TCE) deliveries to NExScI: SOC 9.1 (Q1Q16)1 (Tenenbaum et al. 2014), SOC 9.2 (Q1Q17) aka DR242 (Seader et al. 2015), and SOC 9.3 (Q1Q17) aka DR253 (Twicken et al. 2016). The last few years have seen significant improvements in the SOC science data processing pipeline, leading to higher quality light curves and more sensitive transit searches. The statistical bootstrap analysis results presented here and the numerical results archived at NASAs Exoplanet Science Institute (NExScI) bear witness to these software improvements. This document attempts to introduce and describe the main features and differences between these three data sets as a consequence of the software changes.

Bootstrap

Software for Verifying Image-Correlation Tie Points

A computer program enables assessment of the quality of tie points in the image-correlation processes of the software described in the immediately preceding article. Tie points are computed in mappings between corresponding pixels in the left and right images of a stereoscopic pair. The mappings are sometimes not perfect because image data can be noisy and parallax can cause some points to appear in one image but not the other. The present computer program relies on the availability of a left- right correlation map in addition to the usual right left correlation map. The additional map must be generated, which doubles the processing time. Such increased time can now be afforded in the data-processing pipeline, since the time for map generation is now reduced from about 60 to 3 minutes by the parallelization discussed in the previous article. Parallel cluster processing time, therefore, enabled this better science result. The first mapping is typically from a point (denoted by coordinates x,y) in the left image to a point (x',y') in the right image. The second mapping is from (x',y') in the right image to some point (x",y") in the left image. If (x,y) and(x",y") are identical, then the mapping is considered perfect. The perfect-match criterion can be relaxed by introducing an error window that admits of round-off error and a small amount of noise. The mapping procedure can be repeated until all points in each image not connected to points in the other image are eliminated, so that what remains are verified correlation data.

Klimeck, Gerhard

Optimizing Single Nuclei Sequencing of Brain Samples From Space Flown Mice Across Age and Strain

The NASA GeneLab Sample Processing Laboratory offers high-throughput sequencing services to NASA-funded space biology researchers. Space biology studies have specific challenges such as low sample numbers, introducing susceptibility to batch effects from sample handling. These issues are compounded by complex protocols such as single-nuclei isolation and sequencing, which has recently become an attractive methodology for assessing the cellular diversity within spaceflight samples. High quality single-nuclei sequencing requires reproducible protocols to dissociate tissue and generate clean suspension of intact single nuclei. Producing single-nuclei suspension from brain tissue is particularly challenging due to cell type heterogeneity and the myelin sheath that carries over into the nuclei suspension as debris. Current procedures tend to be time consuming and sometimes include steps that can alter gene expression and create cell-type bias. Commercially available nuclei isolation kits, such as the 10X Genomics nuclei isolation kit, offers a streamlined way to process samples for nuclei isolation, thereby minimizing batch effects and enabling reproducibility. In this study, we report on the performance of the 10X Genomics nuclei isolation kit and Chromium Next GEM Single Cell Multiome ATAC + Gene Expression kit to generate sequencing libraries from space-flown mouse brain samples. Single nuclei sequencing was performed on frozen mouse brain tissue from two spaceflight missions, Rodent Research-10 (RR-10) and RR Reference Mission-2 (RRRM-2). RR-10 mice were female B6129SF2/J, euthanized at 18-19 weeks whereas RRRM-2 mice were female C57BL/6NTac, euthanized at 20 or 37 weeks. Sequencing data was processed using standard GeneLab data processing pipelines. We report evaluation of the performance of the 10X Genomics nuclei isolation kit for spaceflight samples from mouse brain, and evaluation of reproducibility across different mouse strains and age groups. We also report preliminary scientific results including cell type inference, cell clustering, and differentially expressed genes and pathways between spaceflight and ground control samples.

RR-10

Automated X-ray and Optical Analysis of the Virtual Observatory and Grid Computing

We are developing a system to combine the Web Enabled Source Identification with X-Matching (WESIX) web service, which emphasizes source detection on optical images,with the XAssist program that automates the analysis of X-ray data. XAssist is continuously processing archival X-ray data in several pipelines. We have established a workflow in which FITS images and/or (in the case of X ray data) an X-ray field can be input to WESIX. Intelligent services return available data (if requested fields have been processed) or submit job requests to a queue to be performed asynchronously. These services will be available via web services (for non-interactive use by Virtual Observatory portals and applications) and through web applications (written in the Django web application framework). We are adding web services for specific XAssist functionality such as determining .the exposure and limiting flux for a given position on the sky and extracting spectra and images for a given region. We are improving the queuing system in XAssist to allow for "watch lists" to be specified by users, and when X-ray fields in a user's watch list become publicly available they will be automatically added to the queue. XAssist is being expanded to be used as a survey planning 1001 when coupled with simulation software, including functionality for NuStar, eRosita, IXO, and the Wide Field Xray Telescope (WFXT), as part of an end to end simulation/analysis system. We are also investigating the possibility of a dedicated iPhone/iPad app for querying pipeline data, requesting processing, and administrative job control.

Ptak, A.

TESS Science Processing Operations Center Pipeline and Data Products

TESS (Transiting Exoplanet Survey Satellite) launched on 18-4-2018 to conduct a two-year, near all-sky survey for at least 50 nearby exoplanets for which masses can be obtained. TESS just completed surveying the southern hemisphere, identifying hundreds of candidate exoplanet systems and unveiling a plethora of exciting non-exoplanet astrophysics results, such as asteroseismology, asteroids, and supernova. The TESS Science Processing Operations Center (SPOC) at NASA Ames Research Center processes the image data downlinked from TESS every two weeks to generate a variety of data products hosted at the Mikulski Archive for Space Telescopes (MAST). For each approximately 1-month sector, the SPOC calibrates the image data for both 30-minute Full Frame Images (FFIs) and up to 20,000 pre-selected 2-minute target star postage stamps. Simple aperture photometry and systematic error-corrected flux-time series are generated for the 2-minute data. The data products also include co-trending basis vectors (CBVs) and calibration files, such as the Pixel Response Functions (PRF). The archival files are modeled after Kepler's for ease of use, and include Target Pixel Files (TPFs) containing original and calibrated 2-minute image data, Light Curve files (LCs) containing the photometric time series for each 2-minute target, as well as the Data Validation products. New products derived from the FFIs include light curves for the 2-minute targets and CBVs. The TESS Mission is funded by NASA's Science Mission Directorate as an Astrophysics Explorer Mission.

Science Pipeline

TESS Science Processing Operations Center Pipeline and Data Products

TESS launched 18-4-2018 to conduct a two-year, near all-sky survey for at least 50 nearby exoplanets for which masses can be. TESS just completed surveying the southern hemisphere, identifying hundreds of candidate exoplanet systems and unveiling a plethora of exciting non-exoplanet astrophysics results, such as asteroseismology, asteroids, and supernova. The TESS Science Processing Operations Center (SPOC) at NASA Ames Research Center processes the image data downlinked from TESS every two weeks to generate a variety of data products hosted at the Mikulski Archive for Space Telescopes (MAST). For each ~1 month sector, the SPOC calibrates the image data for both 30-min Full Frame Images (FFIs) and up to 20,000 pre-selected 2-min target star postage stamps. Simple aperture photometry and systematic error-corrected flux time series are generated for the 2-min data. The data products also include co-trending basis vectors (CBVs) and calibration files, such as the Pixel Response Functions (PRF). The archival files are modeled after Kepler's for ease of use, and include Target Pixel Files (TPFs) containing original and calibrated 2-min image data, Light Curve files (LCs) containing the photometric time series for each 2-min target, as well as the Data Validation products. New products derived from the FFIs include light curves for the 2-min targets and CBVs. The TESS Mission is funded by NASA's Science Mission Directorate as an Astrophysics Explorer Mission.

Jenkins, Jon M.

TESS Science Processing Operations Center Pipeline and Data Products

TESS launched 18 April 2018 to conduct a two-year, near all-sky survey for at least 50 small, nearby exoplanets for which masses can be ascertained and whose atmospheres can be characterized by ground- and space-based follow-on observations. TESS just completed its survey of the southern hemisphere, identifying >600 candidate exoplanets and unveiling a plethora of exciting non-exoplanet astrophysics results, such as asteroseismology, asteroids, and supernova. The TESS Science Processing Operations Center (SPOC) processes the data downlinked every two weeks to generate a range of data products hosted at the Mikulski Archive for Space Telescopes (MAST). For each sector (~1 month) of observations, the SPOC calibrates the image data for both 30-min Full Frame Images (FFIs) and up to 20,000 pre-selected 2-min target star postage stamps. Data products for the 2-min targets include simple aperture photometry and systematic error-corrected flux time series. The SPOC also conducts searches for transiting exoplanets in the 2-min data for each sector and generates Data Validation time series and associated reports for each transit-like feature identified in the search. Multi-sector searches for exoplanets are conducted periodically to discover longer period planets, including those in the James Webb Continuous Viewing Zone (CVZ), which are observed for up to one year. Data products also include co-trending basis vectors (CBVs) and calibration files, such as the Pixel Response Functions across the field of view of each of TESS's four cameras. To maximize the usability, the TESS science data products are modeled after those for Kepler, including Target Pixel Files and Light Curve files.In this talk, I describe the SPOC pipeline and the chief differences between the TESS and the Kepler pipelines, and the major updates to the SPOC pipeline (4.0) available now to the community at MAST. I also discuss the documentation available to the community to help them in properly interpreting and analyzing the TESS data products.The TESS Mission is funded by NASA's Science Mission Directorate as an Astrophysics Explorer Mission.

Jenkins, Jon M.

NASA GeneLab RNASeq Consensus Pipeline: A Nextflow Implementation

The NASA GeneLab project (genelab.nasa.gov) seeks to accelerate space biology research through cataloging and democratizing omics data. Since raw omics data is largely inaccessible to non-bioinformaticians, GeneLab works with the scientific community to develop standard processing pipelines to generate and publish processed data. Unlike raw data, processed data has greater immediate value to a wide range of users with varying technical backgrounds and computational capabilities. Standardizing processing workflows is essential to match the pace of raw data generation, ensure reproducibility, and enable standardized processed data for comparison across datasets. Previously, GeneLab developed a standardized pipeline for processing RNAseq data, referred to as the ‘GeneLab RNAseq Consensus Pipeline (RCP)’, in collaboration with GeneLab’s Analysis Working Groups. The work presented here is a Nextflow implementation of GeneLab’s RCP that automates and accelerates data processing of RNASeq datasets hosted on GeneLab. In addition to the core data processing, the workflow also includes staging of GeneLab raw data and a robust verification and validation (V&V) program that runs after each processing step to identify errors in real-time, stop additional downstream computation, and preserve computational resources. The workflow, including the staging and V&V functionality, is open source for others to reuse and modify at https://github.com/nasa/GeneLab_Data_Processing/tree/master/RNAseq.

Jonathan Dejesus Oribello

NASA GeneLab RNASeq Consensus Pipeline: A Nextflow Implementation

The NASA GeneLab project (genelab.nasa.gov) seeks to accelerate space biology research through cataloging and democratizing omics data. Since raw omics data is largely inaccessible to non-bioinformaticians, GeneLab works with the scientific community to develop standard processing pipelines to generate and publish processed data. Unlike raw data, processed data has greater immediate value to a wide range of users with varying technical backgrounds and computational capabilities. Standardizing processing workflows is essential to match the pace of raw data generation, ensure reproducibility, and enable standardized processed data for comparison across datasets. Previously, GeneLab developed a standardized pipeline for processing RNAseq data, referred to as the ‘GeneLab RNAseq Consensus Pipeline (RCP)’, in collaboration with GeneLab’s Analysis Working Groups. The work presented here is a Nextflow implementation of GeneLab’s RCP that automates and accelerates data processing of RNASeq datasets hosted on GeneLab. In addition to the core data processing, the workflow also includes staging of GeneLab raw data and a robust verification and validation (V&V) program that runs after each processing step to identify errors in real-time, stop additional downstream computation, and preserve computational resources. The workflow, including the staging and V&V functionality, is open source for others to reuse and modify at https://github.com/nasa/GeneLab_Data_Processing/tree/master/RNAseq.

Jonathan D Oribello

Surface Biology & Geology Pathfinder Data Analysis Pipeline

NASA's future global orbital mission, currently in development as the Surface Biology and Geology (SBG) Designated Observable study, will acquire relatively high resolution solar-reflected spectroscopy and thermal infrared observations. Innovative processes must be utilized for handling the high volume of data anticipated to be collected, which is anticipated to exceed 100 terabytes/day, greater than NASA's total extant airborne hyperspectral data collection. Collecting, processing/re-processing, disseminating, and exploiting this volume of data presents new challenges. To begin addressing them, NASA is drawing upon the expertise developed from its astrophysics programs to address Earth science and applications. Specifically, NASA is adapting the science processing operations technology developed for the Kepler and TESS planet-hunting missions for imaging spectroscopy data processing. This technology development has been the foundation for the remarkable scientific successes of Kepler and TESS. The Kepler/TESS data processing technology provides a scalable architecture for robust, repeatable, and replicable science and application products while enabling the Earth science community to develop, test, and implement new algorithms. Our effort to leverage this existing capability has begun by ingesting data and applying workflows from the EO-1/Hyperion 17-year mission archive that provides globally sampled visible through shortwave infrared spectra that are representative of SBG data types and volumes. This pathfinding data processing system will help define the solutions to processing SBG data volumes and will enable the scientific community to interact with the data and processing pipeline to create new science products.

Jenkins, Jon

First Results of Venus Express Spacecraft Observations with Wettzell

The ESA Venus Express spacecraft was observed at X-band with the Wettzell radio telescope in October-December 2009 in the framework of an assessment study of the possible contribution of the European VLBI Network to the upcoming ESA deep space missions. A major goal of these observations was to develop and test the scheduling, data capture, transfer, processing, and analysis pipeline. Recorded data were transferred from Wettzell to Metsahovi for processing, and the processed data were sent from Mets ahovi to JIVE for analysis. A turnover time of 24 hours from observations to analysis results was achieved. The high dynamic range of the detections allowed us to achieve a milliHz level of spectral resolution accuracy and to extract the phase of the spacecraft signal carrier line. Several physical parameters can be determined from these observational results with more observational data collected. Among other important results, the measured phase fluctuations of the carrier line at different time scales can be used to determine the influence of the solar wind plasma density fluctuations on the accuracy of the astrometric VLBI observations.

Calves, Guifre Molera

Enabling Model Organism and Commercial Astronaut Data Access Through the NASA Open Science Data Repository

NASA’s Open Science Data Repository (OSDR) brings together omics data from NASA’s GeneLab project and non-omics data, including physiological, phenotypic, imaging, and behavioral data from NASA’s Ames Life Sciences Data Archive (ALSDA) collected from decades of space biology research, providing open and FAIR (findable, accessible, interoperable, and reusable) access of these precious data to scientists world-wide. This rich source of meticulously curated metadata and data from spaceflight and analog studies has been mined by the scientific community resulting in dozens of high impact scientific publications that reveals a complex network of molecular and physiological effects of spaceflight across living systems, from microbes to plants, to mammals. Understanding how these effects translate to the human condition is critical as we move deeper into the era of commercial space travel. However, the integration of data, specifically omics data, from astronauts is particularly challenging due to their sensitive nature. OSDR has risen to this challenge by developing a mechanism to control access to identifiable levels of omics data, such as raw sequence data, while enabling public access to processed, unidentifiable, data and associated metadata that will allow the scientific community to interrogate human astronaut data alongside data from model organisms to begin answering these critical questions. The 2021 SpaceX Inspiration4 (I4) mission collected a comprehensive atlas of biological measurements from four civilian astronauts, providing a wealth of data to characterize the effects of spaceflight on the human body. These data include both non-omics and omics assays such as direct RNA sequencing (RNA-seq), single nuclei ATAC-seq and RNA-seq, metagenomics, proteomics, and comprehensive metabolic and cytokine panels, all of which have been integrated into the OSDR system across no less than 9 studies. Each study has been carefully curated using community-backed OSDR standards for sample and assay level metadata ensuring these data are findable and accessible. In addition to hosting both raw and processed data from the principal investigator team for each assay type, the GeneLab team plans to re-process the I4 omics data using GeneLab’s standard processing pipelines. The GeneLab processed data outputs will allow for comparisons across studies on OSDR and enable visualization of these data through the OSDR data visualization platform thereby enabling data reusability and interoperability. Here we describe the robust privacy and security protocols implemented by OSDR to safeguard sensitive health data from astronauts while facilitating metadata and processed data sharing for research purposes. We further provide a road map for navigating the vast amount of data provided for each I4 study on the OSDR, including experimental design, associated experiments, payloads, and missions, data generation and analysis protocols, and associated scientific articles. Additionally, we illustrate how to interrogate the standardized metadata provided in the sample and assay tables as well as various means to download and access the data including programmatically through the GeneLab Open API (GLOpenAPI). The open access of datasets in NASA’s OSDR provides a unique opportunity for the scientific community, as well as citizen scientists and students, to continue using OSDR resources to further unlock profound insights into the consequences of space travel on the human body. Through implementation of security measures to protect sensitive human data, the OSDR seeks to strengthen the science exchange between the Biological and Physical Sciences Program and the Human Research Program, per recommendation 4-1 of the 2023-2032 Decadal Survey, and encourage further sharing and dissemination of astronaut data to provide the scientific community with the resources needed to lay the groundwork for developing targeted mitigation strategies to help withstand the rigors of long-duration spaceflight.

Amanda Marie Saravia-butler

Enabling Model Organism and Commercial Astronaut Data Access Through the NASA Open Science Data Repository

NASA’s Open Science Data Repository (OSDR) brings together omics data from NASA’s GeneLab project and non-omics data, including physiological, phenotypic, imaging, and behavioral data from NASA’s Ames Life Sciences Data Archive (ALSDA) collected from decades of space biology research, providing open and FAIR (findable, accessible, interoperable, and reusable) access of these precious data to scientists world-wide. This rich source of meticulously curated metadata and data from spaceflight and analog studies has been mined by the scientific community resulting in dozens of high impact scientific publications that reveals a complex network of molecular and physiological effects of spaceflight across living systems, from microbes to plants, to mammals. Understanding how these effects translate to the human condition is critical as we move deeper into the era of commercial space travel. However, the integration of data, specifically omics data, from astronauts is particularly challenging due to their sensitive nature. OSDR has risen to this challenge by developing a mechanism to control access to identifiable levels of omics data, such as raw sequence data, while enabling public access to processed, unidentifiable, data and associated metadata that will allow the scientific community to interrogate human astronaut data alongside data from model organisms to begin answering these critical questions. The 2021 SpaceX Inspiration4 (I4) mission collected a comprehensive atlas of biological measurements from four civilian astronauts, providing a wealth of data to characterize the effects of spaceflight on the human body. These data include both non-omics and omics assays such as direct RNA sequencing (RNA-seq), single nuclei ATAC-seq and RNA-seq, metagenomics, proteomics, and comprehensive metabolic and cytokine panels, all of which have been integrated into the OSDR system across no less than 9 studies. Each study has been carefully curated using community-backed OSDR standards for sample and assay level metadata ensuring these data are findable and accessible. In addition to hosting both raw and processed data from the principal investigator team for each assay type, the GeneLab team plans to re-process the I4 omics data using GeneLab’s standard processing pipelines. The GeneLab processed data outputs will allow for comparisons across studies on OSDR and enable visualization of these data through the OSDR data visualization platform thereby enabling data reusability and interoperability. Here we describe the robust privacy and security protocols implemented by OSDR to safeguard sensitive health data from astronauts while facilitating metadata and processed data sharing for research purposes. We further provide a road map for navigating the vast amount of data provided for each I4 study on the OSDR, including experimental design, associated experiments, payloads, and missions, data generation and analysis protocols, and associated scientific articles. Additionally, we illustrate how to interrogate the standardized metadata provided in the sample and assay tables as well as instructions for how to download and access the data. The I4 datasets described here re present the first ever comprehensive collection of commercial astronaut data.

Amanda M Saravia-Butler

Modular on-board adaptive imaging

Feature extraction involves the transformation of a raw video image to a more compact representation of the scene in which relevant information about objects of interest is retained. The task of the low-level processor is to extract object outlines and pass the data to the high-level process in a format that facilitates pattern recognition tasks. Due to the immense computational load caused by processing a 256x256 image, even a fast minicomputer requires a few seconds to complete this low-level processing. It is, therefore, necessary to consider hardware implementation of these low-level functions to achieve real-time processing speeds. The considered project had the objective to implement a system in which the continuous feature extraction process is not affected by the dynamic changes in the scene, varying lighting conditions, or object motion relative to the cameras. Due to the high bandwidth (3.5 MHz) and serial nature of the TV data, a pipeline processing scheme was adopted as the overall architecture of this system. Modularity in the system is achieved by designing circuits that are generic within the overall system.

Eskenazi, R.

Parallel processing in a host plus multiple array processor system for radar

Host plus multiple array processor architecture is demonstrated to yield a modular, fast, and cost-effective system for radar processing. Software methodology for programming such a system is developed. Parallel processing with pipelined data flow among the host, array processors, and discs is implemented. Theoretical analysis of performance is made and experimentally verified. The broad class of problems to which the architecture and methodology can be applied is indicated.

Barkan, B. Z.