Search NASASearch

SEARCH · Search NASA

Results for “pipeline data processing”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 73 records · Page 4

Intelligent experiments through real-time AI: Fast Data Processing and Autonomous Detector Control for sPHENIX and future EIC detectors

This R&D project, initiated by the DOE Nuclear Physics AI-Machine Learning initiative in 2022, leverages AI to address data processing challenges in high-energy nuclear experiments (RHIC, LHC, and future EIC). Our focus is on developing a demonstrator for real-time processing of high-rate data streams from sPHENIX experiment tracking detectors. The limitations of a 15 kHz maximum trigger rate imposed by the calorimeters can be negated by intelligent use of streaming technology in the tracking system. The approach efficiently identifies low momentum rare heavy flavor events in high-rate p+p collisions (3MHz), using Graph Neural Network (GNN) and High Level Synthesis for Machine Learning (hls4ml). Success at sPHENIX promises immediate benefits, minimizing resources and accelerating the heavy-flavor measurements. The approach is transferable to other fields. For the EIC, we develop a DIS-electron tagger using Artificial Intelligence - Machine Learning (AI-ML) algorithms for real-time identification, showcasing the transformative potential of AI and FPGA technologies in high-energy nuclear and particle experiments real-time data processing pipelines.

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND

SITCOMTN-149: An Interim Report on the LSSTComCam On-Sky Campaign

From 24 October to 11 December 2024, the NSF-DOE Vera C. Rubin Observatory conducted an on-sky campaign using the engineering LSST Commissioning Camera (LSSTComCam) to test the end-to-end functionality of hardware and software, as well as operational procedures. This interim report provides a preliminary technical overview of our understanding of the integrated system performance based tests and analyses conducted during the LSSTComCam on-sky campaign. The objectives are to synthesize what we have learned about the system in a timely way to inform on-going commissioning efforts, and to inform the Rubin science community on the progress of the LSSTComCam on-sky campaign. The report is organized into sections to describe major activities during the campaign, as well as multiple aspects of the demonstrated system and science performance. All of the results presented here are to be understood as work in progress using engineering data and the initial versions of the data processing pipelines; the report is a living document that will be updated as analyses are refined.

79 ASTRONOMY AND ASTROPHYSICS

Optimizing Single Nuclei Sequencing of Brain Samples From Space Flown Mice Across Age and Strain

The NASA GeneLab Sample Processing Laboratory offers high-throughput sequencing services to NASA-funded space biology researchers. Space biology studies have specific challenges such as low sample numbers, introducing susceptibility to batch effects from sample handling. These issues are compounded by complex protocols such as single-nuclei isolation and sequencing, which has recently become an attractive methodology for assessing the cellular diversity within spaceflight samples. High quality single-nuclei sequencing requires reproducible protocols to dissociate tissue and generate clean suspension of intact single nuclei. Producing single-nuclei suspension from brain tissue is particularly challenging due to cell type heterogeneity and the myelin sheath that carries over into the nuclei suspension as debris. Current procedures tend to be time consuming and sometimes include steps that can alter gene expression and create cell-type bias. Commercially available nuclei isolation kits, such as the 10X Genomics nuclei isolation kit, offers a streamlined way to process samples for nuclei isolation, thereby minimizing batch effects and enabling reproducibility. In this study, we report on the performance of the 10X Genomics nuclei isolation kit and Chromium Next GEM Single Cell Multiome ATAC + Gene Expression kit to generate sequencing libraries from space-flown mouse brain samples. Single nuclei sequencing was performed on frozen mouse brain tissue from two spaceflight missions, Rodent Research-10 (RR-10) and RR Reference Mission-2 (RRRM-2). RR-10 mice were female B6129SF2/J, euthanized at 18-19 weeks whereas RRRM-2 mice were female C57BL/6NTac, euthanized at 20 or 37 weeks. Sequencing data was processed using standard GeneLab data processing pipelines. We report evaluation of the performance of the 10X Genomics nuclei isolation kit for spaceflight samples from mouse brain, and evaluation of reproducibility across different mouse strains and age groups. We also report preliminary scientific results including cell type inference, cell clustering, and differentially expressed genes and pathways between spaceflight and ground control samples.

RR-10

Automated X-ray and Optical Analysis of the Virtual Observatory and Grid Computing

We are developing a system to combine the Web Enabled Source Identification with X-Matching (WESIX) web service, which emphasizes source detection on optical images,with the XAssist program that automates the analysis of X-ray data. XAssist is continuously processing archival X-ray data in several pipelines. We have established a workflow in which FITS images and/or (in the case of X ray data) an X-ray field can be input to WESIX. Intelligent services return available data (if requested fields have been processed) or submit job requests to a queue to be performed asynchronously. These services will be available via web services (for non-interactive use by Virtual Observatory portals and applications) and through web applications (written in the Django web application framework). We are adding web services for specific XAssist functionality such as determining .the exposure and limiting flux for a given position on the sky and extracting spectra and images for a given region. We are improving the queuing system in XAssist to allow for "watch lists" to be specified by users, and when X-ray fields in a user's watch list become publicly available they will be automatically added to the queue. XAssist is being expanded to be used as a survey planning 1001 when coupled with simulation software, including functionality for NuStar, eRosita, IXO, and the Wide Field Xray Telescope (WFXT), as part of an end to end simulation/analysis system. We are also investigating the possibility of a dedicated iPhone/iPad app for querying pipeline data, requesting processing, and administrative job control.

Ptak, A.

TESS Science Processing Operations Center Pipeline and Data Products

TESS (Transiting Exoplanet Survey Satellite) launched on 18-4-2018 to conduct a two-year, near all-sky survey for at least 50 nearby exoplanets for which masses can be obtained. TESS just completed surveying the southern hemisphere, identifying hundreds of candidate exoplanet systems and unveiling a plethora of exciting non-exoplanet astrophysics results, such as asteroseismology, asteroids, and supernova. The TESS Science Processing Operations Center (SPOC) at NASA Ames Research Center processes the image data downlinked from TESS every two weeks to generate a variety of data products hosted at the Mikulski Archive for Space Telescopes (MAST). For each approximately 1-month sector, the SPOC calibrates the image data for both 30-minute Full Frame Images (FFIs) and up to 20,000 pre-selected 2-minute target star postage stamps. Simple aperture photometry and systematic error-corrected flux-time series are generated for the 2-minute data. The data products also include co-trending basis vectors (CBVs) and calibration files, such as the Pixel Response Functions (PRF). The archival files are modeled after Kepler's for ease of use, and include Target Pixel Files (TPFs) containing original and calibrated 2-minute image data, Light Curve files (LCs) containing the photometric time series for each 2-minute target, as well as the Data Validation products. New products derived from the FFIs include light curves for the 2-minute targets and CBVs. The TESS Mission is funded by NASA's Science Mission Directorate as an Astrophysics Explorer Mission.

Science Pipeline

TESS Science Processing Operations Center Pipeline and Data Products

TESS launched 18-4-2018 to conduct a two-year, near all-sky survey for at least 50 nearby exoplanets for which masses can be. TESS just completed surveying the southern hemisphere, identifying hundreds of candidate exoplanet systems and unveiling a plethora of exciting non-exoplanet astrophysics results, such as asteroseismology, asteroids, and supernova. The TESS Science Processing Operations Center (SPOC) at NASA Ames Research Center processes the image data downlinked from TESS every two weeks to generate a variety of data products hosted at the Mikulski Archive for Space Telescopes (MAST). For each ~1 month sector, the SPOC calibrates the image data for both 30-min Full Frame Images (FFIs) and up to 20,000 pre-selected 2-min target star postage stamps. Simple aperture photometry and systematic error-corrected flux time series are generated for the 2-min data. The data products also include co-trending basis vectors (CBVs) and calibration files, such as the Pixel Response Functions (PRF). The archival files are modeled after Kepler's for ease of use, and include Target Pixel Files (TPFs) containing original and calibrated 2-min image data, Light Curve files (LCs) containing the photometric time series for each 2-min target, as well as the Data Validation products. New products derived from the FFIs include light curves for the 2-min targets and CBVs. The TESS Mission is funded by NASA's Science Mission Directorate as an Astrophysics Explorer Mission.

Jenkins, Jon M.

TESS Science Processing Operations Center Pipeline and Data Products

TESS launched 18 April 2018 to conduct a two-year, near all-sky survey for at least 50 small, nearby exoplanets for which masses can be ascertained and whose atmospheres can be characterized by ground- and space-based follow-on observations. TESS just completed its survey of the southern hemisphere, identifying >600 candidate exoplanets and unveiling a plethora of exciting non-exoplanet astrophysics results, such as asteroseismology, asteroids, and supernova. The TESS Science Processing Operations Center (SPOC) processes the data downlinked every two weeks to generate a range of data products hosted at the Mikulski Archive for Space Telescopes (MAST). For each sector (~1 month) of observations, the SPOC calibrates the image data for both 30-min Full Frame Images (FFIs) and up to 20,000 pre-selected 2-min target star postage stamps. Data products for the 2-min targets include simple aperture photometry and systematic error-corrected flux time series. The SPOC also conducts searches for transiting exoplanets in the 2-min data for each sector and generates Data Validation time series and associated reports for each transit-like feature identified in the search. Multi-sector searches for exoplanets are conducted periodically to discover longer period planets, including those in the James Webb Continuous Viewing Zone (CVZ), which are observed for up to one year. Data products also include co-trending basis vectors (CBVs) and calibration files, such as the Pixel Response Functions across the field of view of each of TESS's four cameras. To maximize the usability, the TESS science data products are modeled after those for Kepler, including Target Pixel Files and Light Curve files.In this talk, I describe the SPOC pipeline and the chief differences between the TESS and the Kepler pipelines, and the major updates to the SPOC pipeline (4.0) available now to the community at MAST. I also discuss the documentation available to the community to help them in properly interpreting and analyzing the TESS data products.The TESS Mission is funded by NASA's Science Mission Directorate as an Astrophysics Explorer Mission.

Jenkins, Jon M.

TomoPyUI : a user-friendly tool for rapid tomography alignment and reconstruction

The management and processing of synchrotron and neutron computed tomography data can be a complex, labor-intensive and unstructured process. Users devote substantial time to both manually processing their data ( i.e. organizing data/metadata, applying image filters etc. ) and waiting for the computation of iterative alignment and reconstruction algorithms to finish. In this work, we present a solution to these problems: TomoPyUI , a user interface for the well known tomography data processing package TomoPy . This highly visual Python software package guides the user through the tomography processing pipeline from data import, preprocessing, alignment and finally to 3D volume reconstruction. The TomoPyUI systematic intermediate data and metadata storage system improves organization, and the inspection and manipulation tools (built within the application) help to avoid interrupted workflows. Notably, TomoPyUI operates entirely within a Jupyter environment. Herein, we provide a summary of these key features of TomoPyUI , along with an overview of the tomography processing pipeline, a discussion of the landscape of existing tomography processing software and the purpose of TomoPyUI , and a demonstration of its capabilities for real tomography data collected at SSRL beamline 6-2c.

73 NUCLEAR PHYSICS AND RADIATION PHYSICS

NASA GeneLab RNASeq Consensus Pipeline: A Nextflow Implementation

The NASA GeneLab project (genelab.nasa.gov) seeks to accelerate space biology research through cataloging and democratizing omics data. Since raw omics data is largely inaccessible to non-bioinformaticians, GeneLab works with the scientific community to develop standard processing pipelines to generate and publish processed data. Unlike raw data, processed data has greater immediate value to a wide range of users with varying technical backgrounds and computational capabilities. Standardizing processing workflows is essential to match the pace of raw data generation, ensure reproducibility, and enable standardized processed data for comparison across datasets. Previously, GeneLab developed a standardized pipeline for processing RNAseq data, referred to as the ‘GeneLab RNAseq Consensus Pipeline (RCP)’, in collaboration with GeneLab’s Analysis Working Groups. The work presented here is a Nextflow implementation of GeneLab’s RCP that automates and accelerates data processing of RNASeq datasets hosted on GeneLab. In addition to the core data processing, the workflow also includes staging of GeneLab raw data and a robust verification and validation (V&V) program that runs after each processing step to identify errors in real-time, stop additional downstream computation, and preserve computational resources. The workflow, including the staging and V&V functionality, is open source for others to reuse and modify at https://github.com/nasa/GeneLab_Data_Processing/tree/master/RNAseq.

Jonathan Dejesus Oribello

NASA GeneLab RNASeq Consensus Pipeline: A Nextflow Implementation

The NASA GeneLab project (genelab.nasa.gov) seeks to accelerate space biology research through cataloging and democratizing omics data. Since raw omics data is largely inaccessible to non-bioinformaticians, GeneLab works with the scientific community to develop standard processing pipelines to generate and publish processed data. Unlike raw data, processed data has greater immediate value to a wide range of users with varying technical backgrounds and computational capabilities. Standardizing processing workflows is essential to match the pace of raw data generation, ensure reproducibility, and enable standardized processed data for comparison across datasets. Previously, GeneLab developed a standardized pipeline for processing RNAseq data, referred to as the ‘GeneLab RNAseq Consensus Pipeline (RCP)’, in collaboration with GeneLab’s Analysis Working Groups. The work presented here is a Nextflow implementation of GeneLab’s RCP that automates and accelerates data processing of RNASeq datasets hosted on GeneLab. In addition to the core data processing, the workflow also includes staging of GeneLab raw data and a robust verification and validation (V&V) program that runs after each processing step to identify errors in real-time, stop additional downstream computation, and preserve computational resources. The workflow, including the staging and V&V functionality, is open source for others to reuse and modify at https://github.com/nasa/GeneLab_Data_Processing/tree/master/RNAseq.

Jonathan D Oribello

Surface Biology & Geology Pathfinder Data Analysis Pipeline

NASA's future global orbital mission, currently in development as the Surface Biology and Geology (SBG) Designated Observable study, will acquire relatively high resolution solar-reflected spectroscopy and thermal infrared observations. Innovative processes must be utilized for handling the high volume of data anticipated to be collected, which is anticipated to exceed 100 terabytes/day, greater than NASA's total extant airborne hyperspectral data collection. Collecting, processing/re-processing, disseminating, and exploiting this volume of data presents new challenges. To begin addressing them, NASA is drawing upon the expertise developed from its astrophysics programs to address Earth science and applications. Specifically, NASA is adapting the science processing operations technology developed for the Kepler and TESS planet-hunting missions for imaging spectroscopy data processing. This technology development has been the foundation for the remarkable scientific successes of Kepler and TESS. The Kepler/TESS data processing technology provides a scalable architecture for robust, repeatable, and replicable science and application products while enabling the Earth science community to develop, test, and implement new algorithms. Our effort to leverage this existing capability has begun by ingesting data and applying workflows from the EO-1/Hyperion 17-year mission archive that provides globally sampled visible through shortwave infrared spectra that are representative of SBG data types and volumes. This pathfinding data processing system will help define the solutions to processing SBG data volumes and will enable the scientific community to interact with the data and processing pipeline to create new science products.

Jenkins, Jon

First Results of Venus Express Spacecraft Observations with Wettzell

The ESA Venus Express spacecraft was observed at X-band with the Wettzell radio telescope in October-December 2009 in the framework of an assessment study of the possible contribution of the European VLBI Network to the upcoming ESA deep space missions. A major goal of these observations was to develop and test the scheduling, data capture, transfer, processing, and analysis pipeline. Recorded data were transferred from Wettzell to Metsahovi for processing, and the processed data were sent from Mets ahovi to JIVE for analysis. A turnover time of 24 hours from observations to analysis results was achieved. The high dynamic range of the detections allowed us to achieve a milliHz level of spectral resolution accuracy and to extract the phase of the spacecraft signal carrier line. Several physical parameters can be determined from these observational results with more observational data collected. Among other important results, the measured phase fluctuations of the carrier line at different time scales can be used to determine the influence of the solar wind plasma density fluctuations on the accuracy of the astrometric VLBI observations.

Calves, Guifre Molera

Enabling Model Organism and Commercial Astronaut Data Access Through the NASA Open Science Data Repository

NASA’s Open Science Data Repository (OSDR) brings together omics data from NASA’s GeneLab project and non-omics data, including physiological, phenotypic, imaging, and behavioral data from NASA’s Ames Life Sciences Data Archive (ALSDA) collected from decades of space biology research, providing open and FAIR (findable, accessible, interoperable, and reusable) access of these precious data to scientists world-wide. This rich source of meticulously curated metadata and data from spaceflight and analog studies has been mined by the scientific community resulting in dozens of high impact scientific publications that reveals a complex network of molecular and physiological effects of spaceflight across living systems, from microbes to plants, to mammals. Understanding how these effects translate to the human condition is critical as we move deeper into the era of commercial space travel. However, the integration of data, specifically omics data, from astronauts is particularly challenging due to their sensitive nature. OSDR has risen to this challenge by developing a mechanism to control access to identifiable levels of omics data, such as raw sequence data, while enabling public access to processed, unidentifiable, data and associated metadata that will allow the scientific community to interrogate human astronaut data alongside data from model organisms to begin answering these critical questions. The 2021 SpaceX Inspiration4 (I4) mission collected a comprehensive atlas of biological measurements from four civilian astronauts, providing a wealth of data to characterize the effects of spaceflight on the human body. These data include both non-omics and omics assays such as direct RNA sequencing (RNA-seq), single nuclei ATAC-seq and RNA-seq, metagenomics, proteomics, and comprehensive metabolic and cytokine panels, all of which have been integrated into the OSDR system across no less than 9 studies. Each study has been carefully curated using community-backed OSDR standards for sample and assay level metadata ensuring these data are findable and accessible. In addition to hosting both raw and processed data from the principal investigator team for each assay type, the GeneLab team plans to re-process the I4 omics data using GeneLab’s standard processing pipelines. The GeneLab processed data outputs will allow for comparisons across studies on OSDR and enable visualization of these data through the OSDR data visualization platform thereby enabling data reusability and interoperability. Here we describe the robust privacy and security protocols implemented by OSDR to safeguard sensitive health data from astronauts while facilitating metadata and processed data sharing for research purposes. We further provide a road map for navigating the vast amount of data provided for each I4 study on the OSDR, including experimental design, associated experiments, payloads, and missions, data generation and analysis protocols, and associated scientific articles. Additionally, we illustrate how to interrogate the standardized metadata provided in the sample and assay tables as well as various means to download and access the data including programmatically through the GeneLab Open API (GLOpenAPI). The open access of datasets in NASA’s OSDR provides a unique opportunity for the scientific community, as well as citizen scientists and students, to continue using OSDR resources to further unlock profound insights into the consequences of space travel on the human body. Through implementation of security measures to protect sensitive human data, the OSDR seeks to strengthen the science exchange between the Biological and Physical Sciences Program and the Human Research Program, per recommendation 4-1 of the 2023-2032 Decadal Survey, and encourage further sharing and dissemination of astronaut data to provide the scientific community with the resources needed to lay the groundwork for developing targeted mitigation strategies to help withstand the rigors of long-duration spaceflight.

Amanda Marie Saravia-butler

Enabling Model Organism and Commercial Astronaut Data Access Through the NASA Open Science Data Repository

NASA’s Open Science Data Repository (OSDR) brings together omics data from NASA’s GeneLab project and non-omics data, including physiological, phenotypic, imaging, and behavioral data from NASA’s Ames Life Sciences Data Archive (ALSDA) collected from decades of space biology research, providing open and FAIR (findable, accessible, interoperable, and reusable) access of these precious data to scientists world-wide. This rich source of meticulously curated metadata and data from spaceflight and analog studies has been mined by the scientific community resulting in dozens of high impact scientific publications that reveals a complex network of molecular and physiological effects of spaceflight across living systems, from microbes to plants, to mammals. Understanding how these effects translate to the human condition is critical as we move deeper into the era of commercial space travel. However, the integration of data, specifically omics data, from astronauts is particularly challenging due to their sensitive nature. OSDR has risen to this challenge by developing a mechanism to control access to identifiable levels of omics data, such as raw sequence data, while enabling public access to processed, unidentifiable, data and associated metadata that will allow the scientific community to interrogate human astronaut data alongside data from model organisms to begin answering these critical questions. The 2021 SpaceX Inspiration4 (I4) mission collected a comprehensive atlas of biological measurements from four civilian astronauts, providing a wealth of data to characterize the effects of spaceflight on the human body. These data include both non-omics and omics assays such as direct RNA sequencing (RNA-seq), single nuclei ATAC-seq and RNA-seq, metagenomics, proteomics, and comprehensive metabolic and cytokine panels, all of which have been integrated into the OSDR system across no less than 9 studies. Each study has been carefully curated using community-backed OSDR standards for sample and assay level metadata ensuring these data are findable and accessible. In addition to hosting both raw and processed data from the principal investigator team for each assay type, the GeneLab team plans to re-process the I4 omics data using GeneLab’s standard processing pipelines. The GeneLab processed data outputs will allow for comparisons across studies on OSDR and enable visualization of these data through the OSDR data visualization platform thereby enabling data reusability and interoperability. Here we describe the robust privacy and security protocols implemented by OSDR to safeguard sensitive health data from astronauts while facilitating metadata and processed data sharing for research purposes. We further provide a road map for navigating the vast amount of data provided for each I4 study on the OSDR, including experimental design, associated experiments, payloads, and missions, data generation and analysis protocols, and associated scientific articles. Additionally, we illustrate how to interrogate the standardized metadata provided in the sample and assay tables as well as instructions for how to download and access the data. The I4 datasets described here re present the first ever comprehensive collection of commercial astronaut data.

Amanda M Saravia-Butler

The Dynamic Networks Experiments: Virtual Experiments to Quantify Gains in Nuclear Explosion Monitoring

We describe an ongoing series of virtual experiments conducted collaboratively by four United States National Laboratories: Sandia National Laboratories, Los Alamos National Laboratory, Lawrence Livermore National Laboratory, and Pacific Northwest National Laboratory. These Dynamic Network Experiments (DNEs) provide an experimental framework to evaluate the potential impact of new research tools on nuclear explosion monitoring. The second DNE (DNE2), completed in 2024, exploited waveform data (seismic, infrasound, and electromagnetic) that was recorded by multi-modal sensors within and near the Nevada National Security Site and synthetic radionuclide signatures over multiple time periods. During the execution of DNE2, we processed and analyzed data through a multi-stage event processing pipeline that ingested raw data, performed quality control, detected signals, built events from these signals, located these events, and characterized the events’ source types and sizes. For each stage and over the entire event processing pipeline, we evaluated performance changes by comparing the performance of new data processing methods, models, and algorithms against a baseline. We also performed an additional execution phase to assess event processing pipeline function, speed, and efficiency against that of an expert analyst, including computational and manual efforts. Finally, we assessed the impact and effort of modern computing infrastructure on the monitoring pipeline. This paper describes key elements of the DNEs, from formulation through execution, as demonstrated in DNE2. The DNEs introduce several novel concepts to quantitatively measure the potential impact of new methods on explosion monitoring, including the collaborative design of multi-modal datasets, performance and logistical metrics, and integrated analyses.

42 ENGINEERING

Modular on-board adaptive imaging

Feature extraction involves the transformation of a raw video image to a more compact representation of the scene in which relevant information about objects of interest is retained. The task of the low-level processor is to extract object outlines and pass the data to the high-level process in a format that facilitates pattern recognition tasks. Due to the immense computational load caused by processing a 256x256 image, even a fast minicomputer requires a few seconds to complete this low-level processing. It is, therefore, necessary to consider hardware implementation of these low-level functions to achieve real-time processing speeds. The considered project had the objective to implement a system in which the continuous feature extraction process is not affected by the dynamic changes in the scene, varying lighting conditions, or object motion relative to the cameras. Due to the high bandwidth (3.5 MHz) and serial nature of the TV data, a pipeline processing scheme was adopted as the overall architecture of this system. Modularity in the system is achieved by designing circuits that are generic within the overall system.

Eskenazi, R.

Parallel processing in a host plus multiple array processor system for radar

Host plus multiple array processor architecture is demonstrated to yield a modular, fast, and cost-effective system for radar processing. Software methodology for programming such a system is developed. Parallel processing with pipelined data flow among the host, array processors, and discs is implemented. Theoretical analysis of performance is made and experimentally verified. The broad class of problems to which the architecture and methodology can be applied is indicated.

Barkan, B. Z.

The Kepler Science Operations Center Pipeline Framework Extensions

The Kepler Science Operations Center (SOC) is responsible for several aspects of the Kepler Mission, including managing targets, generating on-board data compression tables, monitoring photometer health and status, processing the science data, and exporting the pipeline products to the mission archive. We describe how the generic pipeline framework software developed for Kepler is extended to achieve these goals, including pipeline configurations for processing science data and other support roles, and custom unit of work generators that control how the Kepler data are partitioned and distributed across the computing cluster. We describe the interface between the Java software that manages the retrieval and storage of the data for a given unit of work and the MATLAB algorithms that process these data. The data for each unit of work are packaged into a single file that contains everything needed by the science algorithms, allowing these files to be used to debug and evolve the algorithms offline.

Klaus, Todd C.