Search NASA⌕ Search

SEARCH · Search NASA

Results for “data pipeline”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 163 records · Page 9

The Kepler End-to-End Model: Creating High-Fidelity Simulations to Test Kepler Ground Processing

The Kepler mission is designed to detect the transit of Earth-like planets around Sun-like stars by observing 100,000 stellar targets. Developing and testing the Kepler ground-segment processing system, in particular the data analysis pipeline, requires high-fidelity simulated data. This simulated data is provided by the Kepler End-to-End Model (ETEM). ETEM simulates the astrophysics of planetary transits and other phenomena, properties of the Kepler spacecraft and the format of the downlinked data. Major challenges addressed by ETEM include the rapid production of large amounts of simulated data, extensibility and maintainability.

Bryson, Stephen T.↗

First Results of Venus Express Spacecraft Observations with Wettzell

The ESA Venus Express spacecraft was observed at X-band with the Wettzell radio telescope in October-December 2009 in the framework of an assessment study of the possible contribution of the European VLBI Network to the upcoming ESA deep space missions. A major goal of these observations was to develop and test the scheduling, data capture, transfer, processing, and analysis pipeline. Recorded data were transferred from Wettzell to Metsahovi for processing, and the processed data were sent from Mets ahovi to JIVE for analysis. A turnover time of 24 hours from observations to analysis results was achieved. The high dynamic range of the detections allowed us to achieve a milliHz level of spectral resolution accuracy and to extract the phase of the spacecraft signal carrier line. Several physical parameters can be determined from these observational results with more observational data collected. Among other important results, the measured phase fluctuations of the carrier line at different time scales can be used to determine the influence of the solar wind plasma density fluctuations on the accuracy of the astrometric VLBI observations.

Calves, Guifre Molera↗

Data Science and the Knowledge Discovery Adventure

This talk will cover the important steps involved in the data science and knowledge discovery process: • Initial fact gathering (interview domain experts, review reports, articles, state-of-the-art) • Identify the problem (prediction, classification, statistical analysis, etc.) • Survey supporting data sources • Understand the data (numerical, categorical, text, sampling rate, data quality issues, etc.) • Selecting relevant features and sources • Acquire the data (set up agreements with the data stewards, APIs to download, etc.) • Merge data sources (temporal, spatial, common key, other ontologies...) • Feature Engineering (non linear domain knowledge or physics-based relationships) • Build data processing pipeline (may need to tap into data stream, develop parallel processing algorithm, federated learning etc.) • Build model and test (tune hyper-parameters, cross validation.) • Analyze/Validate results (do the results make sense. Does it answer the original question). • Deploy/Publish (Monitor and assess benefits)

Data science↗

Automated X-ray and Optical Analysis of the Virtual Observatory and Grid Computing

We are developing a system to combine the Web Enabled Source Identification with X-Matching (WESIX) web service, which emphasizes source detection on optical images,with the XAssist program that automates the analysis of X-ray data. XAssist is continuously processing archival X-ray data in several pipelines. We have established a workflow in which FITS images and/or (in the case of X ray data) an X-ray field can be input to WESIX. Intelligent services return available data (if requested fields have been processed) or submit job requests to a queue to be performed asynchronously. These services will be available via web services (for non-interactive use by Virtual Observatory portals and applications) and through web applications (written in the Django web application framework). We are adding web services for specific XAssist functionality such as determining .the exposure and limiting flux for a given position on the sky and extracting spectra and images for a given region. We are improving the queuing system in XAssist to allow for "watch lists" to be specified by users, and when X-ray fields in a user's watch list become publicly available they will be automatically added to the queue. XAssist is being expanded to be used as a survey planning 1001 when coupled with simulation software, including functionality for NuStar, eRosita, IXO, and the Wide Field Xray Telescope (WFXT), as part of an end to end simulation/analysis system. We are also investigating the possibility of a dedicated iPhone/iPad app for querying pipeline data, requesting processing, and administrative job control.

Ptak, A.↗

GL4U: Using Space Biology Omics Data to Provide Bioinformatics Training for Students and Educators

NASA’s GeneLab project provides researchers open access to space-relevant multi-omics data via the Open Science Data Repository (OSDR) that can be mined to understand the effects of spaceflight on biological systems. To maximize the number of scientists who understand and utilize GeneLab data and data processing pipelines, GeneLab created GeneLab for Colleges and Universities (GL4U). GL4U provides space biology-relevant training in bioinformatics to the next generation of scientists through direct (training students) and indirect (training educators) approaches. The GL4U pilot programs were conducted in June 2021 (direct training) and 2022 (indirect training). During the pilots, students and educators at Historically Black Colleges and Universities (HBCUs) and Minority Serving Institutions (MSIs) participated in a week-long (direct training) or two-week-long (indirect training) bootcamp consisting of space biology-specific lectures and hands-on instruction using Jupyter Notebooks to analyze space biology RNA sequencing data from OSDR. During the educator pilot, participants received materials, training, and the necessary compute resources to enable them to run the bootcamp at their home institutions, thereby extending the reach of this initiative. In July 2023, GeneLab is partnering with JPL to expand GL4U to include amplicon sequencing (Amp-Seq) analysis training. During the GL4U Amp-Seq bootcamp, student and educator participants will receive training on how to analyze and interpret Amp-Seq data using the NASA GeneLab data processing pipeline. All bootcamp material, including instructions for requesting compute resources, will be made publicly available on GitHub for educators to teach the GL4U content in subsequent semesters. GL4U provides undergraduate students from underrepresented groups the opportunity to learn about NASA and Space Biology, and to enhance their career prospects by gaining hands-on experience analyzing omics data, a skillset that is highly applicable and marketable in the life sciences. We present results from pre- and post-training surveys completed by all participants of the Amp-Seq bootcamp.

Amanda M Saravia-Butler↗

Data for An End-to-End Pipeline for Succinic Acid Production at an Industrially Relevant Scale Using Issatchenkia orientalis

Microbial production of succinic acid (SA) at an industrially relevant scale has been hindered by high downstream processing costs arising from neutral pH fermentation for over three decades. Here, we metabolically engineer the acid-tolerant yeast Issatchenkia orientalis for SA production, attaining the highest titers in sugar-based media at low pH (pH 3) in fed-batch fermentations, i.e. 109.5 g/L in minimal medium and 104.6 g/L in sugarcane juice medium. We further perform batch fermentation using sugarcane juice medium in a pilot-scale fermenter (300×) and achieve 63.1 g/L of SA, which can be directly crystallized with a yield of 64.0%. Finally, we simulate an end-to-end low-pH SA production pipeline, and techno-economic analysis and life cycle assessment indicate our process is financially viable and can reduce greenhouse gas emissions by 34–90% relative to fossil-based production processes. We expect I. orientalis can serve as a general industrial platform for production of organic acids.

Metabolomics↗

Pipeline for Applications-Based Data Discovery

From disaster response and mitigation to monitoring water quality or protecting wildlife habitat, satellite Earth observation data can be applied in countless ways to meet pressing needs and benefit society. The crucial first step toward successful data application is data discovery. Potential users often know exactly what data they need--what Earth feature or phenomenon they need to observe, how frequently, and at what resolution or level of accuracy--but may still struggle to discover the existing observations that meet their needs. We have developed a pipeline to connect applications-based users to specific satellites and data collections within NASA's Earth observation program of record that are highly relevant to their data needs. This pipeline combines available information on satellite and instrument measurement characteristics with an innovative machine learning-based approach that identifies instruments that are most relevant to the feature or phenomenon of interest.

Katrina S Virts↗

Pipeline synthetic aperture radar data compression utilizing systolic binary tree-searched architecture for vector quantization

A system for data compression utilizing systolic array architecture for Vector Quantization (VQ) is disclosed for both full-searched and tree-searched. For a tree-searched VQ, the special case of a Binary Tree-Search VQ (BTSVQ) is disclosed with identical Processing Elements (PE) in the array for both a Raw-Codebook VQ (RCVQ) and a Difference-Codebook VQ (DCVQ) algorithm. A fault tolerant system is disclosed which allows a PE that has developed a fault to be bypassed in the array and replaced by a spare at the end of the array, with codebook memory assignment shifted one PE past the faulty PE of the array.

Chang, Chi-Yung↗

Kepler: A Search for Terrestrial Planets - Kepler Data Characterization Handbook

The Kepler Data Characteristics Handbook (KDCH) provides a description of all phenomena identified in the Kepler data throughout the mission, and an explanation for how these characteristics are handled by the final version of the Kepler Data Processing Pipeline (SOC 9.3).The KDCH complements the Kepler Data Release Notes (KDRNs), which document phenomena and processing unique to a data release. The original motivation for this separation into static, explanatory text and a more journalistic set of figures and tables in the KDRN was for the user to become familiar with the Data Characteristics Handbook, then peruse the short Notes for a new quarter, referring back to the Handbook when necessary. With the completion of the Kepler mission and the final Data Release 25, both the KDCH and the DRN encompass the entire Kepler mission, so the distinction between them is in the level of exposition, not the extent of the time interval discussed.

Kepler↗

Presearch Data Conditioning in the Kepler Science Operations Center Pipeline

We describe the Presearch Data Conditioning (PDC) software component and its context in the Kepler Science Operations Center (SOC) pipeline. The primary tasks of this component are to correct systematic and other errors, remove excess flux due to aperture crowding, and condition the raw flux light curves for over 160,000 long cadence (~thirty minute) and 512 short cadence (~one minute) targets across the focal plane array. Long cadence corrected flux light curves are subjected to a transiting planet search in a subsequent pipeline module. We discuss the science algorithms for long and short cadence PDC: identification and correction of unexplained (i.e., unrelated to known anomalies) discontinuities; systematic error correction; and excess flux removal. We discuss the propagation of uncertainties from raw to corrected flux. Finally, we present examples of raw and corrected flux time series for flight data to illustrate PDC performance. Corrected flux light curves produced by PDC are exported to the Multi-mission Archive at Space Telescope [Science Institute] (MAST) and will be made available to the general public in accordance with the NASA/Kepler data release policy.

Twicken, Joseph D.↗

Astro-H Data Analysis, Processing and Archive

Astro-H (Hitomi) is an X-ray Gamma-ray mission led by Japan with international participation, launched on February 17, 2016. The payload consists of four different instruments (SXS, SXI, HXI and SGD) that operate simultaneously to cover the energy range from 0.3 keV up to 600 keV. This paper presents the analysis software and the data processing pipeline created to calibrate and analyze the Hitomi science data along with the plan for the archive and user support.These activities have been a collaborative effort shared between scientists and software engineers working in several institutes in Japan and USA.

Archive↗

Curation and Dissemination of Complex Multi-Modal Datasets for Radiation Detection, Localization, and Tracking

The PANDAWN sensor network in Chicago, IL, is a state-of-the-art testbed for networked, multi-modal sensing. It integrates AI/data science methods into its operation, from data acquisition to automated data labeling and curation workflows. The curation and dissemination of diverse multi-modal datasets will enable the development of new radiological/nuclear (R/N) detection, localization, and tracking algorithms and methods relevant across the nonproliferation mission space. This article first introduces the PANDAWN sensor network and the features that make it stand out from previous multi-modal data acquisition efforts. We then review the various data streams acquired on the PANDAWN nodes and present the implementation of an automated data curation pipeline that includes the labeling of radiation and contextual data streams. Here, we finally provide a short overview of different studies that leveraged the curated datasets.

Data curation↗

Rapid and Reliable Damage Proxy Map from InSAR Coherence

Future radar satellites will visit SoCal within a day after a disaster event. Data acquisition latency in 2015-2020 is 8 to approx. 15 hours. Data transfer latency that often involves human/agency intervention far exceeds the data acquisition latency. Need interagency cooperation to establish automatic pipeline for data transfer. The algorithm is tested with ALOS PALSAR data of Pasadena, California. Quantitative quality assessment is being pursued: Meeting with Pasadena City Hall computer engineers for a complete list of demolition/construction project 1. Estimate the probability of detection and probability of false alarm 2. Estimate the optimal threshold value.

InSAR↗

Astro-H/Hitomi Data Analysis, Processing, and Archive

Astro-H is the x-ray/gamma-ray mission led by Japan with international participation, launched on February 17, 2016. Soon after launch, Astro-H was renamed Hitomi. The payload consists of four different instruments (SXS, SXI, HXI, and SGD) that operate simultaneously to cover the energy range from 0.3 keV up to 600 keV. On March 27, 2016, JAXA lost contact with the satellite and, on April 28, they announced the cessation of the efforts to restore mission operations. Hitomi collected about one months worth of data with its instruments. This paper presents the analysis software and the data processing pipeline created to calibrate and analyze the Hitomi science data, along with the plan for the archive. These activities have been a collaborative effort shared between scientists and software engineers working in several institutes in Japan and United States.

Angelini, Lorella↗

TPSAS-NF1676L-32493-DND

The Committee on Earth Observation Satellites (CEOS) System Engineering Office (SEO) has supported the Open Data Cube (ODC) initiative to provide a data architecture solution that has value to its global users and increases the impact of EO satellite data. ODC is an open-source platform for processing satellite data. We have developed software products and tools around the core ODC that would help users perform machine learning on EO satellite data. The recent United Nations (UN) Sustainable Development Agenda provides a shared blueprint for peace and prosperity for people and for the planet, considering our current situation and helping to create a plan. The core of this agenda is a set of seventeen Sustainable Development Goals (SDGs), which represent an urgent call for action by all countries - both developed and developing - in a global partnership. The CEOS SEO team has recently developed and released a set of innovative Jupyter notebooks addressing UN SDGs 6.6.1 (spatial extents of water-related ecosystems), 11.3.1 (ratio of land consumption rate to population growth rate), and 15.3.1 (proportion of land that is degraded over total land area). These notebooks empower users by providing features that will assist with streamlining analysis ready data retrieval, processing, and visualization. We have recently incorporated several machine learning techniques in these notebooks. In this paper, we present the lessons learned from our experience on classifying land using supervised and unsupervised machine learning techniques using ODC framework for UN SDGs. We identify the current limitations of ODC to seamlessly support machine learning techniques. We propose features that would help machine learning, specifically within the ODC framework. We propose a thematic indexing/loading of data for both unsupervised learning as well as data annotation/labeling pipeline. Currently, ODC supports machine learning by separating data-management from the analysis process. It works as a mechanism to load cubes of data. ODC does not natively support features that are vital in machine learning such as validation splits, fair/balanced sampling, establishing load size constraints, etc. We believe that our proposed features will empower users by providing features that bring machine learning techniques closed to ODC. Enhancements to ODC to better accommodate machine learning techniques can assist in fulfilling UN SDGs such as 6.3.2, 6.4.2, 6.6.1, 11.3.1, 14.1.1, 15.1.1, 15.3.1, and 15.4.2.

Syed R Rizvi↗

AAM NC ATI TechTalk - Aerograph Architecture v1

Aerograph is NASA’s data management system for Advanced Air Mobility. Its mission is to support AAM research by providing a reliable and secure system that collects, stores, protects, and shares AAM data. Its vision is to provide a system that AAM research scientists, aerospace engineers, data scientists, and analysts trust for obtaining NC data and performing key analyses. The types of data Aerograph manages involves data related to flight test events, including: Aircraft Performance and Characterization (e.g., position reports) Airspace (e.g., operation intent, waypoints, and constraints) Environment (e.g., surface and wind weather) Infrastructure (e.g., surveillance coverage) Derivative Analytical Artifacts (e.g., glide path performance chart, 3D position chart, Integrated Data Product)

Aerograph↗

Comparability of Liquid Chromatography Tandem Mass Spectrometry Analysis of Dissolved Organic Matter across Laboratories

Non-targeted liquid chromatography tandem highresolution mass spectrometry (LC−MS/MS) is increasingly applied for the structure-resolved chemical analysis of dissolved organic matter (DOM). With new developments in MS instrumentation and analysis software, the approach has gained substantial momentum over the past decade. However, achieving high-quality analytical data that is reproducible and comparable across laboratories can be a bottleneck in non-targeted metabolomics and organic matter chemical analysis, especially for data reuse in repository-scale analyses. Understanding the capabilities as well as challenges of comparing LC−MS/MS data from different laboratories is necessary for inferring global trends from public data sets. To illuminate instrumentation factors that drive differences and variability, we used a standardized data analysis pipeline, including classical (CMN) and featurebased molecular networking (FBMN), to analyze data from a ring trial by 24 laboratories on identical sample sets of algal and DOM extracts that were mixed in predefined concentrations and spiked with standards. Our results showed that data sets from similar mass spectrometer types with unified instrument parameters were qualitatively comparable, resolving the same general trends and shared mass spectral features. Interlaboratory comparability was best for high-intensity features, while low-intensity features showed greater detection variability. Our analysis also highlights challenges when comparing data from instruments with different acquisition rates or operating with less standardized methods. Lastly, we provide recommendations for data integration, public data sharing, standardization, and best practices for standardized LC−MS/MS data acquisition, which will be critical for long-term time series and intercomparability of DOM chemical analyses.

DOM↗

WebGeocalc and Cosmographia: Modern Tools to Access OPS SPICE Data

For more than two decades navigation and other ancillary data from most US and international planetary science missions have been packaged using "SPICE" (Spacecraft, Planet, Instrument, Camera-matrix, Events) system data files (a.k.a. SPICE kernels) and, in conjunction with SPICE Toolkit software used by scientists and engineers to compute observation geometry in various ground system tools ranging from mission planning and analysis applications to data production pipelines to science analysis tools. The traditional way for accessing SPICE data is by downloading necessary SPICE kernels to a user’s workstation, installing the SPICE Toolkit software available from NAIF, and writing an application calling APIs from the SPICE Toolkit library to compute numeric geometric parameters of interest. While this approach did and still does provide the greatest flexibility in implementing geometric computations of interest, it proved to be complicated for users with little programming abilities, required data to be always copied to the users’ workstations, and lacked any out-of-the-box visualization capabilities. To address these shortcomings NAIF developed the WebGeocalc (WGC) tool and extended the publicly available Cosmographia program to use SPICE. Employing these two new tools in mission operations enables easier access to SPICE computations and SPICE-based visualizations for a wider variety of mission personnel.

Semenov, Boris V.↗