Search NASASearch

SEARCH · Search NASA

Results for “pipeline data processing”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 73 records · Page 4

Automated Detection and Analysis of Resident Space Objects with the 1.3-Meter Eugene Stansbery-Meter Class Autonomous Telescope

Optical telescopes dedicated to the detection of orbital debris employ large-area detectors that generate a large number of images each night. Such surveys require automated data analysis pipelines that process the images and detect moving objects. We present an overview of the data analysis pipeline employed by the 1.3-meter Eugene Stansbery-Meter Class Autonomous Telescope (ES-MCAT) on Ascension Island, operated by NASA’s Orbital Debris Program Office. The pipeline enfolds the astrometric and photometric calibration of the images, star-trail removal, object detection, correlation over multiple sequential image frames, and orbital parameter estimation. The performance of the pipeline was investigated by means of Monte-Carlo simulations in which simulated object tracks were inserted into ES-MCAT images and then processed by the pipeline. This technique allows one to confidently estimate the completeness for the detection of resident space objects as a function of apparent magnitude and angular velocity. This paper discusses these techniques and provides examples using actual data.

Paul Hickson

Parallel processing in a host plus multiple array processor system for radar

Host plus multiple array processor architecture is demonstrated to yield a modular, fast, and cost-effective system for radar processing. Software methodology for programming such a system is developed. Parallel processing with pipelined data flow among the host, array processors, and discs is implemented. Theoretical analysis of performance is made and experimentally verified. The broad class of problems to which the architecture and methodology can be applied is indicated.

Barkan, B. Z.

The Kepler Science Operations Center Pipeline Framework Extensions

The Kepler Science Operations Center (SOC) is responsible for several aspects of the Kepler Mission, including managing targets, generating on-board data compression tables, monitoring photometer health and status, processing the science data, and exporting the pipeline products to the mission archive. We describe how the generic pipeline framework software developed for Kepler is extended to achieve these goals, including pipeline configurations for processing science data and other support roles, and custom unit of work generators that control how the Kepler data are partitioned and distributed across the computing cluster. We describe the interface between the Java software that manages the retrieval and storage of the data for a given unit of work and the MATLAB algorithms that process these data. The data for each unit of work are packaged into a single file that contains everything needed by the science algorithms, allowing these files to be used to debug and evolve the algorithms offline.

Klaus, Todd C.

Transcriptomics Processing Pipelines for Space Biology: An Open Source and Consensus-Driven Approach

Transcriptomics holds significant value in elucidating the relationship between gene expression, experimental factors, biological factors, and various types of omics data. Enhancing our understanding of these connections is paramount for foundational biology, which plays a pivotal role in devising solutions for challenges pertinent to both space travel and terrestrial life. The NASA GeneLab project, part of the Open Science Data Repository (OSDR.nasa.gov), seeks to accelerate space biology research through cataloging and democratizing ‘omics data, including transcriptomics. Since raw omics data are largely inaccessible to non-bioinformaticians, GeneLab works with the scientific community via the Open Science Analysis Working Groups (AWGs) to develop standard processing pipelines to generate and publish processed data. Unlike raw data, processed data have greater immediate value to diverse users with varying technical backgrounds and computational capabilities. Standardizing processing workflows is essential to match the pace of raw data generation, ensure reproducibility, and enable standardized processed data for comparison across datasets. As of June 2023, transcriptomics studies comprise over half of GeneLab datasets hosted on the OSDR, including data from bulk RNA-seq and Affymetrix or Agilent 1-Channel DNA microarray assays. In collaboration with the AWGs, GeneLab developed consensus processing pipelines for these transcriptomics data types that includes quality control, background correction (microarray only), data normalization and quantification, culminating in the detection and annotation of differentially expressed genes. The work presented here describes Nextflow implementations of GeneLab’s consensus transcriptomics pipelines that automates and accelerates processing of these datasets. In addition to the core data processing, these workflows also include raw data staging and a robust verification and validation program to identify errors in real-time, stop additional downstream computation, and preserve computational resources. These workflows are used to generate GeneLab processed data hosted on the OSDR, and are publicly available as open source software for others to use at: https://github.com/nasa/GeneLab_Data_Processing.

Jonathan Oribello

Open Science for Life in Space: Data Sharing and Tools for Knowledge Discovery

Molecular-omics, physiological-phenotypic-behavioral, and environmental-radiation telemetry data from spaceflight biological and health studies are increasingly being made findable, accessible, interoperable, and reusable for the scientific public. These data, as well as space science-relevant biospecimens, are available through NASA’s Open Science Data Repository (OSDR), which is the new umbrella grouping of NASA GeneLab, the Ames Life Sciences Data Archive (ALSDA), and the NASA Biological Institutional Scientific Collection (NBISC). The quality of data is underpinned by datasets having rich metadata (determined through Analysis Working Group members), processing pipelines to enable data reuse standards, and ontologies specifying terminology semantics (e.g., the Radiation Biology Ontology).

space biology

Ramdb: The NASA Raman Spectral Database (version 1.00).

Given that, in most instances, minimal sample preparation is required and due to its contactless instrument design, Raman spectroscopy is one of the most versatile vibrational spectroscopic techniques for the chemical analysis of environmental and biological specimens. The diversity of applications of Raman spectroscopy ranges anywhere from art [1] to planetary science missions [2]. The advancement in the use of Raman spectroscopy in Solar System missions, notably in post-mission sample return analysis, requires a spectral library holding the broad range of specimens that could be found in Solar System sources. For this purpose, we have initiated the development of a Raman spectral database (Ramdb) at NASA Ames Research Center. Currently, the database includes experimental and theoretical Raman spectra of PAHs [3, 4], as well as laboratory Raman spectra of amino acids, carbon allotropes, minerals, and analogs relevance to Earth Sciences [5], Exobiology [6], Planetary [7], and Astrochemistry [8] to name just a few examples. Ramdb can be found on the web at www.astrochemistry.org/ramdb, where raw and processed Raman spectra can be downloaded in CSV format. The laboratory Raman spectra are measured using a laser Raman spectrometer (JASCO NRS-5500-532QRI). The Raman instrument is equipped with three excitation lasers, with wavelengths of 405, 532, and 785 nm. A clean silicon substrate is used as the internal standard for wavenumber calibration. Powdered samples were prepared (microscopic >10 um, grounded microscopic < 10 um) on glass slides. Some raw data exhibited a background signal arising as a combination of laser-induced fluorescence from the sample. To correct this background, we developed a Python pipeline that uses open-source Python libraries. Ramdb provides both raw and processed (using Python pipeline) data, which includes tabulated Raman shift transitions and other measurement details. The theoretical Raman band positions of PAHs (pyrene monomers and tetramer clusters) were computed using density functional theory (DFT) with the help of the Gaussian 16 suite of programs [9]. In the near future, Ramdb will serve as a repository of Raman spectral data from Laboratory Astrophysics and Planetary Science experiments involving the irradiation of organic compounds under simulated space and planetary conditions. In addition, online and offline tools will be developed for utilising the database for comparison to the user’s sample.

N Punnakayathil

Tracking the Hunga Tonga-Hunga Ha’apai Eruption Stratospheric Aerosol and Trace Gas Plumes Using Machine Learning

On January 15, 2022, the Hunga Tonga-Hunga Ha’apai (hereafter, Hunga Tonga) submarine volcano had an explosive eruption that thrusted ash, gases, and water vapor through the troposphere into the stratosphere and mesosphere. Previous studies manually tracked the aerosol and trace gas plumes over time across different positions in the southern hemisphere. Using data retrieved from low earth orbiting satellite instruments (e.g., OMPS, OMI, and CALIPSO), this research demonstrates how open-source machine learning (ML) models, like Meta’s Segment Anything Model (SAM), with prompt engineering can perform automatic plume tracking following the Hunga Tonga eruption. This extensible methodology, and modular data processing and modeling pipeline using NASA Earthdata and Openscapes, establishes a framework for systematically and rapidly studying extreme events, including volcanic eruptions and large-scale wildfires. By combining advanced machine learning techniques, such as SAM’s zero-shot learning, with large volumes of remote sensing data, this work demonstrates how AI and open science can accelerate research and generate actionable results. The tools and technologies presented here can help translate earth science to action from NASA’s current and future Earth observing satellite missions (e.g., the Atmosphere Observing System (AOS)), and assist researchers and stakeholders in understanding, mapping, and responding to natural disasters and extreme events in a changing world.

David M. Giles

Evaluating Machine Learning Approaches to Plume Tracking

On July 15, 2022, the Hunga Tonga-Hunga Ha’apai (HTHH) submarine volcano erupted, propelling trace gasses and ash through the troposphere and up into the stratosphere. Previous studies manually tracked the aerosol and trace gas plumes over time across different positions in the southern hemisphere. Using imagery from NASA’s Earth Observing System, including MODIS aerosol products and OMI sulfur dioxide products, this research demonstrates how open-source machine learning (ML) models, like Meta’s Segment Anything Model (SAM), can perform automatic plume tracking following the Hunga Tonga eruption. This extensible methodology, and modular data processing and modeling pipeline, establishes a framework for systematically and rapidly studying natural disasters, including additional volcanic eruptions and large-scale wildfires. By combining advanced machine learning techniques, such as SAM’s zero-shot learning, with large volumes of NASA’s Earth Observation and remote sensing data, this work shows how AI and open science can accelerate research and generate actionable results, even for unprecedented events. The tools and technologies presented here can help translate earth science to action from NASA’s current and future Earth observing satellite missions, and assist researchers and stakeholders in understanding, mapping, and responding to natural disasters in a changing world.

machine learning

A Data Exploration Tool for Large Sets of Spectra

We present an exploration tool for very large spectrum data sets such as the SDSS (Sloan Digital Sky Survey), LAMOST (Large Sky Area Multi-Object Fiber Spectroscopic Telescope), and 4MOST (4-meter Multi-Object Spectroscopic Telescope) data sets. The tool works in two stages: the first uses batch processing and the second runs interactively. The latter employs the NASA hyperwall, a configuration of 128 workstation displays (8 by 16 array) controlled by a parallelized software suite running on NASA's Pleiades supercomputer. The stellar subset of the Sloan Digital Sky Survey, DR10, was chosen to show how the our tool may be used. In stage one, SDSS files for 569,740 stars are processed through our data pipeline. The pipeline fits each spectrum using an iterative continuum algorithm, distinguishing emission from absorption and handling molecular absorption bands correctly. It then measures 1659 discrete atomic and molecular spectral features that were carefully preselected based on their likelihood of being visible at some spectral type. The depths relative to the local continuum at each feature wavelength are determined for each spectrum: these depths, the local S/N (signal to noise ratio) level, and DR10-supplied variables such as magnitudes, colors, positions, and radial velocities are the basic measured quantities used on the hyperwall. In stage two, each hyperwall panel is used to display a 2-D scatter plot showing the depth of feature A vs the depth of feature B for all of the stars. A and B change from panel to panel. The relationships between the various (A,B) strengths and any distinctive clustering are immediately apparent when examining and inter-comparing the different panels on the hyperwall. The interactive software allows the user to select the stars in any interesting region of any 2-D plot on the hyperwall, immediately rendering the same stars on all the other 2-D plots in a unique color. The process may be repeated multiple times, each selection displaying a distinctive color on all the plots. At any time, the spectra of the selected stars may be examined in detail on a connected workstation display. We illustrate how our approach allows us to quickly isolate and examine such interesting stellar subsets as EMP (Extremely Metal‐Poor) stars, CV (Cataclymic Variable) stars and C (Carbon)-rich stars.

Data Exploration

Kepler Planet Detection Metrics: Pixel-Level Transit Injection Tests of Pipeline Detection Efficiency for Data Release 25

This document describes the results of the fourth pixel-level transit injection experiment, which was designed to measure the detection efficiency of both the Kepler pipeline (Jenkins 2002, 2010; Jenkins et al. 2017) and the Robovetter (Coughlin 2017). Previous transit injection experiments are described in Christiansen et al. (2013, 2015a,b, 2016).In order to calculate planet occurrence rates using a given Kepler planet catalogue, produced with a given version of the Kepler pipeline, we need to know the detection efficiency of that pipeline. This can be empirically determined by injecting a suite of simulated transit signals into the Kepler data, processing the data through the pipeline, and examining the distribution of successfully recovered transits. This document describes the results for the pixel-level transit injection experiment performed to accompany the final Q1-Q17 Data Release 25 (DR25) catalogue (Thompson et al. 2017)of the Kepler Objects of Interest. The catalogue was generated using the SOC pipeline version 9.3 and the DR25 Robovetter acting on the uniformly processed Q1-Q17 DR25 light curves (Thompson et al. 2016a) and assuming the Q1-Q17 DR25 Kepler stellar properties (Mathur et al. 2017).

Pixel-Level Transit Injection

The Physics of Cooling Flow Clusters with Central Radio Sources

Central galaxies in rich clusters are the sites of cluster cooling flows, with large masses of gas cooling through part of the X-ray band. Many of these galaxies host powerful radio sources. These sources can displace and compress the X-ray gas leading to enhanced cooling and star formation. We observed the bright cooling flow Abell 2626 with a strangely distorted central radio source. We wished to understand the interaction of radio and X-ray thermal plasma, and to determine the dynamical nature of this cluster. One aim was to constrain the source of additional pressure in radio "holes" in the X-ray emission needed to support overlying shells of X-ray gas. We also aimed to study the problem of the lack of kT < 1-2 keV gas in cooling flows by searching for abundance inhomogeneities, heating from the radio source, and excess absorption. We also have a Chandra observation of this cluster. There were problems with the pipeline processing of this data due to a telemetry dropout. We are publishing the Chandra and XMM data together. Delays with the Chandra data have slowed up the publication. At the center of the cluster, there is a complex interaction of the odd, Z-shaped radio source, and the X-ray plasma. However, there are no clear radio bubbles. Also, the cluster SO galaxy IC 5337, which is projected 1.5 arcmin west of the cluster center, has unusual tail-like structures in both the radio and X-ray. It appears to be falling into the cluster center. There is a hot, probably shocked region of gas to the southwest, which is apparently due to the merger of a subcluster in this part of the system. There is also a merging subcluster to the northeast. The axes of these two mergers agrees with a supercluster filament structure.

Sarazin, Craig L.

Parallel integrated frame synchronizer chip

A parallel integrated frame synchronizer which implements a sequential pipeline process wherein serial data in the form of telemetry data or weather satellite data enters the synchronizer by means of a front-end subsystem and passes to a parallel correlator subsystem or a weather satellite data processing subsystem. When in a CCSDS mode, data from the parallel correlator subsystem passes through a window subsystem, then to a data alignment subsystem and then to a bit transition density (BTD)/cyclical redundancy check (CRC) decoding subsystem. Data from the BTD/CRC decoding subsystem or data from the weather satellite data processing subsystem is then fed to an output subsystem where it is output from a data output port.

Ghuman, Parminder Singh

Intensity Conserving Spectral Fitting

The detailed shapes of spectral line profiles provide valuable information about the emitting plasma, especially when the plasma contains an unresolved mixture of velocities, temperatures, and densities. As a result of finite spectral resolution, the intensity measured by a spectrometer is the average intensity across a wavelength bin of non-zero size. It is assigned to the wavelength position at the center of the bin. However, the actual intensity at that discrete position will be different if the profile is curved, as it invariably is. Standard fitting routines (spline, Gaussian, etc.) do not account for this difference, and this can result in significant errors when making sensitive measurements. Detection of asymmetries in solar coronal emission lines is one example. Removal of line blends is another. We have developed an iterative procedure that corrects for this effect. It can be used with any fitting function, but we employ a cubic spline in a new analysis routine called Intensity Conserving Spline Interpolation (ICSI). As the name implies, it conserves the observed intensity within each wavelength bin, which ordinary fits do not. Given the rapid convergence, speed of computation, and ease of use, we suggest that ICSI be made a standard component of the processing pipeline for spectroscopic data.

Klimchuk, J. A.

NASA's Unsteady Pressure-Sensitive Paint Research and Operational Capability Developments

In the last three years, several advancements have been made to produce a new state-of-the-art capability in the field of Aerosciences. NASA’s Aerosciences Evaluations and Test Capabilities (AETC) Portfolio Office has funded a multi-year project to produce the unsteady Pressure-Sensitive Paint (uPSP) technology as an operational capability in key ground test facilities at NASA. The research and development has primarily been conducted at NASA Ames Research Center’s (ARC) Unitary Plan Wind Tunnel (UPWT) 11-by 11-ft Transonic Wind Tunnel (TWT). The NASA ARC UPWT is one of the ground test facilities under NASA AETC’s Portfolio Office. AETC’s goals are to provide the tools to deliver the technology innovations and breakthroughs necessary to address increasingly complex research and development challenges. AETC’s integrated approach will consider the complimentary high-end compute capabilities necessary to advance analysis in conjunction with ground experimental capabilities. The uPSP Capability Challenge Project is a demonstration of several different technologies: 1) the unsteady Pressure-Sensitive Paint (uPSP) technology, and 2) Project: Red Rover, establishing a secure, reliable, fast connection between experimental and computation facilities, leveraging NASA’s computational resources within the High-End Compute Capability (HECC) Project for processing, storing, and sharing data efficiently. This project demonstrates the technical diversity and technical inclusion need to advance the field of Aerosciences. The approach to combine subject matter experts in experimental methods, optical methods, production wind tunnel testing, network engineering, high-end computing, signal processing, grid generation, and visualization while establishing the required infrastructure for subject matter experts to have access to the data while the wind tunnel test is being conducted. The most recent advancements for the uPSP technology have focused on three key areas: development of data products, robust processing pipeline, operational efficiencies and uncertainty quantification.

buffet

From Pixels to Planets

The Kepler Mission was launched in 2009 as NASAs first mission capable of finding Earth-size planets in the habitable zone of Sun-like stars. Its telescope consists of a 1.5-m primary mirror and a 0.95-m aperture. The 42 charge-coupled devices in its focal plane are read out every half hour, compressed, and then downlinked monthly. After four years, the second of four reaction wheels failed, ending the original mission. Back on earth, the Science Operations Center developed the Science Pipeline to analyze about 200,000 target stars in Keplers field of view, looking for evidence of periodic dimming suggesting that one or more planets had crossed the face of its host star. The Pipeline comprises several steps, from pixel-level calibration, through noise and artifact removal, to detection of transit-like signals and the construction of a suite of diagnostic tests to guard against false positives. The Kepler Science Pipeline consists of a pipeline infrastructure written in the Java programming language, which marshals data input to and output from MATLAB applications that are executed as external processes. The pipeline modules, which underwent continuous development and refinement even after data started arriving, employ several analytic techniques, many developed for the Kepler Project. Because of the large number of targets, the large amount of data per target and the complexity of the pipeline algorithms, the processing demands are daunting. Some pipeline modules require days to weeks to process all of their targets, even when run on NASA's 128-node Pleiades supercomputer. The software developers are still seeking ways to increase the throughput. To date, the Kepler project has discovered more than 4000 planetary candidates, of which more than 1000 have been independently confirmed or validated to be exoplanets. Funding for this mission is provided by NASAs Science Mission Directorate.

supercomputers

Orchestrator Telemetry Processing Pipeline

Orchestrator is a software application infrastructure for telemetry monitoring, logging, processing, and distribution. The architecture has been applied to support operations of a variety of planetary rovers. Built in Java with the Eclipse Rich Client Platform, Orchestrator can run on most commonly used operating systems. The pipeline supports configurable parallel processing that can significantly reduce the time needed to process a large volume of data products. Processors in the pipeline implement a simple Java interface and declare their required input from upstream processors. Orchestrator is programmatically constructed by specifying a list of Java processor classes that are initiated at runtime to form the pipeline. Input dependencies are checked at runtime. Fault tolerance can be configured to attempt continuation of processing in the event of an error or failed input dependency if possible, or to abort further processing when an error is detected. This innovation also provides support for Java Message Service broadcasts of telemetry objects to clients and provides a file system and relational database logging of telemetry. Orchestrator supports remote monitoring and control of the pipeline using browser-based JMX controls and provides several integration paths for pre-compiled legacy data processors. At the time of this reporting, the Orchestrator architecture has been used by four NASA customers to build telemetry pipelines to support field operations. Example applications include high-volume stereo image capture and processing, simultaneous data monitoring and logging from multiple vehicles. Example telemetry processors used in field test operations support include vehicle position, attitude, articulation, GPS location, power, and stereo images.

Powell, Mark

Investigation into Cloud Computing for More Robust Automated Bulk Image Geoprocessing

Geospatial resource assessments frequently require timely geospatial data processing that involves large multivariate remote sensing data sets. In particular, for disasters, response requires rapid access to large data volumes, substantial storage space and high performance processing capability. The processing and distribution of this data into usable information products requires a processing pipeline that can efficiently manage the required storage, computing utilities, and data handling requirements. In recent years, with the availability of cloud computing technology, cloud processing platforms have made available a powerful new computing infrastructure resource that can meet this need. To assess the utility of this resource, this project investigates cloud computing platforms for bulk, automated geoprocessing capabilities with respect to data handling and application development requirements. This presentation is of work being conducted by Applied Sciences Program Office at NASA-Stennis Space Center. A prototypical set of image manipulation and transformation processes that incorporate sample Unmanned Airborne System data were developed to create value-added products and tested for implementation on the "cloud". This project outlines the steps involved in creating and testing of open source software developed process code on a local prototype platform, and then transitioning this code with associated environment requirements into an analogous, but memory and processor enhanced cloud platform. A data processing cloud was used to store both standard digital camera panchromatic and multi-band image data, which were subsequently subjected to standard image processing functions such as NDVI (Normalized Difference Vegetation Index), NDMI (Normalized Difference Moisture Index), band stacking, reprojection, and other similar type data processes. Cloud infrastructure service providers were evaluated by taking these locally tested processing functions, and then applying them to a given cloud-enabled infrastructure to assesses and compare environment setup options and enabled technologies. This project reviews findings that were observed when cloud platforms were evaluated for bulk geoprocessing capabilities based on data handling and application development requirements.

Brown, Richard B.

Updates in Developing a Prototype Science Pipeline and Full-Volume, Global Hyperspectral Synthetic Data Sets for NASA’s Earth System Observatory’s Upcoming Surface, Biology and Geology Mission

The Surface Biology and Geology (SBG) mission recently passed mission confirmation review and has entered phase A – design and development. SBG will acquire high resolution solar-reflected spectroscopy and thermal infrared observations at a data rate of ~2.5 TB/day and generate products at ~40 TB/day. Given that the per-day volume is greater than NASA’s total extant airborne hyperspectral data collection, collecting, processing, disseminating, and exploiting the SBG data present new challenges. To meet these challenges, we have developed a prototype science pipeline and a full-volume global hyperspectral synthetic data set to help prepare for SBG’s flight (see poster GC42D-0730). Our science pipeline is based on the science processing technology developed for NASA’s Kepler and TESS planet-hunting missions. The pipeline infrastructure, Ziggy, provides a scalable architecture for robust, repeatable, and replicable science and application products that can be run on a range of systems from a laptop to the cloud or a supercomputer. Ziggy is compliant with NASA Procedural Requirement (NPR) 7150.2C, is at a technical readiness level (TRL) of 7 and has been released to github.com/nasa/ziggy. We integrated Ziggy with EO-1/Hyperion workflows to build a prototype pipeline and ingested the 17-year mission archive that provides globally sampled visible through shortwave infrared spectra that are representative of SBG data types and volumes. We fully implemented the first stage and processed the entire 55 TB Hyperion data set from the raw data (Level 0) to top-of-the-atmosphere radiance (Level 1R). We are currently evaluating the ISOFIT atmospheric correction module to convert the L1R data to surface reflectance (Level 2) before reprocessing the full data set to L2. Crosschecks are being performed with RadCalNet as well as with coincident observations by AVIRIS. We are also investigating modern methods for georectifying the Hyperion scenes. Finally, we describe an analysis of the cost to conduct forward processing and reprocessing campaigns for SBG on HECC with dedicated compute and storage resources using the resurrected Hyperion pipeline as a proxy for full-volume SBG data. The analysis demonstrates that SBG L0 data can be processed to L2 on HECC with full reprocessing campaigns every two years for ~$2.6M over a 7-year lifespan. Moreover, 69% of the system capacity would be available for other activities, possibly enabling future open-source science activities, including algorithm development, L3+ processing, .etc.

ESD