Search NASASearch

SEARCH · Search NASA

Results for “pipeline data processing”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 91 records · Page 5

Automated Detection and Analysis of Resident Space Objects with the 1.3-Meter Eugene Stansbery-Meter Class Autonomous Telescope

Optical telescopes dedicated to the detection of orbital debris employ large-area detectors that generate a large number of images each night. Such surveys require automated data analysis pipelines that process the images and detect moving objects. We present an overview of the data analysis pipeline employed by the 1.3-meter Eugene Stansbery-Meter Class Autonomous Telescope (ES-MCAT) on Ascension Island, operated by NASA’s Orbital Debris Program Office. The pipeline enfolds the astrometric and photometric calibration of the images, star-trail removal, object detection, correlation over multiple sequential image frames, and orbital parameter estimation. The performance of the pipeline was investigated by means of Monte-Carlo simulations in which simulated object tracks were inserted into ES-MCAT images and then processed by the pipeline. This technique allows one to confidently estimate the completeness for the detection of resident space objects as a function of apparent magnitude and angular velocity. This paper discusses these techniques and provides examples using actual data.

Paul Hickson

Parallel processing in a host plus multiple array processor system for radar

Host plus multiple array processor architecture is demonstrated to yield a modular, fast, and cost-effective system for radar processing. Software methodology for programming such a system is developed. Parallel processing with pipelined data flow among the host, array processors, and discs is implemented. Theoretical analysis of performance is made and experimentally verified. The broad class of problems to which the architecture and methodology can be applied is indicated.

Barkan, B. Z.

The Kepler Science Operations Center Pipeline Framework Extensions

The Kepler Science Operations Center (SOC) is responsible for several aspects of the Kepler Mission, including managing targets, generating on-board data compression tables, monitoring photometer health and status, processing the science data, and exporting the pipeline products to the mission archive. We describe how the generic pipeline framework software developed for Kepler is extended to achieve these goals, including pipeline configurations for processing science data and other support roles, and custom unit of work generators that control how the Kepler data are partitioned and distributed across the computing cluster. We describe the interface between the Java software that manages the retrieval and storage of the data for a given unit of work and the MATLAB algorithms that process these data. The data for each unit of work are packaged into a single file that contains everything needed by the science algorithms, allowing these files to be used to debug and evolve the algorithms offline.

Klaus, Todd C.

A machine-learning-driven data labeling pipeline for scientific analysis in MLExchange

This study introduces a novel labeling pipeline to accelerate the labeling process of scientific data sets by using artificial intelligence (AI)-guided tagging techniques. This pipeline includes a set of interconnected web-based graphical user interfaces (GUIs), where Data Clinic and MLCoach enable the preparation of machine learning (ML) models for data reduction and classification, respectively, while Label Maker is used for label assignment. Throughout this pipeline, data can be accessed through a direct connection to a file system or through Tiled for access through Hypertext Transfer Protocol (HTTP). Our experimental results present three use cases where this labeling pipeline has been instrumental for the study of large X-ray scattering data sets in the area of pattern recognition, the remote analysis of resonant soft X-ray scattering data and the fine-tuning process of foundation models. These use cases highlight the labeling capabilities of this pipeline, including the ability to label large data sets in a short period of time, to perform remote data analysis while minimizing data movement and to enhance the fine-tuning process of complex ML models with human involvement.

Chavez, Tanny (ORCID:0000000193172896)

Transcriptomics Processing Pipelines for Space Biology: An Open Source and Consensus-Driven Approach

Transcriptomics holds significant value in elucidating the relationship between gene expression, experimental factors, biological factors, and various types of omics data. Enhancing our understanding of these connections is paramount for foundational biology, which plays a pivotal role in devising solutions for challenges pertinent to both space travel and terrestrial life. The NASA GeneLab project, part of the Open Science Data Repository (OSDR.nasa.gov), seeks to accelerate space biology research through cataloging and democratizing ‘omics data, including transcriptomics. Since raw omics data are largely inaccessible to non-bioinformaticians, GeneLab works with the scientific community via the Open Science Analysis Working Groups (AWGs) to develop standard processing pipelines to generate and publish processed data. Unlike raw data, processed data have greater immediate value to diverse users with varying technical backgrounds and computational capabilities. Standardizing processing workflows is essential to match the pace of raw data generation, ensure reproducibility, and enable standardized processed data for comparison across datasets. As of June 2023, transcriptomics studies comprise over half of GeneLab datasets hosted on the OSDR, including data from bulk RNA-seq and Affymetrix or Agilent 1-Channel DNA microarray assays. In collaboration with the AWGs, GeneLab developed consensus processing pipelines for these transcriptomics data types that includes quality control, background correction (microarray only), data normalization and quantification, culminating in the detection and annotation of differentially expressed genes. The work presented here describes Nextflow implementations of GeneLab’s consensus transcriptomics pipelines that automates and accelerates processing of these datasets. In addition to the core data processing, these workflows also include raw data staging and a robust verification and validation program to identify errors in real-time, stop additional downstream computation, and preserve computational resources. These workflows are used to generate GeneLab processed data hosted on the OSDR, and are publicly available as open source software for others to use at: https://github.com/nasa/GeneLab_Data_Processing.

Jonathan Oribello

Open Science for Life in Space: Data Sharing and Tools for Knowledge Discovery

Molecular-omics, physiological-phenotypic-behavioral, and environmental-radiation telemetry data from spaceflight biological and health studies are increasingly being made findable, accessible, interoperable, and reusable for the scientific public. These data, as well as space science-relevant biospecimens, are available through NASA’s Open Science Data Repository (OSDR), which is the new umbrella grouping of NASA GeneLab, the Ames Life Sciences Data Archive (ALSDA), and the NASA Biological Institutional Scientific Collection (NBISC). The quality of data is underpinned by datasets having rich metadata (determined through Analysis Working Group members), processing pipelines to enable data reuse standards, and ontologies specifying terminology semantics (e.g., the Radiation Biology Ontology).

space biology

A Performant, Scalable Processing Pipeline for High‐Quality and FAIR Environmental Sensor Data

High-resolution environmental monitoring is necessary to record, understand, and predict biogeochemical and ecological changes particularly in coastal systems but brings significant challenges in processing and making rapidly available the resulting data. The COMPASS-FME project established a network of coastal observational sites across the Chesapeake Bay and western Lake Erie regions extensively instrumented with soil, vegetation, and weather sensors logging data every 15 min. Our data processing framework, written in R and completely open source, prioritizes rapid model-experiment iteration and makes biogeochemical data rapidly available for quality assurance/quality control, analysis, and model ingestion. This pipeline is distinguished by a standardized and modular approach to data curation, extensive metadata and documentation, and its high performance. These attributes combine to make biogeochemical data rapidly accessible across COMPASS-FME and the broader community. Flexible, powerful, and reproducible approaches to handling high-volume environmental data are crucial for accelerating biogeosciences research.

Pennington, Stephanie C. [Pacific Northwest Nation

EV-ELM (Electric Vehicle Policies with the Energy Language Model) [SWR-25-156]

Electric Vehicle Policies with the Energy Language Model (EV-ELM) leverages previous work using Large Language Models (LLMs) to find, download, and parse policy information related to energy infrastructure. In this application, we use LLMs to find policy documents related to the permitting and installation of electric vehicle charging infrastructure. This software contains the code to find, download, and parse these documents, while a related data record in the Open Energy Data Initiative (OEDI) will include the resulting output dataset that can be used for downstream analysis. The EV-ELM repository contains code for the EV-ELM project, which focuses on retrieving and processing EV permitting processes using large language models. The project is composed of two pipelines: (1) a web scraping pipeline for discovering and downloading EV permitting documents, and (2) a document parsing and extraction pipeline that processes the downloaded files to produce structured data. The web scraping pipeline is designed to extract relevant information from various websites, while the document parsing pipeline processes and analyzes the extracted documents to derive meaningful insights. Both pipelines depend on the NLR elm repository, which provides essential tools and functionalities for handling and processing the data. The web scraping pipeline is a modified version of the ordinance_gpt example within the elm repository. It has been adapted to fit the specific requirements of the EV-ELM project, ensuring that it effectively captures and processes the necessary information related to EV permitting.

Olson, Reid [National Laboratory of the Rockies (N

Ramdb: The NASA Raman Spectral Database (version 1.00).

Given that, in most instances, minimal sample preparation is required and due to its contactless instrument design, Raman spectroscopy is one of the most versatile vibrational spectroscopic techniques for the chemical analysis of environmental and biological specimens. The diversity of applications of Raman spectroscopy ranges anywhere from art [1] to planetary science missions [2]. The advancement in the use of Raman spectroscopy in Solar System missions, notably in post-mission sample return analysis, requires a spectral library holding the broad range of specimens that could be found in Solar System sources. For this purpose, we have initiated the development of a Raman spectral database (Ramdb) at NASA Ames Research Center. Currently, the database includes experimental and theoretical Raman spectra of PAHs [3, 4], as well as laboratory Raman spectra of amino acids, carbon allotropes, minerals, and analogs relevance to Earth Sciences [5], Exobiology [6], Planetary [7], and Astrochemistry [8] to name just a few examples. Ramdb can be found on the web at www.astrochemistry.org/ramdb, where raw and processed Raman spectra can be downloaded in CSV format. The laboratory Raman spectra are measured using a laser Raman spectrometer (JASCO NRS-5500-532QRI). The Raman instrument is equipped with three excitation lasers, with wavelengths of 405, 532, and 785 nm. A clean silicon substrate is used as the internal standard for wavenumber calibration. Powdered samples were prepared (microscopic >10 um, grounded microscopic < 10 um) on glass slides. Some raw data exhibited a background signal arising as a combination of laser-induced fluorescence from the sample. To correct this background, we developed a Python pipeline that uses open-source Python libraries. Ramdb provides both raw and processed (using Python pipeline) data, which includes tabulated Raman shift transitions and other measurement details. The theoretical Raman band positions of PAHs (pyrene monomers and tetramer clusters) were computed using density functional theory (DFT) with the help of the Gaussian 16 suite of programs [9]. In the near future, Ramdb will serve as a repository of Raman spectral data from Laboratory Astrophysics and Planetary Science experiments involving the irradiation of organic compounds under simulated space and planetary conditions. In addition, online and offline tools will be developed for utilising the database for comparison to the user’s sample.

N Punnakayathil

Untargeted, tandem mass spectrometry (LC/MS-MS) metaproteomes from soil samples in control and warming plots in Blodgett Forest, CA (2014-2021)

The pathways of carbon transport and loss through and from soils—soil organic matter (SOM) depolymerization to dissolved organic carbon and mineralization to carbon dioxide (CO2)—are fundamentally driven by microbial activity, which is strongly regulated by environmental conditions. As part of Lawrence Berkeley National Laboratory (LBNL) Terrestrial Ecosystem Science (TES) Belowground Biogeochemistry Science Focus Area (SFA), we have established a novel whole-soil long-term warming experiment at the University of California (UC) Blodgett Forest Research Station (Sierra Nevada) in 2014, where we study the role of biogeochemical, microbial and geochemical process interactions in SOM decomposition and stabilization. This package contains soil metaproteomics data in the context of site specific metagenomes from soil depth profiles in three paired control and warming plots from a temperate mixed forest in Northern California. Each paired plot had been subjected to experimental warming since June 2014 to simulate a predicted climate change scenario for northern California. These metaproteomes were collected in 2018 after 4.5 years of warming from five depth intervals (0-10 cm, 10-30 cm, 30-45 cm, 45-60 cm, 60-80 cm). For protein identification, the collected spectra were searched following a target-decoy search strategy against a database of metagenome predicted proteins (covering 96 samples from 2014 to 2021) representing the complete sequence diversity at the site. Data was searched with mass spectrometry database search tool (MS-GF+) using Pacific Northwest National Laboratory (PNNL)'s Data Management System (DMS) Processing pipeline. The metagenomes are published as part of another data package. Raw metaproteomic data and the data products from MS-GF+ are deposited in the Mass Spectrometry Interactive Virtual Environment (MassIVE) database under accession no. MSV000097826. Here we present a dataset that includes spectral counts for the detected proteins across samples (EMSL50964_BrodieAllMAGs_Globals_SC.txt), the sequences of the detected proteins, and sample metadata file that contains site information for the soil metaproteome samples.

Belowground Biogeochemistry Science Focus Area

Tracking the Hunga Tonga-Hunga Ha’apai Eruption Stratospheric Aerosol and Trace Gas Plumes Using Machine Learning

On January 15, 2022, the Hunga Tonga-Hunga Ha’apai (hereafter, Hunga Tonga) submarine volcano had an explosive eruption that thrusted ash, gases, and water vapor through the troposphere into the stratosphere and mesosphere. Previous studies manually tracked the aerosol and trace gas plumes over time across different positions in the southern hemisphere. Using data retrieved from low earth orbiting satellite instruments (e.g., OMPS, OMI, and CALIPSO), this research demonstrates how open-source machine learning (ML) models, like Meta’s Segment Anything Model (SAM), with prompt engineering can perform automatic plume tracking following the Hunga Tonga eruption. This extensible methodology, and modular data processing and modeling pipeline using NASA Earthdata and Openscapes, establishes a framework for systematically and rapidly studying extreme events, including volcanic eruptions and large-scale wildfires. By combining advanced machine learning techniques, such as SAM’s zero-shot learning, with large volumes of remote sensing data, this work demonstrates how AI and open science can accelerate research and generate actionable results. The tools and technologies presented here can help translate earth science to action from NASA’s current and future Earth observing satellite missions (e.g., the Atmosphere Observing System (AOS)), and assist researchers and stakeholders in understanding, mapping, and responding to natural disasters and extreme events in a changing world.

David M. Giles

Evaluating Machine Learning Approaches to Plume Tracking

On July 15, 2022, the Hunga Tonga-Hunga Ha’apai (HTHH) submarine volcano erupted, propelling trace gasses and ash through the troposphere and up into the stratosphere. Previous studies manually tracked the aerosol and trace gas plumes over time across different positions in the southern hemisphere. Using imagery from NASA’s Earth Observing System, including MODIS aerosol products and OMI sulfur dioxide products, this research demonstrates how open-source machine learning (ML) models, like Meta’s Segment Anything Model (SAM), can perform automatic plume tracking following the Hunga Tonga eruption. This extensible methodology, and modular data processing and modeling pipeline, establishes a framework for systematically and rapidly studying natural disasters, including additional volcanic eruptions and large-scale wildfires. By combining advanced machine learning techniques, such as SAM’s zero-shot learning, with large volumes of NASA’s Earth Observation and remote sensing data, this work shows how AI and open science can accelerate research and generate actionable results, even for unprecedented events. The tools and technologies presented here can help translate earth science to action from NASA’s current and future Earth observing satellite missions, and assist researchers and stakeholders in understanding, mapping, and responding to natural disasters in a changing world.

machine learning

A Data Exploration Tool for Large Sets of Spectra

We present an exploration tool for very large spectrum data sets such as the SDSS (Sloan Digital Sky Survey), LAMOST (Large Sky Area Multi-Object Fiber Spectroscopic Telescope), and 4MOST (4-meter Multi-Object Spectroscopic Telescope) data sets. The tool works in two stages: the first uses batch processing and the second runs interactively. The latter employs the NASA hyperwall, a configuration of 128 workstation displays (8 by 16 array) controlled by a parallelized software suite running on NASA's Pleiades supercomputer. The stellar subset of the Sloan Digital Sky Survey, DR10, was chosen to show how the our tool may be used. In stage one, SDSS files for 569,740 stars are processed through our data pipeline. The pipeline fits each spectrum using an iterative continuum algorithm, distinguishing emission from absorption and handling molecular absorption bands correctly. It then measures 1659 discrete atomic and molecular spectral features that were carefully preselected based on their likelihood of being visible at some spectral type. The depths relative to the local continuum at each feature wavelength are determined for each spectrum: these depths, the local S/N (signal to noise ratio) level, and DR10-supplied variables such as magnitudes, colors, positions, and radial velocities are the basic measured quantities used on the hyperwall. In stage two, each hyperwall panel is used to display a 2-D scatter plot showing the depth of feature A vs the depth of feature B for all of the stars. A and B change from panel to panel. The relationships between the various (A,B) strengths and any distinctive clustering are immediately apparent when examining and inter-comparing the different panels on the hyperwall. The interactive software allows the user to select the stars in any interesting region of any 2-D plot on the hyperwall, immediately rendering the same stars on all the other 2-D plots in a unique color. The process may be repeated multiple times, each selection displaying a distinctive color on all the plots. At any time, the spectra of the selected stars may be examined in detail on a connected workstation display. We illustrate how our approach allows us to quickly isolate and examine such interesting stellar subsets as EMP (Extremely Metal‐Poor) stars, CV (Cataclymic Variable) stars and C (Carbon)-rich stars.

Data Exploration

Kepler Planet Detection Metrics: Pixel-Level Transit Injection Tests of Pipeline Detection Efficiency for Data Release 25

This document describes the results of the fourth pixel-level transit injection experiment, which was designed to measure the detection efficiency of both the Kepler pipeline (Jenkins 2002, 2010; Jenkins et al. 2017) and the Robovetter (Coughlin 2017). Previous transit injection experiments are described in Christiansen et al. (2013, 2015a,b, 2016).In order to calculate planet occurrence rates using a given Kepler planet catalogue, produced with a given version of the Kepler pipeline, we need to know the detection efficiency of that pipeline. This can be empirically determined by injecting a suite of simulated transit signals into the Kepler data, processing the data through the pipeline, and examining the distribution of successfully recovered transits. This document describes the results for the pixel-level transit injection experiment performed to accompany the final Q1-Q17 Data Release 25 (DR25) catalogue (Thompson et al. 2017)of the Kepler Objects of Interest. The catalogue was generated using the SOC pipeline version 9.3 and the DR25 Robovetter acting on the uniformly processed Q1-Q17 DR25 light curves (Thompson et al. 2016a) and assuming the Q1-Q17 DR25 Kepler stellar properties (Mathur et al. 2017).

Pixel-Level Transit Injection

The Physics of Cooling Flow Clusters with Central Radio Sources

Central galaxies in rich clusters are the sites of cluster cooling flows, with large masses of gas cooling through part of the X-ray band. Many of these galaxies host powerful radio sources. These sources can displace and compress the X-ray gas leading to enhanced cooling and star formation. We observed the bright cooling flow Abell 2626 with a strangely distorted central radio source. We wished to understand the interaction of radio and X-ray thermal plasma, and to determine the dynamical nature of this cluster. One aim was to constrain the source of additional pressure in radio "holes" in the X-ray emission needed to support overlying shells of X-ray gas. We also aimed to study the problem of the lack of kT < 1-2 keV gas in cooling flows by searching for abundance inhomogeneities, heating from the radio source, and excess absorption. We also have a Chandra observation of this cluster. There were problems with the pipeline processing of this data due to a telemetry dropout. We are publishing the Chandra and XMM data together. Delays with the Chandra data have slowed up the publication. At the center of the cluster, there is a complex interaction of the odd, Z-shaped radio source, and the X-ray plasma. However, there are no clear radio bubbles. Also, the cluster SO galaxy IC 5337, which is projected 1.5 arcmin west of the cluster center, has unusual tail-like structures in both the radio and X-ray. It appears to be falling into the cluster center. There is a hot, probably shocked region of gas to the southwest, which is apparently due to the merger of a subcluster in this part of the system. There is also a merging subcluster to the northeast. The axes of these two mergers agrees with a supercluster filament structure.

Sarazin, Craig L.

Parallel integrated frame synchronizer chip

A parallel integrated frame synchronizer which implements a sequential pipeline process wherein serial data in the form of telemetry data or weather satellite data enters the synchronizer by means of a front-end subsystem and passes to a parallel correlator subsystem or a weather satellite data processing subsystem. When in a CCSDS mode, data from the parallel correlator subsystem passes through a window subsystem, then to a data alignment subsystem and then to a bit transition density (BTD)/cyclical redundancy check (CRC) decoding subsystem. Data from the BTD/CRC decoding subsystem or data from the weather satellite data processing subsystem is then fed to an output subsystem where it is output from a data output port.

Ghuman, Parminder Singh

Intensity Conserving Spectral Fitting

The detailed shapes of spectral line profiles provide valuable information about the emitting plasma, especially when the plasma contains an unresolved mixture of velocities, temperatures, and densities. As a result of finite spectral resolution, the intensity measured by a spectrometer is the average intensity across a wavelength bin of non-zero size. It is assigned to the wavelength position at the center of the bin. However, the actual intensity at that discrete position will be different if the profile is curved, as it invariably is. Standard fitting routines (spline, Gaussian, etc.) do not account for this difference, and this can result in significant errors when making sensitive measurements. Detection of asymmetries in solar coronal emission lines is one example. Removal of line blends is another. We have developed an iterative procedure that corrects for this effect. It can be used with any fitting function, but we employ a cubic spline in a new analysis routine called Intensity Conserving Spline Interpolation (ICSI). As the name implies, it conserves the observed intensity within each wavelength bin, which ordinary fits do not. Given the rapid convergence, speed of computation, and ease of use, we suggest that ICSI be made a standard component of the processing pipeline for spectroscopic data.

Klimchuk, J. A.

NASA's Unsteady Pressure-Sensitive Paint Research and Operational Capability Developments

In the last three years, several advancements have been made to produce a new state-of-the-art capability in the field of Aerosciences. NASA’s Aerosciences Evaluations and Test Capabilities (AETC) Portfolio Office has funded a multi-year project to produce the unsteady Pressure-Sensitive Paint (uPSP) technology as an operational capability in key ground test facilities at NASA. The research and development has primarily been conducted at NASA Ames Research Center’s (ARC) Unitary Plan Wind Tunnel (UPWT) 11-by 11-ft Transonic Wind Tunnel (TWT). The NASA ARC UPWT is one of the ground test facilities under NASA AETC’s Portfolio Office. AETC’s goals are to provide the tools to deliver the technology innovations and breakthroughs necessary to address increasingly complex research and development challenges. AETC’s integrated approach will consider the complimentary high-end compute capabilities necessary to advance analysis in conjunction with ground experimental capabilities. The uPSP Capability Challenge Project is a demonstration of several different technologies: 1) the unsteady Pressure-Sensitive Paint (uPSP) technology, and 2) Project: Red Rover, establishing a secure, reliable, fast connection between experimental and computation facilities, leveraging NASA’s computational resources within the High-End Compute Capability (HECC) Project for processing, storing, and sharing data efficiently. This project demonstrates the technical diversity and technical inclusion need to advance the field of Aerosciences. The approach to combine subject matter experts in experimental methods, optical methods, production wind tunnel testing, network engineering, high-end computing, signal processing, grid generation, and visualization while establishing the required infrastructure for subject matter experts to have access to the data while the wind tunnel test is being conducted. The most recent advancements for the uPSP technology have focused on three key areas: development of data products, robust processing pipeline, operational efficiencies and uncertainty quantification.

buffet