Search NASA⌕ Search

SEARCH · Search NASA

Results for “data pipeline”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 307 records · Page 17

Feasibility study, software design, layout and simulation of a two-dimensional fast Fourier transform machine for use in optical array interferometry

The NASA-Cornell Univ.-Worcester Polytechnic Institute Fast Fourier Transform (FFT) chip based on the architecture of the systolic FFT computation as presented by Boriakoff is implemented into an operating device design. The kernel of the system, a systolic inner product floating point processor, was designed to be assembled into a systolic network that would take incoming data streams in pipeline fashion and provide an FFT output at the same rate, word by word. It was thoroughly simulated for proper operation, and it has passed a comprehensive set of tests showing no operational errors. The black box specifications of the chip, which conform to the initial requirements of the design as specified by NASA, are given. The five subcells are described and their high level function description, logic diagrams, and simulation results are presented. Some modification of the Read Only Memory (ROM) design were made, since some errors were found in it. Because a four stage pipeline structure was used, simulating such a structure is more difficult than an ordinary structure. Simulation methods are discussed. Chip signal protocols and chip pinout are explained.

Boriakoff, Valentin↗

Telecommunications issues of intelligent database management for ground processing systems in the EOS era

Future NASA earth science missions, including the Earth Observing System (EOS), will be generating vast amounts of data that must be processed and stored at various locations around the world. Here we present a stepwise-refinement of the intelligent database management (IDM) of the distributed active archive center (DAAC - one of seven regionally-located EOSDIS archive sites) architecture, to showcase the telecommunications issues involved. We develop this architecture into a general overall design. We show that the current evolution of protocols is sufficient to support IDM at Gbps rates over large distances. We also show that network design can accommodate a flexible data ingestion storage pipeline and a user extraction and visualization engine, without interference between the two.

Touch, Joseph D.↗

NASA Tech Briefs, December 2005

Topics covered include: Video Mosaicking for Inspection of Gas Pipelines; Shuttle-Data-Tape XML Translator; Highly Reliable, High-Speed, Unidirectional Serial Data Links; Data-Analysis System for Entry, Descent, and Landing; Hybrid UV Imager Containing Face-Up AlGaN/GaN Photodiodes; Multiple Embedded Processors for Fault-Tolerant Computing; Hybrid Power Management; Magnetometer Based on Optoelectronic Microwave Oscillator; Program Predicts Time Courses of Human/ Computer Interactions; Chimera Grid Tools; Astronomer's Proposal Tool; Conservative Patch Algorithm and Mesh Sequencing for PAB3D; Fitting Nonlinear Curves by Use of Optimization Techniques; Tool for Viewing Faults Under Terrain; Automated Synthesis of Long Communication Delays for Testing; Solving Nonlinear Euler Equations With Arbitrary Accuracy; Self-Organizing-Map Program for Analyzing Multivariate Data; Tool for Sizing Analysis of the Advanced Life Support System; Control Software for a High-Performance Telerobot; Java Radar Analysis Tool; Architecture for Verifiable Software; Tool for Ranking Research Options; Enhanced, Partially Redundant Emergency Notification System; Close-Call Action Log Form; Task Description Language; Improved Small-Particle Powders for Plasma Spraying; Bonding-Compatible Corrosion Inhibitor for Rinsing Metals; Wipes, Coatings, and Patches for Detecting Hydrazines; Rotating Vessels for Growing Protein Crystals; Oscillating-Linear-Drive Vacuum Compressor for CO2; Mechanically Biased, Hinged Pairs of Piezoelectric Benders; Apparatus for Precise Indium-Bump Bonding of Microchips; Radiation Dosimetry via Automated Fluorescence Microscopy; Multistage Magnetic Separator of Cells and Proteins; Elastic-Tether Suits for Artificial Gravity and Exercise; Multichannel Brain-Signal-Amplifying and Digitizing System; Ester-Based Electrolytes for Low-Temperature Li-Ion Cells; Hygrometer for Detecting Water in Partially Enclosed Volumes; Radio-Frequency Plasma Cleaning of a Penning Malmberg Trap; Reduction of Flap Side Edge Noise - the Blowing Flap; and Preventing Accidental Ignition of Upper-Stage Rocket Motors.

Source record↗

Intensity Conserving Spectral Fitting

The detailed shapes of spectral line profiles provide valuable information about the emitting plasma, especially when the plasma contains an unresolved mixture of velocities, temperatures, and densities. As a result of finite spectral resolution, the intensity measured by a spectrometer is the average intensity across a wavelength bin of non-zero size. It is assigned to the wavelength position at the center of the bin. However, the actual intensity at that discrete position will be different if the profile is curved, as it invariably is. Standard fitting routines (spline, Gaussian, etc.) do not account for this difference, and this can result in significant errors when making sensitive measurements. Detection of asymmetries in solar coronal emission lines is one example. Removal of line blends is another. We have developed an iterative procedure that corrects for this effect. It can be used with any fitting function, but we employ a cubic spline in a new analysis routine called Intensity Conserving Spline Interpolation (ICSI). As the name implies, it conserves the observed intensity within each wavelength bin, which ordinary fits do not. Given the rapid convergence, speed of computation, and ease of use, we suggest that ICSI be made a standard component of the processing pipeline for spectroscopic data.

Klimchuk, J. A.↗

Evidence for an Optical Low-Frequency Quasi-Periodic Oscillation in the Kepler Light Curve of an Active Galaxy

We report evidence for a quasi-periodic oscillation (QPO) in the optical light curve of KIC 9650712, a narrow-line Seyfert 1 galaxy in the original Kepler field. After the development and application of a pipeline for Kepler data specific to active galactic nuclei (AGNs), one of our sample of 21 AGNs selected by infrared photometry and X-ray flux demonstrates a peak in the power spectrum at log ν = −6.58 Hz, corresponding to a temporal period of t = 44 days. We note that although the power spectrum is well fit by a model consisting of a Lorentzian and a single power law, alternative continuum models cannot be ruled out. From optical spectroscopy, we measure the black hole mass of this AGN as log (M(sub BH)/solar mass) = 8.17. We find that this frequency lies along a correlation between low-frequency QPOs and black hole mass from stellar and intermediate mass black holes to AGNs, similar to the known correlation in high-frequency QPOs.

quasi-periodic oscillation (QPO)↗

BRAKER3: Fully automated genome annotation using RNA-seq and protein evidence with GeneMark-ETP, AUGUSTUS, and TSEBRA

Gene prediction has remained an active area of bioinformatics research for a long time. Still, gene prediction in large eukaryotic genomes presents a challenge that must be addressed by new algorithms. The amount and significance of the evidence available from transcriptomes and proteomes vary across genomes, between genes, and even along a single gene. User-friendly and accurate annotation pipelines that can cope with such data heterogeneity are needed. The previously developed annotation pipelines BRAKER1 and BRAKER2 use RNA-seq or protein data, respectively, but not both. A further significant performance improvement integrating all three data types was made by the recently released GeneMark-ETP. We here present the BRAKER3 pipeline that builds on GeneMark-ETP and AUGUSTUS, and further improves accuracy using the TSEBRA combiner. BRAKER3 annotates protein-coding genes in eukaryotic genomes using both short-read RNA-seq and a large protein database, along with statistical models learned iteratively and specifically for the target genome. We benchmarked the new pipeline on genomes of 11 species under an assumed level of relatedness of the target species proteome to available proteomes. BRAKER3 outperforms BRAKER1 and BRAKER2. The average transcript-level F1-score is increased by about 20 percentage points on average, whereas the difference is most pronounced for species with large and complex genomes. BRAKER3 also outperforms other existing tools, MAKER2, Funannotate, and FINDER. The code of BRAKER3 is available on GitHub and as a ready-to-run Docker container for execution with Docker or Singularity. Overall, BRAKER3 is an accurate, easy-to-use tool for eukaryotic genome annotation.

59 BASIC BIOLOGICAL SCIENCES↗

The Atacama Cosmology Telescope: Data Characterization and Map Making

We present a description of the data reduction and mapmaking pipeline used for the 2008 observing season of the Atacama Cosmology Telescope (ACT). The data presented here at 148 GHz represent 12% or the 90 TB collected by ACT from 2007 to 2010. In 2008 we observed for 136 days, producing a total of 142h of data (11 TB for the 148 GHz band only), with a daily average of 10.5 h of observation. From these, 108.5 h were devoted to 850 sq deg stripe (11.2 h by 9 deg.1) centered on a declination of -52 deg.7, while 175 h were devoted to a 280 square deg stripe (4.5 h by 4 deg.8) centered at the celestial equator. We discuss sources of statistical and systematic noise, calibration, telescope pointing and data selection. Out of 1260 survey hours and 1024 detectors per array, 816 h and 593 effective detectors remain after data selection for this frequency band, yielding a 38 % survey efficiency. The total sensitivity in 2008, determined from the noise level between 5 Hz and 20 Hz in the time-ordered data stream (TOD), is 32 muK square root of s in CMB units. Atmospheric brightness fluctuations constitute the main contaminant in the data and dominate the detector and noise covariance at low frequencies in the TOD. The maps were made by solving the lease squares problem using the Preconditioned Conjugate Gradient method, incorporating the details of the detector and noise correlations. Cross-correlation with WMAP sky maps as well as analysis from simulations reveal the our maps are unbiased at l > 300. This paper accompanies the public release of the 148 GHz southern stripe maps from 2008. The techniques described here will be applied to future maps and data releases.

Duenner, Rolando↗

RadLab and the Environmental Data Application Dashboard: Graphical and Programming Interfaces for Interrogation of Space Telemetry Data

Sensors on the International Space Station (ISS) and multiple spacecraft elsewhere in Earth orbit and in deep space continuously monitor and collect environmental data, transmitting this information back to Earth. These data include ionizing radiation and, on the ISS, CO2, relative humidity levels, and temperature, and are of great importance to space biology research. Ionizing radiation in particular has been established in ground-based experiments as being correlated with increased risk of carcinogenesis and cardiovascular and neurological effects. Looking ahead to future long duration crewed missions beyond low Earth orbit, the ability to study how factors including CO2 levels, light cycle, temperature modulate the response to ionizing radiation and microgravity is essential. To date, access to these data has been fragmented across space agencies, spacecraft, and databases. To address this issue, NASA’s Open Science Data Repository (osdr.nasa.gov) has developed two Web applications: the Environmental Data Application (EDA) and a radiation-specific RadLab. Each consists of an API (application programming interface) and an associated GUI (graphical user interface) that provide single points of access to the data. To date, OSDR has focused on the sensors from payloads and radiation detectors located on the ISS. The Web applications process telemetry information and associated data, such as spacecraft location and orientation, from multiple international databases. The applications’ request syntax enables users to interrogate these data by craft, sensor type, time range, radiation type (galactic cosmic rays, solar particle events, the contribution of the South Atlantic Anomaly), facilitating arbitrary comparisons of original source data at varying time resolutions. The applications provide programmatic access for use in computational pipelines and GUIs for data visualization and exploration, making these data FAIR (Findable, Accessible, Interoperable, and Reusable), complementing the biological data contained in OSDR, and providing the space science community with a valuable resource for scientific analyses.

radiation↗

UBW (USLCI-Brightway2) [SWR-25-169]

Life cycle inventory (LCI) data are critical for robust life cycle assessment (LCA), yet many widely used datasets such as the U.S. Life Cycle Inventory (USLCI) are not natively compatible with advanced modeling frameworks like Brightway2. This work presents an automated pipeline to transform USLCI data into a fully functional Brightway2 project. The workflow performs systematic data cleaning, resolves duplicate process and exchange identifiers, and applies allocation to multi-output processes. Technosphere and biosphere flows are harmonized through unit conversions and a bridge mapping to the biosphere3 database, with comprehensive logging of missing flows and cutoff issues. The resulting Brightway2 database is validated using matrix diagnostics to ensure consistency of the technosphere, and is benchmarked via life cycle impact assessment (LCIA) methods such as ReCiPe and IPCC GWP. Outputs include reproducible CSV exports of corrected processes, elementary flows, characterization factors, and LCIA results, alongside backup utilities for project sharing. This pipeline lowers barriers for integrating USLCI data into open-source LCA workflows, enabling reproducible, validated LCA inventories within the Brightway 2 framework.

Ghosh, Tapajyoti [National Laboratory of the Rocki↗

mphys-surrogate-model

This repository contains python scripts for building and studying reduced-order-modeling representations of droplet coalescence for eventual use in atmospheric models. The included data are generated from high-fidelity superdroplet methods and are utilized by machine learning pipelines to build data-driven models of droplet size distributions that evolve under coalescence. This repository further includes scripts to determine prediction (uncertainty) intervals on the data-driven model products based on conformal prediction.

Katona, JonasE [Lawrence Livermore National Labora↗

ICED: An Integrated CGRA Framework Enabling DFVS-Aware Acceleration

oarse-grained reconfigurable arrays (CGRAs) are a promising solution to enable energy-efficient acceleration of applications from different domains. By leveraging reconfiguration at the functional level, they can adapt to significantly different computational patterns. Existing CGRA mapping approaches extract instruction-level parallelism, exploit loop-pipelining opportunities, guarantee the data dependency, and target high throughput of a given loop. However, the recurrence data-dependency in the DFG and the mismatch between required and available computing/communication resources complicate the mapping, and might lead to significant unbalances in the utilization of the CGRA's tiles. This results in wasted power for tiles with low utilization. Applying dynamic voltage and frequency scaling (DVFS) can potentially solve this challenge and improve energy efficiency by adjusting voltage and frequency of different tiles independently. CGRAs have also been successful in accelerating data-dependent streaming applications. However, in these applications, the execution time of each kernel in the pipeline might dynamically vary depending on the characteristics of the input. This also leads to under-utilization of resources for the dynamically changing kernels that do not limit the application throughput. DVFS can also improve energy efficiency for these applications by dynamically changing the voltage and frequency levels of tiles that host non performance-constraining kernels. This paper proposes ICEDTEA -- an integrated DVFS-aware framework to map applications on CGRAs that support power islands. ICEDTEA proposes a CGRA architecture supporting DVFS islands at varying granularity (from a single tile to a group of tiles) and the related DVFS-aware compilation and mapping toolchain. ICEDTEA is the first work that introduces DVFS support for spatio-temporal CGRAs at power-island levels. The experimental evaluation shows that ICEDTEA improves average utilization by 2.3$\times$ and energy-efficiency by 1.32$\times$ over a conventional CGRA. With streaming applications, ICEDTEA improves energy efficiency by 1.12$\times$ over a state-of-the-art CGRA that introduces partial dynamic reconfiguration to adapt to variations in kernels' throughput.

Tan, Cheng↗

Untargeted, tandem mass spectrometry (LC/MS-MS) metaproteomes from soil samples in control and warming plots in Blodgett Forest, CA (2014-2021)

The pathways of carbon transport and loss through and from soils—soil organic matter (SOM) depolymerization to dissolved organic carbon and mineralization to carbon dioxide (CO2)—are fundamentally driven by microbial activity, which is strongly regulated by environmental conditions. As part of Lawrence Berkeley National Laboratory (LBNL) Terrestrial Ecosystem Science (TES) Belowground Biogeochemistry Science Focus Area (SFA), we have established a novel whole-soil long-term warming experiment at the University of California (UC) Blodgett Forest Research Station (Sierra Nevada) in 2014, where we study the role of biogeochemical, microbial and geochemical process interactions in SOM decomposition and stabilization. This package contains soil metaproteomics data in the context of site specific metagenomes from soil depth profiles in three paired control and warming plots from a temperate mixed forest in Northern California. Each paired plot had been subjected to experimental warming since June 2014 to simulate a predicted climate change scenario for northern California. These metaproteomes were collected in 2018 after 4.5 years of warming from five depth intervals (0-10 cm, 10-30 cm, 30-45 cm, 45-60 cm, 60-80 cm). For protein identification, the collected spectra were searched following a target-decoy search strategy against a database of metagenome predicted proteins (covering 96 samples from 2014 to 2021) representing the complete sequence diversity at the site. Data was searched with mass spectrometry database search tool (MS-GF+) using Pacific Northwest National Laboratory (PNNL)'s Data Management System (DMS) Processing pipeline. The metagenomes are published as part of another data package. Raw metaproteomic data and the data products from MS-GF+ are deposited in the Mass Spectrometry Interactive Virtual Environment (MassIVE) database under accession no. MSV000097826. Here we present a dataset that includes spectral counts for the detected proteins across samples (EMSL50964_BrodieAllMAGs_Globals_SC.txt), the sequences of the detected proteins, and sample metadata file that contains site information for the soil metaproteome samples.

Belowground Biogeochemistry Science Focus Area↗

RadLab: Graphical and Programming Interfaces for Interrogation of Space Telemetry Data

Sensors on multiple spacecraft in and beyond low Earth orbit continuously monitor and collect space radiation data and transmit it back to Earth. These data are of vast importance to space biology research, as ionizing radiation affects living organisms—astronauts and non-human experiment subjects alike—placing them at higher risk of carcinogenesis, degenerative diseases, and radiation sickness. Therefore, knowledge of the biological effects of space radiation is essential for planning future crewed missions beyond low Earth orbit. The RadLab project, initiated by GeneLab and ALSDA (the Open Science Data Repository; OSDR) and sponsored by the NASA Human Research Program, is a new effort aimed at connecting dosimetry data from radiation detectors located on the International Space Station (ISS), as well as other spacecraft. To date, access to these data has been fragmented across space agencies and databases; to address this issue, we have developed an application programming interface (API) and an associated graphical user interface (GUI) designed to provide a single point of access to the data. As of now, OSDR has focused on the detectors located on the ISS, with the long-term goal to establish a self-sustained portal receiving continuous updates through APIs connecting to multiple radiation databases of varying scope, as well as individual investigator contributions. The RadLab API implements a request syntax enabling users to query data by craft, sensor type, timespan, etc, allowing for arbitrary combinations of original source data, thus providing programmatic access for use in computational pipelines, while the GUI facilitates data visualization and exploration, making these data FAIR (Findable, Accessible, Interoperable, and Reusable), complementing the biological data contained in OSDR, and providing the space science community with a valuable resource for scientific analyses.

radiation↗

Dark Energy Survey Year 6 Results: Weak Lensing and Galaxy Clustering Cosmological Analysis Framework

We present the methodology for the weak lensing and galaxy clustering analyses of the Dark Energy Survey (DES) Year 6 data set. In this work, we design and validate the analysis pipeline for the cosmic shear, galaxy clustering plus galaxy$-$galaxy lensing ($2 \times 2$pt), and the joint analysis in the $3 \times 2$pt. Our framework accounts for key theoretical uncertainties, such as baryonic feedback and galaxy bias, incorporating both linear and non-linear models. We apply scale cuts in regimes where theoretical modeling becomes unreliable. The robustness of the pipeline is validated using mock data and simulations, confirming unbiased cosmological constraints and highlighting the importance of posterior projection effects in the validation process. As a result, we deliver robust and validated analysis pipelines for cosmic shear, $2 \times 2$pt, and $3 \times 2$pt in $Λ$CDM and $w$CDM scenarios, including a well-defined set of scales suitable for real data analysis, a robust prescription for theoretical systematics, and the theoretical covariance of the signal. This comprehensive methodology also lays the groundwork for future galaxy surveys such as the Vera C. Rubin Observatory Legacy Survey of Space and Time.

Sanchez-Cid, D. [Zurich U.; Madrid, CIEMAT; Madrid↗

Iterative solution of large, sparse linear systems on a static data flow architecture - Performance studies

The applicability of static data flow architectures to the iterative solution of sparse linear systems of equations is investigated. An analytic performance model of a static data flow computation is developed. This model includes both spatial parallelism, concurrent execution in multiple PE's, and pipelining, the streaming of data from array memories through the PE's. The performance model is used to analyze a row partitioned iterative algorithm for solving sparse linear systems of algebraic equations. Based on this analysis, design parameters for the static data flow architecture as a function of matrix sparsity and dimension are proposed.

Reed, D. A.↗

An analysis of parameter compression and Full-Modeling techniques with Velocileptors for DESI 2024 and beyond

In anticipation of forthcoming data releases of current and future spectroscopic surveys, we present the validation tests and analysis of systematic effects within velocileptors modeling pipeline when fitting mock data from the AbacusSummit N-body simulations. We compare the constraints obtained from parameter compression methods to the direct fitting (Full-Modeling) approaches of modeling the galaxy power spectra, and show that the ShapeFit extension to the traditional template method is consistent with the Full-Modeling method within the standard ΛCDM parameter space. We show the dependence on scale cuts when fitting the different redshift bins using the ShapeFit and Full-Modeling methods. We test the ability to jointly fit data from multiple redshift bins as well as joint analysis of the pre-reconstruction power spectrum with the post-reconstruction BAO correlation function signal. We further demonstrate the behavior of the model when opening up the parameter space beyond ΛCDM and also when combining likelihoods with external datasets, namely the Planck CMB priors. Finally, we describe different parametrization options for the galaxy bias, counterterm, and stochastic parameters, and employ the halo model in order to physically motivate suitable priors that are necessary to ensure the stability of the perturbation theory.

79 ASTRONOMY AND ASTROPHYSICS↗

The European Southern Observatory-MIDAS table file system

The new and substantially upgraded version of the Table File System in MIDAS is presented as a scientific database system. MIDAS applications for performing database operations on tables are discussed, for instance, the exchange of the data to and from the TFS, the selection of objects, the uncertainty joins across tables, and the graphical representation of data. This upgraded version of the TFS is a full implementation of the binary table extension of the FITS format; in addition, it also supports arrays of strings. Different storage strategies for optimal access of very large data sets are implemented and are addressed in detail. As a simple relational database, the TFS may be used for the management of personal data files. This opens the way to intelligent pipeline processing of large amounts of data. One of the key features of the Table File System is to provide also an extensive set of tools for the analysis of the final results of a reduction process. Column operations using standard and special mathematical functions as well as statistical distributions can be carried out; commands for linear regression and model fitting using nonlinear least square methods and user-defined functions are available. Finally, statistical tests of hypothesis and multivariate methods can also operate on tables.

Peron, M.↗

Orchestrator Telemetry Processing Pipeline

Orchestrator is a software application infrastructure for telemetry monitoring, logging, processing, and distribution. The architecture has been applied to support operations of a variety of planetary rovers. Built in Java with the Eclipse Rich Client Platform, Orchestrator can run on most commonly used operating systems. The pipeline supports configurable parallel processing that can significantly reduce the time needed to process a large volume of data products. Processors in the pipeline implement a simple Java interface and declare their required input from upstream processors. Orchestrator is programmatically constructed by specifying a list of Java processor classes that are initiated at runtime to form the pipeline. Input dependencies are checked at runtime. Fault tolerance can be configured to attempt continuation of processing in the event of an error or failed input dependency if possible, or to abort further processing when an error is detected. This innovation also provides support for Java Message Service broadcasts of telemetry objects to clients and provides a file system and relational database logging of telemetry. Orchestrator supports remote monitoring and control of the pipeline using browser-based JMX controls and provides several integration paths for pre-compiled legacy data processors. At the time of this reporting, the Orchestrator architecture has been used by four NASA customers to build telemetry pipelines to support field operations. Example applications include high-volume stereo image capture and processing, simultaneous data monitoring and logging from multiple vehicles. Example telemetry processors used in field test operations support include vehicle position, attitude, articulation, GPS location, power, and stereo images.

Powell, Mark↗