Search NASA⌕ Search

SEARCH · Search NASA

Results for “data pipeline”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 181 records · Page 10

The Nasa SRA Process as It Relates to Open-Source Workflows Developed for GeneLab Data Processing

To release open, standards-compliant processed data sets in the Open Science Data Repository (OSDR), the GeneLab Data Processing team works with the scientific community through the OSDR Analysis Working Groups to design and build open-source data processing pipelines. Once baselined internally, these pipelines are wrapped into workflows and published on the NASA GeneLab Data Processing public GitHub repository along with detailed instructions for installation and use. Each workflow must be approved through NASA's Software Release Authorization (SRA) process prior to publishing. However, the SRA process lacks sufficient documentation and clarity regarding which forms are applicable for new open-source software that utilizes publicly available 3rd party tools, and the SRA process can take several months to complete, making sharing software outside of NASA cumbersome and in contradiction with the concept of Open Science. Furthermore, the SRA process was designed as a one-size fits all approach and thus many of the questions asked are not applicable to our open-source workflows. Here we describe the software provided on the NASA GeneLab Data Processing GitHub repository, summarize our experiences with the SRA process to release these software, and propose a more stream-lined approach for review of open-source projects.

Software Release Authorization↗

The NASA SRA Process as it Relates to Open-Source Workflows Developed for GeneLab Data Processing

To release open, standards-compliant processed data sets in the Open Science Data Repository (OSDR), the GeneLab Data Processing team works with the scientific community through the OSDR Analysis Working Groups to design and build open-source data processing pipelines. Once baselined internally, these pipelines are wrapped into workflows and published on the NASA GeneLab Data Processing public GitHub repository along with detailed instructions for installation and use. Each workflow must be approved through NASA's Software Release Authorization (SRA) process prior to publishing. However, the SRA process lacks sufficient documentation and clarity regarding which forms are applicable for new open-source software that utilizes publicly available 3rd party tools, and the SRA process can take several months to complete, making sharing software outside of NASA cumbersome and in contradiction with the concept of Open Science. Furthermore, the SRA process was designed as a one-size fits all approach and thus many of the questions asked are not applicable to our open-source workflows. Here we describe the software provided on the NASA GeneLab Data Processing GitHub repository, summarize our experiences with the SRA process to release these software, and propose a more stream-lined approach for review of open-source projects.

Software Release Authorization↗

Kepler: A Search for Terrestrial Planets - SOC 9.3 DR25 Pipeline Parameter Configuration Reports

This document describes the manner in which the pipeline and algorithm parameters for the Kepler Science Operations Center (SOC) science data processing pipeline were managed. This document is intended for scientists and software developers who wish to better understand the software design for the final Kepler codebase (SOC 9.3) and the effect of the software parameters on the Data Release (DR) 25 archival products.

exoplanet↗

Reinforcement Learning for In-Spill Optimization of the Mu2e Resonant Extraction: Compensating Non-Stationarity

We present design considerations and challenges for the fast machine learning component of a third-order resonant beam extraction regulation system being commissioned to deliver steady beam rates to the mu2e experiment at Fermilab. Dedicated quadrupoles drive the tune toward the 29/3 resonance each spill, extracting beam at kV multiwire septa. The overall Spill Regulation System consists of (1) a “slow” process using ~100-spill averages to adjust the base quad ramp infrequently, (2) a feedforward harmonic content compensator, and (3) the “fast” ML agent reacting during each ongoing spill with on-the-fly additive corrections to the sum of (1) and (2). We have demonstrated improved beam-rate steadying for a fast ML agent compared to a PID controller using a quasi-physical spill simulation, and demonstrated distillation of that simulation into a predictive surrogate model. Current work includes a data-and-training pipeline to generate data-aware surrogates with real-world dynamics, even as the dynamics shift unpredictably. The surrogates are to act as RL environments against which to train our fast ML control agents before deploying them on FPGA in the live system. Further current efforts focus on modeling and controlling beam loss around the storage ring, understanding additional available hardware inputs to the model, and the interplay of these with beam-steadying performance.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

Kepler Data Validation I: Architecture, Diagnostic Tests, and Data Products for Vetting Transiting Planet Candidates

The Kepler Mission was designed to identify and characterize transiting planets in the Kepler Field of View and to determine their occurrence rates. Emphasis was placed on identification of Earth-size planets orbiting in the Habitable Zone of their host stars. Science data were acquired for a period of four years. Long-cadence data with 29.4 min sampling were obtained for approx. 200,000 individual stellar targets in at least one observing quarter in the primary Kepler Mission. Light curves for target stars are extracted in the Kepler Science Data Processing Pipeline, and are searched for transiting planet signatures. A Threshold Crossing Event is generated in the transit search for targets where the transit detection threshold is exceeded and transit consistency checks are satisfied. These targets are subjected to further scrutiny in the Data Validation (DV) component of the Pipeline. Transiting planet candidates are characterized in DV, and light curves are searched for additional planets after transit signatures are modeled and removed. A suite of diagnostic tests is performed on all candidates to aid in discrimination between genuine transiting planets and instrumental or astrophysical false positives. Data products are generated per target and planet candidate to document and display transiting planet model fit and diagnostic test results. These products are exported to the Exoplanet Archive at the NASA Exoplanet Science Institute, and are available to the community. We describe the DV architecture and diagnostic tests, and provide a brief overview of the data products. Transiting planet modeling and the search for multiple planets on individual targets are described in a companion paper. The final revision of the Kepler Pipeline code base is available to the general public through GitHub. The Kepler Pipeline has also been modified to support the Transiting Exoplanet Survey Satellite (TESS) Mission which is expected to commence in 2018.

data analysis↗

TomoPyUI : a user-friendly tool for rapid tomography alignment and reconstruction

The management and processing of synchrotron and neutron computed tomography data can be a complex, labor-intensive and unstructured process. Users devote substantial time to both manually processing their data ( i.e. organizing data/metadata, applying image filters etc. ) and waiting for the computation of iterative alignment and reconstruction algorithms to finish. In this work, we present a solution to these problems: TomoPyUI , a user interface for the well known tomography data processing package TomoPy . This highly visual Python software package guides the user through the tomography processing pipeline from data import, preprocessing, alignment and finally to 3D volume reconstruction. The TomoPyUI systematic intermediate data and metadata storage system improves organization, and the inspection and manipulation tools (built within the application) help to avoid interrupted workflows. Notably, TomoPyUI operates entirely within a Jupyter environment. Herein, we provide a summary of these key features of TomoPyUI , along with an overview of the tomography processing pipeline, a discussion of the landscape of existing tomography processing software and the purpose of TomoPyUI , and a demonstration of its capabilities for real tomography data collected at SSRL beamline 6-2c.

73 NUCLEAR PHYSICS AND RADIATION PHYSICS↗

Processing Raw HST Data With Up-to-Date Calibration Data

On-the-Fly Reprocessing (OTFR) is a collection of data-processing routines that work within the context of the Hubble Space Telescope (HST) pipeline data-flow system. The purpose served by OTFR is to generate, on demand, scientifically useful data products from raw HST data stored in an archive. First, on the basis of the requested final data products, OTFR retrieves the corresponding sets of raw data from the archives. Next, OTFR processes the raw data sets to remove artifacts and to establish proper header and other template information. Finally, the calibration routines appropriate to the specific data sets are invoked to produce the requested data products, and the data products are released to an archive distribution system for transmission to the requesting party. OTFR offers two notable advantages: (1) Inasmuch as calibrated data occupy about 8 times as much storage space as do raw data, by obviating storage of calibrated data, OTFR reduces the storage capacity needed by the archive; and (2) the calibration routines can be updated to give requesters the benefit of the most recent calibrations.

Miller, Warren↗

Generating Land Surface Reflectance for the New Generation of Geostationary Satellite Sensors with the MAIAC Algorithm

The latest generation of geostationary satellite sensors, including the GOES-16/ABI and the Himawari 8/AHI, provide exciting capability to monitor land surface at very high temporal resolutions (5-15 minute intervals) and with spatial and spectral characteristics that mimic the Earth Observing System flagship MODIS. However, geostationary data feature changing sun angles at constant view geometry, which is almost reciprocal to sun-synchronous observations. Such a challenge needs to be carefully addressed before one can exploit the full potential of the new sources of data. Here we take on this challenge with Multi-Angle Implementation of Atmospheric Correction (MAIAC) algorithm, recently developed for accurate and globally robust applications like the MODIS Collection 6 re-processing. MAIAC first grids the top-of- atmosphere measurements to a fixed grid so that the spectral and physical signatures of each grid cell are stacked (“remembered”) over time and used to dramatically improve cloud/shadow/snow detection, which is by far the dominant error source in the remote sensing. It also exploits the changing sun-view geometry of the geostationary sensor to characterize surface BRDF with augmented angular resolution for accurate aerosol retrievals and atmospheric correction. The high temporal resolutions of the geostationary data indeed make the BRDF retrieval much simpler and more robust as compared with sun-synchronous sensors such as MODIS. As a prototype test for the geostationary-data processing pipeline on NASA Earth Exchange (GEONEX), we apply MAIAC to process 18 months of data from Himawari 8/AHI over Australia. We generate a suite of test results, including the input TOA reflectance and the output cloud mask, aerosol optical depth (AOD), and the atmospherically-corrected surface reflectance for a variety of geographic locations, terrain, and land cover types. Comparison with MODIS data indicates a general agreement between the retrieved surface reflectance products. Furthermore, the geostationary results satisfactorily capture the movement of clouds and variations in atmospheric dust/aerosol concentrations, suggesting that high quality land surface and vegetation datasets from the advanced geostationary sensors can help complement and improve the corresponding EOS products.

geostationary satellite sensors↗

Carbon dioxide pipeline network transportation cost model: evaluating economic and geographic factors for efficient carbon capture, storage, and utilization

This study presents a comprehensive pipeline network modeling framework to estimate the CO 2 delivery cost for CO 2 utilization and geologic CO 2 storage across the United States. We developed a Python-based CO 2 pipeline transportation cost model leveraging Argonne National Laboratory’s pipeline engineering expertise and detailed natural gas transmission pipeline cost data across U.S. regions. Using existing road corridors as practical routing guides, the model designs pipeline networks that aggregate CO 2 from one or multiple sources and deliver it to selected destinations. It then minimizes the total transportation cost by optimizing pipeline diameters and incorporating booster pumps. A key contribution is the incorporation of up-to-date, region-specific cost factors with itemized components for materials, labor, miscellaneous construction expenses, and right-of-way acquisition. Results emphasize that regional variation and economies of scale associated with CO 2 pipeline costs are significant and should be explicitly accounted for in screening and planning studies. By combining realistic routing constraints with regionalized cost inputs, the model provides transparent design methodology and location-specific insights into source–destination delivery costs, including the effects of routing complexity along existing road networks. We demonstrate the model with two illustrative case studies – one for CO 2 storage and one for CO 2 utilization – in which the model designs pipeline networks spanning hundreds of miles across the states, collecting CO 2 from multiple sources and delivering it to designated endpoints while minimizing levelized cost of delivery via diameter and compression optimization. The model offers a practical, scalable approach for alternative design option screening and early-stage CO 2 transportation planning.

CCS↗

Exploring Beyond Earth's Atmosphere with Human-Machine Teams

NASA's highly successful Kepler Mission has revolutionized our understanding of the Galaxy. We now know that planets, even Earth-size planets in the habitable zone, are common. With the end of the Kepler Mission we now look to the future with the Transiting Exoplanet Survey Satellite (TESS) which will discover thousands of exoplanets in orbit around the brightest stars in the sky. In a two-year survey, TESS will perform an all-sky search of more than 200,000 stars for temporary drops in brightness caused by planetary transits. With Kepler and TESS, humanity is finally at the verge of studying the masses, sizes, densities, orbits, and atmospheres of a large cohort of small planets, including a sample of rocky worlds in the habitable zones of their host stars which may prove to host life. The massive data sets generated by Kepler and TESS must be meticulously combed for the weakest planetary signals every month. While a daunting and error-prone task for humans, this is an exciting opportunity for the breakthroughs recently seen in machine learning. Specifically, traditional methods for identifying planet transits require extensive data processing pipelines followed by extensive human vetting. This manual process risks loss of information due to the data processing and to inconsistency and biases due to individual human vetters. The latest advancements in machine learning will allow an objective classifier to minimize the losses of information and greatly lessen the burden on the human vetters, in addition to providing assessment of quality and score to each planet candidate, freeing the humans to concentrate on border cases and other more interesting investigations.

Smith, Jeffrey C.↗

Prediction of Aircraft Estimated Time of Arrival Using A Supervised Learning Approach

We present a novel data-driven approach for prediction of the estimated time of arrival (ETA) of aircraft in the terminal area via the implementation of a Random Forest regression model. The model uses data fused from a number of sources (flight track, weather, flight plan information, etc.) and provides predictions for the remaining flight time for aircraft landing at Dallas/Fort Worth (DFW) International Airport. The predictions are made when the aircraft is at a distance of 200-miles from the airport. The results show that the model is able to predict estimated time of arrival to within ± 5 min for 90% of the flights in the test data with the mean absolute error being lower at 145 seconds. This paper covers the entire pipeline of data collection, preprocessing, setup and training of the ML model, and the results obtained for DFW.

Machine learning↗

The DECADE cosmic shear project I: A new weak lensing shape catalog of 107 million galaxies

We present the Dark Energy Camera All Data Everywhere (DECADE) weak lensing dataset: a catalog of 107 million galaxies observed by the Dark Energy Camera (DECam) in the northern Galactic cap. This catalog was assembled from public DECam data including survey and standard observing programs. These data were consistently processed with the Dark Energy Survey Data Management pipeline as part of the DECADE campaign and serve as the basis of the DECam Local Volume Exploration survey (DELVE) Early Data Release 3 (EDR3). We apply the Metacalibration measurement algorithm to generate and calibrate galaxy shapes. After cuts, the resulting cosmology-ready galaxy shape catalog covers a region of $5,\!412 \,\,{\rm deg}^2$ with an effective number density of $4.59\,\, {\rm arcmin}^{-2}$. The coadd images used to derive this data have a median limiting magnitude of $r = 23.6$, $i = 23.2$, and $z = 22.6$, estimated at ${\rm S/N} = 10$ in a 2 arcsecond aperture. We present a suite of detailed studies to characterize the catalog, measure any residual systematic biases, and verify that the catalog is suitable for cosmology analyses. In parallel, we build an image simulation pipeline to characterize the remaining multiplicative shear bias in this catalog, which we measure to be $m = (-2.454 \pm 0.124) \times10^{-2}$ for the full sample. Despite the significantly inhomogeneous nature of the data set, due to it being an amalgamation of various observing programs, we find the resulting catalog has sufficient quality to yield competitive cosmological constraints.

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND ↗

Scalable GPS Data Logging To Support Advanced Fleet Analysis

This highlight details the key takeaways from a project that utilized NLR's Fleet Research, Energy Data, and Insights (FleetREDI) data analysis pipeline. National Laboratory of the Rockies researchers developed and demonstrated low-cost, open-source Arduino data loggers with 3D-printed cases that are compatible with global navigational systems and built with components available ubiquitously worldwide, enabling cost-effective collection and analysis of fleet operational data. Validated on an overseas transit bus fleet, NLR analysis showed that, with sufficient charging opportunities, 90% of observed duty cycles could be accomplished by electric buses with no modifications to operations.

33 ADVANCED PROPULSION SYSTEMS↗

Estimating Electrification Potential for Class 8 Regional-Haul Trucks

This one-page highlight details the key takeaways from a project that utilized NREL's Fleet Research, Energy Data, and Insights (FleetREDI) data analysis pipeline. As part of the North American Council for Freight Efficiency's (NACFE's) Run on Less Depot data workshop, NREL sought to understand how Tesla semi-trucks would perform in real-world regional haul applications. Analysis reveals that the modeled Tesla trucks, with an average efficiency of 1.78 kWh/mi, struggle to achieve full operational coverage using current battery and charging configurations assuming operations remain unchanged. However, in an extreme case where ubiquitous charging exists, 100% EV coverage is possible for the given drive cycles. These findings highlight the trade-off between battery size and charge rate in electrification potential and emphasize the necessity for advancements in charging infrastructure to enable electric trucks for regional haul operations.

ADVANCED PROPULSION SYSTEMS↗

DEVELOP’s Approach to Experiential Learning

The NASA DEVELOP Program addresses environmental decision making needs and geoscience workforce development through 10-week feasibility studies that apply Earth observations to environmental issues at hand. The program builds capacity to use geospatial information in both its participants (students, recent graduates, early career professionals, and transitioning career professionals) and partner organizations (federal agencies, state & local governments, non-profits, and private industry). This is accomplished through a structured project execution model that provides opportunities for participants to have autonomy, learn “on the job,” and gain new skillsets for working with remote sensing data. A pipeline of leadership positions enhances opportunities for individuals engaged in the program to get hands-on experience conducting data analyses, communicating their work, leading technical projects, and building their knowledge bank of Earth-observing satellite capabilities. These skillsets and knowledge are then transferred to partners through the projects. This panel contribution will introduce the DEVELOP model, highlight experiences of participants, and share the program’s insights into good practices for effective experiential learning.

Capacity Building↗

The FRB-searching Pipeline of the Tianlai Cylinder Pathfinder Array

This paper presents the design, calibration, and survey strategy of the Fast Radio Burst (FRB) digital backend and its real-time data processing pipeline employed in the Tianlai Cylinder Pathfinder Array. The array, consisting of three parallel cylindrical reflectors and equipped with 96 dual-polarization feeds, is a radio interferometer array designed for conducting drift scans of the northern celestial semi-sphere. The FRB digital backend enables the formation of 96 digital beams, effectively covering an area of approximately 40 square degrees with the 3 dB beam. Our pipeline demonstrates the capability to conduct an automatic search of FRBs, detecting at quasi-real-time and classifying FRB candidates automatically. The current FRB searching pipeline has an overall recall rate of 88%. During the commissioning phase, we successfully detected signals emitted by four well-known pulsars: PSR B0329+54, B2021+51, B0823+26, and B2020+28. We report the first discovery of an FRB by our array, designated as FRB 20220414A. We also investigate the optimal arrangement for the digitally formed beams to achieve maximum detection rate by numerical simulation.

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND ↗

RTN-117: Image Calibration and Instrument Signal Removal for the First Year of the LSST

The NSF-DOE Vera C. Rubin Observatory Legacy Survey of Space and Time (LSST) requires calibration products that provide uniform, stable, and accurate photometric and astrometric performance across the 3.2 gigapixel focal plane and throughout the 10-year survey. This paper details the algorithms and workflows used to produce instrument calibrations and remove instrumental artifacts for Data Preview 2 (DP2)---the first end-to-end processing demonstration using on-sky data with the LSST Camera (LSSTCam). We describe the verification, acceptance, and certification framework used to assess calibration quality and quantify residual systematics. We show baseline metrics on calibrated science images to evaluate the robustness of the calibration and instrument signature removal (ISR) data processing pipelines for DP2. Finally, we summarize the known limitations observed in DP2 production and outline expected algorithmic improvements for the first public LSST data release (Data Release~1, DR1).

79 ASTRONOMY AND ASTROPHYSICS↗

Joint US-Japan Observations with the Infrared Space Observatory (ISO): Deep Surveys and Observations of High-Z Objects

Several important milestones were passed during the past year of our ISO observing program: (1) Our first ISO data were successfully obtained. ISOCAM data were taken for our primary deep field target in the 'Lockman Hole'. Thirteen hours of integration (taken over 4 contiguous orbits) were obtained in the LW2 filter of a 3 ft x 3 ft region centered on the position of minimum HI column density in the Lockman Hole. The data were obtained in microscanning mode. This is the deepest integration attempted to date (by almost a factor of 4 in time) with ISOCAM. (2) The deep survey data obtained for the Lockman Hole were received by the Japanese P.I. (Yoshi Taniguchi) in early December, 1996 (following release of the improved pipeline formatted data from Vilspa), and a copy was forwarded to Hawaii shortly thereafter. These data were processed independently by the Japan and Hawaii groups during the latter part of December 1996, and early January, 1997. The Hawaii group made use of the U.S. ISO data center at IPAC/Caltech in Pasadena to carry out their data reduction, while the Japanese group used a copy of the ISOCAM data analysis package made available to them through an agreement with the head of the ISOCAM team, Catherine Cesarsky. (3) Results of our LW2 Deep Survey in the Lockman Hole were first reported at the ISO Workshop "Taking ISO to the Limits: Exploring the Faintest Sources in the Infrared" held at the ISO Science Operations Center in Villafranca, Spain (VILSPA) on 3-4 February, 1997. Yoshi Taniguchi gave an invited presentation summarizing the results of the U.S.-Japan team, and Dave Sanders gave an invited talk summarizing the results of the Workshop at the conclusion of the two day meeting. The text of the talks by Taniguchi and Sanders are included in the printed Workshop Proceedings, and are published in full on the Web. By several independent accounts, the U.S.-Japan Deep Survey results were one of the highlights of the Workshop; these data showed conclusively that the ISOCAM S/N continues to decrease as the square root of time for periods as long as 13 hours.

Sanders, David B.↗