Search NASA⌕ Search

SEARCH · Search NASA

Results for “data pipeline”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 253 records · Page 14

Cold Weather Impacts on Electric School Bus Performance in Aurora, Colorado

This brief highlight details the key takeaways from a project that utilized NLR's Fleet Research, Energy Data, and Insights (FleetREDI) data analysis pipeline related to electric school bus (ESB) operation. ESBs using battery energy as their primary heating source have a higher energy consumption rate in cold weather, which fleet managers can account for when planning ESB purchases and making dispatching and charging decisions. Researchers found that electric school buses operate 2-5 times more efficiently than conventional buses, on average. Cold weather can double electric school bus energy demands, but strategies such as thermal pre-conditioning significantly reduce this effect. Understanding these impacts can help fleets plan charging, dispatching, and purchase decisions.

33 ADVANCED PROPULSION SYSTEMS↗

Exploration of Real Time Inference for MI-RR Deblending on GPU/TPU Systems

The Fermilab Main Injector (MI) and Recycler Ring (RR) share a common beam loss monitor (BLM) system, making loss events difficult to attribute to their source machine when beam is present in both simultaneously. The Real-time Edge AI for Distributed Systems (READS) project addresses this by deblending BLM readings in real time using machine learning (ML). The current FPGA based implementation meets the sub-3 ms latency requirement but carries a resource intensive hls4ml development cycle, motivating exploration of GPU based deployment. This paper characterizes inference latency on an NVIDIA Jetson Orin Nano and introduces a packet organization scheme for assembling synchronized event frames from seven distributed BLM DAQ streams. Using a Python based DAQ simulation with injected timing jitter in place of unavailable live beam data, the pipeline achieved an average end to end latency of 0.456 ms (σ = 0.122 ms) across 167,000 test frames, comfortably meeting the timing constraint. Early outliers were attributed to TensorRT warm-up rather than steady state limitations, suggesting GPU based inference is a viable alternative to the existing FPGA implementation.

Yu, Kellen [Fermilab; Cornell U.]↗

Correlative visualization techniques for multidimensional data

Critical to the understanding of data is the ability to provide pictorial or visual representation of those data, particularly in support of correlative data analysis. Despite the advancement of visualization techniques for scientific data over the last several years, there are still significant problems in bringing today's hardware and software technology into the hands of the typical scientist. For example, there are other computer science domains outside of computer graphics that are required to make visualization effective such as data management. Well-defined, flexible mechanisms for data access and management must be combined with rendering algorithms, data transformation, etc. to form a generic visualization pipeline. A generalized approach to data visualization is critical for the correlative analysis of distinct, complex, multidimensional data sets in the space and Earth sciences. Different classes of data representation techniques must be used within such a framework, which can range from simple, static two- and three-dimensional line plots to animation, surface rendering, and volumetric imaging. Static examples of actual data analyses will illustrate the importance of an effective pipeline in data visualization system.

Treinish, Lloyd A.↗

Early-type galaxies: Automated reduction and analysis of ROSAT PSPC data

Preliminary results of early-type galaxies that will be part of a galaxy catalog to be derived from the complete Rosat data base are presented. The stored data were reduced and analyzed by an automatic pipeline. This pipeline is based on a command language scrip. The important features of the pipeline include new data time screening in order to maximize the signal to noise ratio of faint point-like sources, source detection via a wavelet algorithm, and the identification of sources with objects from existing catalogs. The pipeline outputs include reduced images, contour maps, surface brightness profiles, spectra, color and hardness ratios.

Mackie, G.↗

The Mars 2020 Ground Data System Architecture

The Mars 2020 Mission’s primary objective is to collect 20 geographically unique samples during its prime mission of one and a quarter Martian years, or just over 2 Earth years. Mission planners determined the project needed to develop a system that would enable the operations team to analyze engineering and science data, make science decisions, select viable rover targets at a millimeter resolution and validate an uplink bundle for a car sized rover with more complex science instruments than any previous Mars surface mission. All this had to be done within a five hour time frame. Doing this with a small team would be a challenge, but this had to be accomplished by a large team of engineers and scientists located across North America and Europe. Achieving this level of operational efficiency was unheard of in the prime mission. In addition, the mission had another set of requirements that had nothing to do with surface operations; the Mars 2020 Ground Data System (GDS) was also expected to comply with a new set of security requirements to keep up with the ever changing cybersecurity landscape. The Mars 2020 Ground Data System (GDS) is a re-architected version of the Mars Science Laboratory GDS. The primary goal was to integrate the lessons learned from previous Mars surface missions, accommodate a set of new requirements and capabilities required to ensure mission success, and comply with a new set of cybersecurity controls. The new architecture includes several unique qualities including a data lake, language-agnostic system-wide event-based operations, containerization, automated deployment, network segmentation, infrastructure-as-code, API-driven interfaces, and the first Mars surface GDS to operate primarily in the cloud. The new architecture enabled greater access to the system’s data, tighter integration with the operations team, and a higher level of traceability. The availability of the data also enabled a new set of capabilities previously not possible on surface missions. These new capabilities include an autonomous data to information, pipeline for downlink analysis, horizontal scaling of science data processing capabilities, autonomous round trip data tracking of science and engineering data, integration of flight system state into the tactical planning cycle, high fidelity targeting utilizing kinematic data, and hierarchical image and 3d meshes data representations. This paper will introduce the requirements for the Mars 2020 Mission, the heritage architecture, and the rationale for the changes to achieve the new architecture. The paper will continue to describe the fundamental changes made to the GDS architecture, how these changes enabled a more tightly integrated GDS, and the new capabilities that were enabled by the new architecture. The paper will conclude with the lessons learned from the process of rearchitecting a heritage GDS system and from the first 200 days of operations supporting over 800 users from around the world.

Lopez-Roig, Reynaldo↗

Development and Demonstration of a Digital NDE Pipeline for Streamlined Analysis of Ultrasonic Data

Currently when nondestructive evaluation (NDE) is performed on composite structures, the results, although recorded digitally are often manually interpreted and indicated (drawn) on the part being inspected by hand. Following this, a determination must be made on how to disposition that part. This decision could be based on engineering guidelines and best practices, rule-of-thumb, expert opinion or finite-element analysis of the part with some approximate representation of the damage. In order to streamline this process for ultrasonic inspection an effort was undertaken as part of NASA Advanced Composites Project to develop the Digital NDE Pipeline. The Digital NDE Pipeline is an integrated tool suite and associated framework that streamlines the inspection and defect disposition process through model-assisted inspection optimization, automated defect analysis of NDE data, and mapping of NDE data into finite element analysis software. This paper will provide an overview of the Digital NDE Pipeline and provide details of the individual tools developed along with the results of applying these tools to a demonstration test case.

Composites↗

Spaceflight Environmental-Telemetry Data for Biological Science

There is a critical need for better access and visualization of spaceflight environmental telemetry and mission hardware data from sensors including relative humidity, carbon dioxide, oxygen, radiation, airflow, temperature, acceleration, and acoustics. Under the stewardship of the Ames Life Sciences Data Archive (ALSDA) and GeneLab, an effort is underway to consolidate, normalize and provide accessibility of archived mission environmental data and hardware information, with the purpose of providing important context to biological data. This effort is necessary to provide scientific context of its impact upon biological and biomedical data from spaceflight missions and experiments (genomic, metagenomic, gene expression, proteomic, metabolomic, physiological, phenomics, behavioral; tabular, imaging, video). Environmental spaceflight data is derived from dozens of sources, with various formats, and in the past year a pipeline is in development to collect, curate and present this data efficiently. In the upcoming year, a new Data Visualization Portal will utilize the standardized pipeline data to provide easy user access to compare parameters and environmental conditions between missions, locations, subjects, and durations. Environmental and hardware data enables broad accessibility and analytics, without the need for advanced data informatic expertise. Familiarity with the capabilities and limitations of a variety of existing hardware/tools is a strength that could be applied to creation of improved hardware for future ecosystems on the Moon and Mars. The intention is to make biological and environmental telemetry data maximally open-access and FAIR (findable, accessible, interoperable, reusable) for data mining-informatic approaches to support knowledge discovery necessary for low Earth orbit, cis-Lunar, Mars transit, and Mars surface missions.

Danielle K. Lopez↗

Updates in Developing a Prototype Science Pipeline and Full-Volume, Global Hyperspectral Synthetic Data Sets for NASA’s Earth System Observatory’s Upcoming Surface, Biology and Geology Mission

The Surface Biology and Geology (SBG) mission recently passed mission confirmation review and has entered phase A – design and development. SBG will acquire high resolution solar-reflected spectroscopy and thermal infrared observations at a data rate of ~2.5 TB/day and generate products at ~40 TB/day. Given that the per-day volume is greater than NASA’s total extant airborne hyperspectral data collection, collecting, processing, disseminating, and exploiting the SBG data present new challenges. To meet these challenges, we have developed a prototype science pipeline and a full-volume global hyperspectral synthetic data set to help prepare for SBG’s flight (see poster GC42D-0730). Our science pipeline is based on the science processing technology developed for NASA’s Kepler and TESS planet-hunting missions. The pipeline infrastructure, Ziggy, provides a scalable architecture for robust, repeatable, and replicable science and application products that can be run on a range of systems from a laptop to the cloud or a supercomputer. Ziggy is compliant with NASA Procedural Requirement (NPR) 7150.2C, is at a technical readiness level (TRL) of 7 and has been released to github.com/nasa/ziggy. We integrated Ziggy with EO-1/Hyperion workflows to build a prototype pipeline and ingested the 17-year mission archive that provides globally sampled visible through shortwave infrared spectra that are representative of SBG data types and volumes. We fully implemented the first stage and processed the entire 55 TB Hyperion data set from the raw data (Level 0) to top-of-the-atmosphere radiance (Level 1R). We are currently evaluating the ISOFIT atmospheric correction module to convert the L1R data to surface reflectance (Level 2) before reprocessing the full data set to L2. Crosschecks are being performed with RadCalNet as well as with coincident observations by AVIRIS. We are also investigating modern methods for georectifying the Hyperion scenes. Finally, we describe an analysis of the cost to conduct forward processing and reprocessing campaigns for SBG on HECC with dedicated compute and storage resources using the resurrected Hyperion pipeline as a proxy for full-volume SBG data. The analysis demonstrates that SBG L0 data can be processed to L2 on HECC with full reprocessing campaigns every two years for ~$2.6M over a 7-year lifespan. Moreover, 69% of the system capacity would be available for other activities, possibly enabling future open-source science activities, including algorithm development, L3+ processing, .etc.

ESD↗

Virtual Inspection of Advanced Manufacturing via Process-Scale Digital Twins (Abbreviated Report)

Inspection and certification comprise the most significant bottlenecks in advanced manufacturing for NNSA applications, often requiring far more time and resources than the fabrication of the parts themselves. Traditional methods, such as manual review and X-ray computed tomography, are not only slow and costly, but also struggle to provide a clear connection between manufacturing instructions and the final performance of critical components. This gap limits both the agility and assurance needed to support the modernization and safety of the United States nuclear stockpile. In response, our Strategic Initiative established a digital twin framework that integrates realtime process monitoring, automated data analysis, and immersive virtual reality collaboration into a unified inspection pipeline. By leveraging data from sensors, machine instructions, and imaging, we created high-fidelity virtual models of manufactured parts that could be rapidly analyzed and certified. This approach was first demonstrated with Direct Ink Write, and then extended to other manufacturing settings, including conventional (or “subtractive”) manufacturing and to predict the end of life performance of parts per the aging and lifetimes programs. The result is a transformational capability: inspection times have been reduced by a factor of 120,000 without loss of accuracy and while simultaneously improving traceability and confidence in part quality. This framework not only streamlines certification for critical applications, but also positions the national security enterprise to respond more flexibly to emerging challenges, supporting agile manufacturing and digital engineering practices across a broad range of mission-relevant domains.

42 ENGINEERING↗

ISO Key Project: Exploring The Full Range of Quasar/AGN Properties

While most of the work on this program has been completed, as previously reported, the portion of the program dealing with the sub topic of ISO LWS data analysis and reduction for the LWS Extragalactic Science Team and its leader, Dr. Howard Smith, is still active. This program in fact continues to generate results, and newly available computer modeling has extended the value of the datasets, As a result the team has requested and been granted an obtained a no-cost extension to this program, through December 31, 2003. The essence of the proposal is to perform ISO spectroscopic studies, including data analysis and modeling, of star formation regions using an ensemble of archival space-based data from the Infrared Space Observatory's Long Wavelength Spectrometer and Short Wavelength Spectrometer, but including as well some other spectroscopic data bases. Four kinds of regions are considered in the studies: (1) disks around more evolved objects; (2) young, low or high mass pre-main sequence stars in star formation regions; (3) star formation in external, bright IR galaxies; and (4) the galactic center. One prime focus of the program is the OH lines in the far infrared. The program has the following goals: (1) refine the data analysis of ISO observations, to obtain deeper and better SNR results on selected sources. The ISO data itself underwent "pipeline 10" reductions in early 2001, and additional "hands-on data reduction packages" were supplied by the ISO teams in 2001. The Fabry-Perot database in particularly sensitive to noise can slight calibration errors. (2) model the atomic and molecular line shapes, in particular the OH lines, using revised Monte-Carlo techniques developed by the SWAS team at the Center for Astrophysics; (3) attend scientific meetings and workshops; (4) do E&PO activities related to infrared astrophysics and/or spectroscopy.

Wilkes, Belinda↗

ISO Key Project: Exploring the Full Range of Quasar/Agn Properties

While most of the work on this program has been completed, as previously reported, the portion of the program dealing with the subtopic of ISO LWS data analysis and reduction for the LWS Extragalactic Science Team and its leader, Dr. Howard Smith, is still active. This program in fact continues to generate results, and newly available computer modeling has extended the value of the datasets. As a result the team requests a one-year no-cost extension to this program, through 31 December 2004. The essence of the proposal is to perform ISO spectroscopic studies, including data analysis and modeling, of star-formation regions using an ensemble of archival space-based data from the Infrared Space Observatory's Long Wavelength Spectrometer and Short Wavelength Spectrometer, but including as well some other spectroscopic databases. Four kinds of regions are considered in the studies: (1) disks around more evolved objects; (2) young, low or high mass pre-main sequence stars in star-formation regions; (3) star formation in external, bright IR galaxies; and (4) the galactic center. One prime focus of the program is the OH lines in the far infrared. The program has the following goals: 1) Refine the data analysis of ISO observations to obtain deeper and better SNR results on selected sources. The ISO data itself underwent 'pipeline 10' reductions in early 2001, and additional 'hands-on data reduction packages' were supplied by the ISO teams in 2001. The Fabry-Perot database is particularly sensitive to noise and slight calibration errors; 2) Model the atomic and molecular line shapes, in particular the OH lines, using revised Monte-Carlo techniques developed by the SWAS team at the Center for Astrophysics; 3) Attend scientific meetings and workshops; 4) Perform E&PO activities related to infrared astrophysics and/or spectroscopy.

Wilkes, Belinda↗

Explore Full Range of QSO/AGN Properties

The goal of the proposal is to perform ISO spectroscopic studies, including data analysis and modeling, of star formation regions using an ensemble of archival space-based data from the Infrared Space Observatory s Long Wavelength Spectrometer and Short Wavelength Spectrometer, but including as well some other spectroscopic databases. Four kinds of regions are considered in the studies: (1) disks around more evolved objects; (2) young, low or high mass pre-main sequence stars in star formation regions; (3) star formation in external, bright IR galaxies; and (4) the galactic center. One prime focus of the program is the OH lines in the far infrared. The program had the following goals: 1) Refine the data analysis of IS0 observations to obtain deeper and better SNR results on selected sources. The IS0 data itself underwent "pipeline 10" reductions in early 2001, and additional "hands-on data reduction packages" were supplied by the IS0 teams in 2001. The Fabry-Perot database is particularly sensitive to noise and slight calibration errors. 2) Model the atomic and molecular line shapes, in particular the OH lines, using revised monte- carlo techniques developed by the SWAS team at the Center for Astrophysics; 3) Attend scientific meetings and workshops; 4) Do E&PO activities related to infrared astrophysics and/or spectroscopy.

Oliversen, Ronald↗

TESS Data Release Notes:Reprocessing of Sectors 14–19, DR30 & DR33

TESS data release 30 (DR30) provides reprocessed data products of Sector 14 to 19. The updated data products were generated using version 4.0 of the science processing pipeline and conform to the final set of data anomaly flags defined over the last two years of TESS data analysis and pipeline development. Data release 33 (DR33) corresponds to a multisector search for transiting planets in the same reprocessed data. A detailed description of the changes in the data products in DR30 and DR33 is discussed in§2, and a brief list of changes is summarized here: The timestamps for 2 minute cadence and FFI data are more accurate. The differences between reprocessed data and previous data releases are less than 2.0 seconds in all cases. Photometric apertures were increased in size for targets with T mag<11. Three new Data Anomaly Flags were added to mitigate the effects of scattered light:–Cadences with strong scattered light signals or saturation effects that corrupt the calibration data are flagged and removed from analysis (bit 15, value 16384, “Bad Calibration Exclude”).–Scattered light data anomaly flags are customized for each target, and flagged automatically based on the local background level (bit 13, value 4096, ”Scattered light flag”).–Cadences with insufficient targets to derive cotrending basis vectors are flagged and the PDCSAP FLUX light curves are set to NULL at these times (bit 16, value 32768, “Insufficient Targets for Error Correction Exclude”). The planet search of the reprocessed light curves produced a different set of TCEs from the original processed data. Although there is a high degree of overlap between the original and reprocessed data (∼83% of targets produced TCEs in common), new TCEs were produced in DR30 and not every TCE from previous data releases was recovered. The same is true of the multisector search results from DR33 compared to DR28

TESS'↗

A Performant, Scalable Processing Pipeline for High‐Quality and FAIR Environmental Sensor Data

High-resolution environmental monitoring is necessary to record, understand, and predict biogeochemical and ecological changes particularly in coastal systems but brings significant challenges in processing and making rapidly available the resulting data. The COMPASS-FME project established a network of coastal observational sites across the Chesapeake Bay and western Lake Erie regions extensively instrumented with soil, vegetation, and weather sensors logging data every 15 min. Our data processing framework, written in R and completely open source, prioritizes rapid model-experiment iteration and makes biogeochemical data rapidly available for quality assurance/quality control, analysis, and model ingestion. This pipeline is distinguished by a standardized and modular approach to data curation, extensive metadata and documentation, and its high performance. These attributes combine to make biogeochemical data rapidly accessible across COMPASS-FME and the broader community. Flexible, powerful, and reproducible approaches to handling high-volume environmental data are crucial for accelerating biogeosciences research.

Pennington, Stephanie C. [Pacific Northwest Nation↗

DoCeph: DPU-Offloaded Messaging in Ceph for Reduced Host CPU Utilization

Ceph is a widely used distributed object store, but its messenger layer imposes substantial CPU overhead on the host. To address this limitation, we propose DoCeph, a DPU-offloaded storage architecture for Ceph that disaggregates the system by offloading the communication-intensive messaging component to the DPU while retaining the storage backend on the host. The DPU efficiently manages communication, using lightweight RPC for metadata operations and DMA for data transfer. Moreover, DoCeph introduces a pipelining technique that overlaps data transmission with buffer preparation, mitigating hardware-imposed transfer size limitations. We implemented DoCeph on a Ceph cluster with NVIDIA BlueField-3 DPUs. Evaluation results indicate that DoCeph cuts host CPU usage by up to 92% while sustaining stable throughput and providing larger performance benefits for object writes over 1 MB.

Park, Kuri [Sogang University]↗

Infrared Spectroscopy of Star Formation in Galactic and Extragalactic Regions

In this program we proposed to perform a series of spectroscopic studies, including data analysis and modeling, of star formation regions using an ensemble of archival space-based data from the Infrared Space Observatory's Long Wavelength Spectrometer and Short Wavelength Spectrometer, and to take advantage of other spectroscopic databases including the first results from SIRTF. Our emphasis has been on star formation in external, bright IR galaxies, but other areas of research have included young, low or high mass pre-main sequence stars in star formation regions, and the galactic center. The OH lines in the far infrared were proposed as one key focus of this inquiry, because the Principal Investigator (H. Smith) had a full set of OH IR lines from IS0 observations. It was planned that during the proposed 2-1/2 year timeframe of the proposal other data (including perhaps from SIRTF) would become available, and we intended to be responsive to these and other such spectroscopic data sets. The program has the following goals: 1) Refine the data analysis of IS0 observations to obtain deeper and better SNR results on selected sources. The IS0 data itself underwent pipeline 10 reductions in early 2001, and the more 'hands-on data reduction packages' have been released. The IS0 Fabry-Perot database is particularly sensitive to noise and can have slight calibration errors, and improvements are anticipated. We plan to build on these deep analysis tools and contribute to their development. Model the atomic and molecular line shapes, in particular the OH lines, using revised montecarlo techniques developed by the Submillimeter Wave Astronomy Satellite (SWAS) team at the Center for Astrophysics. 2) 3) Use newly acquired space-based SIRTF or SOFIA spectroscopic data as they become available, and contribute to these observing programs as appropriate. 4) Attend scientific meetings and workshops. 5) E&PO activities, especially as related to infrared astrophysics and/or spectroscopy.

Smith, Howard A.↗

A Detection Pipeline for Galactic Binaries in LISA Data

The Galaxy is suspected to contain hundreds of millions of binary white dwarf systems, a large fraction of which will have sufficiently small orbital period to emit gravitational radiation in band for space-based gravitational wave detectors such as the Laser Interferometer Space Antenna (LISA). LISA's main science goal is the detection of cosmological events (supermassive black hole mergers) etc.) however the gravitational signal from the galaxy will be the dominant contribution to the data - including instrumental noise - over approximately two decades in frequency. The catalogue of detectable binary systems will serve as an unparalleled means of studying the Galaxy. Furthermore, to maximize the scientific return from the mission, the data must be "cleansed" of the galactic foreground. We will present an algorithm that can accurately resolve and subtract greater than or equal to 10000 of these sources from simulated data supplied by the Mock LISA Data Challenge Task Force. Using the time evolution of the gravitational wave frequency, we will reconstruct the position of the recovered binaries and show how LISA will sample the entire compact binary population in the Galaxy.

Littenberg, Tyson B.↗