Search NASA⌕ Search

SEARCH · Search NASA

Results for “data pipeline”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 199 records · Page 11

Electrification Analysis: All Aboard America!

This one-page highlight details the key takeaways from a project that utilized NREL's Fleet Research, Energy Data, and Insights (FleetREDI) data analysis pipeline, an electrification analysis for the Bustang motorcoach fleet operated by All Aboard America! Holdings Inc. (AAA). NREL installed logging devices and collected operational data on nine 40-foot Bustang motorcoaches operating on fixed routes from May 2022 through August 2022. The analysis determined that partial fleet electrification may be feasible with electrified motorcoach options currently on the market. While this fleet faces significant challenges to electrification given current market options due to demanding range requirements and relatively limited charging opportunities, vehicles operating on the shorter, lower-grade routes along the I-25 corridor show more immediately available electrification potential. Increases in available battery capacity and the availability of fast-charging locations along I-70 routes are likely critical for electrification of the full fleet.

AAA↗

An XMM-Newton Monitoring Campaign of the Accretion Flow in IGRJ16318-4848

This grant is associated to a successful XMM-Newton-AO3 observational proposal to monitor the spectrum of the X-ray loud component of the recently discovered binary system IGR J16138-4848, to study the conditions of the accretion flows (and their evolution) in binary system. All four EPIC-PN and MOS observations of the target have now been performed (the last one of the 4, only 3 months ago). The four observations were logarithmically spaced, so to cover timescales from days to months. Data from all four pointings have now been reduced, using the XMM-Newton data reduction pipeline, and spectra and lightcurves from the target have been extracted. For the first three observations we have already performed the observation-by-observation data analysis, by fitting the single EPIC spectra with spectral models that include an intrinsic continuum power law (reduced at low energy by neutral absorption), a 6.4 keV iron emission line (detected in all spectra with varying intensity) and a Compton-reflection component. A Compton reflection component is also detected in all spectra, although at lower significance. The analysis of the fourth and last observation of our monitoring campaign has just recently begun. Next, we will (1) stack together the four observations of IGR J16138-4848, to obtain high-accuracy estimates of the average spectral parameters of this object; and then (2) proceed to the time-evolving analysis, of the three spectral parameters: (a) Gamma (the slope of the intrinsic continuum), (b) W(FeK), the equivalent width of the 6.4 keV Iron emission line, and (c) R, the relative amount of Compton reflection. Through this time-resolved spectroscopic analysis we hope to constrain (a) the physical state of the accreting matter and its relation with the X-ray output, and (b) the evolution of the accretion flow geometry, distribution and covering factor.

Mushotzky, Richard↗

Enabling Earth Science Through Cloud Computing

Cloud Computing holds tremendous potential for missions across the National Aeronautics and Space Administration. Several flight missions are already benefiting from an investment in cloud computing for mission critical pipelines and services through faster processing time, higher availability, and drastically lower costs available on cloud systems. However, these processes do not currently extend to general scientific algorithms relevant to earth science missions. The members of the Airborne Cloud Computing Environment task at the Jet Propulsion Laboratory have worked closely with the Carbon in Arctic Reservoirs Vulnerability Experiment (CARVE) mission to integrate cloud computing into their science data processing pipeline. This paper details the efforts involved in deploying a science data system for the CARVE mission, evaluating and integrating cloud computing solutions with the system and porting their science algorithms for execution in a cloud environment.

science data system↗

The PICWidget

The Plug-in Image Component Widget (PICWidget) is a software component for building digital imaging applications. The component is part of a methodology described in GIS Methodology for Planning Planetary-Rover Operations (NPO-41812), which appears elsewhere in this issue of NASA Tech Briefs. Planetary rover missions return a large number and wide variety of image data products that vary in complexity in many ways. Supported by a powerful, flexible image-data-processing pipeline, the PICWidget can process and render many types of imagery, including (but not limited to) thumbnail, subframed, downsampled, stereoscopic, and mosaic images; images coregistred with orbital data; and synthetic red/green/blue images. The PICWidget is capable of efficiently rendering images from data representing many more pixels than are available at a computer workstation where the images are to be displayed. The PICWidget is implemented as an Eclipse plug-in using the Standard Widget Toolkit, which provides a straightforward interface for re-use of the PICWidget in any number of application programs built upon the Eclipse application framework. Because the PICWidget is tile-based and performs aggressive tile caching, it has flexibility to perform faster or slower, depending whether more or less memory is available.

Norris, Jeffrey↗

Photometer Performance Assessment in Kepler Science Data Processing

This paper describes the algorithms of the Photometer Performance Assessment (PPA) software component in the science data processing pipeline of the Kepler mission. The PPA performs two tasks: One is to analyze the health and performance of the Kepler photometer based on the long cadence science data down-linked via Ka band approximately every 30 days. The second is to determine the attitude of the Kepler spacecraft with high precision at each long cadence. The PPA component is demonstrated to work effectively with the Kepler flight data.

Li, Jie↗

Data Validation in the Kepler Science Operations Center Pipeline

We present an overview of the Data Validation (DV) software component and its context within the Kepler ScienceOperations Center (SOC) pipeline and overall Kepler Science mission. The SOC pipeline performs a transiting planetsearch on the corrected light curves for over 150,000 targets across the focal plane array. We discuss the DV strategy forautomated validation of Threshold Crossing Events (TCEs) generated in the transiting planet search. For each TCE, atransiting planet model is fitted to the target light curve. A multiple planet search is conducted by repeating the transitingplanet search on the residual light curve after the model flux has been removed; if an additional detection occurs, aplanet model is fitted to the new TCE. A suite of automated tests are performed after all planet candidates have beenidentified. We describe a centroid motion test to determine the significance of the motion of the target photocenterduring transit and to estimate the coordinates of the transit source within the photometric aperture; a series of eclipsingbinary discrimination tests on the parameters of the planet model fits to all transits and the sequences of odd and eventransits; and a statistical bootstrap to assess the likelihood that the TCE would have been generated purely by chancegiven the target light curve with all transits removed.

photometry↗

Data Validation in the Kepler Science Operations Center Pipeline

We present an overview of the Data Validation (DV) software component and its context within the Kepler Science Operations Center (SOC) pipeline and overall Kepler Science mission. The SOC pipeline performs a transiting planet search on the corrected light curves for over 150,000 targets across the focal plane array. We discuss the DV strategy for automated validation of Threshold Crossing Events (TCEs) generated in the transiting planet search. For each TCE, a transiting planet model is fitted to the target light curve. A multiple planet search is conducted by repeating the transiting planet search on the residual light curve after the model flux has been removed; if an additional detection occurs, a planet model is fitted to the new TCE. A suite of automated tests are performed after all planet candidates have been identified. We describe a centroid motion test to determine the significance of the motion of the target photocenter during transit and to estimate the coordinates of the transit source within the photometric aperture; a series of eclipsing binary discrimination tests on the parameters of the planet model fits to all transits and the sequences of odd and even transits; and a statistical bootstrap to assess the likelihood that the TCE would have been generated purely by chance given the target light curve with all transits removed. Keywords: photometry, data validation, Kepler, Earth-size planets

Wu, Hayley↗

Trust Not Verify? The Critical Need for Data Curation Standards in Materials Informatics

The importance of data curation has been recognized in multiple areas of research; however, the discussion of this important issue is only beginning to emerge in materials science. In this Perspective, we highlight the benefits of using the standardized data curation protocols in materials science and discuss current gaps in accurate and reproducible data reporting using case studies drawn from high-impact materials science papers and well-known databases such as the Crystallography Open Database (COD) and the Cambridge Structural Database (CSD). We argue that both experimental and computational materials scientists need to embrace a culture of rigorous data curation as part of modern research data management. We propose a sample data curation pipeline for materials chemistry and illustrate its use by creating two new materials chemistry databases. Here, we hope that this perspective will serve to catalyze further discussion and promote the continuous development of rigorous data curation practices within the materials science research community. We posit that adherence to best practices of data curation will promote and enhance the reliability, reproducibility, and integrity of materials research and enable the development of reliable AI and machine learning models that critically depend on the use of quality data.

Chemical structure↗

Growth Curve Parameterization of Metabolic Activity of Yeast Cells for BioSentinel

The goal of the BioSentinel small satellite payload is to measure the effect of deep space radiation on the growth and metabolic activity of yeast cells. Raw test data is generated by fluidics cards containing yeast cells rehydrated at different periods, with metabolic activity measured by the reduction of alamarBlue. Each card well has a sensor array that measures the amount of red, green, and infrared light transmitted through the yeast culture. This illumination data is then converted to absorbance values, which are further converted into concentrations. The ultimate objective is to convert these concentrations into biologically-relevant metrics that can be compared against one another to determine changes due to differential radiation exposure. Beginning with IR absorbance data (corresponding to cell density) from ground studies, three parameters from a sigmoidal growth curve were extracted and analyzed: 𝜆 (lag phase), 𝜇 (max growth rate), and A (max cell growth). The data was fit to the Gompertz model of microbial growth using non-linear regression (Minitab), as the fit error was reduced compared to the simpler logistic growth curve. Graphs showed that the data contained a discrepancy (drift) in the lag phase that is attributable to a slow, constant loss of moisture. Correcting this discrepancy by fitting the first 25 hours of the data to a power function and subtracting these values from the absorbance readings obtained a better statistical fit to the growth curve in the lag phase. A power fit was selected over a linear fit because it reflected the effects of constant volume loss. This correction to the BioSentinel data analysis pipeline will enable quantitative statistical analysis of the effect of different levels of deep space radiation on yeast cells. Future work includes automation of drift correction and curve modeling to extract these parameters directly from data.

Growth Curve↗

Validating automated resonance evaluation with synthetic data

The integrity and precision of nuclear data are crucial for a broad spectrum of applications, from national security and nuclear reactor design to medical diagnostics, where the associated uncertainties can significantly impact outcomes. A substantial portion of uncertainty in nuclear data originates from the subjective biases in the evaluation process, a crucial phase in the nuclear data production pipeline. Recent advancements indicate that automation of certain routines can mitigate these biases, thereby standardizing the evaluation process and enhancing reproducibility. This research aims to provide a methodology, framework, and metrics for the validation of automated nuclear data evaluation software leveraging high-quality synthetic data that closely mimic real experimental observables. An introduced error metric provides a scale and intuitive measure of the evaluation quality by quantifying the estimate’s accuracy and performance across the specified energy range. Synthetic data provides access to experimental observables and underlying resonance parameters, enabling comparison of different evaluations. The methodology is demonstrated using Ta-181 isotope data in the resolved resonance region. The Automated Resonance Identification Subroutine (ARIS), which operates without prior resonance information, was used to test and showcase the framework’s capabilities utilizing the proposed error metrics. The results demonstrate the effectiveness of the proposed approach and framework for optimizing software parameters and testing hypotheses through “what-if” controlled experiments, such as modifying assumptions about experimental conditions or average resonance parameters.

22 GENERAL STUDIES OF NUCLEAR REACTORS↗

GL4U: Bioinformatics training for students and educators using space omics data

NASA’s GeneLab project provides researchers open access to space-relevant experiment multi-omics data that can be mined to understand the effects of spaceflight on biological systems. To maximize the number of scientists who understand and utilize GeneLab data and data processing pipelines, GeneLab has created GeneLab for Colleges and Universities (GL4U). GL4U provides space biology-relevant training in bioinformatics to the next generation of scientists through direct and indirect approaches. The GeneLab team plans to host two annual data processing bootcamps, one for college-level students (direct) and one for college educators (indirect – training of trainers), in which participants learn to analyze GeneLab’s space-relevant omics data. The GL4U direct training pilot program was conducted in June 2021. During the pilot, students participated in a week-long bootcamp consisting of space biology-specific lectures and hands-on instruction using Jupyter Notebooks to analyze RNA sequence data. This pilot demonstrated the capacity of GL4U for training young scientists and encouraging data re-use. During the educator pilot, scheduled for June 2022, educators will receive materials and training to enable them to run the bootcamp at their home institutions or alternatively to adapt the content to implement within existing courses, thereby extending the reach of this initiative.

Amanda Marie Saravia-butler↗

GL4U: Using Space Omics Data to Provide Bioinformatics Training for Students and Educators

NASA’s GeneLab project provides researchers open access to space-relevant experiment multi-omics data that can be mined to understand the effects of spaceflight on biological systems. To maximize the number of scientists who understand and utilize GeneLab data and data processing pipelines, GeneLab has created GeneLab for Colleges and Universities (GL4U). GL4U provides space biology-relevant training in bioinformatics to the next generation of scientists through direct and indirect approaches. The GeneLab team plans to host two annual data processing bootcamps, one for college-level students (direct) and one for college educators (indirect – training of trainers), in which participants learn to analyze GeneLab’s space-relevant omics data. The GL4U direct training pilot program was conducted in June 2021. During the pilot, students participated in a week-long bootcamp consisting of space biology-specific lectures and hands-on instruction using Jupyter Notebooks to analyze RNA sequence data. This pilot demonstrated the capacity of GL4U for training young scientists and encouraging data re-use. In June 2022, GL4U partnered with Jet Propulsion Laboratory’s (JPL) Planetary Protection Center of Excellence to conduct the indirect training pilot program by training educators at historically black colleges and universities (HBCUs) and minority serving institutions (MSIs). During the educator pilot, participants received materials, training, and will be provided the necessary compute resources to enable them to run the bootcamp at their home institutions or alternatively to adapt the content to implement within existing courses, thereby extending the reach of this initiative. The GL4U training program provides undergraduate students from underrepresented groups the opportunity to learn about NASA and Space Biology, and to enhance their career prospects by gaining hands-on experience analyzing omics data, a skillset that is highly applicable and marketable in the life sciences. Pre- and post-bootcamp surveys were completed by all participants and show the overwhelming success of the bootcamps.

Amanda M. Saravia-Butler↗

GL4U: Bioinformatics Training for Students and Educators Using Space Omics Data

NASA’s GeneLab project provides researchers open access to space-relevant experiment multi-omics data that can be mined to understand the effects of spaceflight on biological systems. To maximize the number of scientists who understand and utilize GeneLab data and data processing pipelines, GeneLab has created GeneLab for Colleges and Universities (GL4U). GL4U provides space biology-relevant training in bioinformatics to the next generation of scientists through direct and indirect approaches. The GeneLab team plans to host two annual data processing bootcamps, one for college-level students (direct) and one for college educators (indirect – training of trainers), in which participants learn to analyze GeneLab’s space-relevant omics data. The GL4U direct training pilot program was conducted in June 2021. During the pilot, students participated in a week-long bootcamp consisting of space biology-specific lectures and hands-on instruction using Jupyter Notebooks to analyze RNA sequence data. This pilot demonstrated the capacity of GL4U for training young scientists and encouraging data re-use. During the educator pilot, scheduled for June 2022, educators will receive materials and training to enable them to run the bootcamp at their home institutions or alternatively to adapt the content to implement within existing courses, thereby extending the reach of this initiative.

Amanda Marie Saravia-butler↗

Distributed Lunar Data Platform with Advanced Machine Learning Capabilities in Support of Lunar Science and Exploration

The United States 2020 Space Policy directive declares that NASA, in cooperation with private industry, will “extend human economic activity into deep space by establishing a permanent human presence on the Moon”. This goal will require advanced data management, as well as analysis, modeling and representation of lunar information in order to prepare for Artemis human missions, lunar science investigations and exploration. To meet this requirement, we conceptualize and present an implementation strategy for a distributed platform for lunar data retrieval, inferencing and analysis, which will be based on federated learning and the NASA Celestial Mapping System (CMS). In addition to demonstrating the imperative of enabling lunar-borne data to remain in-situ but still accessible, this presentation will also include examples of how third parties could contribute both datasets and new functionality into this platform using an AI-based data import pipeline and a plug-in architecture respectively.

Artificial Intelligence↗

The Kepler Data Processing Handbook: A Field Guide to Prospecting for Habitable Worlds

The Kepler telescope hurtled into orbit in March 2009, initiating NASA's first mission to discover Earth-size planets orbiting Sun-like stars. Kepler simultaneously collected data for approximately 165,000 target stars at a time over its four-year mission, identifying over 4700 planet candidates, over 2300 confirmed or validated planets, and over 2100 eclipsing binaries. While Kepler was designed to discover exoplanets, the long-term, ultrahigh photometric precision measurements it achieved made it a premier observational facility for stellar astrophysics, especially in the field of asteroseismology, and for variable stars, such as RR Lyrae. The Kepler Science Operations Center (SOC) was developed at NASA Ames Research Center to process the data acquired by Kepler from pixel-level calibrations all the way to identifying transiting planet signatures and subjecting them to a suite of diagnostic tests to establish or break confidence in their planetary nature. Detecting small, rocky planets transiting Sun-like stars presents a variety of daunting challenges, including achieving an unprecedented photometric precision of ~20 ppm on 6.5-hour timescales, and supporting the science operations, management, processing, and repeated reprocessing of the accumulating data stream. A newly revised and expanded version of the Kepler Data Processing Handbook (KDPH) has been released to support the legacy archival products. The KDPH details the theory, design and performance of the algorithms supporting each data processing step. This paper presents an overview of the KDPH and features illustrations of several key algorithms in the Kepler Science Data Processing Pipeline. Kepler was selected as the 10th mission of the Discovery Program. Funding for this mission is provided by NASA, Science Mission Directorate.

high performance computing↗

Characterization of contaminants in the Lyman-alpha forest auto-correlation with DESI

Baryon Acoustic Oscillations can be measured with sub-percent precision above redshift two with the Lyman-α (Lyα) forest auto-correlation and its cross-correlation with quasar positions. This is one of the key goals of the Dark Energy Spectroscopic Instrument (DESI) which started its main survey in May 2021. We present in this paper a study of the contaminants to the Lyα forest which are mainly caused by correlated signals introduced by the spectroscopic data processing pipeline as well as astrophysical contaminants due to foreground absorption in the intergalactic medium. Notably, an excess signal caused by the sky background subtraction noise is present in the Lyα auto-correlation in the first line-of-sight separation bin. We use synthetic data to isolate this contribution, we also characterize the effect of spectro-photometric calibration noise, and propose a simple model to account for both effects in the analysis of the Lyα forest. We then measure the auto-correlation of the quasar flux transmission fraction of low redshift quasars, where there is no Lyα forest absorption but only its contaminants. We demonstrate that we can interpret the data with a two-component model: data processing noise and triply ionized Silicon and Carbon auto-correlations. This result can be used to improve the modeling of the Lyα auto-correlation function measured with DESI.

79 ASTRONOMY AND ASTROPHYSICS↗

A Framework for Compressing Unstructured Scientific Data via Serialization

We present a general framework for compressing unstructured scientific data with known local connectivity. A common application is simulation data defined on arbitrary finite element meshes. The framework employs a greedy topology preserving reordering of original nodes which allows for seamless integration into existing data processing pipelines. This reordering process depends solely on mesh connectivity and can be performed offline for optimal efficiency. However, the algorithm’s greedy nature also supports on-the-fly implementation. The proposed method is compatible with any compression algorithm that leverages spatial correlations within the data. The effectiveness of this approach is demonstrated on a large-scale real dataset using several compression methods, including MGARD, SZ, and ZFP.

Reshniak, Viktor [ORNL] (ORCID:0000000315454462)↗

Kepler Planet Detection Metrics: Statistical Bootstrap Test

This document describes the data produced by the Statistical Bootstrap Test over the final three Threshold Crossing Event (TCE) deliveries to NExScI: SOC 9.1 (Q1Q16)1 (Tenenbaum et al. 2014), SOC 9.2 (Q1Q17) aka DR242 (Seader et al. 2015), and SOC 9.3 (Q1Q17) aka DR253 (Twicken et al. 2016). The last few years have seen significant improvements in the SOC science data processing pipeline, leading to higher quality light curves and more sensitive transit searches. The statistical bootstrap analysis results presented here and the numerical results archived at NASAs Exoplanet Science Institute (NExScI) bear witness to these software improvements. This document attempts to introduce and describe the main features and differences between these three data sets as a consequence of the software changes.

Bootstrap↗