Search NASA⌕ Search

SEARCH · Search NASA

Results for “data pipeline”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 271 records · Page 15

Search for anisotropic gravitational-wave backgrounds using data from Advanced LIGO and Advanced Virgo's first three observing runs

We report results from searches for anisotropic stochastic gravitational-wave backgrounds using data from the first three observing runs of the Advanced LIGO and Advanced Virgo detectors. For the first time, we include Virgo data in our analysis and run our search with a new efficient pipeline called PyStochon data folded over one sidereal day. We use gravitational-wave radiometry(broadband and narrow band) to produce sky maps of stochastic gravitational-wave backgrounds and to search for gravitational waves from point sources. A spherical harmonic decomposition method is employed to look for gravitational-wave emission from spatially-extended sources. Neither technique found evidence of gravitational-wave signals. Hence we derive 95% confidence-level upper limit sky maps on the gravitational-wave energy flux from broadband point sources, ranging from F(α,Θ) < (0.013−7.6)×10^(−8)erg/sq. cm s Hz,and on the (normalized) gravitational-wave energy density spectrum from extended sources, ranging from Ω(α,Θ) < (0.57−9.3)×10^(−9) per sr, depending on direction (Θ) and spectral index (α). These limits improve upon previous limits by factors of 2.9−3.5. We also set 95% confidence level upper limits on the frequency-dependent strain amplitudes of quasimonochromatic gravitational waves coming from three interesting targets, Scorpius X-1, SN1987A and the Galactic Center, with best upper limits range fromh(0) < (1.7−2.1)×10^(−25), a factor of ≥ 2.0 improvement compared to previous stochastic radiometer searches.

R. Abbott↗

Automated Data Accountability for Missions in Mars Rover Data

As the Mars Curiosity Rover transmits data to the JPL Ground Data System (GDS), it frequently observes data loss and corruption, requiring re-transmits from the rover and Ground Data System Analysts (GDSA) to monitor the downlink process. As new missions are launched, the GDSA team redistributes analysts to these new missions, causing shortages in previous missions. The GDSA team can significantly benefit from the automation and optimization of the downlink process of telemetry data. In fact, there is a need for a better understanding of why the data is corrupted, so that the GDSA team can best determine the root cause of the issues in the GDS. This paper presents machine learning and deep learning based approaches to automate and optimize the detection of data loss. We first created a pipeline to automatically accumulate data from the telemetry databases (MAROS, Telemetry Data Storage, and GDS Elastic Search Database) in the downlink process. With our newly created datasets, we perform feature selection to supplement the GDSA understanding of the downlink process and provide supplemental analysis on the importance of different features. We implement various machine learning and deep learning based models, including support vector machines, ensemble methods, and deep neural networks and evaluate their accuracies in identifying whether a downlink process is complete or incomplete. We utilize fast hyperparameter optimization methods that allow our models to quickly be re-trained, allowing them to quickly be tuned and optimized on daily incoming data in real time. This hyperparameter optimization also allows our methods to be quickly integrated into other JPL missions. Our results show that our best-performing machine learning and deep learning based models outperform the existing GDSA detection software by 6 accuracy points and can aid analysts by providing insights into the data accountability problem. Since these various machine learning and deep learning approaches vary significantly in interpretability, we provide a discussion on the tradeoffs between their performance and trustworthiness in helping detect issues in data transmission.

Divsalar, Dariush↗

Enabling Model Organism and Commercial Astronaut Data Access Through the NASA Open Science Data Repository

NASA’s Open Science Data Repository (OSDR) brings together omics data from NASA’s GeneLab project and non-omics data, including physiological, phenotypic, imaging, and behavioral data from NASA’s Ames Life Sciences Data Archive (ALSDA) collected from decades of space biology research, providing open and FAIR (findable, accessible, interoperable, and reusable) access of these precious data to scientists world-wide. This rich source of meticulously curated metadata and data from spaceflight and analog studies has been mined by the scientific community resulting in dozens of high impact scientific publications that reveals a complex network of molecular and physiological effects of spaceflight across living systems, from microbes to plants, to mammals. Understanding how these effects translate to the human condition is critical as we move deeper into the era of commercial space travel. However, the integration of data, specifically omics data, from astronauts is particularly challenging due to their sensitive nature. OSDR has risen to this challenge by developing a mechanism to control access to identifiable levels of omics data, such as raw sequence data, while enabling public access to processed, unidentifiable, data and associated metadata that will allow the scientific community to interrogate human astronaut data alongside data from model organisms to begin answering these critical questions. The 2021 SpaceX Inspiration4 (I4) mission collected a comprehensive atlas of biological measurements from four civilian astronauts, providing a wealth of data to characterize the effects of spaceflight on the human body. These data include both non-omics and omics assays such as direct RNA sequencing (RNA-seq), single nuclei ATAC-seq and RNA-seq, metagenomics, proteomics, and comprehensive metabolic and cytokine panels, all of which have been integrated into the OSDR system across no less than 9 studies. Each study has been carefully curated using community-backed OSDR standards for sample and assay level metadata ensuring these data are findable and accessible. In addition to hosting both raw and processed data from the principal investigator team for each assay type, the GeneLab team plans to re-process the I4 omics data using GeneLab’s standard processing pipelines. The GeneLab processed data outputs will allow for comparisons across studies on OSDR and enable visualization of these data through the OSDR data visualization platform thereby enabling data reusability and interoperability. Here we describe the robust privacy and security protocols implemented by OSDR to safeguard sensitive health data from astronauts while facilitating metadata and processed data sharing for research purposes. We further provide a road map for navigating the vast amount of data provided for each I4 study on the OSDR, including experimental design, associated experiments, payloads, and missions, data generation and analysis protocols, and associated scientific articles. Additionally, we illustrate how to interrogate the standardized metadata provided in the sample and assay tables as well as various means to download and access the data including programmatically through the GeneLab Open API (GLOpenAPI). The open access of datasets in NASA’s OSDR provides a unique opportunity for the scientific community, as well as citizen scientists and students, to continue using OSDR resources to further unlock profound insights into the consequences of space travel on the human body. Through implementation of security measures to protect sensitive human data, the OSDR seeks to strengthen the science exchange between the Biological and Physical Sciences Program and the Human Research Program, per recommendation 4-1 of the 2023-2032 Decadal Survey, and encourage further sharing and dissemination of astronaut data to provide the scientific community with the resources needed to lay the groundwork for developing targeted mitigation strategies to help withstand the rigors of long-duration spaceflight.

Amanda Marie Saravia-butler↗

Enabling Model Organism and Commercial Astronaut Data Access Through the NASA Open Science Data Repository

NASA’s Open Science Data Repository (OSDR) brings together omics data from NASA’s GeneLab project and non-omics data, including physiological, phenotypic, imaging, and behavioral data from NASA’s Ames Life Sciences Data Archive (ALSDA) collected from decades of space biology research, providing open and FAIR (findable, accessible, interoperable, and reusable) access of these precious data to scientists world-wide. This rich source of meticulously curated metadata and data from spaceflight and analog studies has been mined by the scientific community resulting in dozens of high impact scientific publications that reveals a complex network of molecular and physiological effects of spaceflight across living systems, from microbes to plants, to mammals. Understanding how these effects translate to the human condition is critical as we move deeper into the era of commercial space travel. However, the integration of data, specifically omics data, from astronauts is particularly challenging due to their sensitive nature. OSDR has risen to this challenge by developing a mechanism to control access to identifiable levels of omics data, such as raw sequence data, while enabling public access to processed, unidentifiable, data and associated metadata that will allow the scientific community to interrogate human astronaut data alongside data from model organisms to begin answering these critical questions. The 2021 SpaceX Inspiration4 (I4) mission collected a comprehensive atlas of biological measurements from four civilian astronauts, providing a wealth of data to characterize the effects of spaceflight on the human body. These data include both non-omics and omics assays such as direct RNA sequencing (RNA-seq), single nuclei ATAC-seq and RNA-seq, metagenomics, proteomics, and comprehensive metabolic and cytokine panels, all of which have been integrated into the OSDR system across no less than 9 studies. Each study has been carefully curated using community-backed OSDR standards for sample and assay level metadata ensuring these data are findable and accessible. In addition to hosting both raw and processed data from the principal investigator team for each assay type, the GeneLab team plans to re-process the I4 omics data using GeneLab’s standard processing pipelines. The GeneLab processed data outputs will allow for comparisons across studies on OSDR and enable visualization of these data through the OSDR data visualization platform thereby enabling data reusability and interoperability. Here we describe the robust privacy and security protocols implemented by OSDR to safeguard sensitive health data from astronauts while facilitating metadata and processed data sharing for research purposes. We further provide a road map for navigating the vast amount of data provided for each I4 study on the OSDR, including experimental design, associated experiments, payloads, and missions, data generation and analysis protocols, and associated scientific articles. Additionally, we illustrate how to interrogate the standardized metadata provided in the sample and assay tables as well as instructions for how to download and access the data. The I4 datasets described here re present the first ever comprehensive collection of commercial astronaut data.

Amanda M Saravia-Butler↗

The Dynamic Networks Experiments: Virtual Experiments to Quantify Gains in Nuclear Explosion Monitoring

We describe an ongoing series of virtual experiments conducted collaboratively by four United States National Laboratories: Sandia National Laboratories, Los Alamos National Laboratory, Lawrence Livermore National Laboratory, and Pacific Northwest National Laboratory. These Dynamic Network Experiments (DNEs) provide an experimental framework to evaluate the potential impact of new research tools on nuclear explosion monitoring. The second DNE (DNE2), completed in 2024, exploited waveform data (seismic, infrasound, and electromagnetic) that was recorded by multi-modal sensors within and near the Nevada National Security Site and synthetic radionuclide signatures over multiple time periods. During the execution of DNE2, we processed and analyzed data through a multi-stage event processing pipeline that ingested raw data, performed quality control, detected signals, built events from these signals, located these events, and characterized the events’ source types and sizes. For each stage and over the entire event processing pipeline, we evaluated performance changes by comparing the performance of new data processing methods, models, and algorithms against a baseline. We also performed an additional execution phase to assess event processing pipeline function, speed, and efficiency against that of an expert analyst, including computational and manual efforts. Finally, we assessed the impact and effort of modern computing infrastructure on the monitoring pipeline. This paper describes key elements of the DNEs, from formulation through execution, as demonstrated in DNE2. The DNEs introduce several novel concepts to quantitatively measure the potential impact of new methods on explosion monitoring, including the collaborative design of multi-modal datasets, performance and logistical metrics, and integrated analyses.

42 ENGINEERING↗

Open Science for Life in Space: Data Sharing and Tools for Knowledge Discovery

Molecular-omics, physiological-phenotypic-behavioral, and environmental-radiation telemetry data from spaceflight biological and health studies are increasingly being made findable, accessible, interoperable, and reusable for the scientific public. These data, as well as space science-relevant biospecimens, are available through NASA’s Open Science Data Repository (OSDR), which is the new umbrella grouping of NASA GeneLab, the Ames Life Sciences Data Archive (ALSDA), and the NASA Biological Institutional Scientific Collection (NBISC). The quality of data is underpinned by datasets having rich metadata (determined through Analysis Working Group members), processing pipelines to enable data reuse standards, and ontologies specifying terminology semantics (e.g., the Radiation Biology Ontology).

space biology↗

Benefits and Limits of Phasing Alleles for Network Inference of Allopolyploid Complexes

Abstract Accurately reconstructing the reticulate histories of polyploids remains a central challenge for understanding plant evolution. Although phylogenetic networks can provide insights into relationships among polyploid lineages, inferring networks may be hindered by the complexities of homology determination in polyploid taxa. We use simulations to show that phasing alleles from allopolyploid individuals can improve phylogenetic network inference under the multispecies coalescent by obtaining the true network with fewer loci compared with haplotype consensus sequences or sequences with heterozygous bases represented as ambiguity codes. Phased allelic data can also improve divergence time estimates for networks, which is helpful for evaluating allopolyploid speciation hypotheses and proposing mechanisms of speciation. To achieve these outcomes in empirical data, we present a novel pipeline that leverages a recently developed phasing algorithm to reliably phase alleles from polyploids. This pipeline is especially appropriate for target enrichment data, where the depth of coverage is typically high enough to phase entire loci. We provide an empirical example in the North American Dryopteris fern complex that demonstrates insights from phased data as well as the challenges of network inference. We establish that our pipeline (PATÉ: Phased Alleles from Target Enrichment data) is capable of recovering a high proportion of phased loci from both diploids and polyploids. These data may improve network estimates compared with using haplotype consensus assemblies by accurately inferring the direction of gene flow, but statistical nonidentifiability of phylogenetic networks poses a barrier to inferring the evolutionary history of reticulate complexes.

Evolutionary Biology↗

Comparison of Ring Diagrams Based on the Doppler Shifts of Synthetic Data Obtained with Bisector Method and SDO/HMI Pipeline

Ring diagrams are cross-sections of three-dimensional spatiotemporal power spectra of solar oscillations. The rings reveal information about sub-surface flows and represent an important tool for helioseismology. Ring diagrams can be constructed using Doppler velocity or intensity maps of spectral lines. How the velocities are computed is an important factor for accuracy of information we can retrieve on subsurface flows. In our work, the ring diagrams are generated from Doppler shift data of synthesized Fe I 6173 Å line. We compare ring diagrams computed by two methods–HMI line-of-sight pipeline and the bisector of Fe line. Fe I line is synthesized for StellarBox 3D Radiative hydrodynamic simulations under LTE assumption. We aim to answer the following questions: 1.How do power spectra obtained from velocities computed with the HMI pipeline and bisector of the Fe I 6173Å compare? 2.How do the power spectral density retrieved with each method vary with heliocentric angle? 3.What is the effect of changing resolution on the power spectral density in ring diagrams obtained with the two methods?

SMD↗

Numerical Simulation on Liquid Hydrogen Chill-Down Process of Vertical Pipeline

In order to improve the cryogenic propellant management technologies for a liquid hydrogen rocket with high specific impulse, JAXA, the University of Tokyo, and the NASA Glenn Research Center have jointly organized a multi-agency model validation collaboration project. As part of this project, JAXA's boiling simulation was validated with NASA's experimental data on vertical pipeline chill-down. Simulation results were in good agreement with the experimental data obtained using an improved boiling model to reproduce the spray flow. This activity achieved liquid hydrogen turbo-pump simulation at JAXA for grasping the boiling flow phenomenon from engine cut-off to re-ignition. This joint research resulted in an international cooperative relationship for discussing the cryogenic propellant management technologies necessary to develop next-generation liquid rockets.

Umemura, Yutaka↗

Optimizing Distributed Training on Frontier for Large Language Models

Large language models (LLMs) have demonstrated remarkable success as foundational models, benefiting various downstream applications through fine-tuning. Loss scaling studies have demonstrated the superior performance of larger LLMs compared to their smaller counterparts. Nevertheless, training LLMs with billions of parameters poses significant challenges and requires considerable computational resources. For example, training a one trillion parameter GPT-style model on 20 trillion tokens requires a staggering 120 million exaflops. This research explores efficient distributed training strategies to extract this computation from Frontier, the world's first exascale supercomputer. We enable and investigate various model and data parallel training techniques, such as tensor parallelism, pipeline parallelism, and sharded data parallelism, to facilitate training a trillion-parameter model on Frontier. We empirically assess these techniques and their associated parameters to determine their impact on memory footprint, communication latency, and GPU's computational efficiency. We analyze the complex interplay among these techniques and find a strategy to combine them to achieve high throughput through hyperparameter tuning. We have identified efficient strategies for training large LLMs of varying sizes through empirical analysis and hyperparameter tuning. For 22 Billion, 175 Billion, and 1 Trillion parameters, we achieved GPU throughputs of 38.38%, 36.14%, and 31.96%, respectively. For the training of the 175 Billion parameter model and the 1 Trillion parameter model, we achieved 100% weak scaling efficiency on 1024 and 3072 Mi250X GPUs, respectively. We also achieved strong scaling efficiencies of 89% and 87% for these two models. We trained these models only tens of iterations instead of training till completion.

Yin, Junqi↗

The Evolution of Stellar Dynamos; Survey for Low Mass Members of NGC2232; An X-Ray Survey of the Open Cluster CR140; Towards a Better Understanding of the Rotation-Activity Relation for Solar-Type Members of the Pleiades

This grant was originally awarded to Dr. Charles Prosser, who died tragically in a car accident in Tucson in 1998. We had hoped to finish the work Charles had started, which involved analysis of ROSAT data for three programs (observations of the clusters NGC2232, Crl4O and the Pleiades) and also analysis of optical data for each cluster in order to allow interpretation of the ROSAT observations. The Pleiades portion of the program was completed during the past year, and a paper published. We have obtained optical imaging of the other two clusters, and those data are being analyzed. Dr. Brian Patten intends to complete analysis of the ROSAT observations and to combine those data with the optical photometry, but progress on those efforts has been slow due to the press of other work (Dr. Patten is responsible for the pipeline processing of data from SWAS). We intend to publish those results as soon as we can, but it will now be completed without further support from this grant.

Stauffer, John R.↗

Names Don't Fly: Smart Filters for Profanity Detection and Classification in User-Generated Content

Generally, names associate with a person’s identity. But what if in the pretext of a legitimate name and given the opportunity, users of software provide names to online web forms that carry along offensive language, slurs, and other profanity that is then sent to Mars ? The answer is simple: they don’t fly. In this paper,we perform model explorations to detect and classify inappropriate content in the names submitted from people across the world to ‘Send Your Names to MARS’ public engagement campaign.We propose a novel pipeline approach, that can effectively overcome the issues of lack of negative samples, noisy labels by gathering expert knowledge over time with human(s) in the loop and data augmentation, and achieve high accuracy in classifying inappropriate names with very little or no context. We describe cloud-based infrastructure to deploy our application and run predictions on large-scale data through our pipeline and achieve significant speedup over offline processes, with enhanced reliability and security.

Soderstrom, Tomas↗

Adaptive Cybersecurity for Distributed Energy Resources (AdCyDER): Online Reinforcement Learning with Stackelberg-Optimized Defenses — Pipeline Architecture, Evaluation Methodology, and Findings from a Synthetic-Data Evaluation

This report documents the design and evaluation of an integrated online-learning pipeline developed within the AdCyDER project for Distributed Energy Resource (DER) cybersecurity. The pipeline couples a Reinforcement Learning (RL) attack classifier — which produces an attack-type probability distribution — with a Stackelberg game-theoretic (GT) defense selector that consumes those distributions alongside SME-encoded priors over (defense, attack) effectiveness pairings and perdefense costs to choose grid-health-preserving defenses. The objective is not attack classification per se but production of distributions that drive effective defense selection through the Stackelberg layer, learned from delayed grid-health feedback rather than labeled attack data. AdCyDER as a whole is broader than the work presented here; this report covers the specific RL/GT loop integration and its evaluation. We present the integrated pipeline (SCADA telemetry with Fronius inverter physics, Suricata IDS, time-windowed aggregation, per-facility LSTM classifier, Stackelberg optimizer, OpenC2 actuators), an experimental campaign of 28 eight-hour iterations across three baseline modes, and a pipeline-ordered diagnostic protocol. The protocol identifies two distinct failure modes within the loop: paired supervised ceilings on the same features establish that the deployed online RL classifier (macro F1 ≈ 0.07) sits at least 4.7× below a same-architecture supervised LSTM (≈ 0.34) and 10–11× below a linear feature-signal ceiling (≈ 0.70–0.79 depending on per-facility isolation), localizing the dominant failure to the training procedure; and the reward signal driving online updates carries weak directional coupling with classifier correctness in the methodology-expected direction (multi-lens convergent: top-decile P(true) records produce more frequent state changes and slightly larger improvements, top-vs-bot Cohen’s 𝑑 ≈ −0.19), but at effect magnitudes too small to drive gradient-based learning at the campaign sample size. The original learning hypothesis is not supported by the data. The primary contributions are the diagnostic methodology — proposed as a transferable falsification protocol for online RL/GT defense pipelines learning from delayed environmental reward — and the open, reproducible experimental infrastructure. We outline reward reformulation as the highest-priority aspirational next step given the underpowered-but-aligned Q6 reading, with hardware-in-the-loop evaluation as the broadest scope-expansion option.

Blakely, Benjamin [Argonne National Laboratory (AN↗

Testing the LSST Difference Image Analysis Pipeline Using Synthetic Source Injection Analysis

Abstract We evaluate the performance of the Legacy Survey of Space and Time Science Pipelines Difference Image Analysis (DIA) on simulated images. By adding synthetic sources to galaxies on images, we trace the recovery of injected synthetic sources to evaluate the pipeline on images from the Dark Energy Science Collaboration Data Challenge 2. The pipeline performs well, with efficiency and flux accuracy consistent with the signal-to-noise ratio of the input images. We explore different spatial degrees of freedom for the Alard–Lupton polynomial-Gaussian image subtraction kernel and analyze for trade-offs in efficiency versus artifact rate. Increasing the kernel spatial degrees of freedom reduces the artifact rate without loss of efficiency. The flux measurements with different kernel spatial degrees of freedom are consistent. We also here provide a set of DIA flags that substantially filter out artifacts from the DIA source table. We explore the morphology and possible origins of the observed remaining subtraction artifacts and suggest that given the complexity of these artifact origins, a convolution kernel with a set of flexible bases with spatial variation may be needed to yield further improvements.

Liu, S. (ORCID:0000000244612143)↗

GeoNEX: A Geostationary Earth Observatory

The latest generation of geostationary satellites (Himawari 8/9, GOES-16/17, FY-4, GK-2A) carries sensors that closely mimic the spatial and spectral characteristics of widely used polar-orbiting, global monitoring sensors such as MODIS and VIIRS. When combined, data from various currently operating/planned geostationary platforms provide a geo-ring of hyper-temporal (5-10 minutes), multispectral observations at spatial resolutions as high as 500 m. These high frequency observations offer exciting new possibilities for monitoring our planet, including better retrievals of geophysical variables by overcoming cloud cover, enabling studies of diurnally varying phenomena in the atmosphere, land, and the oceans, and support operational decision-making in agriculture, hydrology and disaster management. The NASA Earth Exchange (NEX) team, in collaboration with scientists from JAXA, KARI, NOAA and other international institutions, created the GeoNEX (www.nasa.gov/geonex) pipeline to integrate data from all available geostationary platforms and produce and distribute spatially, temporally, and radiometrically consistent data for the earth science community. We envision various institutions adapting the Geo component (e.g., GeoNOAA, GeoKARI, GeoChiba, GeoJAXA, GeoCMA) and customizing the pipeline and downstream products to serve the local/regional research and applied science communities. To facilitate collaborative work among the partners, we have established the OpenNEX platform on the public cloud. OpenNEX provides researchers, developers, educators, and ordinary users with easy access to an integrated Earth science computational and data platform, enabling citizen scientists and application developers to realize the full value of GeoNEX data assets and software tools.

GeoNEX↗

ASDC’s Python-Based Metadata Extraction Pipeline for Suborbital Campaigns

The FAIRness of data products, especially findability and accessibility depend on rich metadata which, when extracted, can allow for proper curation. Over the past few years, the Atmospheric Science Data Center (ASDC) suborbital science support team has developed a metadata extraction pipeline to ensure the required metadata can be retrieved systematically, effectively, and efficiently to ensure the data can be used by a broad community. The development of a pipeline has presented many, but necessary, challenges to support archival and distribution of ASDC’s 30+ suborbital missions. Though sufficient metadata is provided by instrument scientists, the metadata may not be readily machine actionable due to different formats and templates. Further complicating metadata extraction, our team has found that the nature of metadata can be quite diverse given the difference in measurement types, instruments, and measurement platforms. A metadata extraction pipeline has been developed to provide an efficient, plugin-in based, method for adding new parsers, a configuration system that lets non-developers customize how files are processed, and a system for identifying and logging metadata quality issues to ensure they are readily found and addressed. The metadata extraction pipeline identifies critical pieces of metadata that are needed to promote data FAIRness, including location, file revision, measurement start/end datetime and can be easily modified to extract further information (such as variables). Given the wide-ranging datasets, the pipeline has been modified to accommodate multiple file formats, including multiple versions of ICARTT (International Consortium for Atmospheric Research on Transport and Transformation), HDF (Hierarchical Data Format), netCDF (network Common Data Form), and multiple versions of the Ames File Format. The pipeline also supports building metadata for file formats that cannot have metadata easily extracted from them, such as PDF (Portable Document Format) and GIF (Graphics Interchange Format). The pipeline has allowed our team to maintain a consistent flow of data and metadata to archival and distribution services, ensuring the ASDC meets the needs of the suborbital science community. This presentation will highlight the ASDC’s suborbital metadata extraction pipeline, its development, how it’s been modified to support data FAIRness, and plans for maintaining the pipeline and adding new features.

Abraham Porter↗

The Kepler Science Operations Center Pipeline Framework Extensions

The Kepler Science Operations Center (SOC) is responsible for several aspects of the Kepler Mission, including managing targets, generating on-board data compression tables, monitoring photometer health and status, processing the science data, and exporting the pipeline products to the mission archive. We describe how the generic pipeline framework software developed for Kepler is extended to achieve these goals, including pipeline configurations for processing science data and other support roles, and custom unit of work generators that control how the Kepler data are partitioned and distributed across the computing cluster. We describe the interface between the Java software that manages the retrieval and storage of the data for a given unit of work and the MATLAB algorithms that process these data. The data for each unit of work are packaged into a single file that contains everything needed by the science algorithms, allowing these files to be used to debug and evolve the algorithms offline.

Klaus, Todd C.↗

GeoNEX: A geostationary earth observatory at NASA Earth eXchange: Earth monitoring from operational geostationary satellite systems

The latest generation of geostationary satellites (Himawari 8/9, GOES-16/17, FY-4, GK-2A) carries sensors that closely mimic the spatial and spectral characteristics of widely used polar-orbiting, global monitoring sensors such as MODIS and VIIRS. When combined, data from various currently operating/planned geostationary platforms provide a geo-ring of hyper-temporal (5-10 minutes), multispectral observations at spatial resolutions as high as 500 m. These high frequency observations offer exciting new possibilities for monitoring our planet, including better retrievals of geophysical variables by overcoming cloud cover, enabling studies of diurnally varying phenomena in the atmosphere, land, and the oceans, and support operational decision-making in agriculture, hydrology and disaster management. The NASA Earth Exchange (NEX) team, in collaboration with scientists from JAXA, KARI, NOAA and other international institutions, created the GeoNEX (www.nasa.gov/geonex) pipeline to integrate data from all available geostationary platforms and produce and distribute spatially, temporally, and radiometrically consistent data for the earth science community. We envision various institutions adapting the Geo component (e.g., GeoNOAA, GeoKARI, GeoChiba, GeoJAXA, GeoCMA) and customizing the pipeline and downstream products to serve the local/regional research and applied science communities.

Ramakrishna R Nemani↗