Search NASA⌕ Search

SEARCH · Search NASA

Results for “open datasets”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 271 records · Page 15

Global 3D Data Visualization and Analysis Platform with Advanced Machine Learning Capabilities in Support of Lunar Exploration

Introduction: The science goals for NASA’s Artemis program include: a) Understanding the character and origin of lunar polar volatiles, b) Conducting experimental science in the lunar environment and c) Investigating and mitigating exploration risks. The permanently shadowed regions (PSRs) on the Lunar south pole are expected to host large quantities of water-ice and volatiles that are important for sustainable Lunar exploration. There are several missions such as onboard Korea Pathfinder Lunar Orbiter (KPLO: Korean name Danuri) with onboard ShadowCam camera, Astrobotic Peregrine Mission One [4], and other efforts underway to obtain high resolution topographic, minerals, volatiles and other information on the moon. We envision a need in immediate future for platforms to integrate these data sets, provide rendering and visualization capabilities in the context of a 3D Lunar globe for easier information access and analysis. NASA's Celestial Mapping System (CMS) is developed to address the need for 3D tools for planetary science investigations, mission planning, in-situ operations, in a 3D-first design constructed around a unified view of a planetary globe. At present CMS provides many critical functionalities that include: 1) equipment planning and optimized placement on Lunar surface 2) line-of-sight (LOS) analysis 3) powerful measurement tools based on 3D terrain with realistic 3D models to represent rovers, astronauts and equipment 4) visualization of derived mapping products (e.g. resource maps), and 5) a data engine for hosting new observations that are not available in other contemporary lunar data tools. Planetary Data Ingestion: CMS can consume and analyze data from locally hosted and external third party sources. It is compatible with Open Geospatial Consortium (OGC) data and file standards and currently integrates datasets from the Astrogeology Science Center of USGS. This includes global and local data acquired from NASA (LRO, Clementine, Lunar Orbiter) and JAXA (SELENE/Kaguya), with capability of integrating more datasets. In addition, users can specify other WMS-hosted data endpoints, which CMS can then query and stream data from automatically. To set-up an automated process for ingestion and accurate rendering, visualization and analysis of external 3rd party planetary datasets within CMS, we initiated the process of ingesting unique dataset of super-enhanced images of the permanently shadowed regions (PSRs) at the lunar poles which were produced by the Hyper-effective nOise Removal U-net Software (HORUS) tool. This tool was developed to enhance the extremely low-light images of the interior of PSRs and provide the ability to see within these regions at and discern surface features (i.e. boulders and craters) down to 3 meters in size. We focused on the Nobile region on the Lunar south pole, selected site for VIPER mission and stitched several images to create a high-resolution map within one of the PSR of Nobile crater. Figure 1 shows the dark PSR zone form the original NAC layer of LRO as the base layer (left image) and the illuminated areas within that crater (center) which was created by ingesting and merging several of HORUS generated images. At present we employ a semi-automated process to ensure spatial accuracy and merger of several overlapping zones. However, we are in the process of completely automating this process by employing AI based techniques that would rank, sort, and stack the images based on their information density. The georectification of the images would employ selected features. Analysis on Ingested Planetary Datasets: Once an external planetary data-set is successfully ingested, georectified and merged seamlessly as a data-layer; CMS’ numerous analysis tools can be used on this data. A Line of Sight (LOS) tool has been developed for CMS which analyzes terrain profiles and obstructions to determine visibility for remote observers. Figure 1 (right image) shows the viewshed analysis on the same PSR in the Nobile region. The yellow pin shows the observer location outside the PSR. The yellow area shows the visible part of PSR. The obstructed area with no visibility for the observer is shown in red. The Measurements tool allows the user to take area and distance measurements of features on the terrain using various shapes. Measurement type can be specified in a number of ways: Line, Path, Polygon, Circle, Ellipse, Square, Rectangle or Freehand. Once the shape is specified, elevation information can then be extracted along each of these shapes. Figure 2 (left) shows the measurements performed on a crater n illuminated PSR in Nobile region. The equipment placement tool allows the user to place a 3D equipment model at a desired location and analyze its coverage area. The equipment placement tool is coupled with LOS to determine the coverage. Figure 2 (right) shows an equipment placed on the Lunar terrain and it’s coverage area. The red rays are blocked sight lines and the green rays are non-obstructed sight lines with the cyan lines showing the point of intersection with the terrain. More details are provided in the video demonstrations in Reference 5. Overcoming Polar Distortions: 3D geospatial applications exhibit significant distortions in polar imagery due to several reasons: 1) distortions in the source imagery, 2) incompatible tessellation algorithms at the poles, and 3) map projections. We are leveraging new tessellation algorithms and reprojecting data using projections that are better suited for Lunar poles. The goal is to seamlessly switch to polar projections while maintaining 3D view and navigation.

Maps↗

Beyond Fair: Engagement, Data Usability, and Open Community Productivity through the NASA Open Science Data Repository

The FAIR principle (findable, accessible, interoperable, and reusable) governs the storage and sharing of NASA space biology and health data[1]. These guiding principles maximize reuse of data and the reproducibility of scientific findings. The NASA Open Science Data Repository (OSDR; an expansion of NASA GeneLab) was built on the FAIR principles and houses over 500 studies and close to 1000 datasets from decades of space life sciences experiments. OSDR embodies the FAIR principles through data governance that includes mediated, embargoed, and fully open access data. The FAIR data governance principles were recently proposed to be expanded to encompass a FAIREST framework for assessing research data repositories (FAIR + Engagement, Social connections, and Trust)[2]. FAIREST emphasizes the importance of data repositories engaging with the scientific community and gaining the trust of researchers regarding data quality. Trust also refers to the TRUST principles developed for assessment of digital repositories: Transparency, Responsibility, User Focus, Sustainability, Technology[3]. We present the “Open Science for Life in Space” Analysis Working Groups (AWGs) as evidence regarding the power of engagement, social connections, and trust which has enhanced OSDR’s capabilities and productivity. AWG members engage in two main activities. One, members provide feedback on OSDR scientific standards for data ingestion, curation, and reuse (study, subject and assay metadata; processing pipelines; dataset formats and uniformed structures for machine-readability). Two, AWG members collaborate to mine-reuse OSDR data to conduct scientific analysis. With nearly 800 active members, the AWGs have resulted in 32 publications re-using OSDR data and contributed many papers in two major special issues in Cell (2020) and Nature (2024). AWGs also serve as networking groups, facilitate social connections between researchers at all levels of experience, and also have a social online ‘Forum’ used to keep members informed on projects and opportunities. This community-centric, productive, and trustworthy data culture has resulted in a broader effect with international space agencies, academics, and the commercial space sector wanting to submit their data to OSDR. Ten studies of Inspiration 4 data were recently publicly released by OSDR, as were some JAXA human data. Coming up soon in OSDR are data submissions from the European Space Agency, Virgin Galactic PIs, and SpaceX Polaris Dawn. A major benefit of OSDR is the array of standardized and uniformly formatted data (which was developed through AWG member consensus), from which visualization tools, analysis tools, and machine learning models can be built or trained. This talk will cover the Multi-Study Visualization Tool, the Environmental Data Application, RadLab, and a UCSF-NSF funded knowledge graph biomedical health discovery tool ‘SPOKE’ currently being integrated with OSDR. OSDR also provides training programs in bioinformatics and machine learning to improve the scientific community’s awareness of data availability and to boost their ability to perform data analysis. The increasing engagement of the scientific community and the public with technologies powered by artificial intelligence (AI) heightens the need for data analysis to be transparent. The AI for Life in Space initiative leverages the data products provided in OSDR to train AI models, with an emphasis on explainable and trustworthy AI, which would not be possible without FAIR data and metadata. Overall, here we will demonstrate the importance for NASA life sciences data repositories to adhere to the FAIREST framework, by providing examples and success stories from different aspects of OSDR.

data↗

Analyzing EOSDIS Dataset Research Outputs using Knowledge Graphs and Large Language Models

Datasets, unlike publications, can be updated over time, with each new version receiving a DOI but not always being linked to previous ones. This complicates tracking citations across a dataset’s lifecycle. We address this by integrating dataset versions and citations into a knowledge graph (KG), which helps trace dataset citations and analyze dataset usage in applied research. To categorize publications from various journals, we fine-tuned NASA IMPACT INDUS Large Language Model (LLM) on a labeled publication set, assigning publications to one of twenty applied research areas. By linking datasets to these research areas, we improved dataset searchability and discovery through these domains.

open-source↗

MOSAIC-CONUS: A Multimodal, Multi-Temporally Paired Dataset for Earth Sciences

Earth embeddings—vector representations of geographic locations indexed in space and time—are emerging as a unifying interface for geospatial AI. However, their quality depends not only on model design, but on how multimodal Earth observation (EO) data are spatially indexed, temporally aligned, and cross-modally associated during pretraining. We introduce MOSAIC-CONUS (Multimodal Observations with Spatially Aligned Imagery, Urban Points of Interest, In-Situ Measurements and Text Captions), a large-scale EO dataset over the contiguous United States, organized around 250,000 stratified point indices that serve as stable spatial keys across seven modalities: active radar, passive optical imagery, lidar-derived elevation, land cover, functional context, hydrometeorological measurements, and textual summaries. Unlike existing EO datasets, MOSAIC-CONUS introduces four contributions not jointly addressed in prior work: 1. an open-source, large-scale multimodal EO corpus structured around point-indexed data designed to support Earth embedding learning; 2. explicit radar-optical pairing tables spanning twelve temporal alignment regimes, formalizing cross-sensor alignment as a controllable variable for analyzing how temporal mismatch across modalities influences learned embeddings quality; 3. a benchmark suite spanning cross-modal retrieval, annual nightlights regression, and basin-held-out streamflow prediction, positioning MOSAIC-CONUS as a benchmark-ready resource for multimodal AI systems; and 4. a language-based embedding layer through co-registered textual summaries, enabling Earth embeddings to function as a queryable interface for agentic AI systems. The dataset and pairing protocols are publicly released.

54 ENVIRONMENTAL SCIENCES↗

Toward the Neutrino Discovery Platform: An Auditable, Uncertainty-Bearing Toolchain for MINERvA Open-Data Cross-Section Analysis

The Neutrino Discovery Platform (NDP) aims to accelerate DUNE-era science by making the neutrino program's existing datasets analyzable through fast, reproducible, and auditable workflows. We report a working version of two of its layers, data curation and agentic orchestration, built and tested end to end on MINERvA open data. The guiding lesson throughout is that a cross section is a measurement, and not just a plotted shape, only if it carries a defensible systematic-uncertainty budget, a trustworthy unfolding, and a reproducible record. Using a single medium-energy playlist pair from the MINERvA open-data release (about $2.05\times10^{17}$ protons on target of data), we first reproduced the shapes of two published charged-current inclusive $\nu_\mu$ measurements through a complete extraction ladder: selection, background subtraction, D'Agostini unfolding, efficiency correction, and flux normalization. These shape-level reproductions ran and tracked the published results, but they lacked the systematic-uncertainty machinery that defines a MINERvA cross section. To supply it, we vendored and built the MINERvA Analysis Toolkit and developed a many-universe systematic-uncertainty tool that produces a portable covariance artifact, a parallel event-loop runner, and a per-run auditability harness. Validated against a published covariance release, the toolchain reproduces the released statistical, flux, and muon-energy-scale terms and shows that they account for roughly 63\% of the total variance, with the remainder unreleased. Using this same infrastructure, we then performed a measurement of our own design, the hadronic recoil-energy distribution of low-energy ($E_\nu<2.5$~GeV) charged-current inclusive events, and found data/simulation shape agreement of $\chi^2/\mathrm{ndf}=1.26$. Together these results show that the platform supports original physics and not only reproductions.

Breaux, Auto [Tulane U. (main)]↗

Connecting Users and Applications with Po.daac Hosted GHRSST Data

The 80+ GHRSST public datasets represent a rich resource for sea surface temperature research and applications given their time series length, resolution, spatial coverage, varying measurement types and processing levels, and availability in the full spectrum of PO.DAAC tools and services ecosystem. The PO.DAAC has created a publicly accessible recipe suite for the user community to perform straightforward yet powerful computations on GHRSST data using python recipes, Jupyter notebooks, R, Matlab, and the NCO programming language. These recipes include numerical computations for regional and global SST trends, anomaly derivations, EOF analysis, climate signal reproduction, and ocean phenology. For example, one recipe reproduces a famous SST based warming figure from the Fourth National Climate Assessment (USA) while another focuses on quantifying the regional changes in ocean SST phenology. Most are python-based while some contain hybrid calls and leverage the NCO programming interface too. All are available on the PO.DAAC user forum (https://podaac.jpl.nasa.gov/forum/) and/or via the open source NASA GitHub repository (https://github.com/nasa/podaac_tools_and_services). Several are available in the Jupyter notebook framework including podaacypy (https://github.com/nasa/podaacpy), a recipe for GHRSST granule metadata discovery and application, and more recently a Jupyter notebook developed to support data analysis and visualization of a cloud-based Zarr formatted Level 4 MUR dataset in the AWS Open Data Registry. Throughout the summer of 2020, the PO.DAAC intends to add and migrate more of its numerical recipes to the Jupyter notebook framework and publish them on its open source GitHub repository.

Gentemann, Chelle↗

Observational ozone datasets over the global oceans and polar regions (version 2024)

Studying tropospheric ozone over the remote areas of the planet, such as the open oceans and the polar regions, is crucial to understand the role of ozone as a global climate forcer and regulator of atmospheric oxidative capacity. A focus on the pristine oceanic and polar regions complements the available land-based datasets and provides insights into key photochemical and depositional loss processes that control the concentrations and spatiotemporal variability in ozone as well as the physicochemical mechanisms driving these patterns. However, an assessment of the role of ozone over the oceanic and polar regions has been hampered by a lack of comprehensive observational datasets. Here, we present the first comprehensive collection of ozone data over the oceans and the polar regions. The overall dataset consists of 77 ship cruises/buoy-based observations and 48 aircraft-based campaigns. The dataset, consisting of more than 630 000 independent ozone measurement data points covering the period from 1977 to 2022 and an altitude range from the surface to 5000 m (with a focus on the lowest 2000 m), allows systematic analyses of the spatiotemporal distribution and long-term trends over the 11 defined ocean/polar regions. The datasets from ships, buoys, and aircraft are complemented by ozonesonde data from 29 launch sites or field campaigns and by 21 non-polar and 17 polar ground-based station datasets. The datasets contain information on how long the observed air masses were isolated from land, as estimated by backward trajectories from the individual observation points. To extract observations representative of oceanic conditions, we recommend using a subset of the data with an isolation time of 72 h or longer, from the analysis with coincident radon observations. These filtered oceanic and polar data showed typically flat diurnal cycles at high latitudes, whereas daytime decreases in ozone (11 %–16 %) were observed at lower latitudes. The ship/buoy- and aircraft-based datasets presented here will supplement the land-based ones in the TOAR-II (Tropospheric Ozone Assessment Report Phase II) database to provide a fully global assessment of tropospheric ozone. The described dataset is available at https://doi.org/10.17596/0004044 (Kanaya et al., 2025).

Kanaya, Yugo [Japan Agency for Marine-Earth Scienc↗

Low Latency Flux and Concentration Datasets in Support of Greenhouse Gas Monitoring Based on NASA's GEOS Modeling and Data Assimilation System

We present efforts to develop space-based greenhouse gas monitoring systems that can provide low latency information and traceability to independent observations. Through support from its Carbon Monitoring System program, NASA has developed the capability to assimilate XCO2 retrievals from the Orbiting Carbon Observatory, 2 (OCO-2) into the Goddard Earth Observing System (GEOS) Constituent Data Assimilation System (CoDAS) to create gap-filled, three-dimensional (3D) estimates of CO2 mixing ratio. When OCO-2 data are not available, concentration fields are further informed by a bottom-up flux package based on remotely sensed fire radiative power, nighttime lights, and vegetation reflectance combined with estimates of atmospheric growth rate based on surface in situ data. The 3D nature of this dataset supports evaluation with independent aircraft data, helping to ensure transparency of remotely sensed data products. These quasi-operational data are currently produced 2-3 months behind real time and are distributed via NASA and international dashboard services to a variety of end users. In this presentation, we provide an overview of the system as well as remaining data gaps and modeling challenges. We also highlight the application of this dataset for detecting emissions anomalies associated with COVID-19 and comparing against independent emissions estimates. Finally, we highlight a new NASA initiative called the Earth Information System (EIS), which aims to support open science and applications by leveraging emerging cloud computing capabilities to increase access to NASA’s greenhouse gas datasets, opportunities for co-development, and transparency in methods for analysis and flux attribution.

Lesley Ott↗

Open Science for Life in Space: Bioimaging, Data Sharing, and Tools for Knowledge Discovery

Precious space-flown biological experiments have both multi-omic and phenotypic data which NASA strives to make maximally open access for reuse. Currently a number of these space-relevant bioimaging datasets are being reused for AI/ML approaches. NASA Ames Life Science Data Archive and NASA GeneLab are working to make all current and future bioimaging data even more accessible and reusable. Standards for collection and curation are being implemented to enable scientists worldwide access to these data for further discovery and use.

data science↗

AstroAmpSeq: Microbial Bioinformatics Education with NASA GeneLab’s Amplicon Pipeline

The prevalence and importance of large sequencing datasets in microbiology has led to a movement to share microbial ecology experimental data through open-access databases. This is particularly true of experiments that are difficult to replicate, such as those conducted in the spaceflight environment and shared via NASA GeneLab. It is now possible and indeed valuable for students to access and re-analyze these shared datasets for educational and research purposes. To provide students with experience utilizing microbial bioinformatics tools, GeneLab for Colleges and Universities (GL4U) has designed AstroAmpSeq, a week-long, virtually implemented project-based learning (PBL) minicourse to instruct undergraduate students on 16S amplicon sequencing. AstroAmpSeq was created to be accessible to students without prior bioinformatics or microbial ecology experience. During the minicourse students work in teams to process, analyze, and visualize a subsample of GeneLab dataset GLDS-280 using GeneLab’s standard amplicon processing pipeline, which is based in R. Students develop a hypothesis related to the dataset then generate and analyze figures to evaluate their hypothesis. Formative assessment of student learning is determined via pre- and post-evaluations, peer feedback, and self-reflection. Project and presentation rubrics serve as a summative assessment of student learning. GL4U AstroAmpSeq not only meets American Society for Microbiology Curriculum Guidelines, but also incites student interest in research by an inquiry-based approach and can be made part of a larger semester-long curriculum. GL4U AstroAmpSeq raises awareness of space microbiology and bioinformatics as a field and career path among undergraduates. Further, by using a GeneLab dataset and nesting microbiology techniques into the real-world application of space biology, AstroAmpSeq enforces deeper and longer-lasting student learning.

microbiology↗

SetGo: Metadata Readiness for Scientific AI Datasets

Scientific datasets intended for AI use require both computational readiness for model training and metadata readiness for discovery, sharing, and reuse. The Readiness Engine for Data Integration (REDI) addresses computational readiness, but no corresponding tool evaluates whether a dataset’s metadata are sufficiently complete, governed, and standards-compliant for publication and agent-based consumption. Existing FAIR assessors operate only on published repository records, and no single system covers FAIR compliance, licensing, provenance, governance, reproducibility, and catalog readiness together. We present SetGo, an open-source Python toolkit that assesses and repairs metadata readiness across these six dimensions before a dataset is published or archived. Applied to four scientific corpora, SetGo surfaces deficiencies that general-purpose tools do not detect: ERA5 climate metadata scores 4% on ACDD 1.3 compliance; materials datasets fail OPTIMADE species-definition requirements; and PDB-derived proteomics data carries licensing terms incompatible with standard SPDX identifiers. Guided enrichment raises overall FAIR scores from 52–57% to 81–91%, and a single setgo publish command pushes to Hugging Face Hub, CKAN, or OpenMetadata with ML Commons Croissant 1.0 metadata sidecars. To support interactive and automated workflows, SetGo integrates with coding agents powered by large language models (LLMs) through a /setgo skill that enables natural-language execution of the full assess–enrich–publish loop, with user involvement limited to supplying missing metadata values.

Wilkinson, Sean [ORNL] (ORCID:0000000214437479)↗

EC-Bench: A Benchmark for Enzyme Commission Number Prediction

Enzymes are proteins that catalyze specific biochemical reactions in cells. Enzyme Commission (EC) numbers are used to annotate enzymes in a four-level hierarchy that classifies enzymes based on the specific chemical reactions they catalyze. Accurate EC number prediction is essential for understanding enzyme functions. Despite the availability of numerous methods for predicting EC numbers from protein sequences, there is no unified framework for evaluating and studying such methods systematically. This gap limits the ability of the community to identify the most effective approaches for enzyme annotation. We introduce EC-Bench, a benchmark for EC number prediction, consisting of 1) an initial representative set of existing methods (including homology-based, deep learning, contrastive learning, and language model methods), 2) existing and novel accuracy and efficiency performance metrics, and 3) selected datasets to allow for comprehensive comparative study. EC-Bench is open-source and provides a framework for researchers to not only compare among existing methods objectively under uniform conditions, but also to introduce and effectively evaluate performance of new methods in a comparative framework. To demonstrate the utility of EC-Bench, we perform extensive experimentation to compare the existing EC number prediction methods and establish their advantages and disadvantages in a variety of prediction tasks, namely “exact EC number prediction”, “EC number completion” and (partial or additional) “EC number recommendation”. We find wide variation in the performance of different methods, but also subtle but potentially useful differences in the performance of different methods across tasks and for different parts of the EC hierarchy.

59 BASIC BIOLOGICAL SCIENCES↗

RectifHydPlus Data Pipeline

The RectifHydPlus Data Pipeline is an open source and fully reproducible data processing pipeline for creating RectifHydPlus—a dataset of historical monthly net electricity generation for all US hydropower plants (>10MW). The pipeline is coded in R, applying tidyverse libraries and code principles, and using the targets data pipeline framework. All data inputs to the RectifHydPlus Data Pipeline are available from public sources. References to all data inputs, as well as instructions for running the RectifHydPlus Data Pipeline, are available on the GitLab code repository: https://code.ornl.gov/turnersw/rectifhydplus

Turner, SeanWilliam Donald [Oak Ridge National Lab↗

RectifHydPlus Data Pipeline v1.1.0

The RectifHydPlus Data Pipeline is an open source and fully reproducible data processing pipeline for creating RectifHydPlus—a dataset of historical monthly net electricity generation for all US hydropower plants (>10MW). The pipeline is coded in R, applying tidyverse libraries and code principles, and using the targets data pipeline framework. All data inputs to the RectifHydPlus Data Pipeline are available from public sources. References to all data inputs, as well as instructions for running the RectifHydPlus Data Pipeline, are available on the GitLab code repository: https://code.ornl.gov/turnersw/rectifhydplus

Turner, SeanWilliam Donald [Oak Ridge National Lab↗

A Real-Time Satellite-Based Icing Detection System

Aircraft icing is one of the most dangerous weather conditions for general aviation. Currently, model forecasts and pilot reports (PIREPS) constitute much of the database available to pilots for assessing the icing conditions in a particular area. Such data are often uncertain or sparsely available. Improvements in the temporal and areal coverage of icing diagnoses and prognoses would mark a substantial enhancement of aircraft safety in regions susceptible to heavy supercooled liquid water clouds. The use of 3.9 microns data from meteorological satellite imagers for diagnosing icing conditions has long been recognized (e.g., Ellrod and Nelson, 1996) but to date, no explicit physically based methods have been implemented. Recent advances in cloud detection and cloud property retrievals using operational satellite imagery open the door for real-time objective applications of those satellite datasets for a variety of weather phenomena. Because aircraft icing is related to cloud macro- and microphysical properties (e.g., Cober et al. 1995), it is logical that the cloud properties from satellite data would be useful for diagnosing icing conditions. This paper describes the a prototype realtime system for detecting aircraft icing from space.

Minnis, Patrick↗

TPSAS-NF1676L-29062-DND

After more than three decades of research, the role of polar stratospheric clouds (PSCs) in stratospheric ozone depletion is well established. However, important questions remain unanswered that have limited our understanding of PSC processes and how to accurately represent them in global models, calling into question our prognostic capabilities for future ozone loss in a changing climate. A more complete picture of PSC processes on vortex-wide scales is emerging from a suite of contemporary satellite missions: the Michelson Interferometer for Passive Atmospheric Sounding (MIPAS) on Envisat (2002-2012), the Microwave Limb Sounder (MLS) on Aura (2004-present), and the Cloud-Aerosol Lidar with Orthogonal Polarization (CALIOP) on CALIPSO (2006-present). These datasets have motivated numerous research activities that both extend and challenge our present knowledge of PSC processes and modeling capabilities. The SPARC PSC initiative was organized in January 2015 to address key questions related to PSCs and their representation in global models with the following main objectives: identify key PSC parameters required by global models; identify strengths and limitations of the PSC datasets; define a methodology to obtain the key PSC properties required by models from the observational datasets; develop a state of the art PSC climatology; and identify remaining open science questions. In this presentation, we describe the PSCi activity, key findings, and remaining open questions.

Michael C Pitts↗

NASA GeneLab: The NASA Systems Biology Platform for Spaceomics Repository, Analysis and Visualization

At NASA Ames Research Center, the GeneLab Open Science Project is on a mission to gather all large -omics datasets relevant to space biology research. These datasets come from various organisms flown in multiple space habitats such as the International Space Station or the Space Shuttle, in addition to mimicking space-like conditions on ground. Researchers and citizen scientists all around the world have used the data and the analytical tools put together by the GeneLab team to start deciphering new biological impact of microgravity, space ionizing radiation and other space stressors.

GeneLab↗

Object Detection and Pose Estimation Public Competition Sample Data

This data was created with open-source tooling for the purpose of running a public competition. This dataset is a sample of what can be generated for the commissioned competition runner to be able to set up their testing infrastructure. It contains images of publicly available spacecraft models in a Blender scene composed of a light source, a background image, and the spacecraft model with bounding box labels.

James Berck↗