Search NASASearch

Engineering topics

Lauren M Sanders

Publications and source records attributed to Lauren M Sanders.

At least 19 records

NASA Open Science Data Repository: Open Science for Life in Space

Space biology and health data are critical for the success of deep space missions and sustainable human presence off-world. At the core of effectively managing biomedical risks is the commitment to open science principles, which ensure that data are findable, accessible, interoperable, reusable, reproducible and maximally open. The 2021 integration of the Ames Life Sciences Data Archive with GeneLab to establish the NASA Open Science Data Repository significantly enhanced access to a wide range of life sciences, biomedical-clinical, and mission telemetry data alongside existing ‘omics data from GeneLab. This paper describes the new database, its architecture, and new data streams supporting diverse data types and enhancing data submission, retrieval, and analysis. Features include the Biological Data Management Environment for improved data submission, a new user interface, controlled data access, an enhanced API, and comprehensive public visualization tools for environmental telemetry, radiation dosimetry data, and ‘omics analyses. By fostering global collaboration through its Analysis Working Groups and training programs, the Open Science Data Repository promotes widespread engagement in space biology, ensuring transparency and inclusivity in research. It supports the global scientific community in advancing our understanding of spaceflight's impact on biological systems, ensuring humans will thrive in future deep space missions.

OSDR

Tracking Community Building in Open Science

Open Science is enabled by a vibrant community of researchers who regularly engage with the data, from its production to its organization, curation, archiving, dissemination, analysis, and publication. This presentation will examine community building in open science. The NASA Open Science Data Repository (OSDR) makes data available to the public following the FAIR (Findability, Accessibility, Interoperability, and Reusability) principles. OSDR takes open science further with the OS Analysis Working Groups (AWGs) that facilitate community development and promotion. The primary activity of each AWG is to establish and validate analytical processes to generate higher-order data from data housed in OSDR. There are a number of these groups on various topics, including the Animal AWG, Plant AWG, Microbial AWG, Multi-Omics AWG, AI/ML AWG, and the Ames Life Sciences Data Archive (ALSDA) AWG. The international volunteers participating in these AWGs come from academia, citizen science initiatives, industry, and government. They include researchers, principal investigators, professors, trained hobbyists, and students from various domains and disciplines. Anyone may request to join the AWGs, and membership requests are vetted monthly by the group organizers before granting admission. Core to membership is demonstrated expertise through records of training, integrity, work in the professed domain(s), and good community standing. Regular virtual meetings are held for each AWG, with a varying cadence depending on the group's needs and goals. AWG communities share their expertise in research including cutting edge tools, software, frameworks, data formats, and libraries accelerating research collectively. This collaborative approach helps community members cross technology gaps and identify emerging challenges. These diverse communities encompass a wide range of individuals hailing from various sectors within the Science Mission Directorate and beyond. They serve as a means to promote and enhance transparency, accessibility, and inclusion. An annual AWG Symposium brings contributors together in person. Participation in AWGs can be synchronous or asynchronous, with some groups performing most of their work in off hours. Participants gain valuable skills and connections that allow them to add value to their communities and new organizations that they join, resulting in an expanded return on investment for the space life science community. Open science is increasingly a federal mandate and initiatives like NASA's Transform to Open Science and instruments like the Decadal Survey of Biological and Physical Sciences in Space demonstrate the need to carefully consider best practices in this domain. Here, we present greater detail about the makeup and participation metrics of the various AWGs affiliated with OSDR and details of successful peer-reviewed publication campaigns.

Christina M Johnson

Open Science for Life in Space: Bioimaging, Data Sharing, and Tools for Knowledge Discovery

Precious space-flown biological experiments have both multi-omic and phenotypic data which NASA strives to make maximally open access for reuse. Currently a number of these space-relevant bioimaging datasets are being reused for AI/ML approaches. NASA Ames Life Science Data Archive and NASA GeneLab are working to make all current and future bioimaging data even more accessible and reusable. Standards for collection and curation are being implemented to enable scientists worldwide access to these data for further discovery and use.

data science

Open Science for Life in Space: Data Sharing and Tools for Knowledge Discovery

The fast-growing array of space biological data, which in the past was simply archived after minimal analysis, holds great potential if it can be reorganized and formatted for Open Science. Organizing the data for such analysis is a challenge because of its diverse nature (molecular, cellular, tissue, whole organism, behavior; tabular, imagery). Open Science is the concept that the more people have access to scientifically curated data, the more knowledge will be gained. This led NASA to start the development of GeneLab in 2015. GeneLab houses spaceflight and space-analog multi-omics datasets from plant, rodent, small animal, and microbial experiments. The success and knowledge gained from GeneLab led to a new alliance of NASA “Open Science Data Repositories” (OSDR), which include the Ames Life Sciences Data Archive (ALSDA) and the NASA Biological Institutional Scientific Collection (NBISC). Both are adopting the GeneLab data system, so data are more findable, accessible, interoperable, and reusable (FAIR). OSDR systems provide users the ability to upload, download, search, share, analyze, and visualize. Open Science also needs strong confidence in the data, which is gained through building science communities. With ~400 current members, GeneLab and ALSDA formed Analysis Working Groups (AWGs) to provide feedback on processing pipelines, metadata curation standards (for ‘omics and phenotypic-physiological-behavioral assays), and to collaborate in effectively reusing data. The AWG also led to the development of the Radiation Biology Ontology (RBO), ensuring radiation metadata are efficiently captured, connected, and interoperable. Feedback from the AWG provided design input toward the new single point-of-entry data submission portal for all investigators to submit, curate, and share their research data. Space biological data is now maximally open access, collected-curated with rich metadata, and formatted for interoperability to enable systems biology, meta-analysis, knowledge graphs, machine learning, modeling, and other reuse approaches. With potential for further federation of OSDR for data mining with traditional biological and medical databases (NIH, NCI, EBI, etc.), a new era for space biology has begun to support the knowledge discovery necessary for Lunar and Martian missions.

Ryan T Scott

Benchmark Models for Classification of Radiation Type Induced in Immune Cells

NASA Biological and Physical Sciences and the Science Mission Directorate have published a benchmark dataset of mouse immune cells subjected to radiation-induced DNA damage. The dataset comprises ML-ready microscopic imagery of said cells, including labels indicating radiation type and dose. The machine learning team at NASA Interagency Implementation and Advanced Concept Team (IMPACT) created multiple benchmark models. Initially, we conducted a preliminary analysis using thresholding. The algorithm used thresholds on average brightness of the available images to classify them into their respective radiation type. We also tested machine learning approaches. Convolutional Neural Networks (CNN) emerged as the best-performing model. This poster presents the benchmark scores obtained by the models.

Vishal Perekadan

RadLab and the Environmental Data Application Dashboard: Graphical and Programming Interfaces for Interrogation of Space Telemetry Data

Sensors on the International Space Station (ISS) and multiple spacecraft elsewhere in Earth orbit and in deep space continuously monitor and collect environmental data, transmitting this information back to Earth. These data include ionizing radiation and, on the ISS, CO2, relative humidity levels, and temperature, and are of great importance to space biology research. Ionizing radiation in particular has been established in ground-based experiments as being correlated with increased risk of carcinogenesis and cardiovascular and neurological effects. Looking ahead to future long duration crewed missions beyond low Earth orbit, the ability to study how factors including CO2 levels, light cycle, temperature modulate the response to ionizing radiation and microgravity is essential. To date, access to these data has been fragmented across space agencies, spacecraft, and databases. To address this issue, NASA’s Open Science Data Repository (osdr.nasa.gov) has developed two Web applications: the Environmental Data Application (EDA) and a radiation-specific RadLab. Each consists of an API (application programming interface) and an associated GUI (graphical user interface) that provide single points of access to the data. To date, OSDR has focused on the sensors from payloads and radiation detectors located on the ISS. The Web applications process telemetry information and associated data, such as spacecraft location and orientation, from multiple international databases. The applications’ request syntax enables users to interrogate these data by craft, sensor type, time range, radiation type (galactic cosmic rays, solar particle events, the contribution of the South Atlantic Anomaly), facilitating arbitrary comparisons of original source data at varying time resolutions. The applications provide programmatic access for use in computational pipelines and GUIs for data visualization and exploration, making these data FAIR (Findable, Accessible, Interoperable, and Reusable), complementing the biological data contained in OSDR, and providing the space science community with a valuable resource for scientific analyses.

radiation

Optimizing a Small RNAseq Analysis Pipeline for NASA GeneLab Using Open-Source Tools and Libraries

Small RNA sequencing (small RNAseq) is a powerful tool for studying the regulation of gene expression in various organisms. Small RNAseq has been leveraged in space biology research to study how expression of small RNAs, e.g. micro RNAs (miRNAs), small interfering RNAs (siRNAs), and piwi-interacting RNAs (piRNAs), change upon exposure to the space environment. NASA GeneLab currently hosts small RNAseq raw data derived from space-relevant experiments on the Open Science Data Repository (OSDR). To maximize the accessibility of these data to the scientific community, in addition to hosting raw data, which is only interpretable by bioinformaticians, GeneLab plans to process all small RNAseq datasets and make those processed data available to the scientific community via the OSDR. In this study, we present the development of the GeneLab standardized pipeline for processing small RNAseq datasets. Using human, plant, and synthetic small RNAseq datasets, we interrogate various open-source software and publicly available databases to evaluate their accuracy and reproducibility in each step of the pipeline. For quality control and adapter detection and trimming, we evaluated TrimGalore!, FASTX, SeqKit, and DNApi methods to optimize alignment to reference genomes. We compared BWA, Bowtie, and Bowtie2 to determine the optimal alignment tool. For each alignment tool we also assessed various reference databases, including Ensembl reference genomes and different types of small RNA reference databases, including genome, hairpin, and miRNA references from the miRbase and MirGeneDB databases. To quantify the aligned data, we compared SAMtools, HTSeq, and RSEM for counting alignment events from each alignment tool used. Finally, we evaluated various tools, including DESeq2 and EdgeR, for data normalization and subsequent differential expression analysis. We will present the results from our comparative analyses for each pipeline step and propose a consensus pipeline for processing small RNAseq data derived from various organisms exposed to the space environment.

SmallRNAseq, NASA GeneLab, quality control, adapte

Transcriptomics Processing Pipelines for Space Biology: An Open Source and Consensus-Driven Approach

Transcriptomics holds significant value in elucidating the relationship between gene expression, experimental factors, biological factors, and various types of omics data. Enhancing our understanding of these connections is paramount for foundational biology, which plays a pivotal role in devising solutions for challenges pertinent to both space travel and terrestrial life. The NASA GeneLab project, part of the Open Science Data Repository (OSDR.nasa.gov), seeks to accelerate space biology research through cataloging and democratizing ‘omics data, including transcriptomics. Since raw omics data are largely inaccessible to non-bioinformaticians, GeneLab works with the scientific community via the Open Science Analysis Working Groups (AWGs) to develop standard processing pipelines to generate and publish processed data. Unlike raw data, processed data have greater immediate value to diverse users with varying technical backgrounds and computational capabilities. Standardizing processing workflows is essential to match the pace of raw data generation, ensure reproducibility, and enable standardized processed data for comparison across datasets. As of June 2023, transcriptomics studies comprise over half of GeneLab datasets hosted on the OSDR, including data from bulk RNA-seq and Affymetrix or Agilent 1-Channel DNA microarray assays. In collaboration with the AWGs, GeneLab developed consensus processing pipelines for these transcriptomics data types that includes quality control, background correction (microarray only), data normalization and quantification, culminating in the detection and annotation of differentially expressed genes. The work presented here describes Nextflow implementations of GeneLab’s consensus transcriptomics pipelines that automates and accelerates processing of these datasets. In addition to the core data processing, these workflows also include raw data staging and a robust verification and validation program to identify errors in real-time, stop additional downstream computation, and preserve computational resources. These workflows are used to generate GeneLab processed data hosted on the OSDR, and are publicly available as open source software for others to use at: https://github.com/nasa/GeneLab_Data_Processing.

Jonathan Oribello

Open Science for Plants in Space: Data Sharing, Standards, and Informatics for Reuse and Knowledge Discovery

Upcoming deep space missions will rely on plants for crew and ecosystem health. Open access space biology data enables scientists to examine the biological responses of plants to ionizing radiation, altered gravity, low atmospheric pressure, elevated CO2, altered photoperiods and many other abiotic stressors. Open Science is the practice of making research available to all, while respecting diverse cultures, to foster collaborations with equity. NASA has declared 2023 as the ‘Year of Open Science’ and created a 5-year Transform to Open Science (TOPS) initiative designed to rapidly transform the agency toward an inclusive culture of open science. NASA’s Open Science Data Repository (OSDR) within the Biological and Physical Sciences Division provides access to data from space-relevant biological experiments. OSDR combines two databases, GeneLab and Ames Life Sciences Data Archive (ALSDA) to maximize access to standardized ‘omics (e.g., transcriptomics, proteomics) and phenotypic data (e.g., microscopy, biomass), respectively. GeneLab started in 2014 with the creation of the first space-relevant FAIR (Findable, Accessible, Interoperable, Reusable) biological ‘omics repository, providing detailed metadata on investigation, sample, and assay levels. The addition of ALSDA to OSDR expands plant data analysis capabilities across both phenotypic and ‘omics data. Today, OSDR hosts 62+ plant datasets and has enabled 58 peer-reviewed publications. Most of these publications were collaboration efforts under the OSDR Analysis Working Groups (AWGs). AWGs provide great opportunities for investigators to collaborate with community members and set new standards for space-relevant data and metadata. The AWGs welcome any ASGSR members interested in contributing plant expertise for space biology, and to serve as subject matter experts as we establish the framework for modern plant data archiving. Investigators are encouraged to submit their space-relevant plant datasets to OSDR and visit the site to learn about the tools OSDR has to offer (osdr.nasa.gov/bio).

FAIR

Enabling Model Organism and Commercial Astronaut Data Access Through the NASA Open Science Data Repository

NASA’s Open Science Data Repository (OSDR) brings together omics data from NASA’s GeneLab project and non-omics data, including physiological, phenotypic, imaging, and behavioral data from NASA’s Ames Life Sciences Data Archive (ALSDA) collected from decades of space biology research, providing open and FAIR (findable, accessible, interoperable, and reusable) access of these precious data to scientists world-wide. This rich source of meticulously curated metadata and data from spaceflight and analog studies has been mined by the scientific community resulting in dozens of high impact scientific publications that reveals a complex network of molecular and physiological effects of spaceflight across living systems, from microbes to plants, to mammals. Understanding how these effects translate to the human condition is critical as we move deeper into the era of commercial space travel. However, the integration of data, specifically omics data, from astronauts is particularly challenging due to their sensitive nature. OSDR has risen to this challenge by developing a mechanism to control access to identifiable levels of omics data, such as raw sequence data, while enabling public access to processed, unidentifiable, data and associated metadata that will allow the scientific community to interrogate human astronaut data alongside data from model organisms to begin answering these critical questions. The 2021 SpaceX Inspiration4 (I4) mission collected a comprehensive atlas of biological measurements from four civilian astronauts, providing a wealth of data to characterize the effects of spaceflight on the human body. These data include both non-omics and omics assays such as direct RNA sequencing (RNA-seq), single nuclei ATAC-seq and RNA-seq, metagenomics, proteomics, and comprehensive metabolic and cytokine panels, all of which have been integrated into the OSDR system across no less than 9 studies. Each study has been carefully curated using community-backed OSDR standards for sample and assay level metadata ensuring these data are findable and accessible. In addition to hosting both raw and processed data from the principal investigator team for each assay type, the GeneLab team plans to re-process the I4 omics data using GeneLab’s standard processing pipelines. The GeneLab processed data outputs will allow for comparisons across studies on OSDR and enable visualization of these data through the OSDR data visualization platform thereby enabling data reusability and interoperability. Here we describe the robust privacy and security protocols implemented by OSDR to safeguard sensitive health data from astronauts while facilitating metadata and processed data sharing for research purposes. We further provide a road map for navigating the vast amount of data provided for each I4 study on the OSDR, including experimental design, associated experiments, payloads, and missions, data generation and analysis protocols, and associated scientific articles. Additionally, we illustrate how to interrogate the standardized metadata provided in the sample and assay tables as well as instructions for how to download and access the data. The I4 datasets described here re present the first ever comprehensive collection of commercial astronaut data.

Amanda M Saravia-Butler

Tracking Community Building in Open Science

Open Science is enabled by a vibrant community of researchers who regularly engage with the data, from its production to its organization, curation, archiving, dissemination, analysis, and publication. This presentation will examine community building in open science. The NASA Open Science Data Repository (OSDR) makes data available to the public following the FAIR (Findability, Accessibility, Interoperability, and Reusability) principles. OSDR takes open science further with the OS Analysis Working Groups (AWGs) that facilitate community development and promotion. The primary activity of each AWG is to establish and validate analytical processes to generate higher-order data from data housed in OSDR. There are a number of these groups on various topics, including the Animal AWG, Plant AWG, Microbial AWG, Multi-Omics AWG, AI/ML AWG, and the Ames Life Sciences Data Archive (ALSDA) AWG. The international volunteers participating in these AWGs come from academia, citizen science initiatives, industry, and government. They include researchers, principal investigators, professors, trained hobbyists, and students from various domains and disciplines. Anyone may request to join the AWGs, and membership requests are vetted monthly by the group organizers before granting admission. Core to membership is demonstrated expertise through records of training, integrity, work in the professed domain(s), and good community standing. Regular virtual meetings are held for each AWG, with a varying cadence depending on the group's needs and goals. AWG communities share their expertise in research including cutting edge tools, software, frameworks, data formats, and libraries accelerating research collectively. This collaborative approach helps community members cross technology gaps and identify emerging challenges. These diverse communities encompass a wide range of individuals hailing from various sectors within the Science Mission Directorate and beyond. They serve as a means to promote and enhance transparency, accessibility, and inclusion. An annual AWG Symposium brings contributors together in person. Participation in AWGs can be synchronous or asynchronous, with some groups performing most of their work in off hours. Participants gain valuable skills and connections that allow them to add value to their communities and new organizations that they join, resulting in an expanded return on investment for the space life science community. Open science is increasingly a federal mandate and initiatives like NASA's Transform to Open Science and instruments like the Decadal Survey of Biological and Physical Sciences in Space demonstrate the need to carefully consider best practices in this domain. Here, we present greater detail about the makeup and participation metrics of the various AWGs affiliated with OSDR and details of successful peer-reviewed publication campaigns.

Christina M Johnson

Feature Selection in High-Dimensional Space with Applications to Gene Expression Data

Recent years have seen rapid growth in high-dimensional datasets. Most existing machine learning (ML) algorithms fail in high-dimensional settings where many features could be redundant. A critical process of feature selection is thus applied in such a setting that helps in identifying the most relevant features while removing redundant ones. With the increase in high dimensionality, one is also faced with problems of efficiency and interpretation in performing such selection methods. Therefore, this paper proposes a “novel” feature selection framework that uses an ensemble of interpretable ML algorithms to perform feature selection and the ranking of final features. Finally, this framework is applied to a gene expression dataset obtained through collaboration with the National Aeronautics and Space Administration (NASA)’s Biological and Physical Sciences (BPS) team and helps identify important and relevant genes contributing to specific target attributes through classification tasks.

Nishan Pantha

RadLab: A Comprehensive Database and Graphical and Programming Interfaces for Space Radiation Data

RadLab, a component of the NASA Open Science Data Repository (OSDR), is a database of radiation measurements from multiple instruments and spacecraft that provides visual and programmatic interfaces for interrogation and retrieval of these data. The attributes of data available through RadLab include spacecraft, types of radiation sensing instruments, locations within the spacecraft (e.g. ISS modules), associated celestial bodies, trajectories, and spacecraft coordinates; the primary type of data is the absorbed dose rate, as well as flux and dose equivalent rate where available. The application programming interface (API) implements a request syntax for retrieval of timestamped data filtered by various combinations of such attributes; the graphical user interface (GUI) extends this functionality with visualizations (time series plots, comparison plots, geospatial visualizations) which provide easy means to assess data availability, iteratively refine search parameters, interactively inspect the data, and export target data subsets. Datasets are continuously being added to the RadLab database as part of the rolling release process. Investigators from multiple countries, including the US, Canada, Germany, Bulgaria, Hungary, Italy, Japan, Russia and the Czech Republic, have committed to provide data from their instruments in and beyond low Earth orbit. The current release contains datasets provided by US and international collaborators and includes readings from multiple modules of the ISS, the BioSentinel CubeSat, Chang’e 4, the Lunar Reconnaissance Orbiter, the ExoMars Orbiter, and the Curiosity rover. Datasets are associated with respective RadLab knowledgebase articles which include instrument descriptions and provide bibliographical references. RadLab aims to provide a comprehensive, dynamic compendium of space radiation data, enabling the scientific community to perform analyses of data from multiple detectors and to determine the radiation environment of research missions and experiments. Some of its applications include inference of absorbed radiation dose for NASA GeneLab payloads, and training predictive models as part of the 2024 FDL-X challenge. The platform is actively expanding and seeking additional data, with plans to also cover past (e.g. Shuttle, Mir) and future (e.g. Artemis) missions. The RadLab Working Group has been created to aid in this process as well as to foster collaborations among data contributors and users, to develop standards for data harmonization, and to guide the development of the platform, with the goal to establish the use of RadLab in space radiation research and to advance our understanding of the radiation environment in outer space.

Kirill Grigorev

RadLab: A Comprehensive Database and Graphical and Programming Interfaces for Biologically Relevant Space Radiation Data

RadLab, a new component of the NASA Open Science Data Repository (OSDR), comprises a database of radiation measurements relevant to space biology, and visual and programmatic interfaces for interrogation and retrieval of these data. The attributes of data available through RadLab include spacecraft, types of radiation sensing instruments, locations within the spacecraft (e.g. modules of the ISS), associated celestial bodies, trajectories, and spacecraft coordinates. The application programming interface (API) implements a request syntax for retrieval of timestamped data filtered by various combinations of such attributes; the graphical user interface (GUI) extends this functionality with visualizations, such as spacecraft schematics, time series plots, geospatial visualizations, and provides easy means to iteratively refine search parameters, inspect the data on the fly, and download target subsets. The release of RadLab currently available to the public contains datasets provided by US and international collaborators and focuses on data recorded on the ISS. Investigators from multiple countries, including the US, Canada, Germany, Bulgaria, Hungary, Italy, Japan, Russia and the Czech Republic, have committed to provide data from their instruments in and beyond low Earth orbit; RadLab will also soon expand to include past (e.g. Shuttle and Mir) and future (e.g. Artemis) data. RadLab will provide a comprehensive, dynamic compendium of space radiation data, enabling the scientific community to perform analyses of data from multiple detectors and to determine the radiation environment of research missions and experiments. The RadLab Working Group has been formed to foster collaborations among data contributors and users, to identify data sources, to put in place standards for data harmonization, and to guide the development of the platform, with the goal to establish the use of RadLab in space radiation research and to advance our understanding of the space radiation environment in human habitats.

database

Beyond Fair: Engagement, Data Usability, and Open Community Productivity through the NASA Open Science Data Repository

The FAIR principle (findable, accessible, interoperable, and reusable) governs the storage and sharing of NASA space biology and health data[1]. These guiding principles maximize reuse of data and the reproducibility of scientific findings. The NASA Open Science Data Repository (OSDR; an expansion of NASA GeneLab) was built on the FAIR principles and houses over 500 studies and close to 1000 datasets from decades of space life sciences experiments. OSDR embodies the FAIR principles through data governance that includes mediated, embargoed, and fully open access data. The FAIR data governance principles were recently proposed to be expanded to encompass a FAIREST framework for assessing research data repositories (FAIR + Engagement, Social connections, and Trust)[2]. FAIREST emphasizes the importance of data repositories engaging with the scientific community and gaining the trust of researchers regarding data quality. Trust also refers to the TRUST principles developed for assessment of digital repositories: Transparency, Responsibility, User Focus, Sustainability, Technology[3]. We present the “Open Science for Life in Space” Analysis Working Groups (AWGs) as evidence regarding the power of engagement, social connections, and trust which has enhanced OSDR’s capabilities and productivity. AWG members engage in two main activities. One, members provide feedback on OSDR scientific standards for data ingestion, curation, and reuse (study, subject and assay metadata; processing pipelines; dataset formats and uniformed structures for machine-readability). Two, AWG members collaborate to mine-reuse OSDR data to conduct scientific analysis. With nearly 800 active members, the AWGs have resulted in 32 publications re-using OSDR data and contributed many papers in two major special issues in Cell (2020) and Nature (2024). AWGs also serve as networking groups, facilitate social connections between researchers at all levels of experience, and also have a social online ‘Forum’ used to keep members informed on projects and opportunities. This community-centric, productive, and trustworthy data culture has resulted in a broader effect with international space agencies, academics, and the commercial space sector wanting to submit their data to OSDR. Ten studies of Inspiration 4 data were recently publicly released by OSDR, as were some JAXA human data. Coming up soon in OSDR are data submissions from the European Space Agency, Virgin Galactic PIs, and SpaceX Polaris Dawn. A major benefit of OSDR is the array of standardized and uniformly formatted data (which was developed through AWG member consensus), from which visualization tools, analysis tools, and machine learning models can be built or trained. This talk will cover the Multi-Study Visualization Tool, the Environmental Data Application, RadLab, and a UCSF-NSF funded knowledge graph biomedical health discovery tool ‘SPOKE’ currently being integrated with OSDR. OSDR also provides training programs in bioinformatics and machine learning to improve the scientific community’s awareness of data availability and to boost their ability to perform data analysis. The increasing engagement of the scientific community and the public with technologies powered by artificial intelligence (AI) heightens the need for data analysis to be transparent. The AI for Life in Space initiative leverages the data products provided in OSDR to train AI models, with an emphasis on explainable and trustworthy AI, which would not be possible without FAIR data and metadata. Overall, here we will demonstrate the importance for NASA life sciences data repositories to adhere to the FAIREST framework, by providing examples and success stories from different aspects of OSDR.

data

RadLab: A Comprehensive Database and Graphical and Programming Interfaces for Space Radiation Data

RadLab, a component of the NASA Open Science Data Repository (OSDR), is a database of radiation measurements from multiple instruments and spacecraft that provides visual and programmatic interfaces for interrogation and retrieval of these data. The attributes of data available through RadLab include spacecraft, types of radiation sensing instruments, locations within the spacecraft (e.g. ISS modules), associated celestial bodies, trajectories, and spacecraft coordinates; the primary type of data is the absorbed dose rate, as well as flux and dose equivalent rate where available. The application programming interface (API) implements a request syntax for retrieval of timestamped data filtered by various combinations of such attributes; the graphical user interface (GUI) extends this functionality with visualizations (time series plots, comparison plots, geospatial visualizations) which provide easy means to assess data availability, iteratively refine search parameters, interactively inspect the data, and export target data subsets. Datasets are continuously being added to the RadLab database as part of the rolling release process. Investigators from multiple countries, including the US, Canada, Germany, Bulgaria, Hungary, Italy, Japan, Russia and the Czech Republic, have committed to provide data from their instruments in and beyond low Earth orbit. The current release contains datasets provided by US and international collaborators and includes readings from multiple modules of the ISS, the BioSentinel CubeSat, Chang’e 4, the Lunar Reconnaissance Orbiter, the ExoMars Orbiter, and the Curiosity rover. Datasets are associated with respective RadLab knowledgebase articles which include instrument descriptions and provide bibliographical references. RadLab aims to provide a comprehensive, dynamic compendium of space radiation data, enabling the scientific community to perform analyses of data from multiple detectors and to determine the radiation environment of research missions and experiments. Some of its applications include inference of absorbed radiation dose for NASA GeneLab payloads, and training predictive models as part of the 2024 FDL-X challenge. The platform is actively expanding and seeking additional data, with plans to also cover past (e.g. Shuttle, Mir) and future (e.g. Artemis) missions. The RadLab Working Group has been created to aid in this process as well as to foster collaborations among data contributors and users, to develop standards for data harmonization, and to guide the development of the platform, with the goal to establish the use of RadLab in space radiation research and to advance our understanding of the radiation environment in outer space.

Kirill Grigorev