Search NASA⌕ Search

SEARCH · Search NASA

Results for “Omics, GeneLab”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 127 records · Page 7

GeneLab: NASA's Open Access, Collaborative Platform for Systems Biology and Space Medicine

NASA is investing in GeneLab1 (http:genelab.nasa.gov), a multi-year effort to maximize utilization of the limited resources to conduct biological and medical research in space, principally aboard the International Space Station (ISS). High-throughput genomic, transcriptomic, proteomic or other omics analyses from experiments conducted on the ISS will be stored in the GeneLab Data Systems (GLDS), an open-science information system that will also include a biocomputation platform with collaborative science capabilities, to enable the discovery and validation of molecular networks.

Berrios, Daniel C.↗

GeneLab: NASA's Open Access, Collaborative Platform for Systems Biology and Space Medicine

NASA is investing in GeneLab1 (http:genelab.nasa.gov), a multi-year effort to maximize utilization of the limited resources to conduct biological and medical research in space, principally aboard the International Space Station (ISS). High-throughput genomic, transcriptomic, proteomic or other omics analyses from experiments conducted on the ISS will be stored in the GeneLab Data Systems (GLDS), an open-science information system that will also include a biocomputation platform with collaborative science capabilities, to enable the discovery and validation of molecular networks.

Berrios, Daniel C.↗

Cross Kingdom Analysis of Data Within the GeneLab Repository Identifies a Potential Conserved Response of Life to the Stress Associated with Spaceflight

It is important to determine the health risks and potential survival for astronauts associated with long-term space missions. This entails not only understanding the impact the space environment will have on humans, but also how it will affect other organisms needed for humans to survive in space such as plants. In addition, it has been reported in the literature that hundreds of genes seem to be conserved and/or transferred between different organisms from bacteria, archaea, fungi, microorganisms, and plants to animals. Since space travel involves humans in a closed environment over a long period of time, we hypothesize that potential conserved biological factors will occur between the different organisms in that environment possibly due to transfer of genes. Determining the conserved factors that are commonly being regulated in space can shed insight into possible universal master regulators and also determine the symbiotic relationship between the organisms in space. Utilizing NASA's GeneLab Data Repository (a rapidly expanding, curated clustering of spaceflight-related ‘omics-level datasets for all organisms), we were able to uncover a novel pathway and factors that were commonly shared between humans, mice, plants, C. Elegans, and drosophilas. Through ChIP-Seq enrichment analysis techniques utilizing various GeneLab datasets from each species that were flown in space, we found the following factors to be conserved across all species: oxidative stress, DNA damage (through GABPA/NRFs and NFY), SIX5, GTF2B and glutamine synthetase. Such commonalities would likely reflect the effects of factors such as microgravity and the increased radiation exposure inherent in spaceflight on basic physical processes shared by all biological systems at the cellular level. Differences between organismal responses revealed by GeneLab's data should also help understand the unique reactions to life in space that arise from the very different lifestyles of microbes, animals and plants.

Barker, Richard↗

Systemic Microgravity Response: Utilizing GeneLab to Develop Hypotheses for Spaceflight Risks

Biological risks associated with microgravity are a major concern for long-term space travel. Although determination of risk has been a focus for NASA research, data examining systemic (i.e., multi- or pan-tissue) responses to space flight are sparse. To perform our analysis, we utilized the NASA GeneLab database which is a publicly available repository containing a wide array of omics results from experiments conducted with: i) with different flight conditions (space shuttle (STS) missions vs. International Space Station (ISS); ii) a variety of tissues; and 3) assays that measure epigenetic, transcriptional, and protein expression changes. Meta-analysis of the transcriptomic data from 7 different murine and rat data sets, examining tissues such as liver, kidney, adrenal gland, thymus, mammary gland, skin, and skeletal muscle (soleus, extensor digitorum longus, tibialis anterior, quadriceps, and gastrocnemius) revealed for the first time, the existence of potential master regulators coordinating systemic responses to microgravity in rodents. We identified p53, TGF1 and immune related pathways as the highly prevalent pan-tissue signaling pathways that are affected by microgravity. Some variability in the degree of change in their expression across species, strain and time of flight was also observed. Interestingly, while certain skeletal muscle (gastrocnemius and soleus) exhibited an overall down-regulation of these genes, some other muscle types such as the extensor digitorum longus, tibialis anterior and quadriceps, showed an up-regulated expression, indicative of potential compensatory mechanisms to prevent microgravity-induced atrophy. Key genes isolated by unbiased systems analyses displayed a major overlap between tissue types and flight conditions and established TGF1 to be the most connected gene across all data sets. Finally, a set of microgravity responsive miRNA signature was identified and based on their predicted functional state and subsequent impact on health, a theoretical health risk score was calculated. The genes and miRNAs identified from our analyses can be targeted for future research involving efficient countermeasure design. Our study thus exemplifies the utility of GeneLab data repository to aid in the process of performing novel hypothesis based spaceflight research aimed at elucidating the global impact of environmental stressors at multiple biological scales.

GeneLab↗

Systemic Response to Microgravity: Utilizing GeneLab Datasets to Identify Molecular Targets for Future Hypotheses-Driven Spaceflight Studies

Biological risks associated with microgravity are a major concern for long-term space travel. Although determination of risk has been a focus for NASA research, data examining systemic (i.e., multi- or pan-tissue) responses to space flight are sparse. To perform our analysis, we utilized the NASA GeneLab database which is a publicly available repository containing a wide array of omics results from experiments conducted with: i) with different flight conditions (space shuttle (STS) missions vs. International Space Station (ISS); ii) a variety of tissues; and 3) assays that measure epigenetic, transcriptional, and protein expression changes. Meta-analysis of the transcriptomic data from 7 different murine and rat data sets, examining tissues such as liver, kidney, adrenal gland, thymus, mammary gland, skin, and skeletal muscle (soleus, extensor digitorum longus, tibialis anterior, quadriceps, and gastrocnemius) revealed for the first time, the existence of potential master regulators coordinating systemic responses to microgravity in rodents. We identified p53, TGF(beta)1 and immune related pathways as the highly prevalent pan-tissue signaling pathways that are affected by microgravity. Some variability in the degree of change in their expression across species, strain and time of flight was also observed. Interestingly, while certain skeletal muscle (gastrocnemius and soleus) exhibited an overall down-regulation of these genes, some other muscle types such as the extensor digitorum longus, tibialis anterior and quadriceps, showed an up-regulated expression, indicative of potential compensatory mechanisms to prevent microgravity-induced atrophy. Key genes isolated by unbiased systems analyses displayed a major overlap between tissue types and flight conditions and established TGF(beta)1 to be the most connected gene across all data sets. Finally, a set of microgravity responsive miRNA signature was identified and based on their predicted functional state and subsequent impact on health, a theoretical health risk score was calculated. The genes and miRNAs identified from our analyses can be targeted for future research involving efficient countermeasure design. Our study thus exemplifies the utility of GeneLab data repository to aid in the process of performing novel hypothesis based spaceflight research aimed at elucidating the global impact of environmental stressors at multiple biological scales.

GeneLab↗

The NASA GeneLab Project

"NASA's GeneLab Project has evolved from being a simple repository that hosts multi-omics datasets generated from spaceflight experiments, to a complete solution for analysis and visualization of spaceflight related omics. We will show how Omics can help elucidate the impact of spaceflight factors (e.g. CO2) on organisms, tissues and cells."

OMICS↗

Open Science for Plants in Space: Improvements in NASA's Open Science Data Repository

Upcoming deep space missions will rely on plants for crew and ecosystem health. Open access space biology data enables scientists to examine the biological responses of plants to ionizing radiation, altered gravity, elevated CO2, and many other abiotic stressors. NASA has declared 2023 as the ‘Year of Open Science’ and created a 5-year Transform to Open Science (TOPS) initiative designed to rapidly transform the agency toward an inclusive culture of open science. NASA’s Open Science Data Repository (OSDR) combines two databases, GeneLab and Ames Life Sciences Data Archive (ALSDA) to maximize access to standardized ‘omics (e.g., transcriptomics, proteomics) and phenotypic data (e.g., microscopy, biomass), respectively. Current OSDR standards include the ISA (Investigation-Study-Assay) experiment model, assay metadata configurations, and standardized terminology and ontologies. In 2024 OSDR will include a new suite of features for improved FAIR compliance including downloadable plant metadata templates, data submission tools and overall improved AI-readiness of plant datasets. AI/ML methods can be helpful tools to overcome the inherent challenges of space biology research (small sample size, sparse and heterogeneous data etc.). However these methods are built on an assumption of normalized and well-curated data. OSDR’s new curation tools will improve users ability to leverage ML and AI methods to model space biology data and better understand the complex effects of spaceflight on living systems across hierarchical biological levels. We look forward to sharing our advances with the spaceflight community.

FAIR↗

The GeneLab Buffet: A Bioinformatic MATRIX of MANGO and TOAST

The GeneLab data repository provides an unparalleled resource for exploring how spaceflight affects organisms with omics-level insights. However, two major interlinked challenges to capitalizing on the information within these data are their vast breadth and the often-specialized expertise that has been required in the past for their analysis. How do you compare responses within and between studies, especially if you are a non-bioinformatics specialist? This presentation will discuss how Space Biology data can be accessed using software to help provide these data resources to address research questions and generate new hypotheses. The presentation will cover a wide range of the available space life science tools but will focus on TOAST, MANGO, the MATRIX, RadBioApp and other interactive relational databases (https://genelab.nasa.gov/external-vis-apps). These exploration environments have been developed to search the GeneLab data repository for new insights that inform how model organisms respond to microgravity, radiation and other factors associated with spaceflight. The presentation will be interactive, and participants will have the opportunity to ask questions and learn more about the data viz and modeling tools that are available to them.

AstroBotany↗

Increasing the Statistical Rigor of Cross-Species Differential Expression Analysis

Microgravity inflicts substantial, but undercharacterized, pressure on organisms that induces metabolic responses such as increased microbial virulence and antibiotic resistance, altered organ weights in developing rats, and loss of bone tissue in astronauts. Numerous studies have analyzed the effects of microgravity on specific organisms, tissues, or test conditions, but these projects are necessarily limited by the small sample size of space research. Increasing the sample size of spaceflight studies is non-trivial; however, pooling data from numerous studies can greatly increase the statistical rigor of comparative analyses. The GeneLab houses datasets from 73 spaceflight studies that performed transcription profiling assays. These data encompass a diverse array of organisms ranging from Escherichia coli to Mus musculus to Homo sapiens and comprise studies analyzing ionizing radiation, mammalian pregnancy, etc. Collectively, the GeneLab database contains a large quantity of transcription assays and RNA sequence data analyzing Differential Gene Expression (DGE) between microand normogravity. Xspecies, a cross-species analysis method for DGE developed by Kristiansson, et al. in 2012, identifies homologous genes between species that are universally up- or downregulated in response to test conditions. Previous work by an intern at GeneLab applied Xspecies to 19 datasets containing seven different species and identified 14 homologous groups differentially expressed under spaceflight conditions including several heat shock proteins and cytoskeletal components. Unfortunately, these results may be biased by the disproportionate number of studies on Arabidopsis thaliana (5) and Mus musculus (6) and the results are not normalized by evolutionary distances. Here, we present modifications to the Xspecies algorithm that permits incorporation of multi-omic data and normalizes data for effect size, directionality, and evolutionary distances. We then apply this algorithm to all currently available GeneLab studies

Xspecies↗

NASA's GeneLab Phase II: Federated Search and Data Discovery

GeneLab is currently being developed by NASA to accelerate 'open science' biomedical research in support of the human exploration of space and the improvement of life on earth. Phase I of the four-phase GeneLab Data Systems (GLDS) project emphasized capabilities for submission, curation, search, and retrieval of genomics, transcriptomics and proteomics ('omics') data from biomedical research of space environments. The focus of development of the GLDS for Phase II has been federated data search for and retrieval of these kinds of data across other open-access systems, so that users are able to conduct biological meta-investigations using data from a variety of sources. Such meta-investigations are key to corroborating findings from many kinds of assays and translating them into systems biology knowledge and, eventually, therapeutics.

exobiology↗

NASAs GeneLab Phase II: Federated Search and Data Discovery

GeneLab is currently being developed by NASA to accelerate open science biomedical research in support of the human exploration of space and the improvement of life on earth. Phase I of the four-phase GeneLab Data Systems (GLDS) project emphasized capabilities for submission, curation, search, and retrieval of genomics, transcriptomics and proteomics (omics) data from biomedical research of space environments. The focus of development of the GLDS for Phase II has been federated data search for and retrieval of these kinds of data across other open-access systems, so that users are able to conduct biological meta-investigations using data from a variety of sources. Such meta-investigations are key to corroborating findings from many kinds of assays and translating them into systems biology knowledge and, eventually, therapeutics.

genome↗

Expanding Biological Repository Data Available for Sharing and Knowledge Discovery

Biology has developed next-generation data science and alternative analytical approaches with methodologies which require principal investigator (PI) experimental assay data be re-used. This new approach involves mining multiple datasets at once from various hierarchical organizations of biological complexity, while concurrently evaluating how experimental factors affect endpoints of standard assays. The purpose of the NASA Ames Life Sciences Data Archive (ALSDA) is to collect, curate, and make findable, accessible, interoperable, and reusable (FAIR) all non-human space-relevant biological data. These data include mission metadata, subject metadata, assay metadata (parameters), raw and processed assay data, assay imagery, and subject-experienced telemetry (radiation, temperature, humidity, acoustics, vibrations). ALSDA has transformed to bring current biological repository data and all future collected data into this new scientific data mining reality. It has integrated into the ‘NASA Open Science’ group of projects to facilitate a suite of new tools and workflows to improve data accessibility and reusability by implementing data management plans, automating data submission agreements, and adopting the single-point-of-entry data submission portal, originally developed by NASA GeneLab. These systems required ALSDA to develop science assay configurations for the submission portal, capturing essential assay parameters according to established norms in each sub-field within biology. The submission portal expedites data collection by enhancing ease of PI data submission, providing a user interface and specificity for which data is to be submitted. ALSDA datasets are curated to maintain rich metadata, accuracy of datasets, data transparency, provenance, and additionally ensure data are machine-readable (e.g., R and Python languages). ALSDA integration with GeneLab and its analysis portals enable higher-order physiological-level datasets be mined in conjunction with -omics datasets. As ALSDA physiological-level datasets are published (micro-computed tomography, histology, intraocular pressure, hormonal assays, immunostaining, ultrasonography), the merging of hierarchical organizations of biological complexity from spaceflight will enable new knowledge discovery approaches.

Ryan T Scott↗

Machine Learning Approaches to Increasing Value of Spaceflight Omics Databases

The number of spaceflight bioscience mission opportunities is too small to allow all relevant biological and environmental parameters to be experimentally identified. Simulated spaceflight experiments in ground-based facilities (GBFs), such as clinostats, are each suitable only for particular investigations -- a rotating-wall vessel may be 'simulated microgravity' for cell differentiation (hours), but not DNA repair (seconds) -- and introduce confounding stimuli, such as motor vibration and fluid shear effects. This uncertainty over which biological mechanisms respond to a given form of simulated space radiation or gravity, as well as its side effects, limits our ability to baseline spaceflight data and validate mission science. Machine learning techniques autonomously identify relevant and interdependent factors in a data set given the set of desired metrics to be evaluated: to automatically identify related studies, compare data from related studies, or determine linkages between types of data in the same study. System-of-systems (SoS) machine learning models have the ability to deal with both sparse and heterogeneous data, such as that provided by the small and diverse number of space biosciences flight missions; however, they require appropriate user-defined metrics for any given data set. Although machine learning in bioinformatics is rapidly expanding, the need to combine spaceflight/GBF mission parameters with omics data is unique. This work characterizes the basic requirements for implementing the SoS approach through the System Map (SM) technique, a composite of a dynamic Bayesian network and Gaussian mixture model, in real-world repositories such as the GeneLab Data System and Life Sciences Data Archive. The three primary steps are metadata management for experimental description using open-source ontologies, defining similarity and consistency metrics, and generating testing and validation data sets. Such approaches to spaceflight and GBF omics data may soon enable unique insight into which measured phenomena correlate to biological mechanisms that are truly affected by spaceflight conditions; which are most likely to be confounded by other variables; and which are insufficiently characterized, significantly increasing existing and future science return from ISS and spaceflight missions.

Gentry, Diana↗

The NASA Open Science Data Repository: Biomedical Data, Analysis Tools, and Informatic Collaborations

Increased biomedical risks and challenges associated with deep space missions require knowledge discovery, health countermeasures, and biomedical support capabilities. Maximally open-access and reusable data is needed by developers, scientists, and engineers to develop these systems. The NASA Open Science Data Repository (OSDR) is a maximally open access and FAIR database (ie., findable, accessible, interoperable, and reusable), and meets various scientific, technical, and operational needs. It offers users and submitters the ability to upload, download, search, share, analyze, cite, and visualize data across ‘omics, physiological, phenotypic, payload, hardware, behavioral, bioimaging, video, and environmental monitoring telemetry datasets. OSDR is an expanded database, based upon the successes of NASA GeneLab. OSDR has >460 studies with datasets covering model organisms to non-NASA human astronauts. There are ~12 datasets from the Inspiration 4 (I4) mission, spanning metagenomics, comprehensive metabolic panels, clonal hematopoiesis, spatial transcriptomics, proteomics, and cytokine panels. In the interest of data privacy, two I4 datasets with raw files relating to the epitranscriptome, and a new request feature is live in OSDR (with a backend review process established) developed from industry norms. OSDR is collecting and curating biomedical human data from a new sub-orbital research flight and is open to more space life science/biomedical submissions from the international and commercial sectors. OSDR also recently began a collaboration with the European Space Agency (ESA) to collect and curate >200 terabytes of human and model organism data. The OSDR submission portal is designed to ingest and curate ~25 ‘omics and ~50 physiological-phenotypic-imaging assay data types. Tools available for OSDR users include: 1) an Environmental Data Application to compare radiation, CO2, relative humidity, temperature, and other telemetry across missions and subjects, 2) the RadLab database, a collaboration between NASA, ESA, the German and Italian Space Agencies, and the Bulgarian Academy of Sciences, and 3) a Multi-study visualization tool which enables users to look across and combine ‘omics datasets. There are ~600 volunteer OSDR Analysis Working Group (AWG) members providing feedback on scientific data/metadata standards and collaborating to mine-reuse OSDR in research. OSDR/GeneLab has enabled ~60 publications reusing data as of October 2023.

space biology↗

Evaluation of Correction Methods for NASA GeneLab Transcriptomic Datasets

Conducting space biology experiments aboard the International Space Station, particularly those utilizing complex model organisms like mice, is expensive and difficult due to limited crew availability, hardware, and space. As a result, sample numbers from these studies are low, reducing the statistical power of any one experiment. Aggregating spaceflight datasets serves as a method to increase sample numbers, allowing for novel insights through bioinformatic analysis of ‘omics data from merged datasets. However, aggregating datasets can introduce unwanted variation including 1) differences in sample handling, processing, and sequencing platforms between datasets (technical variation) as well as 2) differences in experimental design between datasets. In the present study, NASA GeneLab-hosted RNAseq datasets from mouse liver tissues were used to evaluate several statistical methods to correct for this unwanted variation through two approaches, reference-based and standard. The following correction algorithms were applied with (reference-based) and/or without (standard) considering Universal Mouse RNA Reference samples: ComBat and ComBat_seq from the SVA package, median polish, empirical Bayes, and ANOVA-based algorithms from the MBatch package, and negative binomial regression normalization in the DESeq2 package. For each approach, after the correction algorithm was applied, differential gene expression (DGE) analysis of flight and ground control samples was performed with the combined data. The robustness of each tool was evaluated using BatchQC to determine statistical differences between datasets before and after correction, Principal Component Analysis to evaluate global gene expression in samples before and after correction, and by comparing DGE analysis of individual datasets and combined datasets before and after correction. The results showed that the reference-based approach introduced several additional (and likely artificial) DEGs when compared with the respective standard approach. Of the methods tested, standard ComBat and DESeq2 were identified as the most robust correction methods for combining spaceflight mouse liver RNAseq datasets hosted on GeneLab.

GeneLab↗

NASA GeneLab Concept of Operations

NASA's GeneLab aims to greatly increase the number of scientists that are using data from space biology investigations on board ISS, emphasizing a systems biology approach to the science. When completed, GeneLab will provide the integrated software and hardware infrastructure, analytical tools and reference datasets for an assortment of model organisms. GeneLab will also provide an environment for scientists to collaborate thereby increasing the possibility for data to be reused for future experimentation. To maximize the value of data from life science experiments performed in space and to make the most advantageous use of the remaining ISS research window, GeneLab will apply an open access approach to conducting spaceflight experiments by generating, and sharing the datasets derived from these biological studies in space.Onboard the ISS, a wide variety of model organisms will be studied and returned to Earth for analysis. Laboratories on the ground will analyze these samples and provide genomic, transcriptomic, metabolomic and proteomic data. Upon receipt, NASA will conduct data quality control tasks and format raw data returned from the omics centers into standardized, annotated information sets that can be readily searched and linked to spaceflight metadata. Once prepared, the biological datasets, as well as any analysis completed, will be made public through the GeneLab Space Bioinformatics System webb as edportal. These efforts will support a collaborative research environment for spaceflight studies that will closely resemble environments created by the Department of Energy (DOE), National Center for Biotechnology Information (NCBI), and other institutions in additional areas of study, such as cancer and environmental biology. The results will allow for comparative analyses that will help scientists around the world take a major leap forward in understanding the effect of microgravity, radiation, and other aspects of the space environment on model organisms. These efforts will speed the process of scientific sharing, iteration, and discovery.

Space Life Science↗

GL4U: GeneLab for Colleges and Universities

GeneLab for Colleges and Universities (GL4U) will provide space biology-relevant training in bioinformatics to the next generation of scientists through direct and indirect approaches. The GeneLab (GL) team will host two annual data processing bootcamps, one for college-level students (direct) and one for college educators (indirect – Training of Trainers), in which participants learn to analyze space-relevant omics data hosted on GL. The first bootcamp took place in early June 2021 with about 30 SJSU undergraduate students and covered space biology-specific lectures and hands-on instruction using Jupyter Notebooks (JNs) for RNA sequence (RNAseq) data analysis. All training materials including the enclosed files listed below will be made publicly available on GitHub. RNAseq Bootcamp Lectures (attached in combined file): Introduction to NASA, Space Biology, GeneLab, and the Command Line: NASA_GL_CL_Intro_FINAL.pdf - DRAFT from initial submission NASA_SB_GL_CL_Intro_FULL.pdf - FINAL version presented during the bootcamp - only minor edits from the draft version RNAseq and Data Processing Overview: RNAseq_Overview_FINAL.pdf - DRAFT from initial submission RNAseq_Overview_FULL.pdf - FINAL version presented during the bootcamp - only minor edits from the draft version Overview of the Statistics Used for RNAseq Data Analysis: SJSU_Statistics_Intro_Lecture_FINAL.pdf - DRAFT from initial submission Statistics_Overview_FULL.pdf - FINAL version presented during the bootcamp - only minor edits from the draft version Completed JNs in HTML format (attached in combined file): Unix_Intro_JN_06-2021_completed.html R_Intro_JN_06-2021_completed.html RNAseq_fastq_to_counts_JN_06-2021_completed.html RNAseq_DGE_JN_06-2021_completed.html RNAseq Bootcamp Recordings (attached): GL4U_RNAseq_Bootcamp_June_2021_Pilot_Day1_Part_1_of_5.mp4 GL4U_RNAseq_Bootcamp_June_2021_Pilot_Day1_Part_2_of_5.mp4 GL4U_RNAseq_Bootcamp_June_2021_Pilot_Day1_Part_3_of_5.mp4 *There were issues with the part 4 recording so that is not available GL4U_RNAseq_Bootcamp_June_2021_Pilot_Day1_Part_5_of_5.mp4 GL4U_RNAseq_Bootcamp_June_2021_Pilot_Day2_Part_1_of_3.mp4 GL4U_RNAseq_Bootcamp_June_2021_Pilot_Day2_Part_2_of_3.mp4 GL4U_RNAseq_Bootcamp_June_2021_Pilot_Day2_Part_3_of_3.mp4 GL4U_RNAseq_Bootcamp_June_2021_Pilot_Day3_Part_1_of_4.mp4 GL4U_RNAseq_Bootcamp_June_2021_Pilot_Day3_Part_2_of_4.mp4 GL4U_RNAseq_Bootcamp_June_2021_Pilot_Day3_Part_3_of_4.mp4 GL4U_RNAseq_Bootcamp_June_2021_Pilot_Day3_Part_4_of_4.mp4 GL4U_RNAseq_Bootcamp_June_2021_Pilot_Day4_Part_1_of_4.mp4 GL4U_RNAseq_Bootcamp_June_2021_Pilot_Day4_Part_2_of_4.mp4 GL4U_RNAseq_Bootcamp_June_2021_Pilot_Day4_Part_3_of_4.mp4 GL4U_RNAseq_Bootcamp_June_2021_Pilot_Day4_Part_4_of_4.mp4 GL4U_RNAseq_Bootcamp_June_2021_Pilot_Day5_Part_1_of_4.mp4 GL4U_RNAseq_Bootcamp_June_2021_Pilot_Day5_Part_2_of_4.mp4 GL4U_RNAseq_Bootcamp_June_2021_Pilot_Day5_Part_3_of_4.mp4 GL4U_RNAseq_Bootcamp_June_2021_Pilot_Day5_Part_4_of_4.mp4

GeneLab↗

Combining RNA-SEQ Datasets from NASA GENELAB: An Evaluation of Correction Methods

Background: Conducting space biology experiments aboard the International Space Station, particularly those utilizing complex model organisms like mice, is expensive and difficult due to limited crew availability, hardware, and space. As a result, sample numbers from these studies are low, reducing the statistical power of any one experiment. Aggregating spaceflight datasets serves as a method to increase sample numbers, allowing for novel insights through bioinformatic analysis of ‘omics data from merged datasets. However, aggregating datasets can introduce unwanted variation including 1) differences in sample handling, processing, and sequencing platforms between datasets (technical variation) as well as 2) differences in experimental design between datasets. Methods: In the present study, NASA GeneLab-hosted RNAseq datasets from mouse liver tissues were used to evaluate several statistical methods to correct for this unwanted variation through two approaches, reference-based and standard. The following correction algorithms were applied with (reference-based) and/or without (standard) considering Universal Mouse RNA Reference samples: ComBat and ComBat_seq from the SVA package, the median polish, empirical Bayes, and ANOVA-based algorithms from the MBatch package, and negative binomial regression normalization in the DESeq2 package. For each approach, after the correction algorithm was applied, differential gene expression (DGE) analysis of flight and ground control samples was performed with the combined data. The robustness of each tool was evaluated using BatchQC to determine statistical differences between datasets before and after correction, Principal Component Analysis to evaluate global gene expression in samples before and after correction, and by comparing DGE analysis of individual datasets and combined datasets before and after correction. Results: The results showed that the reference-based approach introduced several additional (and likely artificial) differentially expressed genes when compared with the respective standard approach. Conclusions: Of the methods tested, standard ComBat_seq and DESeq2 were identified as the most robust correction methods for combining spaceflight mouse liver RNAseq datasets hosted on GeneLab.

Finsam Samson↗