Search NASA⌕ Search

SEARCH · Search NASA

Results for “Bioinformatics”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 163 records · Page 9

Machine Learning Approaches to Increasing Value of Spaceflight Omics Databases

The number of spaceflight bioscience mission opportunities is too small to allow all relevant biological and environmental parameters to be experimentally identified. Simulated spaceflight experiments in ground-based facilities (GBFs), such as clinostats, are each suitable only for particular investigations -- a rotating-wall vessel may be 'simulated microgravity' for cell differentiation (hours), but not DNA repair (seconds) -- and introduce confounding stimuli, such as motor vibration and fluid shear effects. This uncertainty over which biological mechanisms respond to a given form of simulated space radiation or gravity, as well as its side effects, limits our ability to baseline spaceflight data and validate mission science. Machine learning techniques autonomously identify relevant and interdependent factors in a data set given the set of desired metrics to be evaluated: to automatically identify related studies, compare data from related studies, or determine linkages between types of data in the same study. System-of-systems (SoS) machine learning models have the ability to deal with both sparse and heterogeneous data, such as that provided by the small and diverse number of space biosciences flight missions; however, they require appropriate user-defined metrics for any given data set. Although machine learning in bioinformatics is rapidly expanding, the need to combine spaceflight/GBF mission parameters with omics data is unique. This work characterizes the basic requirements for implementing the SoS approach through the System Map (SM) technique, a composite of a dynamic Bayesian network and Gaussian mixture model, in real-world repositories such as the GeneLab Data System and Life Sciences Data Archive. The three primary steps are metadata management for experimental description using open-source ontologies, defining similarity and consistency metrics, and generating testing and validation data sets. Such approaches to spaceflight and GBF omics data may soon enable unique insight into which measured phenomena correlate to biological mechanisms that are truly affected by spaceflight conditions; which are most likely to be confounded by other variables; and which are insufficiently characterized, significantly increasing existing and future science return from ISS and spaceflight missions.

Gentry, Diana↗

One Step Closer to Mars with Aquaponics: Cultivating Citizen Science in K12 Schools

The Microbial Ecology and Biogeochemistry Research Laboratory at NASA Ames Research Center focuses primarily on the nutrient cycling and diversity of complex microbial communities. NASA is interested in the composition and functioning of microbial mat communities as these processes fundamentally shape the form and function of these analogs for the earliest forms of life on Earth (3.6 billion years ago), and likely will on other planets as well. Aquaponics systems are supported by microbial communities who perform many complex ecosystem services, including cycling nitrogen. Microbes are integral to the stability and productivity of aquaponics systems, which are analogous to microbial communities in food production systems that are essential for building efficient life support systems for long-distance space travel. Students at Meadow Park Middle School created 10 parallel aquaponics systems and took temporal microbial samples to characterize whether any macro-ecology variables impacted or changed the microbial diversity of these systems. Students additionally created a website so that other classrooms can pursue similar projects in their own schools (https://go.nasa.gov/2uJhxmF). Our lab at NASA Ames has sequenced water samples from each of the 10 tanks at 3 timepoints using a MinION sequencer. MPMS students will be involved in the analysis of the bioinformatics data generated through this collaboration. Our ongoing collaboration aims to collect and analyze data in the classroom setting that has utility for research scientists, while involving students as collaborators in the research process.

Kolattukudy, Maria↗

GeneLab Analysis Working Group Kick-Off Meeting

Goals to achieve for GeneLab AWG - GL vision - Review of GeneLab AWG charter Timeline and milestones for 2018 Logistics - Monthly Meeting - Workshop - Internship - ASGSR Introduction of team leads and goals of each group Introduction of all members Q/A Three-tier Client Strategy to Democratize Data Physiological changes, pathway enrichment, differential expression, normalization, processing metadata, reproducibility, Data federation/integration with heterogeneous bioinformatics external databases The GLDS currently serves over 100 omics investigations to the biomedical community via open access. In order to expand the scope of metadata record searches via the GLDS, we designed a metadata warehouse that collects and updates metadata records from external systems housing similar data. To demonstrate the capabilities of federated search and retrieval of these data, we imported metadata records from three open-access data systems into the GLDS metadata warehouse: NCBI's Gene Expression Omnibus (GEO), EBI's PRoteomics IDEntifications (PRIDE) repository, and the Metagenomics Analysis server (MG-RAST). Each of these systems defines metadata for omics data sets differently. One solution to bridge such differences is to employ a common object model (COM) to which each systems' representation of metadata can be mapped. Warehoused metadata records are then transformed at ETL to this single, common representation. Queries generated via the GLDS are then executed against the warehouse, and matching records are shown in the COM representation (Fig. 1). While this approach is relatively straightforward to implement, the volume of the data in the omics domain presents challenges in dealing with latency and currency of records. Furthermore, the lack of a coordinated has been federated data search for and retrieval of these kinds of data across other open-access systems, so that users are able to conduct biological meta-investigations using data from a variety of sources. Such meta-investigations are key to corroborating findings from many kinds of assays and translating them into systems biology knowledge and, eventually, therapeutics.

GeneLab↗

GeneLab: Omics Database for Spaceflight Experiments

Motivation - To curate and organize expensive spaceflight experiments conducted aboard space stations and maximize the scientific return of investment, while democratizing access to vast amounts of spaceflight related omics data generated from several model organisms. Results - The GeneLab Data System (GLDS) is an open access database containing fully coordinated and curated "omics" (genomics, transcriptomics, proteomics, metabolomics) data, detailed metadata and radiation dosimetry for a variety of model organisms. GLDS is supported by an integrated data system allowing federated search across several public bioinformatics repositories. Archived datasets can be queried using full-text search (e.g., keywords, Boolean and wildcards) and results can be sorted in multifactorial manner using assistive filters. GLDS also provides a collaborative platform built on GenomeSpace for sharing files and analyses with collaborators. It currently houses 172 datasets and supports standard guidelines for submission of datasets, MIAME (for microarray), ENCODE Consortium Guidelines (for RNA-seq) and MIAPE Guidelines (for proteomics).

omics↗

Characterization of Plastic Degrading Bacteria from Environmental Samples by Genetic and Biochemical Analysis

Plastic is the major waste-product during NASA space missions, recycling this waste-stream to produce other beneficial materials would decrease upmass. Bacterial called plastisomes have been demonstrated to metabolize non-biodegradable plastics such as polyethylene and polystyrene. Characterization and engineering of these bacteria, and their eventual incorporation as life support systems would enable space flight beyond lower earth orbit. We will utilize molecular techniques to identify and isolate the most productive plastisome. Environmental samples obtained from locations known to be rich in plastic will be cultured in a laboratory defined-media supplemented with plastic as the sole carbon source. Cultures will be monitored for growth over time. Ribosomal DNA will be amplified from cultures that exhibit growth using PCR. These amplified fragments will be sequenced to determine the identity of the consortia in the cultures. We will then perform bioinformatics analysis on the data to identify the plastisomes and generate phylogenetic trees. Morphological and physiological profile of the plastisomes will also be conducted by microscopy and biochemical tests. Our results would reveal a bacterial strain that can break down plastics efficiently. The implication for this project would not only benefit space exploration but also make a major impact towards sustainability development on Earth.

plastic conversion↗

The Evolution of Planetary Protection Implementation on Mars Landed Missions

NASA has developed requirements dedicated to the prevention of forward and backward contamination during space exploration. Historically, international agreements provided guidelines to prevent contamination of the Moon and other celestial bodies, as well as the Earth (e.g., sample return missions). The UN Outer Space Treaty was established in 1967 and the Committee on Space Research (COSPAR) maintains a planetary protection policy complying with Article IX of this treaty. By avoiding forward contamination, the integrity of scientific exploration is preserved. Planetary Protection mission requirements are levied on missions to control contamination. These requirements are dependent on the science of the mission and on the celestial bodies encountered or targeted along the way. Consequently, categories are assigned to missions, and specific implementation plans are developed to meet the planetary protection requirements. NASA missions have evolved over time with increasingly more demanding scientific objectives and more complex flight systems to achieve those objectives and, thus, planetary protection methods and processes used for implementation have become much more intricate, complicated, and challenging. Here, we will portray the evolution of planetary protection implementation at JPL in several important areas throughout the course of NASA sponsored robotic Mars lander or rover missions, starting from Mars Pathfinder through the beginning of Mars 2020. Highlighted in the discussion will be process changes in planetary protection requirements development and flow down. Development and implementation of new and improved methods used in the reduction of spacecraft bioburden will be discussed as well as approaches and challenges that come along with setting up remote laboratories to perform bioassays. The consequences and forward planning of delays on missions will be highlighted as well as lessons learned on the impact of communication and training in achieving planetary protection requirements. The evolution of methods used for the detection of microbial bioburden on spacecraft hardware will be considered. These methods use standard microbiology as well as the adaptation of advances in biotechnology, molecular biology, and bioinformatics. Technical approaches developed for the prevention of contamination and recontamination of hardware during Assembly, Test, and Launch Operations will be discussed.

Kazarians, Gayane A.↗

DNA Damage Response to Low and High-LET in a Large Cohort of Mice and Humans and Latest Advancement in NASA Space Omics

This presentation will first focus on a thorough evaluation of the DNA damage response to both low and high-LET in a cohort of 76 mice primary skin fibroblast derived from 15 different strains or in human blood mononuclear cells derived from 550 healthy donors. In both the human and mice work, we have hypothesized that DNA repair capacity can be used as a marker to evaluate and differentiate individual radiation sensitivity. More specifically, this work is based on the concept that the combined time-dose dependence of radiation-induced foci (RIF) of p53-binding protein 1 (53BP1) following low-LET exposure contains sufficient information to infer sensitivity to any other LET. This work is one of the most extensive studies on the kinetics and possible genetic underpinnings of radiation-induced DNA damage and repair. Results on humans are still preliminary as we are still in the process of collecting and isolating primary blood mononuclear cells from 500 to 800 healthy subjects of European descent, 18-75 years of age, 50/50 male/female distribution. We have analyzed 53BP1+ RIF formation as well as oxidative stress and cell death in primary cells from 192 subjects in response to the same HZE particles as used in mice: 600 MeV/n Fe, 350 MeV/n Ar and 350 MeV/n Si, 1.1 and 3 particles/100m2, 4 and 24 hours after irradiation. The second part of the talk will focus on describing GeneLab: The NASA Systems Biology Platform for Space Omics Repository, Analysis and Visualization. NASA GeneLab is an open-access repository for omics datasets generated by biological experiments conducted in space or experiments relevant to spaceflight (e.g. simulated cosmic radiation, simulated microgravity, bed rest studies). Started as a repository designed to archive precious omics from space experiments, GeneLab has expanded its scope to maximize the intelligibility of the raw data (e.g. RNAseq, microarray, WGBS, metagenome), particularly for users with limited bioinformatics knowledge. As such GeneLab is now providing processed data derived from the raw data covering a large spectrum of omics (genome, epigenome, transcriptome, epitranscriptome, proteome, metabolome), to help users explore important questions: Which genes or proteins are expressed differently in space for various living organisms? What are the consequences arising from these changes? What specifics DNA mutations or epigenetic changes happen in space? What species or genetic features lead to better adaption to such a unique environment? In this presentation, we will report on the current and future objectives for GeneLab, and review recent published studies relating molecular changes observed in various animal models and tissue with microgravity, radiation, circadian rhythm, hydration and carbon dioxide conditions.

DNA repair kinetics↗

GeneLab: A Systems Biology Platform for Omics Analysis

NASA GeneLab is an open-access repository for omics datasets generated by biological experiments conducted in space or experiments relevant to spaceflight (e.g. simulated cosmic radiation, simulated microgravity, bed rest studies). The GeneLab Data Systems (GLDS) version 4.0 will be available on October 1st 2019, and will provide the latest in terms of professional state-of-the-art bioinformatics platform for the space biology and radiation community to upload their data into an omics data commons, to process their data with vetted standard workflows and to compare to existing analyses. Started in 2015 as a repository designed to archive omics data from space experiments, GeneLab has expanded its scope to all ionizing radiation omics experiments conducted on the ground and has put considerable effort in providing carefully characterized radiation metadata on all dataset. GeneLab is also providing processed data derived from the raw data covering a large spectrum of omics (genome, epigenome, transcriptome, epitranscriptome, proteome, metabolome) to help users explore important questions: 1) Which genes or proteins are expressed differently in space for various living organisms? 2) What specific DNA mutations or epigenetic changes happen in space or after exposure to ionizing radiation? and 3) How does genetics affect these responses? Processed data available on GeneLab are derived by standard data analysis workflows vetted by hundreds of scientists who volunteered to join one of the four GeneLab Analysis Working Groups (Animal AWG, Plant AWG, Microbe AWG, Multi-Omics AWG). In this presentation, we will discuss how to bridge the gap between irradiation studies performed on earth and biological experiments conducted in space since the early 1990's. We will discuss how radiation dosimetry was estimated for datasets derived from samples collected during the Space Shuttle era or on the International Space Station. Finally, we will address future strategies regarding dose monitoring in future missions into space, inter-agency efforts to unify data under one umbrella, and knowledge dissemination across the radiation research community and the space biology community.

open-science↗

NASA GeneLab Space Omics Database: Expanding from Space to Ionizing Radiation Data on the Ground

NASA GeneLab is an open-access repository for omics datasets generated by biological experiments conducted in space or ground experiments relevant to spaceflight (e.g. simulated cosmic radiation, simulated microgravity, bed rest studies). The GeneLab Data Systems (GLDS) version 4.0 will be available on October 1st 2019, and will provide a state-of-the-art bioinformatics platform for the space biology and radiation communities to upload their data into an omics data commons, to process their data with vetted standard workflows and to compare with existing analyses. Started in 2015 as a repository designed to archive omics data from space experiments, GeneLab has expanded its scope to all ionizing radiation omics experiments conducted on the ground and has put considerable effort in providing carefully characterized radiation metadata on all datasets. GeneLab is also providing processed data derived from the raw data covering a large spectrum of omics (genome, epigenome, transcriptome, epitranscriptome, proteome, metabolome) to help users explore important questions: 1) Which genes or proteins are expressed differently in space for various living organisms? 2) What specific DNA mutations or epigenetic changes happen in space or after exposure to ionizing radiation? and 3) How does genetics affect these responses? Processed data available on GeneLab are derived by standard data analysis workflows vetted by hundreds of scientists who volunteered to join one of the four GeneLab Analysis Working Groups (Animal AWG, Plant AWG, Microbe AWG, Multi-Omics AWG). In this presentation, we will discuss how to bridge the gap between irradiation studies performed on earth and biological experiments conducted in space since the early 1990's. We will discuss how radiation dosimetry was estimated for datasets derived from samples collected during the Space Shuttle era on the International Space Station and on other orbiting platforms. Finally, we will address future strategies regarding dose monitoring in future missions into space, inter-agency efforts to unify data under one umbrella, and knowledge dissemination across the radiation research community and the space biology community.

open-science↗

GeneLab: Overview of Challenges and Opportunities

The NASA GeneLab project capitalizes on multi-omic technologies to maximize the return on spaceflight experiments. To do this, GeneLab maintains a publicly accessible database (GLDS) that houses spaceflight and spaceflight relevant multi-omics data, and collaborates with NASA principal investigators and projects to generate additional omics data. GeneLab houses more than 200 transcriptomic, proteomic, metabolomic and epigenomic datasets from plant, animal and microbial experiments, with a growing number of these having been produced by the GeneLab sample processing lab. The GLDS contains rich metadata about each experiment and has recently integrated radiation dosimetery data from experiments flown on the Space Shuttle. GeneLab has also recently implemented an effort to present processed data in the GLDS in addition to the raw omics data. The processed data will enable interpretation of the data by a larger group of students, scientists and the general public. Standard pipelines for the transformation of raw data into visualizations were developed by four GeneLab Analysis Working Groups (animals, plants, microbes, multi-omics) comprised of over 100 scientists from NASA and academia. These pipelines are now being used by a group of bioinformatics interns to provide standard basic analysis of the data for incorporation into GLDS.

Galazka, Jonathan M.↗

GeneLab: Open Science for Life in Space

The NASA GeneLab project capitalizes on multi-omic technologies to maximize the return on spaceflight experiments. To do this, GeneLab maintains a publicly accessible database (GLDS) that houses spaceflight and spaceflight relevant multi-omics data, and collaborates with NASA principal investigators and projects to generate additional omics data. GeneLab houses more than 200 transcriptomic, proteomic, metabolomic and epigenomic datasets from plant, animal and microbial experiments, with a growing number of these having been produced by the GeneLab sample processing lab. The GLDS contains rich metadata about each experiment and has recently integrated radiation dosimetery data from experiments flown on the Space Shuttle. GeneLab has also recently implemented an effort to present processed data in the GLDS in addition to the raw omics data. The processed data will enable interpretation of the data by a larger group of students, scientists and the general public. Standard pipelines for the transformation of raw data into visualizations were developed by four GeneLab Analysis Working Groups (animals, plants, microbes, multi-omics) comprised of over 100 scientists from NASA and academia. These pipelines are now being used by a group of bioinformatics interns to provide standard basic analysis of the data for incorporation into GLDS.

Galazka, Jonathan M.↗

Beyond Nanopore Sequencing in Space: Identifying the Unknown

Astronaut Kate Rubins sequenced DNA on the International Space Station (ISS) for the first time in August 2016 (Figure 1A). A 2D sequencing library containing an equal mixture of lambda bacteriophage, Escherichia coli, and Mus musculus was prepared on the ground with a SQK_MAP006 kit and sent to the ISS frozen and loaded into R7.3 flow cells. After a total of 9 on-orbit sequencing runs over 6 months, it was determined that there was no decrease in sequencing performance on-orbit compared to ground controls (1). A total of ~280,000 and ~130,000 reads generated on-orbit and on the ground, respectively, identified 90% of reads that were attributed to 30% lambda bacteriophage, 30% Escherichia coli, and 30% M. musculus (Figure 1B). Extensive bioinformatics analysis determined comparable 2D and 1D read accuracies between flight and ground runs (Figure 1C), and data collected from the ISS were able to construct directed assemblies of E.coli and lambda genomes at 100% and M. musculus mitochondrial genome at 96.7%. These findings validate sequencing as a viable option for potential on-orbit applications such as environmental microbial monitoring and disease diagnosis. Current microbial monitoring of the ISS applies culture-based techniques that provide colony forming unit (CFU) data for air, water, and surface samples. The identity of the cultured microorganisms in unknown until sample return and ground-based analysis, a process that can take up to 60 days. For sequencing to benefit ISS applications, spaceflight-compatible sample preparation techniques are required. Subsequent to the testing of the MinION on-orbit, a sample-to-sequence method was developed using miniPCR™ and basic pipetting, which was only recently proven to be effective in microgravity. The work presented here details the in- flight sample preparation process and the first application of DNA sequencing on the ISS to identify unknown ISS-derived microorganisms.

Stahl, Sarah E.↗

Evaluation of Correction Methods for NASA GeneLab Transcriptomic Datasets

Conducting space biology experiments aboard the International Space Station, particularly those utilizing complex model organisms like mice, is expensive and difficult due to limited crew availability, hardware, and space. As a result, sample numbers from these studies are low, reducing the statistical power of any one experiment. Aggregating spaceflight datasets serves as a method to increase sample numbers, allowing for novel insights through bioinformatic analysis of ‘omics data from merged datasets. However, aggregating datasets can introduce unwanted variation including 1) differences in sample handling, processing, and sequencing platforms between datasets (technical variation) as well as 2) differences in experimental design between datasets such as sex or age of the model organism used. In the present study, NASA GeneLab-hosted RNAseq datasets from rodent liver tissues were used to evaluate several statistical methods to correct for this unwanted variation through two approaches, reference-based and standard. The following correction algorithms were applied with (reference-based) and/or without (standard) considering Universal Mouse RNA Reference samples: ComBat and ComBat_seq from the SVA package, median polish, empirical Bayes, and ANOVA-based algorithms from the MBatch package, and negative binomial regression normalization in the DESeq2 package. For each approach, after the correction algorithm was applied, differential gene expression (DGE) analysis of flight and ground control samples was performed with the combined data. The robustness of each tool was evaluated using BatchQC, to determine statistical differences between datasets before and after correction, Principal Component Analysis, to evaluate global gene expression in samples before and after correction, and by comparing DGE analysis of individual datasets and combined datasets before and after correction. The results showed that the reference-based approach introduced several additional (and likely artificial) DEGs when compared with the standard approach. Thus, the most robust standard correction will be implemented in the GeneLab Visualization 2.0 platform when datasets are combined.

GeneLab, RNA-seq, Batch Correction↗

Evaluation of Correction Methods for NASA GeneLab Transcriptomic Datasets

Conducting space biology experiments aboard the International Space Station, particularly those utilizing complex model organisms like mice, is expensive and difficult due to limited crew availability, hardware, and space. As a result, sample numbers from these studies are low, reducing the statistical power of any one experiment. Aggregating spaceflight datasets serves as a method to increase sample numbers, allowing for novel insights through bioinformatic analysis of ‘omics data from merged datasets. However, aggregating datasets can introduce unwanted variation including 1) differences in sample handling, processing, and sequencing platforms between datasets (technical variation) as well as 2) differences in experimental design between datasets. In the present study, NASA GeneLab-hosted RNAseq datasets from mouse liver tissues were used to evaluate several statistical methods to correct for this unwanted variation through two approaches, reference-based and standard. The following correction algorithms were applied with (reference-based) and/or without (standard) considering Universal Mouse RNA Reference samples: ComBat and ComBat_seq from the SVA package, median polish, empirical Bayes, and ANOVA-based algorithms from the MBatch package, and negative binomial regression normalization in the DESeq2 package. For each approach, after the correction algorithm was applied, differential gene expression (DGE) analysis of flight and ground control samples was performed with the combined data. The robustness of each tool was evaluated using BatchQC to determine statistical differences between datasets before and after correction, Principal Component Analysis to evaluate global gene expression in samples before and after correction, and by comparing DGE analysis of individual datasets and combined datasets before and after correction. The results showed that the reference-based approach introduced several additional (and likely artificial) DEGs when compared with the respective standard approach. Of the methods tested, standard ComBat and DESeq2 were identified as the most robust correction methods for combining spaceflight mouse liver RNAseq datasets hosted on GeneLab.

GeneLab↗

Biospecimen and Data Sharing: NASA Institutional Scientific Collection at Ames Research Center (ISC-ARC), and the Ames Life Sciences Data Archive (ALSDA)

For decades, NASA and international partners have conducted biological experiments in space to understand effects of spaceflight and address potential hazards. To enable spaceflight back to the Moon, and then to Mars and beyond, it is imperative to further understand basic science and health risks associated with spaceflight, along with developing countermeasures. The sending of experiments and organisms into space is a costly endeavor. To maximize scientific return, sharing with the scientific community both space-flown biospecimens and data from completed experiments is essential. New fundamental, applied, and bioinformatic science insights can be gained from specimen and data sharing efforts. Data reuse enables spaceflight health risk modeling, analyzing adverse outcomes across spaceflight hazards, and deep space autonomous support for the flight medical officer.

Data↗

GeneLab

GeneLab collects and enables analysis of spaceflight and ground-based spaceflight simulation genomic data, RNA and protein expression, and metabolic profiles. It interfaces with other existing databases containing spaceflight omic data. The 2011 National Research Council (NRC) Decadal Survey on NASA Life and Physical Sciences called for increased opportunities for multi-investigator spaceflight opportunities and greater use of genomic approaches to meet the needs of NASA researchers. To address these recommendations of the NRC Decadal Survey, the Space Life and Physical Sciences Research and Applications Division of NASA's Human Exploration and Operations Mission Directorate has initiated a transition to an Open Science architecture to increase research opportunities, and has developed the GeneLab Platform based on highly leveraged and integrated bioinformatics analytics. GeneLab is an interactive, open-access resource where scientists can upload, download, store, search, share, transfer, and analyze omics data from spaceflight and corresponding analogue experiments. Users can explore GeneLab datasets in the Data Repository, analyze data using the Analysis Platform, visualize high-order data and create collaborative projects using the Collaborative Workspace. Our primary goal is to maximize the utilization of the valuable biological research conducted aboard the International Space Station (ISS) by collecting genomic, transcriptomic, proteomic, and metabolomics data known as “omics”. By providing a portal linking processed data to flight parameters, GeneLab enables exploration of the molecular network responses of terrestrial biology to the space environment. This allows researchers to understand the complex responses of biological systems to the space environment. This technology development activity was transferred from the Human Exploration and Operations Mission Directorate to the Science Mission Directorate Division of Biological and Physical Sciences (BPS) in October 2020.

GeneLab↗

Combining RNA-SEQ Datasets from NASA GENELAB: An Evaluation of Correction Methods

Background: Conducting space biology experiments aboard the International Space Station, particularly those utilizing complex model organisms like mice, is expensive and difficult due to limited crew availability, hardware, and space. As a result, sample numbers from these studies are low, reducing the statistical power of any one experiment. Aggregating spaceflight datasets serves as a method to increase sample numbers, allowing for novel insights through bioinformatic analysis of ‘omics data from merged datasets. However, aggregating datasets can introduce unwanted variation including 1) differences in sample handling, processing, and sequencing platforms between datasets (technical variation) as well as 2) differences in experimental design between datasets. Methods: In the present study, NASA GeneLab-hosted RNAseq datasets from mouse liver tissues were used to evaluate several statistical methods to correct for this unwanted variation through two approaches, reference-based and standard. The following correction algorithms were applied with (reference-based) and/or without (standard) considering Universal Mouse RNA Reference samples: ComBat and ComBat_seq from the SVA package, the median polish, empirical Bayes, and ANOVA-based algorithms from the MBatch package, and negative binomial regression normalization in the DESeq2 package. For each approach, after the correction algorithm was applied, differential gene expression (DGE) analysis of flight and ground control samples was performed with the combined data. The robustness of each tool was evaluated using BatchQC to determine statistical differences between datasets before and after correction, Principal Component Analysis to evaluate global gene expression in samples before and after correction, and by comparing DGE analysis of individual datasets and combined datasets before and after correction. Results: The results showed that the reference-based approach introduced several additional (and likely artificial) differentially expressed genes when compared with the respective standard approach. Conclusions: Of the methods tested, standard ComBat_seq and DESeq2 were identified as the most robust correction methods for combining spaceflight mouse liver RNAseq datasets hosted on GeneLab.

Finsam Samson↗

Spaceflight Biospecimen Sharing in Support of Science Discovery and Exploration

For decades, NASA and international partners have flown non-human biological experiments in space to understand the effects of spaceflight and address potential biological hazards. Sending organisms into space is a costly endeavor which makes space-flown biological specimens a valuable resource. To enable maximum scientific return, samples not required by the Principal Investigators are harvested and collected mostly by NASA’s Space Biology Biospecimen Sharing Program. These specimens are collected according to well-established SOPs that maintain quality and integrity. The specimens are then preserved, archived, and made available to the international scientific community through NASA’s Institutional Scientific Collection (ISC) at Ames Research Center (ARC). The ISC-ARC biospecimens and descriptive metadata are findable and accessible for request through the Life Sciences Data Archive (LSDA). The NASA ISC-ARC currently stores over 32,000 specimens from Shuttle, International Space Station, and ground-based investigations (spaceflight analog experiments involving either hindlimb unloading, centrifugation, or partial weight-bearing study designs). Tissues are predominantly from mice and rats, though samples are also available from bacteria and quail. The specimens include tissues from many physiological systems including musculoskeletal, neurosensory, reproductive, respiratory, circulatory, and digestive. Tissues are stored at -80°C, -20°C, +4°C, or ambient and preserved in various fixatives. Descriptive metadata is available for all samples. Historically, these tissues have been used for a wide range of analyses, including histology, genomics, and transcriptomics. Plans are underway to expand the ISC-ARC beyond the mostly-rodent contents, to include a space-relevant microbial culture collection including bacteria, fungi, and yeast. This expansion of the ISC-ARC will now involve identifying and standardizing best practices for microbial curations. To ensure safe long-term storage of microbial isolates, a microbiology laboratory will be dedicated for identification, cell culture, and lyophilization. Awarding of tissue to public science investigators has resulted in 33 publications since 2011, with 48 requests being submitted since 2016. Of note, NASA GeneLab has been awarded ISC-ARC biospecimens in the past few years. GeneLab processes the biospecimens to generate various levels of ‘omics’ data, which are published on GeneLab’s open access online platform for bioinformatics analysis and visualization. This has helped a systems biology community grow around the processed-biospecimens’ datasets, resulting in many new publications and insights. Websites: https://www.nasa.gov/ames/research/space-biosciences/isc-bsp ; https://lsda.jsc.nasa.gov/Biospecimen

Ryan T. Scott↗