Search NASA⌕ Search

SEARCH · Search NASA

Results for “RNAseq”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

Optimizing a Small RNAseq Analysis Pipeline for NASA GeneLab Using Open-Source Tools and Libraries

Small RNA sequencing (small RNAseq) is a powerful tool for studying the regulation of gene expression in various organisms. Small RNAseq has been leveraged in space biology research to study how expression of small RNAs, e.g. micro RNAs (miRNAs), small interfering RNAs (siRNAs), and piwi-interacting RNAs (piRNAs), change upon exposure to the space environment. NASA GeneLab currently hosts small RNAseq raw data derived from space-relevant experiments on the Open Science Data Repository (OSDR). To maximize the accessibility of these data to the scientific community, in addition to hosting raw data, which is only interpretable by bioinformaticians, GeneLab plans to process all small RNAseq datasets and make those processed data available to the scientific community via the OSDR. In this study, we present the development of the GeneLab standardized pipeline for processing small RNAseq datasets. Using human, plant, and synthetic small RNAseq datasets, we interrogate various open-source software and publicly available databases to evaluate their accuracy and reproducibility in each step of the pipeline. For quality control and adapter detection and trimming, we evaluated TrimGalore!, FASTX, SeqKit, and DNApi methods to optimize alignment to reference genomes. We compared BWA, Bowtie, and Bowtie2 to determine the optimal alignment tool. For each alignment tool we also assessed various reference databases, including Ensembl reference genomes and different types of small RNA reference databases, including genome, hairpin, and miRNA references from the miRbase and MirGeneDB databases. To quantify the aligned data, we compared SAMtools, HTSeq, and RSEM for counting alignment events from each alignment tool used. Finally, we evaluated various tools, including DESeq2 and EdgeR, for data normalization and subsequent differential expression analysis. We will present the results from our comparative analyses for each pipeline step and propose a consensus pipeline for processing small RNAseq data derived from various organisms exposed to the space environment.

SmallRNAseq, NASA GeneLab, quality control, adapte↗

RNAseq analysis of Cellvibrio japonicus during starch utilization differentiates between genes encoding carbohydrate active enzymes controlled by substrate detection or growth rate

ABSTRACT Bacterial utilization of starch is increasingly of interest as the importance and contributions of animal gut microbiomes become more defined. Consequently, identifying and characterizing the bacterial enzymes responsible for the degradation, transport, and metabolism of starch will enable developments in pharmaceutical, biotechnological, and culinary industries searching for novel prebiotics, carrier molecules, and low glycemic index sweeteners. The current challenge is that bacteria proficient at starch utilization often have hundreds of carbohydrate active enzymes, and it is unclear which are essential for starch utilization using only homology-based bioinformatics or computational methods. Complementary experimental data are also needed, especially to understand the regulation of bacterial starch utilization. We have completed an RNAseq analysis of the Gram-negative bacterium Cellvibrio japonicus and found that it has sophisticated regulation that includes substrate sensing and growth rate components for genes that encode starch-degrading enzymes. Among the 22 genes predicted to encode starch-active enzymes, C. japonicus has 10 alpha-amylases, 4 alpha-glucosidases, 2 pullulnases, and 2 cyclomaltodextrin glucanotransferases, 15 of which were up-regulated during exponential growth on starch and 8 up-regulated in stationary phase. Growth analyses with an enzyme secretion deficient mutant of C. japonicus suggested that secreted amylases are essential for this bacterium to degrade starch. Our approach of coupling a physiological growth assay with transcriptomic data provides a platform to identify targets for further genetic or biochemical analysis that can be broadly applied to other starch-utilizing bacteria. IMPORTANCE Understanding the bacterial metabolism of starch is important as this polysaccharide is a ubiquitous ingredient in foods, supplements, and medicines, all of which influence gut microbiome composition and health. Our RNAseq and growth data set provides a valuable resource to those who want to better understand the regulation of starch utilization in Gram-negative bacteria. These data are also useful as they provide an example of how to approach studying a starch-utilizing bacterium that has many putative amylases by coupling transcriptomic data with growth assays to overcome the potential challenges of functional redundancy. The RNAseq data can also be used as a part of larger meta-analyses to compare how C. japonicus regulates carbohydrate active enzymes, or how this bacterium compares to gut microbiome constituents in terms of starch utilization potential.

59 BASIC BIOLOGICAL SCIENCES↗

NASA GeneLab RNASeq Consensus Pipeline: A Nextflow Implementation

The NASA GeneLab project (genelab.nasa.gov) seeks to accelerate space biology research through cataloging and democratizing omics data. Since raw omics data is largely inaccessible to non-bioinformaticians, GeneLab works with the scientific community to develop standard processing pipelines to generate and publish processed data. Unlike raw data, processed data has greater immediate value to a wide range of users with varying technical backgrounds and computational capabilities. Standardizing processing workflows is essential to match the pace of raw data generation, ensure reproducibility, and enable standardized processed data for comparison across datasets. Previously, GeneLab developed a standardized pipeline for processing RNAseq data, referred to as the ‘GeneLab RNAseq Consensus Pipeline (RCP)’, in collaboration with GeneLab’s Analysis Working Groups. The work presented here is a Nextflow implementation of GeneLab’s RCP that automates and accelerates data processing of RNASeq datasets hosted on GeneLab. In addition to the core data processing, the workflow also includes staging of GeneLab raw data and a robust verification and validation (V&V) program that runs after each processing step to identify errors in real-time, stop additional downstream computation, and preserve computational resources. The workflow, including the staging and V&V functionality, is open source for others to reuse and modify at https://github.com/nasa/GeneLab_Data_Processing/tree/master/RNAseq.

Jonathan Dejesus Oribello↗

NASA GeneLab RNASeq Consensus Pipeline: A Nextflow Implementation

The NASA GeneLab project (genelab.nasa.gov) seeks to accelerate space biology research through cataloging and democratizing omics data. Since raw omics data is largely inaccessible to non-bioinformaticians, GeneLab works with the scientific community to develop standard processing pipelines to generate and publish processed data. Unlike raw data, processed data has greater immediate value to a wide range of users with varying technical backgrounds and computational capabilities. Standardizing processing workflows is essential to match the pace of raw data generation, ensure reproducibility, and enable standardized processed data for comparison across datasets. Previously, GeneLab developed a standardized pipeline for processing RNAseq data, referred to as the ‘GeneLab RNAseq Consensus Pipeline (RCP)’, in collaboration with GeneLab’s Analysis Working Groups. The work presented here is a Nextflow implementation of GeneLab’s RCP that automates and accelerates data processing of RNASeq datasets hosted on GeneLab. In addition to the core data processing, the workflow also includes staging of GeneLab raw data and a robust verification and validation (V&V) program that runs after each processing step to identify errors in real-time, stop additional downstream computation, and preserve computational resources. The workflow, including the staging and V&V functionality, is open source for others to reuse and modify at https://github.com/nasa/GeneLab_Data_Processing/tree/master/RNAseq.

Jonathan D Oribello↗

RNAseq-based transcriptome assembly of Clostridium acetobutylicum for functional genome annotation and discovery

Accurate genome annotations are essential in modern biology and biotechnology, yet they are still largely based on genome sequencing and comparative analyses. We show that the Clostridium acetobutylicum genome annotation can be markedly improved by integrating bioinformatic predictions with RNA sequencing (RNAseq) data. Samples were acquired under butanol, butyrate, and unstressed treatments across various growth conditions. Analysis of an initial assembly revealed errors due to background signals and limitations of assembly algorithms. Hurdles for RNAseq transcriptome mapping include optimizing library complexity and sequencing depth, yet most studies report low sequencing depth and ignore the effect of ribosomal RNA abundance. An integrative analysis was developed to combine motif predictions, single-nucleotide resolution sequencing depth, and library complexity to resolve difficulties in assembly curation. This minimized false positive error and determined gene boundaries, in some cases, to the exact base-pair of prior studies. This will be the first strand-specific transcriptome assembly in a Clostridium organism.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

RNASeq and Fluorescence Analysis of the Response of ERF2 and ERF104 in Arabidopsis thaliana under Simulated Altered Gravity

As NASA moves closer to long-term human space exploration, the need to understand how to sustain life in space is increasingly pressing. Plants are essential to human sustenance, making it important to understand how spaceflight affects plant health. We used differential gene expression analysis to examine GLDS-251 (RNAseq analysis of the response of Arabidopsis thaliana to fractional gravity under blue-light stimulation during spaceflight) from NASA’s GeneLab data repository and found downregulation of ERF2 and ERF104, transcription factors of the ethylene response factor families, that integrate hormonal pathways involved in abiotic stress responses. Downregulation of ERF2 and ERF104 during spaceflight may indicate a dysregulation of the ethylene signaling pathway. Our hypothesis is that altered gravity downregulates the expression of ERF2 and ERF104 in Arabidopsis thaliana, altering the ethylene signaling pathway and affecting the electron transport chain and light-dependent reactions in chloroplast thylakoids. To test this hypothesis, we propose to grow A. thaliana seedlings (wild-type and mutant/knockout of ERF2 and ERF104) in altered gravity conditions to determine the effects on the expression of ERF2, ERF104, and photosynthesis. We anticipate that ERF2 and ERF104 will be underexpressed in altered gravity conditions and result in decreased regulation of the ethylene signaling pathway.

Arabidopsis↗

RNAseq data for P. putida with vanillate

Illumina sequencing reads from RNA sequencing of vanillate-utilizing strains of Pseudomonas putida, described in Evolution and engineering of pathways for aromatic O-demethylation in Pseudomonas putida KT2440 by A. Bleem, et al. (2024)

Adaptive laboratory evolution↗

GL4U: GeneLab for Colleges and Universities

GeneLab for Colleges and Universities (GL4U) will provide space biology-relevant training in bioinformatics to the next generation of scientists through direct and indirect approaches. The GeneLab (GL) team will host two annual data processing bootcamps, one for college-level students (direct) and one for college educators (indirect – Training of Trainers), in which participants learn to analyze space-relevant omics data hosted on GL. The first bootcamp took place in early June 2021 with about 30 SJSU undergraduate students and covered space biology-specific lectures and hands-on instruction using Jupyter Notebooks (JNs) for RNA sequence (RNAseq) data analysis. All training materials including the enclosed files listed below will be made publicly available on GitHub. RNAseq Bootcamp Lectures (attached in combined file): Introduction to NASA, Space Biology, GeneLab, and the Command Line: NASA_GL_CL_Intro_FINAL.pdf - DRAFT from initial submission NASA_SB_GL_CL_Intro_FULL.pdf - FINAL version presented during the bootcamp - only minor edits from the draft version RNAseq and Data Processing Overview: RNAseq_Overview_FINAL.pdf - DRAFT from initial submission RNAseq_Overview_FULL.pdf - FINAL version presented during the bootcamp - only minor edits from the draft version Overview of the Statistics Used for RNAseq Data Analysis: SJSU_Statistics_Intro_Lecture_FINAL.pdf - DRAFT from initial submission Statistics_Overview_FULL.pdf - FINAL version presented during the bootcamp - only minor edits from the draft version Completed JNs in HTML format (attached in combined file): Unix_Intro_JN_06-2021_completed.html R_Intro_JN_06-2021_completed.html RNAseq_fastq_to_counts_JN_06-2021_completed.html RNAseq_DGE_JN_06-2021_completed.html RNAseq Bootcamp Recordings (attached): GL4U_RNAseq_Bootcamp_June_2021_Pilot_Day1_Part_1_of_5.mp4 GL4U_RNAseq_Bootcamp_June_2021_Pilot_Day1_Part_2_of_5.mp4 GL4U_RNAseq_Bootcamp_June_2021_Pilot_Day1_Part_3_of_5.mp4 *There were issues with the part 4 recording so that is not available GL4U_RNAseq_Bootcamp_June_2021_Pilot_Day1_Part_5_of_5.mp4 GL4U_RNAseq_Bootcamp_June_2021_Pilot_Day2_Part_1_of_3.mp4 GL4U_RNAseq_Bootcamp_June_2021_Pilot_Day2_Part_2_of_3.mp4 GL4U_RNAseq_Bootcamp_June_2021_Pilot_Day2_Part_3_of_3.mp4 GL4U_RNAseq_Bootcamp_June_2021_Pilot_Day3_Part_1_of_4.mp4 GL4U_RNAseq_Bootcamp_June_2021_Pilot_Day3_Part_2_of_4.mp4 GL4U_RNAseq_Bootcamp_June_2021_Pilot_Day3_Part_3_of_4.mp4 GL4U_RNAseq_Bootcamp_June_2021_Pilot_Day3_Part_4_of_4.mp4 GL4U_RNAseq_Bootcamp_June_2021_Pilot_Day4_Part_1_of_4.mp4 GL4U_RNAseq_Bootcamp_June_2021_Pilot_Day4_Part_2_of_4.mp4 GL4U_RNAseq_Bootcamp_June_2021_Pilot_Day4_Part_3_of_4.mp4 GL4U_RNAseq_Bootcamp_June_2021_Pilot_Day4_Part_4_of_4.mp4 GL4U_RNAseq_Bootcamp_June_2021_Pilot_Day5_Part_1_of_4.mp4 GL4U_RNAseq_Bootcamp_June_2021_Pilot_Day5_Part_2_of_4.mp4 GL4U_RNAseq_Bootcamp_June_2021_Pilot_Day5_Part_3_of_4.mp4 GL4U_RNAseq_Bootcamp_June_2021_Pilot_Day5_Part_4_of_4.mp4

GeneLab↗

A consensus-based ensemble approach to improve transcriptome assembly

Systems-level analyses, such as differential gene expression analysis, co-expression analysis, and metabolic pathway reconstruction, depend on the accuracy of the transcriptome. Multiple tools exist to perform transcriptome assembly from RNAseq data. However, assembling high quality transcriptomes is still not a trivial problem. This is especially the case for non-model organisms where adequate reference genomes are often not available. Different methods produce different transcriptome models and there is no easy way to determine which are more accurate. Furthermore, having alternative-splicing events exacerbates such difficult assembly problems. While benchmarking transcriptome assemblies is critical, this is also not trivial due to the general lack of true reference transcriptomes. In this study, we first provide a pipeline to generate a set of the simulated benchmark transcriptome and corresponding RNAseq data. Using the simulated benchmarking datasets, we compared the performance of various transcriptome assembly approaches including both de novo and genome-guided methods. The results showed that the assembly performance deteriorates significantly when alternative transcripts (isoforms) exist or for genome-guided methods when the reference is not available from the same genome. To improve the transcriptome assembly performance, leveraging the overlapping predictions between different assemblies, we present a new consensus-based ensemble transcriptome assembly approach, ConSemble. Without using a reference genome, ConSemble using four de novo assemblers achieved an accuracy up to twice as high as any de novo assemblers we compared. When a reference genome is available, ConSemble using four genome-guided assemblies removed many incorrectly assembled contigs with minimal impact on correctly assembled contigs, achieving higher precision and accuracy than individual genome-guided methods. Furthermore, ConSemble using de novo assemblers matched or exceeded the best performing genome-guided assemblers even when the transcriptomes included isoforms. We thus demonstrated that the ConSemble consensus strategy both for de novo and genome-guided assemblers can improve transcriptome assembly. The RNAseq simulation pipeline, the benchmark transcriptome datasets, and the script to perform the ConSemble assembly are all freely available from: http://bioinfolab.unl.edu/emlab/consemble/.

59 BASIC BIOLOGICAL SCIENCES↗

Data for An Orphan Gene BOOSTER Enhances Photosynthetic Efficiency and Plant Productivity

Seeds of Col-0 wild type, sig6 T-DNA mutants (CS877785, ABRC), PRL-1-OE, and sig6 T-DNA mutants transfected with PRL-1 (sig6::PRL-1) were planted in 1/2 MS media. Seedlings growth including chlorophyll development defects were investigated across the genotypes. Four-days-old-post-light exposure seedlings were harvested and performed RNAseq analysis with four biological replicates.

Biomass Analytics↗

Identification of integrated proteomics and transcriptomics signature of alcohol-associated liver disease using machine learning

Distinguishing between alcohol-associated hepatitis (AH) and alcohol-associated cirrhosis (AC) remains a diagnostic challenge. In this study, we used machine learning with transcriptomics and proteomics data from liver tissue and peripheral mononuclear blood cells (PBMCs) to classify patients with alcohol-associated liver disease. The conditions in the study were AH, AC, and healthy controls. We processed 98 PBMC RNAseq samples, 55 PBMC proteomic samples, 48 liver RNAseq samples, and 53 liver proteomic samples. First, we built separate classification and feature selection pipelines for transcriptomics and proteomics data. The liver tissue models were validated in independent liver tissue datasets. Next, we built integrated gene and protein expression models that allowed us to identify combined gene-protein biomarker panels. For liver tissue, we attained 90% nested-cross validation accuracy in our dataset and 82% accuracy in the independent validation dataset using transcriptomic data. We attained 100% nested-cross validation accuracy in our dataset and 61% accuracy in the independent validation dataset using proteomic data. For PBMCs, we attained 83% and 89% accuracy with transcriptomic and proteomic data, respectively. The integration of the two data types resulted in improved classification accuracy for PBMCs, but not liver tissue. We also identified the following gene-protein matches within the gene-protein biomarker panels: CLEC4M-CLC4M, GSTA1-GSTA2 for liver tissue and SELENBP1-SBP1 for PBMCs. In this study, machine learning models had high classification accuracy for both transcriptomics and proteomics data, across liver tissue and PBMCs. The integration of transcriptomics and proteomics into a multi-omics model yielded improvement in classification accuracy for the PBMC data. The set of integrated gene-protein biomarkers for PBMCs show promise toward developing a liquid biopsy for alcohol-associated liver disease.

60 APPLIED LIFE SCIENCES↗

Genetics and Genomics of Pathogen Resistance in Switchgrass (Final Report)

This project was funded by DOE under Grant no. DE-SC0016108. Originally approved for the 2016-2019 period, two no-cost extensions were solicited and approved, which prolonged the lifespan through July 2021. This final report informs on the results obtained so far from the research implemented. The research hinged on integrating genomics (genomic selection, RNAseq, virus-plant interactions) with classical genetics (conventional breeding) to incorporate durable resistance to fungal (rust) and viral (mosaic) diseases in switchgrass (Panicum virgatum) populations being bred for bioenergy. Higher biomass yield, higher quality (low lignin content), and durable disease resistance are key features to make lignocellulosic switchgrass feedstocks economically competitive and sustainable. Genomic selection is being applied on three generations of a switchgrass population derived from crossing two ecotypes (Kanlow as lowland female and Summer as upland male) with differential performance in terms of biomass yield and quality, disease resistance, and winter survivability. Target populations were screened for rust and mosaic in field and/or lab and phenotyped for biomass yield and quality traits. Genetic analyses were applied across generations to capture the joint inheritance of the targeted traits and predict breeding values for parents and progeny with greater accuracy. Parental and a panel of different switchgrass populations were genotyped with the DArTseq technology to develop SNP (0, 1, 2) and in-silico (presence/absence) DArT markers. Rust inoculations techniques were developed and applied successfully on switchgrass. The original populations (Kanlow and Summer) were sequenced with RNAseq to capture the gene expression profiles across sequential time-points and appraise the basis of greater resistance in the Kanlow vs the Summer ecotype. Constructs of PMV and sPMV mosaic virus were assembled and tested first on proso millet to find the best protocol to use later on switchgrass. Results from the preliminary analyses indicate that 1) ample additive genetic variation is available for selection and improving this inter-ecotypic population for yield, quality, and disease traits, 2) significant gains are to be expected with the genetic correlations being favorable between yield and lignin content and between yield and disease ratings, 3) substantial differences exist in the genetic regions controlling rust resistance in the two ecotypes, 4) co-infection with PMV isolates from Nebraska and its satellite from Kansas elicit severe mosaic symptoms, and 5) two different genetic systems are responsible for imparting resistance to rust and virus in switchgrass.

59 BASIC BIOLOGICAL SCIENCES↗

Enhanced Resistance Pines for Improved Renewable Biofuel and Chemical Production (Technical Report)

We completed phenotyping constitutive and inducible oleoresin flow across two seasons, constitutive resin canal number and density and wood terpene content in our ADEPT2 and CCLONES populations. We completed genetic association between 19 oleoresin phenotypes and a total of 523,192 SNP markers from ADEPT2 and 13,883 SNP markers in CCLONES using four mixed linear models. A total of 293 significant SNPs (FDR = 0.20) were identified. We used the MENTOR tool to mine mechanistic connections from a multiplex network constructed from poplar multi-omic data to construct a conceptual model for a subset of these significant SNPs. Our model contains 6 transcriptional regulators in addition to 3 monoterpene synthases. To generate more lines of evidence for these significant SNPs, we completed a time course RNAseq experiment after inducing vascular zone cells to differentiate into new resin canals with a methyl jasmonate treatment, a single nuclei RNAseq that identified differentiating resin canal epithelial cells and are completing analysis for a QTL study in a hybrid pine population. The time course identified 4634 significantly down and 1890 significantly up regulated transcripts after treatment with methyl jasmonate, an inducer of new resin canal formation in the vascular cambial meristem. To analyze this large set of differentially regulated genes, we created a predictive expression network and analyzed it with random walk restart using 6 seed genes coding for transcription factors regulating xylem differentiation in poplar. Of the top ranked 200 transcripts, 119 transcripts were significant differentially expressed supporting these transcripts as potential candidates regulating resin canal formation. Analysis of single nuclei sequencing of shoot tips that contain differentiating resin canals, identified 10 clusters. One cluster was highly enriched in transcripts coding for 9 of the enzymes in the MEP pathway 3 prenyl synthetases, and 3 monoterpene synthases strongly suggesting that this cluster represents resin canal epithelial cells. We are mining the additional transcripts to create a trajectory analysis. In summary, we have identified > 10 novel genes that are strongly supported candidates for further analysis in breeding lines and for genetic engineering over- and under- expressing lines to increase wood terpene content to improve resistance to insect and fungal pathogens while simultaneously increasing terpene supplies for renewable chemicals and biofuels.

59 BASIC BIOLOGICAL SCIENCES↗

Evaluation of Correction Methods for NASA GeneLab Transcriptomic Datasets

Conducting space biology experiments aboard the International Space Station, particularly those utilizing complex model organisms like mice, is expensive and difficult due to limited crew availability, hardware, and space. As a result, sample numbers from these studies are low, reducing the statistical power of any one experiment. Aggregating spaceflight datasets serves as a method to increase sample numbers, allowing for novel insights through bioinformatic analysis of ‘omics data from merged datasets. However, aggregating datasets can introduce unwanted variation including 1) differences in sample handling, processing, and sequencing platforms between datasets (technical variation) as well as 2) differences in experimental design between datasets. In the present study, NASA GeneLab-hosted RNAseq datasets from mouse liver tissues were used to evaluate several statistical methods to correct for this unwanted variation through two approaches, reference-based and standard. The following correction algorithms were applied with (reference-based) and/or without (standard) considering Universal Mouse RNA Reference samples: ComBat and ComBat_seq from the SVA package, median polish, empirical Bayes, and ANOVA-based algorithms from the MBatch package, and negative binomial regression normalization in the DESeq2 package. For each approach, after the correction algorithm was applied, differential gene expression (DGE) analysis of flight and ground control samples was performed with the combined data. The robustness of each tool was evaluated using BatchQC to determine statistical differences between datasets before and after correction, Principal Component Analysis to evaluate global gene expression in samples before and after correction, and by comparing DGE analysis of individual datasets and combined datasets before and after correction. The results showed that the reference-based approach introduced several additional (and likely artificial) DEGs when compared with the respective standard approach. Of the methods tested, standard ComBat and DESeq2 were identified as the most robust correction methods for combining spaceflight mouse liver RNAseq datasets hosted on GeneLab.

GeneLab↗

Combining RNA-SEQ Datasets from NASA GENELAB: An Evaluation of Correction Methods

Background: Conducting space biology experiments aboard the International Space Station, particularly those utilizing complex model organisms like mice, is expensive and difficult due to limited crew availability, hardware, and space. As a result, sample numbers from these studies are low, reducing the statistical power of any one experiment. Aggregating spaceflight datasets serves as a method to increase sample numbers, allowing for novel insights through bioinformatic analysis of ‘omics data from merged datasets. However, aggregating datasets can introduce unwanted variation including 1) differences in sample handling, processing, and sequencing platforms between datasets (technical variation) as well as 2) differences in experimental design between datasets. Methods: In the present study, NASA GeneLab-hosted RNAseq datasets from mouse liver tissues were used to evaluate several statistical methods to correct for this unwanted variation through two approaches, reference-based and standard. The following correction algorithms were applied with (reference-based) and/or without (standard) considering Universal Mouse RNA Reference samples: ComBat and ComBat_seq from the SVA package, the median polish, empirical Bayes, and ANOVA-based algorithms from the MBatch package, and negative binomial regression normalization in the DESeq2 package. For each approach, after the correction algorithm was applied, differential gene expression (DGE) analysis of flight and ground control samples was performed with the combined data. The robustness of each tool was evaluated using BatchQC to determine statistical differences between datasets before and after correction, Principal Component Analysis to evaluate global gene expression in samples before and after correction, and by comparing DGE analysis of individual datasets and combined datasets before and after correction. Results: The results showed that the reference-based approach introduced several additional (and likely artificial) differentially expressed genes when compared with the respective standard approach. Conclusions: Of the methods tested, standard ComBat_seq and DESeq2 were identified as the most robust correction methods for combining spaceflight mouse liver RNAseq datasets hosted on GeneLab.

Finsam Samson↗

Transcriptomics-based Machine Learning Analysis Predicts Space-Exposed Murine Livers

Limited sample sizes, high data dimensionality, and sensitivity to technical and biological variability of next generation sequencing (NGS), has typically limited machine learning (ML) in space studies and further study of radiation effects. However, pooling smaller studies while addressing intra- and inter-study variabilities allows for ML predictive modeling. Here, integration methods were applied to whole transcriptome shotgun sequencing (RNAseq) data from 6 mouse liver GeneLab datasets (GLDS) with a total of 113 spaceflight and ground-control samples to determine top features relevant to spaceflight including the effect of radiation exposure. Data was normalized within each study, then merged and scaled across all datasets. Data dimensionality was reduced using a minimum redundancy maximum relevance (MRMR) methodology. The top MRMR features were used to predict spaceflight vs. ground-control samples using a Random Forest (RF) classifier with 5-fold cross validation (CV). The ML-based gene sets were further compared against differential gene expression results from individual GLDS. CV training using the top 100 MRMR genes show averages of 86% accuracy and 0.95 AUC value on the validation set over 5 folds (Figure 1A). Baseline set analysis on differentially expressed genes (DEGs) identified using padj ≤ 0.05 show 811 or 68 DEGs overlapping between at least 2 or 3 studies, respectively (Figure 1B). Over-representation analysis showed overlapping biological processes related to fatty acid and lipid metabolism. Set analysis between the MRMR features and the DEGs showed 60 or 8 genes overlapping with at least 1 or 2 studies, respectively. MRMR feature selection and ensemble ML methods (e.g. RF) improve performance relative to a Naïve Bayes classifier when NGS data sets are analyzed. A challenge of applying ML methods across heterogeneous NGS data is accounting for signal:noise ratio. Here, signal validation across studies was shown by intersecting sets between top MRMR genes and DEGs from RNASeq analysis. Non-intersecting sets introduce opportunity to explore spaceflight relevant genes and implementing ML methods across existing NGS datasets may overcome sample size limitations. ML coupled with existing analytical methods enhances understanding of disease by revealing common underlying pathways across datasets.

Machine Learning↗