Search NASA⌕ Search

DOE OSTI · 1825743

A consensus-based ensemble approach to improve transcriptome assembly

Abstract

Systems-level analyses, such as differential gene expression analysis, co-expression analysis, and metabolic pathway reconstruction, depend on the accuracy of the transcriptome. Multiple tools exist to perform transcriptome assembly from RNAseq data. However, assembling high quality transcriptomes is still not a trivial problem. This is especially the case for non-model organisms where adequate reference genomes are often not available. Different methods produce different transcriptome models and there is no easy way to determine which are more accurate. Furthermore, having alternative-splicing events exacerbates such difficult assembly problems. While benchmarking transcriptome assemblies is critical, this is also not trivial due to the general lack of true reference transcriptomes. In this study, we first provide a pipeline to generate a set of the simulated benchmark transcriptome and corresponding RNAseq data. Using the simulated benchmarking datasets, we compared the performance of various transcriptome assembly approaches including both de novo and genome-guided methods. The results showed that the assembly performance deteriorates significantly when alternative transcripts (isoforms) exist or for genome-guided methods when the reference is not available from the same genome. To improve the transcriptome assembly performance, leveraging the overlapping predictions between different assemblies, we present a new consensus-based ensemble transcriptome assembly approach, ConSemble. Without using a reference genome, ConSemble using four de novo assemblers achieved an accuracy up to twice as high as any de novo assemblers we compared. When a reference genome is available, ConSemble using four genome-guided assemblies removed many incorrectly assembled contigs with minimal impact on correctly assembled contigs, achieving higher precision and accuracy than individual genome-guided methods. Furthermore, ConSemble using de novo assemblers matched or exceeded the best performing genome-guided assemblers even when the transcriptomes included isoforms. We thus demonstrated that the ConSemble consensus strategy both for de novo and genome-guided assemblers can improve transcriptome assembly. The RNAseq simulation pipeline, the benchmark transcriptome datasets, and the script to perform the ConSemble assembly are all freely available from: http://bioinfolab.unl.edu/emlab/consemble/.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Voshall, Adam, Behera, Sairam, Li, Xiangjun, Yu, Xiao-Hong, Kapil, Kushagra, Deogun, Jitender S., Shanklin, John, Cahoon, Edgar B., Moriyama, Etsuko N.. 2021-10-21. A consensus-based ensemble approach to improve transcriptome assembly. https://doi.org/10.1186/s12859-021-04434-8

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related reports

Soil metagenomics umbrella narrative

Implementing accessible, authentic research experiences in introductory courses is challenging, particularly at institutions serving diverse student populations. To address this gap, we developed and deployed a Course-based Undergraduate Research Experience (CURE) focused on plant-microbe interactions in General Biology II at Northeastern Illinois University (NEIU), a minority-serving institution with a diverse student body. Students grew sugar beets (Beta vulgaris), extracted DNA from the rhizoplane, and used the Department of Energy Systems Biology Knowledgebase (KBase) for bioinformatic analysis to compare microbial relative abundance in fertilized versus unfertilized soil. Over five semesters, the CURE engaged 103 students and leveraged the intuitive KBase platform to make complex sequencing data accessible. Pre/post-course survey data revealed significant increases in student self-assessed research skills, including the ability to explain results and determine the types of data to collect. Furthermore, students reported significant gains in confidence related to experimental design and hypothesis development, alongside a strong increase in familiarity with KBase. Informal faculty feedback indicated high student engagement and appreciation for the real-world connections (e.g. food systems, agriculture, and health). This scalable, low-cost model effectively integrates data science tools into the foundational curriculum, demonstrating a potent strategy for boosting research skills and broadening participation in authentic scientific inquiry among diverse undergraduate students.

59 BASIC BIOLOGICAL SCIENCES↗

Genome-resolved insights into microbial diversity and elemental cycling in Winogradsky columns

We retained 18 MAGs with ≥50% completion and <10% contamination (i.e., at least medium quality). Of these, 10 had >90% completion and <5% contamination; however, only one (Paceibacteria Bin.003_MG) can be described as high-quality, as the others lacked a full suite of 5S, 16S, and 23S rRNA genes. To maximize the diversity of our recovered MAGs, we also retained one MAG (Chromatiaceae Bin.008_AM) with >40% (but less than 50%) completion and <5% contamination, as well as one (Rhodopseudomonas Bin.015_MK) with >90% completion and <20% (but>10%) contamination. Interestingly, significant chimerism was not detected in this MAG (40) , suggesting that the elevated contamination (20%) may instead reflect two closely related strains collapsing into a single bin. Consistent with this, contig coverage was bimodal, with roughly 17% of the assembly at ~115x and the remaining 83% at ~282x, while GC content remained uniform across both groups (~64%), arguing against contamination from a taxonomically distinct source.

59 BASIC BIOLOGICAL SCIENCES↗