Search NASA⌕ Search

SEARCH · Search NASA

Results for “RNA sequencing”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 181 records · Page 10

Transcriptomics-based Machine Learning Analysis Predicts Space-Exposed Murine Livers

Limited sample sizes, high data dimensionality, and sensitivity to technical and biological variability of next generation sequencing (NGS), typically limits machine learning (ML) approaches in spaceflight studies that include radiation effects. However, pooling smaller studies while addressing intra- and inter-study variabilities allows for ML predictive modeling. Here, integration methods were applied to whole transcriptome shotgun sequencing (RNA-seq) data from six mouse liver GeneLab datasets (GLDS) (n ranging from 6 to 39 samples) from with a total of 81 spaceflight and ground-control samples to determine top features (i.e. genes) relevant to spaceflight including the effect of radiation exposure. RNASeq counts were normalized for each study, then merged and scaled across all datasets. Data dimensionality was reduced using a minimum redundancy maximum relevance (MRMR) methodology. Redundancy and relevance were computed using the Pearson correlation and F-statistic, respectively. The top 100 MRMR features were used to predict spaceflight vs. ground-control samples using Random Forest (RF), Support Vector Machine (SVM), and Linear Discriminant Analysis (LDA) classifiers with 5-fold cross validation (CV). Principal component analysis (PCA) on the complete feature set versus the MRMR features shows separation between spaceflight samples and ground controls (Figure 1A). The ML-based gene sets were compared against differential gene expression results obtained with DESeq2 from individual GLDS. Using all features or randomly sampled subsets at matching set sizes with MRMR, a maximum classifier accuracy of 69% was shown on the test set over 5 folds. For all classifiers, CV training using at least the top 30 MRMR genes show minimum 89% accuracy and 0.95 AUC value on the test set over 5 folds (Figure 1B). Baseline set analysis on differentially expressed genes (DEGs) identified using padj ≤ 0.05 show 295 DEGs that overlap at least two studies and 13 DEGs that overlap three studies (Figure 1C). Set analysis between the top 100 MRMR features and the DEGs showed 47 genes that overlap at least one study and 24 genes that overlap two studies. Over-representation analysis showed overlapping biological processes related to fatty acid and lipid metabolism which may indicate these processes in the response to spaceflight stressors. MRMR feature selection for the selected ML methods improve performance relative to a classifier built on all features or randomly sampled subsets. Permutation feature importance within the decorrelated MRMR features showed concordance in feature ranking between ML methods. A challenge of applying ML methods across heterogeneous NGS data is accounting for signal:noise. Here, signal validation across studies was shown by intersecting sets between top MRMR genes and DEGs from DESeq2 analysis. Non-intersecting sets introduce opportunity to explore genes relevant to differentiating space flight exposed groups and implementing ML methods across existing NGS datasets may overcome sample size limitations.

Machine Learning↗

Transcriptomics-based Machine Learning (ML) Analysis Predicts Space-Exposed Murine Livers

Limited sample sizes, high data dimensionality, and sensitivity to technical and biological variability of next generation sequencing (NGS), typically limits machine learning (ML) approaches in spaceflight studies that include radiation effects. However, pooling smaller studies while addressing intra- and inter-study variabilities allows for ML predictive modeling. Here, integration methods were applied to whole transcriptome shotgun sequencing (RNA-seq) data from six mouse liver GeneLab datasets (GLDS) (n ranging from 6 to 39 samples) from with a total of 81 spaceflight and ground-control samples to determine top features (i.e. genes) relevant to spaceflight including the effect of radiation exposure. RNASeq counts were normalized for each study, then merged and scaled across all datasets. Data dimensionality was reduced using a minimum redundancy maximum relevance (MRMR) methodology. Redundancy and relevance were computed using the Pearson correlation and F-statistic, respectively. The top 100 MRMR features were used to predict spaceflight vs. ground-control samples using Random Forest (RF), Support Vector Machine (SVM), and Linear Discriminant Analysis (LDA) classifiers with 5-fold cross validation (CV). Principal component analysis (PCA) on the complete feature set versus the MRMR features shows separation between spaceflight samples and ground controls (Figure 1A). The ML-based gene sets were compared against differential gene expression results obtained with DESeq2 from individual GLDS. Using all features or randomly sampled subsets at matching set sizes with MRMR, a maximum classifier accuracy of 69% on the test set over 5 folds. For all classifiers, CV training using at least the top 30 MRMR genes show minimum 89% accuracy and 0.95 AUC value on the test set over 5 folds (Figure 1B). Baseline set analysis on differentially expressed genes (DEGs) identified using padj ≤ 0.05 show 295 DEGs that overlap at least two studies and 13 DEGs that overlap three studies (Figure 1C). Set analysis between the top 100 MRMR features and the DEGs showed 47 genes that overlap at least one study and 24 genes that overlap two studies. Over-representation analysis showed overlapping biological processes related to fatty acid and lipid metabolism which may indicate these processes in the response to spaceflight stressors. MRMR feature selection for the selected ML methods improve performance relative to a classifier built on all features or randomly sampled subsets. Permutation feature importance within the decorrelated MRMR features showed concordance in feature ranking between ML methods. A challenge of applying ML methods across heterogeneous NGS data is accounting for signal:noise. Here, signal validation across studies was shown by intersecting sets between top MRMR genes and DEGs from DESeq2 analysis. Non-intersecting sets introduce opportunity to explore genes relevant to differentiating space flight exposed groups and implementing ML methods across existing NGS datasets may overcome sample size limitations.

Machine Learning↗

Transcriptomics-based Machine Learning Analysis Predicts Space-Exposed Murine Livers

Limited sample sizes, high data dimensionality, and sensitivity to technical and biological variability of next generation sequencing (NGS), typically limits machine learning (ML) approaches in spaceflight studies that include radiation effects. However, pooling smaller studies while addressing intra- and inter-study variabilities allows for ML predictive modeling. Here, integration methods were applied to whole transcriptome shotgun sequencing (RNA-seq) data from six mouse liver GeneLab datasets (GLDS) (n ranging from 6 to 39 samples) from with a total of 81 spaceflight and ground-control samples to determine top features (i.e. genes) relevant to spaceflight including the effect of radiation exposure. RNASeq counts were normalized for each study, then merged and scaled across all datasets. Data dimensionality was reduced using a minimum redundancy maximum relevance (MRMR) methodology. Redundancy and relevance were computed using the Pearson correlation and F-statistic, respectively. The top 100 MRMR features were used to predict spaceflight vs. ground-control samples using Random Forest (RF), Support Vector Machine (SVM), and Linear Discriminant Analysis (LDA) classifiers with 5-fold cross validation (CV). Principal component analysis (PCA) on the complete feature set versus the MRMR features shows separation between spaceflight samples and ground controls (Figure 1A). The ML-based gene sets were compared against differential gene expression results obtained with DESeq2 from individual GLDS. Using all features or randomly sampled subsets at matching set sizes with MRMR, a maximum classifier accuracy of 69% was shown on the test set over 5 folds. For all classifiers, CV training using at least the top 30 MRMR genes show minimum 89% accuracy and 0.95 AUC value on the test set over 5 folds (Figure 1B). Baseline set analysis on differentially expressed genes (DEGs) identified using padj ≤ 0.05 show 295 DEGs that overlap at least two studies and 13 DEGs that overlap three studies (Figure 1C). Set analysis between the top 100 MRMR features and the DEGs showed 47 genes that overlap at least one study and 24 genes that overlap two studies. Over-representation analysis showed overlapping biological processes related to fatty acid and lipid metabolism which may indicate these processes in the response to spaceflight stressors. MRMR feature selection for the selected ML methods improve performance relative to a classifier built on all features or randomly sampled subsets. Permutation feature importance within the decorrelated MRMR features showed concordance in feature ranking between ML methods. A challenge of applying ML methods across heterogeneous NGS data is accounting for signal:noise. Here, signal validation across studies was shown by intersecting sets between top MRMR genes and DEGs from DESeq2 analysis. Non-intersecting sets introduce opportunity to explore genes relevant to differentiating space flight exposed groups and implementing ML methods across existing NGS datasets may overcome sample size limitations.

Machine Learning↗

Nanopores and nucleic acids: prospects for ultrarapid sequencing

DNA and RNA molecules can be detected as they are driven through a nanopore by an applied electric field at rates ranging from several hundred microseconds to a few milliseconds per molecule. The nanopore can rapidly discriminate between pyrimidine and purine segments along a single-stranded nucleic acid molecule. Nanopore detection and characterization of single molecules represents a new method for directly reading information encoded in linear polymers. If single-nucleotide resolution can be achieved, it is possible that nucleic acid sequences can be determined at rates exceeding a thousand bases per second.

Review↗

The rRNA evolution and procaryotic phylogeny

Studies of ribosomal RNA primary structure allow reconstruction of phylogenetic trees for prokaryotic organisms. Such studies reveal major dichotomy among the bacteria that separates them into eubacteria and archaebacteria. Both groupings are further segmented into several major divisions. The results obtained from 5S rRNA sequences are essentially the same as those obtained with the 16S rRNA data. In the case of Gram negative bacteria the ribosomal RNA sequencing results can also be directly compared with hybridization studies and cytochrome c sequencing studies. There is again excellent agreement among the several methods. It seems likely then that the overall picture of microbial phylogeny that is emerging from the RNA sequence studies is a good approximation of the true history of these organisms. The RNA data allow examination of the evolutionary process in a semi-quantitative way. The secondary structures of these RNAs are largely established. As a result it is possible to recognize examples of local structural evolution. Evolutionary pathways accounting for these events can be proposed and their probability can be assessed.

Fox, G. E.↗

Assembly of catalytic complexes from randomized oligonucleotides

The early evolution of life relied on catalytic RNAs (ribozymes) for central functions. To test whether early catalysts could have assembled from multiple short nucleic acid fragments in random sequence environments, we performed an in vitro selection from a short RNA library in the presence of 256 different DNA 20-nucleotide oligomers. High-throughput sequencing and biochemical analysis showed that most of the selected 1331 RNA sequences required at least one DNA for activity. Representatives for four of six RNA clusters that depended on DNA cofactors were active even when the 256 DNAs were replaced by completely random DNA 20-nucleotide oligomers. The formation of these catalytic complexes and the recruitment of oligonucleotide cofactors from completely random libraries demonstrate an important principle for the emergence of the earliest oligonucleotide catalysts.

Xu Han↗

Mechanistic modeling of in vitro transcription incorporating effects of magnesium pyrophosphate crystallization

The in vitro transcription (IVT) reaction used in the production of messenger RNA vaccines and therapies remains poorly quantitatively understood. Mechanistic modeling of IVT could inform reaction design, scale-up, and control. In this work, we develop a mechanistic model of IVT to include nucleation and growth of magnesium pyrophosphate crystals and subsequent agglomeration of crystals and DNA. To help generalize this model to different constructs, a novel quantitative description is included for the rate of transcription as a function of target sequence length, DNA concentration, and T7 RNA polymerase concentration. The model explains previously unexplained trends in IVT data and quantitatively predicts the effect of adding the pyrophosphatase enzyme to the reaction system. The model is validated on additional literature data showing an ability to predict transcription rates as a function of RNA sequence length.

59 BASIC BIOLOGICAL SCIENCES↗

A minimal complex of KHNYN and zinc-finger antiviral protein binds and degrades single-stranded RNA

Detecting viral infection is a key role of the innate immune system. The genomes of some RNA viruses have a high CpG dinucleotide content relative to most vertebrate cell RNAs, making CpGs a molecular marker of infection. The human zinc-finger antiviral protein (ZAP) recognizes CpG, mediates clearance of the foreign CpG-rich RNA, and causes attenuation of CpG-rich RNA viruses. While ZAP binds RNA, it lacks enzymatic activity that might be responsible for RNA degradation and thus requires interacting cofactors for its function. One of these cofactors, KHNYN, has a predicted nuclease domain. Using biochemical approaches, we found that the KHNYN NYN domain is a single-stranded RNA ribonuclease that does not have sequence specificity and digests RNA with or without CpG dinucleotides equivalently in vitro. We show that unlike most KH domains, the KHNYN KH domain does not bind RNA. Indeed, a crystal structure of the KH region revealed a double-KH domain with a negatively charged surface that accounts for the lack of RNA binding. Rather, the KHNYN C-terminal domain (CTD) interacts with the ZAP RNA-binding domain (RBD) to provide target RNA specificity. We define a minimal complex composed of the ZAP RBD and the KHNYN NYN-CTD and use a fluorescence polarization assay to propose a model for how this complex interacts with a CpG dinucleotide-containing RNA. In the context of the cell, this module would represent the minimum ZAP and KHNYN domains required for CpG-recognition and ribonuclease activity essential for attenuation of viruses with clusters of CpG dinucleotides.

Yeoh, Zoe C. (ORCID:0000000226949068)↗

Chance and necessity in the selection of nucleic acid catalysts

In Tom Stoppard's famous play [Rosencrantz and Guildenstern are Dead], the ill-fated heroes toss a coin 101 times. The first 100 times they do so the coin lands heads up. The chance of this happening is approximately 1 in 10(30), a sequence of events so rare that one might argue that it could only happen in such a delightful fiction. Similarly rare events, however, may underlie the origins of biological catalysis. What is the probability that an RNA, DNA, or protein molecule of a given random sequence will display a particular catalytic activity? The answer to this question determines whether a collection of such sequences, such as might result from prebiotic chemistry on the early earth, is extremely likely or unlikely to contain catalytically active molecules, and hence whether the origin of life itself is a virtually inevitable consequence of chemical laws or merely a bizarre fluke. The fact that a priori estimates of this probability, given by otherwise informed chemists and biologists, ranged from 10(-5) to 10(-50), inspired us to begin to address the question experimentally. As it turns out, the chance that a given random sequence RNA molecule will be able to catalyze an RNA polymerase-like phosphoryl transfer reaction is close to 1 in 10(13), rare enough, to be sure, but nevertheless in a range that is comfortably accessible by experiment. It is the purpose of this Account to describe the recent advances in combinatorial biochemistry that have made it possible for us to explore the abundance and diversity of catalysts existing in nucleic acid sequence space.

Review, Tutorial↗

Developing Asparagaceae1726: An Asparagaceae‐specific probe set targeting 1726 loci for Hyb‐Seq and phylogenomics in the family

Abstract Premise Target sequence capture (Hyb‐Seq) is a cost‐effective sequencing strategy that employs RNA probes to enrich for specific genomic sequences. By targeting conserved low‐copy orthologs, Hyb‐Seq enables efficient phylogenomic investigations. Here, we present Asparagaceae1726—a Hyb‐Seq probe set targeting 1726 low‐copy nuclear genes for phylogenomics in the angiosperm family Asparagaceae—which will aid the often‐challenging delineation and resolution of evolutionary relationships within Asparagaceae. Methods Here we describe and validate the Asparagaceae1726 probe set (https://github.com/bentzpc/Asparagaceae1726) in six of the seven subfamilies of Asparagaceae. We perform phylogenomic analyses with these 1726 loci and evaluate how inclusion of paralogs and bycatch plastome sequences can enhance phylogenomic inference with target‐enriched data sets. Results We recovered at least 82% of target orthologs from all sampled taxa, and phylogenomic analyses resulted in strong support for all subfamilial relationships. Additionally, topology and branch support were congruent between analyses with and without inclusion of target paralogs, suggesting that paralogs had limited effect on phylogenomic inference. Discussion Asparagaceae1726 is effective across the family and enables the generation of robust data sets for phylogenomics of any Asparagaceae taxon. Asparagaceae1726 establishes a standardized set of loci for phylogenomic analysis in Asparagaceae, which we hope will be widely used for extensible and reproducible investigations of diversification in the family.

Plant Sciences↗

A potential role for RNA aminoacylation prior to its role in peptide synthesis

Coded ribosomal peptide synthesis could not have evolved unless its sequence and amino acid–specific aminoacylated tRNA substrates already existed. We therefore wondered whether aminoacylated RNAs might have served some primordial function prior to their role in protein synthesis. Here, we show that specific RNA sequences can be nonenzymatically aminoacylated and ligated to produce amino acid–bridged stem-loop RNAs. We used deep sequencing to identify RNAs that undergo highly efficient glycine aminoacylation followed by loop-closing ligation. The crystal structure of one such glycine-bridged RNA hairpin reveals a compact internally stabilized structure with the same eponymous T-loop architecture that is found in many noncoding RNAs, including the modern tRNA. We demonstrate that the T-loop-assisted amino acid bridging of RNA oligonucleotides enables the rapid template-free assembly of a chimeric version of an aminoacyl-RNA synthetase ribozyme. We suggest that the primordial assembly of amino acid–bridged chimeric ribozymes provides a direct and facile route for the covalent incorporation of amino acids into RNA. A greater functionality of covalently incorporated amino acids could contribute to enhanced ribozyme catalysis, providing a driving force for the evolution of sequence and amino acid–specific aminoacyl-RNA synthetase ribozymes in the RNA World. The synthesis of specifically aminoacylated RNAs, an unlikely prospect for nonenzymatic reactions but a likely one for ribozymes, could have set the stage for the subsequent evolution of coded protein synthesis.

Science & Technology - Other Topics↗

Populus VariantDB v3.2 facilitates CRISPR and functional genomics research

The success of CRISPR genome editing studies depends critically on the precision of guide RNA (gRNA) design. Sequence polymorphisms in outcrossing tree species pose design hazards that can render CRISPR genome editing ineffective. Despite recent advances in tree genome sequencing with haplotype resolution, sequence polymorphism information remains largely inaccessible to various functional genomics research efforts. The Populus VariantDB v3.2 addresses these challenges by providing a user-friendly search engine to query sequence polymorphisms of heterozygous genomes. The database accepts short sequences, such as gRNAs and primers, as input for searching against multiple poplar genomes, including hybrids, with customizable parameters. We provide examples to showcase the utilities of VariantDB in improving the precision of gRNA or primer design. The platform-agnostic nature of the probe search design makes Populus VariantDB v3.2 a versatile tool for the rapidly evolving CRISPR field and other sequence-sensitive functional genomics applications. The database schema is expandable and can accommodate additional tree genomes to broaden its user base.

59 BASIC BIOLOGICAL SCIENCES↗

Emergence of a replicating species from an in vitro RNA evolution reaction

The technique of self-sustained sequence replication allows isothermal amplification of DNA and RNA molecules in vitro. This method relies on the activities of a reverse transcriptase and a DNA-dependent RNA polymerase to amplify specific nucleic acid sequences. We have modified this protocol to allow selective amplification of RNAs that catalyze a particular chemical reaction. During an in vitro RNA evolution experiment employing this modified system, a unique class of "selfish" RNAs emerged and replicated to the exclusion of the intended RNAs. Members of this class of selfish molecules, termed RNA Z, amplify efficiently despite their inability to catalyze the target chemical reaction. Their amplification requires the action of both reverse transcriptase and RNA polymerase and involves the synthesis of both DNA and RNA replication intermediates. The proposed amplification mechanism for RNA Z involves the formation of a DNA hairpin that functions as a template for transcription by RNA polymerase. This arrangement links the two strands of the DNA, resulting in the production of RNA transcripts that contain an embedded RNA polymerase promoter sequence.

Non-NASA Center↗

A new version of the RDP (Ribosomal Database Project)

The Ribosomal Database Project (RDP-II), previously described by Maidak et al. [ Nucleic Acids Res. (1997), 25, 109-111], is now hosted by the Center for Microbial Ecology at Michigan State University. RDP-II is a curated database that offers ribosomal RNA (rRNA) nucleotide sequence data in aligned and unaligned forms, analysis services, and associated computer programs. During the past two years, data alignments have been updated and now include >9700 small subunit rRNA sequences. The recent development of an ObjectStore database will provide more rapid updating of data, better data accuracy and increased user access. RDP-II includes phylogenetically ordered alignments of rRNA sequences, derived phylogenetic trees, rRNA secondary structure diagrams, and various software programs for handling, analyzing and displaying alignments and trees. The data are available via anonymous ftp (ftp.cme.msu. edu) and WWW (http://www.cme.msu.edu/RDP). The WWW server provides ribosomal probe checking, approximate phylogenetic placement of user-submitted sequences, screening for possible chimeric rRNA sequences, automated alignment, and a suggested placement of an unknown sequence on an existing phylogenetic tree. Additional utilities also exist at RDP-II, including distance matrix, T-RFLP, and a Java-based viewer of the phylogenetic trees that can be used to create subtrees.

Non-NASA Center↗

Evolution of early life inferred from protein and ribonucleic acid sequences

The chemical structures of ferredoxin, 5S ribosomal RNA, and c-type cytochrome sequences have been employed to construct a phylogenetic tree which connects all major photosynthesizing organisms: the three types of bacteria, blue-green algae, and chloroplasts. Anaerobic and aerobic bacteria, eukaryotic cytoplasmic components and mitochondria are also included in the phylogenetic tree. Anaerobic nonphotosynthesizing bacteria similar to Clostridium were the earliest organisms, arising more than 3.2 billion years ago. Bacterial photosynthesis evolved nearly 3.0 billion years ago, while oxygen-evolving photosynthesis, originating in the blue-green algal line, came into being about 2.0 billion years ago. The phylogenetic tree supports the symbiotic theory of the origin of eukaryotes.

Dayhoff, M. O.↗

Analysis of in Vitro Evolution Reveals the Underlying Distribution of Catalytic Activity Among Random Sequences

The emergence of catalytic RNA is believed to have been a key event during the origin of life. Understanding how catalytic activity is distributed across random sequences is fundamental to estimating the probability that catalytic sequences would emerge. Here, we analyze the in vitro evolution of triphosphorylating ribozymes and translate their fitnesses into absolute estimates of catalytic activity for hundreds of ribozyme families. The analysis efficiently identified highly active ribozymes and estimated catalytic activity with good accuracy. The evolutionary dynamics follow Fisher’s Fundamental Theorem of Natural Selection and a corollary, permitting retrospective inference of the distribution of fitness and activity in the random sequence pool for the first time. The frequency distribution of rate constants appears to be log-normal, with a surprisingly steep dropoff at higher activity, consistent with a mechanism for the emergence of activity as the product of many independent contributions.

Abe Pressman↗

Evolution of heliobacteria: implications for photosynthetic reaction center complexes

The evolutionary position of the heliobacteria, a group of green photosynthetic bacteria with a photosynthetic apparatus functionally resembling Photosystem I of plants and cyanobacteria, has been investigated with respect to the evolutionary relationship to Gram-positive bacteria and cyanobacteria. On the basis of 16S rRNA sequence analysis, the heliobacteria appear to be most closely related to Gram-positive bacteria, but also an evolutionary link to cyanobacteria is evident. Interestingly, a 46-residue domain including the putative sixth membrane-spanning region of the heliobacterial reaction center protein show rather strong similarity (33% identity and 72% similarity) to a region including the sixth membrane-spanning region of the CP47 protein, a chlorophyll-binding core antenna polypeptide of Photosystem II. The N-terminal half of the heliobacterial reaction center polypeptide shows a moderate sequence similarity (22% identity over 232 residues) with the CP47 protein, which is significantly more than the similarity with the Photosystem I core polypeptides in this region. An evolutionary model for photosynthetic reaction center complexes is discussed, in which an ancestral homodimeric reaction center protein (possibly resembling the heliobacterial reaction center protein) with 11 membrane-spanning regions per polypeptide has diverged to give rise to the core of Photosystem I, Photosystem II, and of the photosynthetic apparatus in green, purple, and heliobacteria.

NASA Program Exobiology↗

Origins of the plant chloroplasts and mitochondria based on comparisons of 5S ribosomal RNAs

In this paper, we provide macromolecular comparisons utilizing the 5S ribosomal RNA structure to suggest extant bacteria that are the likely descendants of chloroplast and mitochondria endosymbionts. The genetic stability and near universality of the 5S ribosomal gene allows for a useful means to study ancient evolutionary changes by macromolecular comparisons. The value in current and future ribosomal RNA comparisons is in fine tuning the assignment of ancestors to the organelles and in establishing extant species likely to be descendants of bacteria involved in presumed multiple endosymbiotic events.

NASA Discipline Exobiology↗