Search NASASearch

SEARCH · Search NASA

Results for “RNA sequencing”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

Structurally complex and highly active RNA ligases derived from random RNA sequences

Seven families of RNA ligases, previously isolated from random RNA sequences, fall into three classes on the basis of secondary structure and regiospecificity of ligation. Two of the three classes of ribozymes have been engineered to act as true enzymes, catalyzing the multiple-turnover transformation of substrates into products. The most complex of these ribozymes has a minimal catalytic domain of 93 nucleotides. An optimized version of this ribozyme has a kcat exceeding one per second, a value far greater than that of most natural RNA catalysts and approaching that of comparable protein enzymes. The fact that such a large and complex ligase emerged from a very limited sampling of sequence space implies the existence of a large number of distinct RNA structures of equivalent complexity and activity.

Non-NASA Center

Equally parsimonious pathways through an RNA sequence space are not equally likely

An experimental system for determining the potential ability of sequences resembling 5S ribosomal RNA (rRNA) to perform as functional 5S rRNAs in vivo in the Escherichia coli cellular environment was devised previously. Presumably, the only 5S rRNA sequences that would have been fixed by ancestral populations are ones that were functionally valid, and hence the actual historical paths taken through RNA sequence space during 5S rRNA evolution would have most likely utilized valid sequences. Herein, we examine the potential validity of all sequence intermediates along alternative equally parsimonious trajectories through RNA sequence space which connect two pairs of sequences that had previously been shown to behave as valid 5S rRNAs in E. coli. The first trajectory requires a total of four changes. The 14 sequence intermediates provide 24 apparently equally parsimonious paths by which the transition could occur. The second trajectory involves three changes, six intermediate sequences, and six potentially equally parsimonious paths. In total, only eight of the 20 sequence intermediates were found to be clearly invalid. As a consequence of the position of these invalid intermediates in the sequence space, seven of the 30 possible paths consisted of exclusively valid sequences. In several cases, the apparent validity/invalidity of the intermediate sequences could not be anticipated on the basis of current knowledge of the 5S rRNA structure. This suggests that the interdependencies in RNA sequence space may be more complex than currently appreciated. If ancestral sequences predicted by parsimony are to be regarded as actual historical sequences, then the present results would suggest that they should also satisfy a validity requirement and that, in at least limited cases, this conjecture can be tested experimentally.

NASA Discipline Exobiology

Long-read RNA sequencing atlas of human microglia isoforms elucidates disease-associated genetic regulation of splicing

Microglia, the innate immune cells of the central nervous system, have been genetically implicated in multiple neurodegenerative diseases. Mapping the genetics of gene expression in human microglia has identified several loci associated with disease-associated genetic variants in microglia-specific regulatory elements. However, identifying genetic effects on splicing is challenging because of the use of short sequencing reads. Here, we present the isoform-centric microglia genomic atlas (isoMiGA), which leverages long-read RNA sequencing to identify 35,879 novel microglia isoforms. We show that these isoforms are involved in stimulation response and brain region specificity. We then quantified the expression of both known and novel isoforms in a multi-ancestry meta-analysis of 555 human microglia short-read RNA sequencing samples from 391 donors, and found associations with genetic risk loci in Alzheimer’s and Parkinson’s disease. We nominate several loci that may act through complex changes in isoform and splice-site usage.

59 BASIC BIOLOGICAL SCIENCES

Phylogenetic origins of the plant mitochondrion based on a comparative analysis of 5S ribosomal RNA sequences

The complete nucleotide sequences of 5S ribosomal RNAs from Rhodocyclus gelatinosa, Rhodobacter sphaeroides, and Pseudomonas cepacia were determined. Comparisons of these 5S RNA sequences show that rather than being phylogenetically related to one another, the two photosynthetic bacterial 5S RNAs share more sequence and signature homology with the RNAs of two nonphotosynthetic strains. Rhodobacter sphaeroides is specifically related to Paracoccus denitrificans and Rc. gelatinosa is related to Ps. cepacia. These results support earlier 16S ribosomal RNA studies and add two important groups to the 5S RNA data base. Unique 5S RNA structural features previously found in P. denitrificans are present also in the 5S RNA of Rb. sphaeroides; these provide the basis for subdivisional signatures. The immediate consequence of obtaining these new sequences is that it is possible to clarify the phylogenetic origins of the plant mitochondrion. In particular, a close phylogenetic relationship is found between the plant mitochondria and members of the alpha subdivision of the purple photosynthetic bacteria, namely, Rb. sphaeroides, P. denitrificans, and Rhodospirillum rubrum.

Villanueva, E.

Identification of characteristic oligonucleotides in the bacterial 16S ribosomal RNA sequence dataset

MOTIVATION: The phylogenetic structure of the bacterial world has been intensively studied by comparing sequences of 16S ribosomal RNA (16S rRNA). This database of sequences is now widely used to design probes for the detection of specific bacteria or groups of bacteria one at a time. The success of such methods reflects the fact that there are local sequence segments that are highly characteristic of particular organisms or groups of organisms. It is not clear, however, the extent to which such signature sequences exist in the 16S rRNA dataset. A better understanding of the numbers and distribution of highly informative oligonucleotide sequences may facilitate the design of hybridization arrays that can characterize the phylogenetic position of an unknown organism or serve as the basis for the development of novel approaches for use in bacterial identification. RESULTS: A computer-based algorithm that characterizes the extent to which any individual oligonucleotide sequence in 16S rRNA is characteristic of any particular bacterial grouping was developed. A measure of signature quality, Q(s), was formulated and subsequently calculated for every individual oligonucleotide sequence in the size range of 5-11 nucleotides and for 15mers with reference to each cluster and subcluster in a 929 organism representative phylogenetic tree. Subsequently, the perfect signature sequences were compared to the full set of 7322 sequences to see how common false positives were. The work completed here establishes beyond any doubt that highly characteristic oligonucleotides exist in the bacterial 16S rRNA sequence dataset in large numbers. Over 16,000 15mers were identified that might be useful as signatures. Signature oligonucleotides are available for over 80% of the nodes in the representative tree.

NASA Discipline Life Sciences Technologies

Single cell RNA sequencing reveals shifts in cell maturity and function of endogenous and infiltrating cell types in response to acute intervertebral disc injury

Intervertebral disc (IVD) degeneration contributes to disabling back pain. Degeneration can be initiated by injury and progressively leads to an irreversible loss of cells and function. IVD function restoration through cell replacement therapies have had limited success due to knowledge gaps in the critical cell populations important for repair. Here, in this study, we used single cell RNA sequencing to identify the transcriptional changes of IVD resident and infiltrating cell populations from Control and Injured coccygeal IVDs extracted from 12-week-old female C57BL/6J mice 7 days post injury. Clustering, gene ontology, and pseudotime trajectory analyses determined transcriptomic divergences with injury, flow cytometry identified they types of infiltrating immune cells, and immunofluorescence was utilized to define mesenchymal stem cell (MSC) localization. We identified 11 distinct clusters that included IVD, immune, vascular cells, and MSCs. Differential gene expression analysis determined that Outer Annulus Fibrosus, Neutrophils, Saa2-High MSCs, Macrophages, and Krt18 + Nucleus Pulposus (NP) cells were the major drivers of transcriptomic differences between Control and Injured cells. Gene ontology revealed that the most upregulated biological pathways were angiogenesis and T cell-related while wound healing and ECM regulation were downregulated. Pseudotime trajectory analyses revealed that IVD injury directed cells towards increased differentiation in all clusters, except for Krt18 + NP cells which remained in a less mature cell state. Saa2-High and Grem1-High MSCs populations shifted towards more differentiated IVD cells profiles with injury and localized distinctly within the IVD. This study revealed novel MSC populations with the potential to be leveraged for future IVD repair studies.

Cartilage

A Standard RNA Sequencing Assay for Space Biology

Given the limited opportunities for biological experimentation in space, it is often desirable to compare results across experiments to gain additional insights into the effects of spaceflight on biological systems. However, this approach is made difficult by a multitude of confounding factors including differences in strain, hardware configuration, and sample processing. To help harmonize datasets, the NASA GeneLab Project has developed consistent sample and data processing protocols for the generation of raw and processed RNA-sequencing data from various mouse tissues. We will present these and discuss how they can be used to make novel discoveries from these precious samples.

Galazka, Jonathan

Dual-RNA-sequencing to elucidate the interactions between sorghum and Colletotrichum sublineola

In warm and humid regions, the productivity of sorghum is significantly limited by the fungal hemibiotrophic pathogen Colletotrichum sublineola , the causal agent of anthracnose, a problematic disease of sorghum ( Sorghum bicolor (L.) Moench) that can result in grain and biomass yield losses of up to 50%. Despite available genomic resources of both the host and fungal pathogen, the molecular basis of sorghum− C. sublineola interactions are poorly understood. By employing a dual-RNA sequencing approach, the molecular crosstalk between sorghum and C. sublineola can be elucidated. In this study, we examined the transcriptomes of four resistant sorghum accessions from the sorghum association panel (SAP) at varying time points post-infection with C. sublineola . Approximately 0.3% and 93% of the reads mapped to the genomes of C. sublineola and Sorghum bicolor , respectively. Expression profiling of in vitro versus in planta C. sublineola at 1-, 3-, and 5-days post-infection (dpi) indicated that genes encoding secreted candidate effectors, carbohydrate-active enzymes (CAZymes), and membrane transporters increased in expression during the transition from the biotrophic to the necrotrophic phase (3 dpi). The hallmark of the pathogen-associated molecular pattern (PAMP)-triggered immunity in sorghum includes the production of reactive oxygen species (ROS) and phytoalexins. The majority of effector candidates secreted by C. sublineola were predicted to be localized in the host apoplast, where they could interfere with the PAMP-triggered immunity response, specifically in the host ROS signaling pathway. The genes encoding critical molecular factors influencing pathogenicity identified in this study are a useful resource for subsequent genetic experiments aimed at validating their contributions to pathogen virulence. This comprehensive study not only provides a better understanding of the biology of C. sublineola but also supports the long-term goal of developing resistant sorghum cultivars.

Vela, Saddie

Space Flown Rodent Liver RNA Sequencing Data for Machine Learning in Space Biology Research

High-throughput nucleic acid sequencing (DNA-seq, RNA-seq) has become widespread in biomedical research due to the growing availability and affordability of these assays. Data analysis has been accelerated in recent years by the adoption of artificial intelligence (AI) and machine learning (ML) techniques by biomedical researchers. In space biology research, RNAseq datasets from space-flown experimental samples are critical for characterizing the gene expression aberrations associated with exposure to spaceflight stressors. However, space biological experiments tend to be very low sample size, so identifying proper AI/ML algorithms for sequencing data analysis is an ongoing challenge since these algorithms typically require large sample size. The NASA Science Mission Directorate (SMD) has started the “Benchmark Initiative for AI/ML”, focused on creating datasets meant for three main applications: 1) scientific benchmarking, which finds the best algorithm for a specific problem; 2) application benchmarking, which measures algorithm performance against a set of parameters; and 3) system benchmarking, which evaluates performance of hardware and software architecture. These scientific benchmarks consist of an AI-ready dataset and a reference implementation on a specific scientific question. In this work, we focused on generating standardized datasets to allow the scientific community to benchmark AI/ML algorithms in the domain of space biology. We present here a standardized, AI-ready, publicly available benchmark dataset for space biology RNA-seq data as a collaboration between the NASA AI4LS (Artificial Intelligence for Life Sciences) working group. and NASA’s SMD. This dataset consists of space-flown and ground control mouse liver found in the NASA GeneLab omics database. However, to amplify the small sample number (n=112 samples) for ML purposes, we employ Gaussian noise and a generative adversarial network to extend this dataset to 6,000 synthetic samples, matching the original gene expression characteristics.

James Casaletto

Ribosomal RNA sequence suggest microsporidia are extremely ancient eukaryotes

A comparative sequence analysis of the 18S small subunit ribosomal RNA (rRNA) of the microsporidium Vairimorpha necatrix is presented. The results show that this rRNA sequence is more unlike those of other eukaryotes than any known eukaryote rRNA sequence. It is concluded that the lineage leading to microsporidia branched very early from that leading to other eukaryotes.

Vossbrinck, C. R.

Experimental investigation of an RNA sequence space

Modern rRNAs are the historic consequence of an ongoing evolutionary exploration of a sequence space. These extant sequences belong to a special subset of the sequence space that is comprised only of those primary sequences that can validly perform the biological function(s) required of the particular RNA. If it were possible to readily identify all such valid sequences, stochastic predictions could be made about the relative likelihood of various evolutionary pathways available to an RNA. Herein an experimental system which can assess whether a particular sequence is likely to have validity as a eubacterial 5S rRNA is described. A total of ten naturally occurring, and hence known to be valid, sequences and two point mutants of unknown validity were used to test the usefulness of the approach. Nine of the ten valid sequences tested positive whereas both mutants tested as clearly defective. The tenth valid sequence gave results that would be interpreted as reflecting a borderline status were the answer not known. These results demonstrate that it is possible to experimentally determine which sequences in local regions of the sequence space are potentially valid 5S rRNAs.

Lee, Youn-Hyung

Single-cell RNA sequencing reveals plasmid constrains bacterial population heterogeneity and identifies a non-conjugating subpopulation

Transcriptional heterogeneity in isogenic bacterial populations can play various roles in bacterial evolution, but its detection remains technically challenging. Here, we use microbial split-pool ligation transcriptomics to study the relationship between bacterial subpopulation formation and plasmid-host interactions at the single-cell level. We find that single-cell transcript abundances are influenced by bacterial growth state and plasmid carriage. Moreover, plasmid carriage constrains the formation of bacterial subpopulations. Plasmid genes, including those with core functions such as replication and maintenance, exhibit transcriptional heterogeneity associated with cell activity. Notably, we identify a cell subpopulation that does not transcribe conjugal plasmid transfer genes, which may help reduce plasmid burden on a subset of cells. Our study advances the understanding of plasmid-mediated subpopulation dynamics and provides insights into the plasmid-bacteria interplay.

59 BASIC BIOLOGICAL SCIENCES

A Pipeline for Assessing the Quality of Rna-Seq Datasets in GeneLab

Transcriptome profiling by RNA sequencing (RNA-seq) is a powerful approach to identify gene expression changes in organisms exposed to unique environments such as spaceflight. One of the challenges of evaluating RNA-seq data both within and across different space-relevant studies is the ability to control for technical differences, including the use of different library preparation kits, sequencing platforms, RNA yield, and person-to-person variation. To help address this issue, the National Institute of Standards and Technology (NIST, nist.gov) initiated a consortium, at the request of industry and academia, to develop a set of controls for gene expression measurements. The result was a set of 92 unlabeled, polyadenylated transcripts that range from 250 – 2,000 nucleotides in length to mimic natural eukaryotic mRNAs. These External RNA Controls Consortium (ERCC) genes can be used in any RNA-seq experiment, by adding known concentrations of the ERCC genes to samples after RNA extraction, to offer a standard measurement for data comparison. At NASA GeneLab, we employ these controls as part of our standard operating procedures for every in-house RNA-seq study to assess the limit of detection, dynamic range, and power of differential expression analysis both within and across experiments. Here we will discuss the use, benefits, and limitations of ERCC genes and other types of controls, such as universal RNA references, to generate quality control information for RNA-seq studies conducted at GeneLab.

GeneLab