Search NASA⌕ Search

SEARCH · Search NASA

Results for “RNA sequencing”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 235 records · Page 13

Creating Benchmark Data for Artificial Intelligence and Machine Learning Space Biology Research

To identify an appropriate AI/ML approach for a specific problem, the best practice is to measure algorithm performance through the benchmarking process. A scientific benchmark consists of an AI-ready dataset and a reference implementation on a specific scientific question. The NASA Science Mission Directorate (SMD) has started the “Benchmark Initiative for AI/ML to create scientific benchmark datasets in three applications: 1) scientific benchmarking, which finds the best algorithm for a specific problem; 2) application benchmarking, which measures algorithm performance against a set of parameters; and 3) system benchmarking, which evaluates performance of hardware and software architecture. Currently, there are no standardized datasets available to benchmark AI/ML algorithms in the domain of space biology. In this work, we constructed two AI/ML-ready biological datasets from experiments in space-flown mice: cellular imaging and RNA-seq. First, radiation-exposed immune cells harbor DNA damage foci that can be fluorescently marked to visualize the amount of damage following exposure to ionizing radiation. However, such large datasets are difficult to analyze visually, due to imaging inconsistencies and human bias, and classical image processing approaches can fail on imaging artifacts. AI/ML are therefore exciting alternative, providing the speed of machines and the accuracy of humans. We have made this dataset available at https://registry.opendata.aws/bps_microscopy/. Second, high-throughput nucleic acid sequencing (DNA-seq, RNA-seq) has become widespread in biomedical research due to the growing availability and affordability of these assays. However, most sequencing datasets suffer from high dimensionality and low sample count. In this work, we used a generative adversarial network to synthesize a standardized, AI-ready, publicly available benchmark dataset for space biology RNA-seq data with sufficient space-flown and ground control mouse liver samples from NASA GeneLab. This dataset is available at https://registry.opendata.aws/bps_rnaseq/. These datasets are now fully open the Space Biology community to test their favorite AI/ML approaches.

James Casaletto↗

Differential expression of members of the annexin multigene family in Arabidopsis

Although in most plant species no more than two annexin genes have been reported to date, seven annexin homologs have been identified in Arabidopsis, Annexin Arabidopsis 1-7 (AnnAt1--AnnAt7). This establishes that annexins can be a diverse, multigene protein family in a single plant species. Here we compare and analyze these seven annexin gene sequences and present the in situ RNA localization patterns of two of these genes, AnnAt1 and AnnAt2, during different stages of Arabidopsis development. Sequence analysis of AnnAt1--AnnAt7 reveals that they contain the characteristic four structural repeats including the more highly conserved 17-amino acid endonexin fold region found in vertebrate annexins. Alignment comparisons show that there are differences within the repeat regions that may have functional importance. To assess the relative level of expression in various tissues, reverse transcription-PCR was carried out using gene-specific primers for each of the Arabidopsis annexin genes. In addition, northern blot analysis using gene-specific probes indicates differences in AnnAt1 and AnnAt2 expression levels in different tissues. AnnAt1 is expressed in all tissues examined and is most abundant in stems, whereas AnnAt2 is expressed mainly in root tissue and to a lesser extent in stems and flowers. In situ RNA localization demonstrates that these two annexin genes display developmentally regulated tissue-specific and cell-specific expression patterns. These patterns are both distinct and overlapping. The developmental expression patterns for both annexins provide further support for the hypothesis that annexins are involved in the Golgi-mediated secretion of polysaccharides.

Non-NASA Center↗

Class 2 CRISPR/Cas compositions and methods of use

Provided are compositions and methods that include one or more of: (1) a Class 2 CRISPR/Cas effector protein, a nucleic acid encoding the effector protein, and/or a modified host cell comprising the effector protein (and/or a nucleic acid encoding the same); (2) a CRISPR/Cas guide RNA that binds to and provides sequence specificity to the Class 2 CRISPR/Cas effector protein, a nucleic acid encoding the CRISPR/Cas guide RNA, and/or a modified host cell comprising the CRISPR/Cas guide RNA (and/or a nucleic acid encoding the same); and (3) a CRISPR/Cas transactivating noncoding RNA (trancRNA), a nucleic acid encoding the CRISPR/Cas trancRNA, and/or a modified host cell comprising the CRISPR/Cas trancRNA (and/or a nucleic acid encoding the same).

Doudna, Jennifer A.↗

Adaptation of Organisms by Resonance of RNA Transcription with the Cellular Redox Cycle

Sequence variation in organisms differs across the genome and the majority of mutations are caused by oxidation, yet its origin is not fully understood. It has also been shown that the reduction-oxidation reaction cycle is the fundamental biochemical cycle that coordinates the timing of all biochemical processes in that cell, including energy production, DNA replication, and RNA transcription. It is shown that the temporal resonance of transcriptome biosynthesis with the oscillating binary state of the reduction-oxidation reaction cycle serves as a basis for non-random sequence variation at specific genome-wide coordinates that change faster than by accumulation of chance mutations. This work demonstrates evidence for a universal, persistent and iterative feedback mechanism between the environment and heredity, whereby acquired variation between cell divisions can outweigh inherited variation.

Stolc, Viktor↗

The case for relationship of the flavobacteria and their relatives to the green sulfur bacteria

Analysis of 16S rRNA sequences suggests, but does not convincingly demonstrate a specific relationship between the eubacterial phylum defined by the flavobacteria and their relatives and that defined by the green sulfur bacteria. Consequently, we have sequenced the 23S rRNA from several representatives of the former group and one of the latter in order to bring more data to bear upon this point. The 23S rRNA data alone strongly suggest a specific relationship between the two phyla, and, together with the 16S rRNA results, provides what we consider now to be a convincing case for this specific relationship.

NASA Discipline Number 52-30↗

Long-read sequencing transcriptome quantification with lr-kallisto

RNA abundance quantification has become routine and affordable thanks to high-throughput “short-read” technologies that provide accurate molecule counts at the gene level. Similarly accurate and affordable quantification of definitive full-length, transcript isoforms has remained a stubborn challenge, despite its obvious biological significance across a wide range of problems. “Long-read” sequencing platforms now produce data-types that can, in principle, drive routine definitive isoform quantification. However some particulars of contemporary long-read datatypes, together with isoform complexity and genetic variation, present bioinformatic challenges. We show here, using ONT data, that fast and accurate quantification of long-read data is possible and that it is improved by exome capture. To perform quantifications we developed lr-kallisto, which adapts the kallisto bulk and single-cell RNA-seq quantification methods for long-read technologies.

Loving, Rebekah K. (ORCID:0000000187250376)↗

Self-assembly and condensation of intermolecular poly(UG) RNA quadruplexes

Abstract Poly(UG) or ‘pUG’ dinucleotide repeats are highly abundant sequences in eukaryotic RNAs. In Caenorhabditis elegans, pUGs are added to RNA 3′ ends to direct gene silencing within Mutator foci, a germ granule condensate. Here, we show that pUG RNAs efficiently self-assemble into gel condensates through quadruplex (G4) interactions. Short pUG sequences form right-handed intermolecular G4s (pUG G4s), while longer pUGs form left-handed intramolecular G4s (pUG folds). We determined a 1.05 Å crystal structure of an intermolecular pUG G4, which reveals an eight stranded G4 dimer involving 48 nucleotides, 7 different G and U quartet conformations, 7 coordinated potassium ions, 8 sodium ions and a buried water molecule. A comparison of the intermolecular pUG G4 and intramolecular pUG fold structures provides insights into the molecular basis for G4 handedness and illustrates how a simple dinucleotide repeat sequence can form complex structures with diverse topologies.

Biochemistry & Molecular Biology↗

Identification of a dual-specificity protein phosphatase that inactivates a MAP kinase from Arabidopsis

Mitogen-activated protein kinases (MAPKs) play a key role in plant responses to stress and pathogens. Activation and inactivation of MAPKs involve phosphorylation and dephosphorylation on both threonine and tyrosine residues in the kinase domain. Here we report the identification of an Arabidopsis gene encoding a dual-specificity protein phosphatase capable of hydrolysing both phosphoserine/threonine and phosphotyrosine in protein substrates. This enzyme, designated AtDsPTP1 (Arabidopsis thaliana dual-specificity protein tyrosine phosphatase), dephosphorylated and inactivated AtMPK4, a MAPK member from the same plant. Replacement of a highly conserved cysteine by serine abolished phosphatase activity of AtDsPTP1, indicating a conserved catalytic mechanism of dual-specificity protein phosphatases from all eukaryotes.

Non-NASA Center↗

Transcription factor IID in the Archaea: sequences in the Thermococcus celer genome would encode a product closely related to the TATA-binding protein of eukaryotes

The first step in transcription initiation in eukaryotes is mediated by the TATA-binding protein, a subunit of the transcription factor IID complex. We have cloned and sequenced the gene for a presumptive homolog of this eukaryotic protein from Thermococcus celer, a member of the Archaea (formerly archaebacteria). The protein encoded by the archaeal gene is a tandem repeat of a conserved domain, corresponding to the repeated domain in its eukaryotic counterparts. Molecular phylogenetic analyses of the two halves of the repeat are consistent with the duplication occurring before the divergence of the archael and eukaryotic domains. In conjunction with previous observations of similarity in RNA polymerase subunit composition and sequences and the finding of a transcription factor IIB-like sequence in Pyrococcus woesei (a relative of T. celer) it appears that major features of the eukaryotic transcription apparatus were well-established before the origin of eukaryotic cellular organization. The divergence between the two halves of the archael protein is less than that between the halves of the individual eukaryotic sequences, indicating that the average rate of sequence change in the archael protein has been less than in its eukaryotic counterparts. To the extent that this lower rate applies to the genome as a whole, a clearer picture of the early genes (and gene families) that gave rise to present-day genomes is more apt to emerge from the study of sequences from the Archaea than from the corresponding sequences from eukaryotes.

NASA Discipline Exobiology↗

In vitro selection of functional nucleic acids

In vitro selection allows rare functional RNA or DNA molecules to be isolated from pools of over 10(15) different sequences. This approach has been used to identify RNA and DNA ligands for numerous small molecules, and recent three-dimensional structure solutions have revealed the basis for ligand recognition in several cases. By selecting high-affinity and -specificity nucleic acid ligands for proteins, promising new therapeutic and diagnostic reagents have been identified. Selection experiments have also been carried out to identify ribozymes that catalyze a variety of chemical transformations, including RNA cleavage, ligation, and synthesis, as well as alkylation and acyl-transfer reactions and N-glycosidic and peptide bond formation. The existence of such RNA enzymes supports the notion that ribozymes could have directed a primitive metabolism before the evolution of protein synthesis. New in vitro protein selection techniques should allow for a direct comparison of the frequency of ligand binding and catalytic structures in pools of random sequence polynucleotides versus polypeptides.

Review↗

Identification of defective illegitimate recombinational repair of oxidatively-induced DNA double-strand breaks in ataxia-telangiectasia cells

Ataxia-telangiectasia (A-T) is an autosomal-recessive lethal human disease. Homozygotes suffer from a number of neurological disorders, as well as very high cancer incidence. Heterozygotes may also have a higher than normal risk of cancer, particularly for the breast. The gene responsible for the disease (ATM) has been cloned, but its role in mechanisms of the disease remain unknown. Cellular A-T phenotypes, such as radiosensitivity and genomic instability, suggest that a deficiency in the repair of DNA double-strand breaks (DSBs) may be the primary defect; however, overall levels of DSB rejoining appear normal. We used the shuttle vector, pZ189, containing an oxidatively-induced DSB, to compare the integrity of DSB rejoining in one normal and two A-T fibroblast cells lines. Mutation frequencies were two-fold higher in A-T cells, and the mutational spectrum was different. The majority of the mutations found in all three cell lines were deletions (44-63%). The DNA sequence analysis indicated that 17 of the 17 plasmids with deletion mutations in normal cells occurred between short direct-repeat sequences (removing one of the repeats plus the intervening sequences), implicating illegitimate recombination in DSB rejoining. The combined data from both A-T cell lines showed that 21 of 24 deletions did not involve direct-repeats sequences, implicating a defect in the illegitimate recombination pathway. These findings suggest that the A-T gene product may either directly participate in illegitimate recombination or modulate the pathway. Regardless, this defect is likely to be important to a mechanistic understanding of this lethal disease.

Non-NASA Center↗

The direct and indirect drivers shaping RNA viral communities in grassland soils

ABSTRACT Recent studies have revealed diverse RNA viral communities in soils. Yet, how environmental factors influence soil RNA viruses remains largely unknown. Here, we recovered RNA viral communities from bulk metatranscriptomes sequenced from grassland soils managed for 5 years under multiple environmental conditions including water content, plant presence, cultivar type, and soil depth. More than half of the unique RNA viral contigs (64.6%) were assigned with putative hosts. About 74.7% of these classified RNA viral contigs are known as eukaryotic RNA viruses suggesting eukaryotic RNA viruses may outnumber prokaryotic RNA viruses by nearly three times in this grassland. Of the identified eukaryotic RNA viruses and the associated eukaryotic species, the most dominant taxa were Mitoviridae with an average relative abundance of 72.4%, and their natural hosts, Fungi with an average relative abundance of 56.6%. Network analysis and structural equation modeling support that soil water content, plant presence, and type of cultivar individually demonstrate a significant positive impact on eukaryotic RNA viral richness directly as well as indirectly on eukaryotic RNA viral abundance via influencing the co-existing eukaryotic members. A significant negative influence of soil depth on soil eukaryotic richness and abundance indirectly impacts soil eukaryotic RNA viral communities. These results provide new insights into the collective influence of multiple environmental and community factors that shape soil RNA viral communities and offer a structured perspective of how RNA virus diversity and ecology respond to environmental changes. IMPORTANCE Climate change has been reshaping the soil environment as well as the residing microbiome. This study provides field-relevant information on how environmental and community factors collectively shape soil RNA communities and contribute to ecological understanding of RNA viral survival under various environmental conditions and virus-host interactions in soil. This knowledge is critical for predicting the viral responses to climate change and the potential emergence of biothreats.

59 BASIC BIOLOGICAL SCIENCES↗

The origin and early evolution of nucleic acid polymerases

The hypothesis that vestiges of the ancestral RNA-dependent RNA polymerase involved in the replication of RNA genomes of Archean cells are present in the eubacterial RNA-polymerase beta-prime subunit and its homologues is discussed. It is shown that, in the DNA-dependent RNA polymerases from three cellular lineages, a very conserved sequence of eight amino acids, also found in a small RNA-binding site previously described for the E. coli polynucleotide phosphorylase and the S1 ribosomal protein, is present. The optimal conditions for the replicase activity of the avian-myeloblastosis-virus reverse transcriptase are presented. The evolutionary significance of the in vitro modifications of substrate and template specificities of RNA polymerases and reverse transcriptases is discussed.

Lazcano, A.↗

CABO-16S—a Combined Archaea, Bacteria, Organelle 16S rRNA database framework for amplicon analysis of prokaryotes and eukaryotes in environmental samples

Abstract Identification of both prokaryotic and eukaryotic microorganisms in environmental samples is currently challenged by the need for additional sequencing to obtain separate 16S and 18S ribosomal RNA (rRNA) amplicons or the constraints imposed by “universal” primers. Organellar 16S rRNA sequences are amplified and sequenced along with prokaryote 16S rRNA and provide an alternative method to identify eukaryotic microorganisms. CABO-16S combines bacterial and archaeal sequences from the SILVA database with 16S rRNA sequences of plastids and other organelles from the PR2 database to enable identification of all 16S rRNA sequences. Comparison of CABO-16S with SILVA 138.2 results in equivalent taxonomic classification of mock communities and increased classification of diverse environmental samples. In particular, identification of phototrophic eukaryotes in shallow seagrass environments, marine waters, and lake waters was increased. The CABO-16S framework allows users to add custom sequences for further classification of underrepresented clades and can be easily updated with future releases of reference databases. Addition of sequences obtained from Sanger sequencing of methane seep sediments and curated sequences of the polyphyletic SEEP-SRB1 clade resulted in differentiation of syntrophic and non-syntrophic SEEP-SRB1 in hydrothermal vent sediments. CABO-16S highlights the benefit of combining and amending existing training sets when studying microorganisms in diverse environments.

Eitel, Eryn M. (ORCID:0009000723919297)↗

Identification of shared viral sequences in peat moss metagenomes reveals elements of a possible Sphagnum core virome

Viruses are an understudied component of plant microbiomes. Identifying viruses that are shared between individual plants, or members of the “core virome”, could reveal stable viral populations with the potential to modulate the composition and function of the microbiome. Here, we examined the virome associated with Sphagnum mosses, a keystone species that has direct influence over the fate of peatland carbon stores. We analyzed bulk metagenomes and metatranscriptomes generated from Sphagnum field samples collected over a ten-month period to identify virus-like sequences shared among plants. Individual Sphagnum samples harbored distinct DNA and RNA viromes where only a small percentage (< 1%) of the total number of identified viral contigs were shared among all samples. Based on taxonomic classification, the shared viral contigs represent bacterial viruses, or phage (Caudoviricetes), as well as viruses of eukaryotes, namely nucleocytoplasmic large DNA viruses (Nucleocytoviricota) and RNA viruses (Riboviria). We linked the shared phage-like contigs to viral regions within sequenced genomes of bacterial taxa that are members of the Sphagnum core microbiome, suggesting that these contigs represent temperate phage or degraded prophage. The putative nucleocytoplasmic large DNA viruses and RNA viruses were phylogenetically diverse and showed sequence similarity to viruses associated with a broad range of hosts and environmental sources. The identification of shared viral contigs suggested that, despite the compositional heterogeneity between samples, Sphagnum mosses may harbor a core virome. Future work validating the presence of the core virome is warranted as it may aid in understanding how persistent viruses impact microbiome ecology and symbiont evolution within this climatically relevant keystone species.

Metagenomics↗

Structural insights into RNase H catalytic mechanism from room-temperature X-ray and neutron crystallography of apo- and RNA/DNA hybrid-bound enzyme

RNase H enzymes are sequence-nonspecific endonucleases that cleave RNA strands in RNA/DNA hybrid duplexes, an enzymatic process essential in DNA replication and repair in both prokaryotes and eukaryotes. Also, RNase H activity of the reverse transcriptase in human immunodeficiency viruses (HIV-1 and HIV-2) is indispensable for the viral replication cycle. RNase H enzymes play an central role in the development of gene therapies and are targets for novel antivirals. It is therefore of great importance to gain a detailed understanding of the RNase H catalytic mechanism to improve drug design. We utilized Bacillus halodurans RNase H1 (BhRNase H1) to shed light on its function and catalytic mechanism. Room-temperature neutron crystallography of the wild-type and inactive D132N mutant enzymes revealed that E109, belonging to the catalytic DEDD motif, can change its protonation state, allowing us to propose its role in the protonation of the leaving O3′ hydroxyl group of RNA. X-ray crystallography has demonstrated the ability of the RNA/DNA duplex to slide along the protein surface upon metal ion binding at site M A , transforming a product mimic into a Michaelis-like complex, which confirms an essential role of the M A metal ion in catalysis.

Enzyme mechanisms↗