Search NASA⌕ Search

SEARCH · Search NASA

Results for “RNA sequencing”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 163 records · Page 9

Proteogenomic characterization of difficult-to-treat breast cancer with tumor cells enriched through laser microdissection

Abstract Background Breast cancer (BC) is the most commonly diagnosed cancer and the leading cause of cancer death among women globally. Despite advances, there is considerable variation in clinical outcomes for patients with non-luminal A tumors, classified as difficult-to-treat breast cancers (DTBC). This study aims to delineate the proteogenomic landscape of DTBC tumors compared to luminal A (LumA) tumors. Methods We retrospectively collected a total of 117 untreated primary breast tumor specimens, focusing on DTBC subtypes. Breast tumors were processed by laser microdissection (LMD) to enrich tumor cells. DNA, RNA, and protein were simultaneously extracted from each tumor preparation, followed by whole genome sequencing, paired-end RNA sequencing, global proteomics and phosphoproteomics. Differential feature analysis, pathway analysis and survival analysis were performed to better understand DTBC and investigate biomarkers. Results We observed distinct variations in gene mutations, structural variations, and chromosomal alterations between DTBC and LumA breast tumors. DTBC tumors predominantly had more mutations inTP53,PLXNB3, Zinc finger genes, and fewer mutations inSDC2,CDH1,PIK3CA,SVIL, andPTEN. Notably, Cytoband 1q21, which contains numerous cell proliferation-related genes, was significantly amplified in the DTBC tumors. LMD successfully minimized stromal components and increased RNA–protein concordance, as evidenced by stromal score comparisons and proteomic analysis. Distinct DTBC and LumA-enriched clusters were observed by proteomic and phosphoproteomic clustering analysis, some with survival differences. Phosphoproteomics identified two distinct phosphoproteomic profiles for high relapse-risk and low relapse-risk basal-like tumors, involving several genes known to be associated with breast cancer oncogenesis and progression, includingKIAA1522,DCK,FOXO3,MYO9B,ARID1A,EPRS,ZC3HAV1, andRBM14. Lastly, an integrated pathway analysis of multi-omics data highlighted a robust enrichment of proliferation pathways in DTBC tumors. Conclusions This study provides an integrated proteogenomic characterization of DTBC vs LumA with tumor cells enriched through laser microdissection. We identified many common features of DTBC tumors and the phosphopeptides that could serve as potential biomarkers for high/low relapse-risk basal-like BC and possibly guide treatment selections.

Oncology↗

Radiation Stability Evaluation of Protein-Based Nanopores for Mars and Europa Missions

Exploration of our Solar System has revealed a number of locations that are now habitable or could have supported life in the past. One approach to finding life involves detection of informational polymers like deoxyribonucleic acid (DNA) and ribonucleic acid (RNA) that are definitive biosignatures for life as we know it. Alternatively, structural variants of DNA and RNA, collectively termed xenonucleic acids (XNAs) have been shown in the laboratory to behave similarly. Nanopore-based sequencers differ from traditional sequencing technologies in that they do not explicitly require synthesis of DNA before or during analysis. Because of this, nanopore sequencers have been used for the direct sequencing of RNA, and could be used for the detection and analysis of other charged polymers. Here we describe results of exposing the MinION hardware, flow cells, and key reagents to ionizing radiation at doses relevant to Mars and Europa missions (10 to 3000 silicon-equivalent gray).

Burton, Aaron S.↗

Co-conservation of rRNA tetraloop sequences and helix length suggests involvement of the tetraloops in higher-order interactions

Terminal loops containing four nucleotides (tetraloops) are common in structural RNAs, and they frequently conform to one of three sequence motifs, GNRA, UNCG, or CUUG. Here we compare available sequences and secondary structures for rRNAs from bacteria, and we show that helices capped by phylogenetically conserved GNRA loops display a strong tendency to be of conserved length. The simplest interpretation of this correlation is that the conserved GNRA loops are involved in higher-order interactions, intramolecular or intermolecular, resulting in a selective pressure for maintaining the lengths of these helices. A small number of conserved UNCG loops were also found to be associated with conserved length helices, consistent with the possibility that this type of tetraloop also takes part in higher-order interactions.

Non-NASA Center↗

RolyPoly (rp) v0.1.0

The Rolypoly pipeline is designed to process raw RNA-seq data and identify potential RNA viral sequences. It is split into several self contained steps: 1. input data filtering and QC, 2. Genome assembly and refinement, 3. Assembly filtering, 4. Mapping to known RNA viral genomes, 5. Searching for RNA viral marker genes. 6. Genome functional and structural annotation. 6. Report preparation and potential downstream analysis The last module, may include taxonomic assignment, host range estimation, and phenotypic prediction. There are many similar software, but they focus on human related viruses, and lack the downstream applications or differ in their sensitivity. The initial user base are non-computational microbial ecologists who wish to better understand the potential RNA viruses in their own generated samples.

Neri, Uri↗

RNA Crystallization

RNA molecules may be crystallized using variations of the methods developed for protein crystallography. As the technology has become available to syntheisize and purify RNA molecules in the quantities and with the quality that is required for crystallography, the field of RNA structure has exploded. The first consideration when crystallizing an RNA is the sequence, which may be varied in a rational way to enhance crystallizability or prevent formation of alternate structures. Once a sequence has been designed, the RNA may be synthesized chemically by solid-state synthesis, or it may be produced enzymatically using RNA polymerase and an appropriate DNA template. Purification of milligram quantities of RNA can be accomplished by HPLC or gel electrophoresis. As with proteins, crystallization of RNA is usually accomplished by vapor diffusion techniques. There are several considerations that are either unique to RNA crystallization or more important for RNA crystallization. Techniques for design, synthesis, purification, and crystallization of RNAs will be reviewed here.

Golden, Barbara L.↗

The western red cedar (Thuja plicata) 8-8' DIRIGENT family displays diverse expression patterns and conserved monolignol coupling specificity

The isolation and characterization of a multigene family of the first class of dirigent proteins (namely that mainly involved in 8-8' coupling leading to (+)-pinoresinol in this case) is reported, this comprising of nine western red cedar (Thuja plicata) DIRIGENT genes (DIR1-9) of 72-99.5% identity to each other. Their corresponding cDNA clones had coding regions for 180-183 amino acids with each having a predicted molecular mass of ca. 20 kDa including the signal peptide. Real time-PCR established that the DIRIGENT isovariants were differentially expressed during growth and development of T. plicata (P < 0.05). The phylogenetic relationships and the rates and patterns of nucleotide substitution suggest that the DIRIGENT gene may have evolved via paralogous expansion at an early stage of vascular plant diversification. Thereafter, western red cedar paralogues have maintained an high homogeneity presumably via a concerted evolutionary mode. This, in turn, is assumed to be the driving force for the differential formation of 8-8'-linked pinoresinol derived (poly)lignans in the needles, stems, bark and branches, as well as for massive accumulation of 8-8'-linked plicatic acid-derived (poly)lignans in heartwood.

Non-NASA Center↗

Circularization of 23S rRNA but not 16S rRNA within archaeal ribosomes

Background Processing of archaeal 16S and 23S rRNAs is believed to involve excision of individual rRNAs from polycistronic precursors, circularization of excised rRNAs, and re-linearization before the incorporation into ribosomes. However, all the knowledge is derived from several isolated species, leaving open the possibility that different processes may occur in other archaeal groups. Results Here, we investigate rRNAs from diverse and mostly uncultivated archaea. Sequencing of total cellular RNA from eight phylum-level lineages indicates that archaeal circular 23S rRNA transcript abundances vastly exceed those of linear counterparts, and linear versions are often undetectable. As the majority of rRNAs derive from mature ribosomes, the data suggest that ribosomes contain circular 23S rRNAs. Thus, we directly sequence RNA extracted from isolated ribosomes of a model archaeon, Methanosarcina acetivorans, and confirm that the 23S rRNAs in the ribosomes are circular. Structural modeling places the 5′ and 3′ ends of the linear precursors of archaeal 23S rRNAs in close proximity to form a GNRA tetraloop (in which N is A, C, G, or U and R is A or G), consistent with their existence as circular molecules. We also confirm the existence of circular 16S rRNA intermediates in transcriptomes of most archaea, yet a circular form is not evident in some distinct archaeal groups, suggesting that certain archaea do not circularize 16S rRNA during processing. Conclusions Our findings uncover unexpected variations in the processing required to generate mature rRNAs and the conformation of functional molecules in archaeal ribosomes.

Archaea↗

RNA Splicing Events in Circulation Distinguish Individuals With and Without New-onset Type 1 Diabetes

Context: Alterations in RNA splicing may influence protein isoform diversity that contributes to or reflects the pathophysiology of certain diseases. Whereas specific RNA splicing events in pancreatic islets have been investigated in models of inflammation in vitro, how RNA splicing in the circulation correlates with or is reflective of type 1 diabetes (T1D) disease pathophysiology in humans remains unexplored. Objective: To use machine learning to investigate if alternative RNA splicing events differ between individuals with and without new-onset T1D and to determine if these splicing events provide insight into T1D pathophysiology. Methods: RNA deep sequencing was performed on whole blood samples from 2 independent cohorts: a training cohort consisting of 12 individuals with new-onset T1D and 12 age- and sex-matched nondiabetic controls and a validation cohort of the same size and demographics. Machine learning analysis was used to identify specific isoforms that could distinguish individuals with T1D from controls. Results: Distinct patterns of RNA splicing differentiated participants with T1D from unaffected controls. Notably, certain splicing events, particularly involving retained introns, showed significant association with T1D. Machine learning analysis using these splicing events as features from the training cohort demonstrated high accuracy in distinguishing between T1D subjects and controls in the validation cohort. Gene Ontology pathway enrichment analysis of the retained intron category showed evidence for a systemic viral response in T1D subjects. Conclusion: Alternative RNA splicing events in whole blood are significantly enriched in individuals with new-onset T1D and can effectively distinguish these individuals from unaffected controls. Further, our findings also suggest that RNA splicing profiles offer the potential to provide insights into disease pathogenesis.

60 APPLIED LIFE SCIENCES↗

Identification of candidate host-specificity genes in Exserohilum turcicum using comparative genomics and transcriptomics

Abstract Exserohilum turcicum causes northern corn leaf blight and sorghum leaf blight. While the same species cause disease in both crops, the strains are host-specific. Here, we report the sequence and de novo annotated assemblies of one sorghum- and one maize-specific E. turcicum strain. The strains were sequenced using the PacBio Sequel II system. The total genome length for both assemblies was between 44 and 45 Mb with N50 of ∼2.5 Mb. Ninety-eight percent of the Benchmarking Universal Single-Copy Orthologs (BUSCO) for both assemblies had complete status. The estimated number of genes was 11,762 and 12,029 in the sorghum- and maize-specific isolates, respectively. Funannotate, EffectorP, SignalP, and transcriptome data were used to create functional annotation of each genome. The whole-genome comparison identified ten large-scale inversions and three translocations between the maize- and sorghum-specific strains, along with homologous genes and gene duplications. RNA was sequenced from the maize- and sorghum-specific isolate 10 days post-inoculation in maize and sorghum and from axenic cultures. Gene expression data from planta and axenic growth experiments were compared for each strain. Candidate host-specificity genes were identified by combining results from whole-genome comparison, synteny analysis, gene annotations, and transcriptome data. Overall, this study identified several candidate host-specificity genes that provide insights into E. turcicum interaction with its hosts.

Krone, Mara J. (ORCID:0000000159006624)↗

Investigating biological nitrogen fixation via single-cell transcriptomics

The extensive use of nitrogen fertilizers has detrimental environmental consequences, and it is essential for society to explore sustainable alternatives. One promising avenue is engineering root nodule symbiosis, a naturally occurring process in certain plant species within the nitrogen-fixing clade, into non-leguminous crops. Advancements in single-cell transcriptomics provide unprecedented opportunities to dissect the molecular mechanisms underlying root nodule symbiosis at the cellular level. This review summarizes key findings from single-cell studies in Medicago truncatula, Lotus japonicus, and Glycine max. We highlight how these studies address fundamental questions about the development of root nodule symbiosis, including the following findings: (i) single-cell transcriptomics has revealed a conserved transcriptional program in root hair and cortical cells during rhizobial infection, suggesting a common infection pathway across legume species; (ii) characterization of determinate and indeterminate nodules using single-cell technologies supports the compartmentalization of nitrogen fixation, assimilation, and transport into distinct cell populations; (iii) single-cell transcriptomics data have enabled the identification of novel root nodule symbiosis genes and provided new approaches for prioritizing candidate genes for functional characterization; and (iv) trajectory inference and RNA velocity analyses of single-cell transcriptomics data have allowed the reconstruction of cellular lineages and dynamic transcriptional states during root nodule symbiosis.

Lotus japonicus↗

Metabolomic and transcriptomic remodeling of bone marrow myeloid cells in response to maternal obesity

Maternal obesity puts the offspring at high risk of developing obesity and cardiometabolic diseases in adulthood. Here, we utilized a mouse model of maternal high-fat diet (HFD)-induced obesity that recapitulates metabolic perturbations seen in humans. We show increased adiposity in the offspring of HFD-fed mothers (Off-HFD) when compared with the offspring of regular diet-fed mothers (Off-RD). We have previously reported significant immune perturbations in the bone marrow of newly weaned Off-HFD. Here, we hypothesized that lipid metabolism is altered in the bone marrow of Off-HFD versus Off-RD. To test this hypothesis, we investigated the lipidomic profile of bone marrow cells collected from 3-week-old Off-RD and Off-HFD. Diacylglycerols (DAGs), triacylglycerols (TAGs), sphingolipids, and phospholipids were remarkably different between the groups, independent of fetal sex. Levels of cholesteryl esters were significantly decreased in Off-HFD, suggesting reduced delivery of cholesterol. These were accompanied by age-dependent progression of mitochondrial dysfunction in bone marrow cells. We subsequently isolated CD11b+ myeloid cells from 3-wk-old mice and conducted metabolomic, lipidomic, and transcriptomic analyses. The lipidomic profiles of myeloid cells were similar to those of bone marrow cells and included increases in DAGs and decreased TAGs. Transcriptomics revealed altered expression of genes related to immune pathways, including macrophage alternative activation, B-cell receptors, and transforming growth factor-β signaling. All told, this study revealed lipidomic, metabolomic, and gene expression abnormalities in bone marrow cells broadly, and in bone marrow myeloid cells particularly, in the newly weaned offspring of mothers with obesity, which might at least partially explain the progression of metabolic and cardiovascular diseases in their adulthood.

RNA sequencing↗

Transcriptomic Analysis of the CAM Species Kalanchoë fedtschenkoi Under Low- and High-Temperature Regimes

Temperature stress is one of the major limiting environmental factors that negatively impact global crop yields. Kalanchoë fedtschenkoi is an obligate crassulacean acid metabolism (CAM) plant species, exhibiting much higher water-use efficiency and tolerance to drought and heat stresses than C 3 or C 4 plant species. Previous studies on gene expression responses to low- or high-temperature stress have been focused on C 3 and C 4 plants. There is a lack of information about the regulation of gene expression by low and high temperatures in CAM plants. To address this knowledge gap, we performed transcriptome sequencing (RNA-Seq) of leaf and root tissues of K. fedtschenkoi under cold (8 °C), normal (25 °C), and heat (37 °C) conditions at dawn (i.e., 2 h before the light period) and dusk (i.e., 2 h before the dark period). Our analysis revealed differentially expressed genes (DEGs) under cold or heat treatment in comparison to normal conditions in leaf or root tissue at each of the two time points. In particular, DEGs exhibiting either the same or opposite direction of expression change (either up-regulated or down-regulated) under cold and heat treatments were identified. In addition, we analyzed gene co-expression modules regulated by cold or heat treatment, and we performed in-depth analyses of expression regulation by temperature stresses for selected gene categories, including CAM-related genes, genes encoding heat shock factors and heat shock proteins, circadian rhythm genes, and stomatal movement genes. Our study highlights both the common and distinct molecular strategies employed by CAM and C 3 /C 4 plants in adapting to extreme temperatures, providing new insights into the molecular mechanisms underlying temperature stress responses in CAM species.

59 BASIC BIOLOGICAL SCIENCES↗

Hydrothermal systems and the emergence of life

The author reviews current thought about life originating in hyperthermophilic microorganisms. Hyperthermophiles obtain food from chemosynthesis of sulfur and have an RNA nucleotide sequence different from bacteria and eucarya. It is postulated that a hyperthermophile may be the common ancestor of all life. Current research efforts focus on the synthesis of organic compounds in hydrothermal systems.

Non-NASA Center↗

Molecular cloning and characterization of a tomato cDNA encoding a systemically wound-inducible bZIP DNA-binding protein

Localized wounding of one leaf in intact tomato (Lycopersicon esculentum Mill.) plants triggers rapid systemic transcriptional responses that might be involved in defense. To better understand the mechanism(s) of intercellular signal transmission in wounded tomatoes, and to identify the array of genes systemically up-regulated by wounding, a subtractive cDNA library for wounded tomato leaves was constructed. A novel cDNA clone (designated LebZIP1) encoding a DNA-binding protein was isolated and identified. This clone appears to be encoded by a single gene, and belongs to the family of basic leucine zipper domain (bZIP) transcription factors shown to be up-regulated by cold and dark treatments. Analysis of the mRNA levels suggests that the transcript for LebZIP1 is both organ-specific and up-regulated by wounding. In wounded wild-type tomatoes, the LebZIP1 mRNA levels in distant tissue were maximally up-regulated within only 5 min following localized wounding. Exogenous abscisic acid (ABA) prevented the rapid wound-induced increase in LebZIP1 mRNA levels, while the basal levels of LebZIP1 transcripts were higher in the ABA mutants notabilis (not), sitiens (sit), and flacca (flc), and wound-induced increases were greater in the ABA-deficient mutants. Together, these results suggest that ABA acts to curtail the wound-induced synthesis of LebZIP1 mRNA.

Non-NASA Center↗

Light-modulated abundance of an mRNA encoding a calmodulin-regulated, chromatin-associated NTPase in pea

A CDNA encoding a 47 kDa nucleoside triphosphatase (NTPase) that is associated with the chromatin of pea nuclei has been cloned and sequenced. The translated sequence of the cDNA includes several domains predicted by known biochemical properties of the enzyme, including five motifs characteristic of the ATP-binding domain of many proteins, several potential casein kinase II phosphorylation sites, a helix-turn-helix region characteristic of DNA-binding proteins, and a potential calmodulin-binding domain. The deduced primary structure also includes an N-terminal sequence that is a predicted signal peptide and an internal sequence that could serve as a bipartite-type nuclear localization signal. Both in situ immunocytochemistry of pea plumules and immunoblots of purified cell fractions indicate that most of the immunodetectable NTPase is within the nucleus, a compartment proteins typically reach through nuclear pores rather than through the endoplasmic reticulum pathway. The translated sequence has some similarity to that of human lamin C, but not high enough to account for the earlier observation that IgG against human lamin C binds to the NTPase in immunoblots. Northern blot analysis shows that the NTPase MRNA is strongly expressed in etiolated plumules, but only poorly or not at all in the leaf and stem tissues of light-grown plants. Accumulation of NTPase mRNA in etiolated seedlings is stimulated by brief treatments with both red and far-red light, as is characteristic of very low-fluence phytochrome responses. Southern blotting with pea genomic DNA indicates the NTPase is likely to be encoded by a single gene.

NASA Discipline Number 40-50↗

Microbial assessment of cabin air quality on commercial airliners

The microbial burdens of 69 cabin air samples collected from commercial airliners were assessed via conventional culture-dependent, and molecular-based microbial enumeration assays. Cabin air samples from each of four separate flights aboard two different carriers were collected via air-impingement. Microbial enumeration techniques targeting DNA, ATP, and endotoxin were employed to estimate total microbial burden. The total viable microbial population ranged from 0 to 3.6 x10 4 cells per 100 liters of air, as assessed by the ATP-assay. When these same samples were plated on R2A minimal medium, anywhere from 2% to 80% of these viable populations were cultivable. Five of the 29 samples examined exhibited higher cultivable counts than ATP derived viable counts, perhaps a consequence of the dormant nature (and thus lower concentration of intracellular ATP) of cells inhabiting these air cabin samples. Ribosomal RNA gene sequence analysis showed these samples to consist of a moderately diverse group of bacteria, including human pathogens. Enumeration of ribosomal genes via quantitative-PCR indicated that population densities ranged from 5 x 10 1 ' to IO 7 cells per 100 liters of air. Each of the aforementioned strategies for assessing overall microbial burden has its strengths and weaknesses; this publication serves as a testament to the power of their use in concert.

microbe↗

Transcriptomics-based Machine Learning Analysis Predicts Space-Exposed Murine Livers

Limited sample sizes, high data dimensionality, and sensitivity to technical and biological variability of next generation sequencing (NGS), typically limits machine learning (ML) approaches in spaceflight studies that include radiation effects. However, pooling smaller studies while addressing intra- and inter-study variabilities allows for ML predictive modeling. Here, integration methods were applied to whole transcriptome shotgun sequencing (RNA-seq) data from six mouse liver GeneLab datasets (GLDS) (n ranging from 6 to 39 samples) from with a total of 81 spaceflight and ground-control samples to determine top features (i.e. genes) relevant to spaceflight including the effect of radiation exposure. RNASeq counts were normalized for each study, then merged and scaled across all datasets. Data dimensionality was reduced using a minimum redundancy maximum relevance (MRMR) methodology. Redundancy and relevance were computed using the Pearson correlation and F-statistic, respectively. The top 100 MRMR features were used to predict spaceflight vs. ground-control samples using Random Forest (RF), Support Vector Machine (SVM), and Linear Discriminant Analysis (LDA) classifiers with 5-fold cross validation (CV). Principal component analysis (PCA) on the complete feature set versus the MRMR features shows separation between spaceflight samples and ground controls (Figure 1A). The ML-based gene sets were compared against differential gene expression results obtained with DESeq2 from individual GLDS. Using all features or randomly sampled subsets at matching set sizes with MRMR, a maximum classifier accuracy of 69% was shown on the test set over 5 folds. For all classifiers, CV training using at least the top 30 MRMR genes show minimum 89% accuracy and 0.95 AUC value on the test set over 5 folds (Figure 1B). Baseline set analysis on differentially expressed genes (DEGs) identified using padj ≤ 0.05 show 295 DEGs that overlap at least two studies and 13 DEGs that overlap three studies (Figure 1C). Set analysis between the top 100 MRMR features and the DEGs showed 47 genes that overlap at least one study and 24 genes that overlap two studies. Over-representation analysis showed overlapping biological processes related to fatty acid and lipid metabolism which may indicate these processes in the response to spaceflight stressors. MRMR feature selection for the selected ML methods improve performance relative to a classifier built on all features or randomly sampled subsets. Permutation feature importance within the decorrelated MRMR features showed concordance in feature ranking between ML methods. A challenge of applying ML methods across heterogeneous NGS data is accounting for signal:noise. Here, signal validation across studies was shown by intersecting sets between top MRMR genes and DEGs from DESeq2 analysis. Non-intersecting sets introduce opportunity to explore genes relevant to differentiating space flight exposed groups and implementing ML methods across existing NGS datasets may overcome sample size limitations.

Machine Learning↗