Search NASASearch

SEARCH · Search NASA

Results for “RNA sequencing”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 37 records · Page 2

Common 5S rRNA variants are likely to be accepted in many sequence contexts

Over evolutionary time RNA sequences which are successfully fixed in a population are selected from among those that satisfy the structural and chemical requirements imposed by the function of the RNA. These sequences together comprise the structure space of the RNA. In principle, a comprehensive understanding of RNA structure and function would make it possible to enumerate which specific RNA sequences belong to a particular structure space and which do not. We are using bacterial 5S rRNA as a model system to attempt to identify principles that can be used to predict which sequences do or do not belong to the 5S rRNA structure space. One promising idea is the very intuitive notion that frequently seen sequence changes in an aligned data set of naturally occurring 5S rRNAs would be widely accepted in many other 5S rRNA sequence contexts. To test this hypothesis, we first developed well-defined operational definitions for a Vibrio region of the 5S rRNA structure space and what is meant by a highly variable position. Fourteen sequence variants (10 point changes and 4 base-pair changes) were identified in this way, which, by the hypothesis, would be expected to incorporate successfully in any of the known sequences in the Vibrio region. All 14 of these changes were constructed and separately introduced into the Vibrio proteolyticus 5S rRNA sequence where they are not normally found. Each variant was evaluated for its ability to function as a valid 5S rRNA in an E. coli cellular context. It was found that 93% (13/14) of the variants tested are likely valid 5S rRNAs in this context. In addition, seven variants were constructed that, although present in the Vibrio region, did not meet the stringent criteria for a highly variable position. In this case, 86% (6/7) are likely valid. As a control we also examined seven variants that are seldom or never seen in the Vibrio region of 5S rRNA sequence space. In this case only two of seven were found to be potentially valid. The results demonstrate that changes that occur multiple times in a local region of RNA sequence space in fact usually will be accepted in any sequence context in that same local region.

NASA Discipline Exobiology

Batch Effect Correction Methods for NASA GeneLab Transcriptomic Datasets

RNA sequencing (RNA-seq) data from space biology experiments promise to yield invaluable insights into the effects of spaceflight on terrestrial biology. However, sample numbers from each study are low due to limited crew availability, hardware, and space. To increase statistical power, spaceflight RNA-seq datasets from different missions are often aggregated together. However, this can introduce technical variation or "batch effects", often due to differences in sample handling, sample processing, and sequencing platforms. Several computational methods have been developed to correct for technical batch effects, thereby reducing their impact on true biological signals. In this study, we combined 7 mouse liver RNA-seq datasets from NASA GeneLab (part of the NASA Open Science Data Repository) to evaluate several common batch effect correction methods (ComBat and ComBat-seq from the sva R package, and Median Polish, Empirical Bayes, and ANOVA from the MBatch R package). We quantitatively evaluated the ability of these methods to correct for technical batch variables in space biology RNA-seq data using the following criteria: BatchQC, principal component analysis, dispersion separability criterion, log fold change correlation, and differential gene expression analysis. Each batch variable / correction method combination was then assessed using a custom scoring approach to identify the optimal correction method for the combined dataset, by geometrically probing the space of all allowable scoring functions to yield an aggregate volume-based scoring measure. Finally, we describe the way in which the GeneLab multi-study analysis and visualization portal will allow users to examine the presence or absence of batch effects using multiple metrics. If the user chooses to perform batch effect correction, the scoring approach described here can be implemented to identify the optimal correction method to use for their specific combined dataset prior to analysis.

Lauren M. Sanders

Protocol for applying a network-enabled gene discovery pipeline to non-model plant species

Identifying upstream regulators of key genes is essential for understanding gene regulatory mechanisms and translating these insights into functional targets. Here, we present a protocol for applying the network-enabled gene discovery pipeline (NEEDLE) to non-model plant species. We describe steps for environment setup, data preparation, computational analysis, expected outputs, and parameter considerations. NEEDLE integrates RNA sequencing (RNA-seq) processing, weighted gene co-expression analysis (WGCNA), Gene Network Inference with Ensemble of trees (GENIE3), and promoter conservation analysis to prioritize candidate transcriptional regulators.

Plant Sciences

Psychosocial experiences are associated with human brain mitochondrial biology

Psychosocial experiences affect brain health and aging trajectories, but the molecular pathways underlying these associations remain unclear. Normal brain function relies on energy transformation by mitochondria oxidative phosphorylation (OxPhos). Two main lines of evidence position mitochondria both as targets and drivers of psychosocial experiences. On the one hand, chronic stress exposure and mood states may alter multiple aspects of mitochondrial biology; on the other hand, functional variations in mitochondrial OxPhos capacity may alter social behavior, stress reactivity, and mood. But are psychosocial exposures and subjective experiences linked to mitochondrial biology in the human brain? By combining longitudinal antemortem assessments of psychosocial factors with postmortem brain (dorsolateral prefrontal cortex) proteomics in older adults, we find that higher well-being is linked to greater abundance of the mitochondrial OxPhos machinery, whereas higher negative mood is linked to lower OxPhos protein content. Combined, positive and negative psychosocial factors explained 18 to 25% of the variance in the abundance of OxPhos complex I, the primary biochemical entry point that energizes brain mitochondria. Moreover, interrogating mitochondrial psychobiological associations in specific neuronal and nonneuronal brain cells with single-nucleus RNA sequencing (RNA-seq) revealed strong cell-type-specific associations for positive psychosocial experiences and mitochondria in glia but opposite associations in neurons. As a result, these “mind-mitochondria” associations were masked in bulk RNA-seq, highlighting the likely underestimation of true psychobiological effect sizes in bulk brain tissues. Thus, self-reported psychosocial experiences are linked to human brain mitochondrial phenotypes.

59 BASIC BIOLOGICAL SCIENCES

Deficiency in transmitter release triggers homeostatic transcriptional changes that increase presynaptic excitability

Weakening of synaptic transmission at theDrosophilalarval neuromuscular junction triggers two forms of homeostatic compensation, one that increases the probability of glutamate release per action potential (P r ) and another that increases motoneuron (MN) activity. We investigated the molecular changes in MNs that underlie the increase in MN activity. RNA sequencing (RNA-seq) analysis on MNs whose glutamate release is weakened by knockdown of components of the MN transmitter release machinery reveals a reduction in expression of a group of genes that encode potassium channels and their positive modulators. These results identify a mechanism of compensation for weakened synaptic transmission by MNs, which engages a transcriptional program in those cells to increase firing and, thereby, ensure sufficient locomotory drive.

Science & Technology - Other Topics

Human Host Cellular Response to HCoV-229E Infection Proteomics (ACS-JM-DP2)

The purpose of this experiment was to evaluate the human host cellular response to wild-type Human coronavirus strain 229E (HCoV-229E) infection. Sample data was obtained for mock and infected immortalized human lung epithelial cells (A549) (MOI 5) nuclear extracts, immortalized human lung fibroblasts cells (MRC5) (MOI5) nuclear extracts, and primary human airway epithelial (HAE) (MOI 3) cells from lung tissue and processed for proteome analysis. Processed datasets are openly accessible from the download button and contain secondary processed proteomic results files and supporting metadata materials. Experimental proteomics samples were prepared using Limited Proteolysis (LiP) methods for Label-free quantification (LFQ) and global proteomic evaluation. Sample data was acquired using a Q-Exactive HF-X mass spectrometer and was processed and compiled using MaxQuant software (v.1.6.17.0). Processed proteomic data downloads include a sample naming key, processed MaxQuant results/parameters, and protein annotated relative abundance files. See corresponding primary data accessions below and Viral Experiment LiP Analysis source code supporting data transparency and reuse. Experimental transcriptomics samples were collected in parallel and processed for RNA sequencing (RNA-Seq) as summarized under ACS-DP1 (https://data.pnnl.gov/group/nodes/dataset/34069).

59 BASIC BIOLOGICAL SCIENCES

NASA GeneLab Multi-study Visualization Portal

NASA GeneLab has helped advance the field of Space Biology by providing a public repository where researchers can store, share, analyze and visualize the results of space flight related omics experiments. The GeneLab data visualization portal allows any user, regardless of bioinformatics knowledge or access to computational resources, to interact with the experimental data, draw their own conclusions, and gain insights about the effects of space on living systems. These tools help democratize scientific research and foster the NASA Open Science initiative. The new multi-study feature of the GeneLab visualization platform allows users to mine study metadata from RNA sequencing (RNA-seq) experiments to identify samples of interest by filtering datasets based on organism, tissue, assay technology type, and/or factor. Once samples are selected from multiple datasets, users can combine and normalize the sample data, then utilize the visualization displays, including Principal Component Analysis (PCA) plots, to assess sample distributions. Finally, users can perform differential gene expression analysis on the combined data and visualize the results through PCA plots, Volcano plots, Pair plots, Heatmap, Ideogram and Gene Set Enrichment Analysis. All user-generated results and visualizations will be available for download. Here, we present a biological study using samples from multiple GeneLab RNA-seq datasets and analyzed using the multi-study visualization platform to demonstrate inter- and intra-study variability, as well as commonly differentially expressed genes between spaceflight and ground control conditions across datasets. This new feature opens a wide range of possibilities and opportunities for further development including combining other assay technology types and integration with batch effect correction techniques and machine learning applications. Overall, this tool allows users to increase the statistical power of individual experiments, validate hypothesis, identify patterns, and opens the door to new and exciting research.

space biology

RNA language models predict mutations that improve RNA function

Structured RNA lies at the heart of many central biological processes, from gene expression to catalysis. RNA structure prediction is not yet possible due to a lack of high-quality reference data associated with organismal phenotypes that could inform RNA function. We present GARNET (Gtdb Acquired RNa with Environmental Temperatures), a new database for RNA structural and functional analysis anchored to the Genome Taxonomy Database (GTDB). GARNET links RNA sequences to experimental and predicted optimal growth temperatures of GTDB reference organisms. Using GARNET, we develop sequence- and structure-aware RNA generative models, with overlapping triplet tokenization providing optimal encoding for a GPT-like model. Leveraging hyperthermophilic RNAs in GARNET and these RNA generative models, we identify mutations in ribosomal RNA that confer increased thermostability to the Escherichia coli ribosome. The GTDB-derived data and deep learning models presented here provide a foundation for understanding the connections between RNA sequence, structure, and function.

59 BASIC BIOLOGICAL SCIENCES

The promising role of proteomes and metabolomes in defining the single-cell landscapes of plants

The plant community has a strong track-record of RNA sequencing technology deployment, which combined with the recent advent of spatial platforms (e.g., 10x genomics), has resulted in an explosion of outstanding single cell and nuclei datasets that can be put in an in situ context within tissues (e.g., a cell atlas)1. In the genomics era, application of proteomics technologies in the plant sciences has always trailed behind that of RNA sequencing technologies, largely due to accessibility, ease-of-use and access to expertise along with depth of analysis benefits. On the other hand, the use of early analytical tools for characterizing small molecules (metabolites) from plant systems predates nucleic acid sequencing and proteomics analysis2, as the search for plant-based natural products has played a significant role in improving human health throughout history. However, the employment of proteomics and metabolomics assays for characterizing plant cell processes now remains significantly behind transcriptional approaches, even though both provide a direct functional readout of cell states and phenotypes.

Anderton, Christopher R. [BATTELLE (PACIFIC NW LAB

RNAseq data for P. putida with vanillate

Illumina sequencing reads from RNA sequencing of vanillate-utilizing strains of Pseudomonas putida, described in Evolution and engineering of pathways for aromatic O-demethylation in Pseudomonas putida KT2440 by A. Bleem, et al. (2024)

Adaptive laboratory evolution

PPI DataHub Project Data Package: S. elongatus PCC 7942 Circadian Control Bioproduction Transcriptomics (PB-DP3)

The purpose of this experiment was to evaluate how circadian clock regulation impacts carbon partitioning between storage, growth, and product synthesis in Synechococcus elongatus PCC 7942 in providing insights to strategies for enhanced bioproduction. Sample data was acquired using a Illumina HiSeq sequencer system and processed for RNA sequencing (RNA-Seq) expression analysis. Transcriptomic differential expression analysis revealed coordinated circadian clock-driven adjustment of the cell cycle and rewiring of energy and carbon metabolism. Processed RNA-Seq datasets are openly accessible from the PNNL DataHub project dataset download page and contain secondary processed RNA-seq results files and supporting metadata materials linked to relevant source code information supporting data transparency and reuse.

59 BASIC BIOLOGICAL SCIENCES

Human Host Cellular Response to HCoV-229E Infection Transcriptomics (ACS-DP1)

The purpose of this experiment was to evaluate the human host cellular response to wild-type Human coronavirus strain 229E (HCoV-229E) infection. Sample data was obtained for mock and infected immortalized human lung epithelial cells (A549) (MOI 5), immortalized human lung fibroblasts cells (MRC5) (MOI5), and primary human airway epithelial (HAE) (MOI 3) cells from lung tissue. Sample data was acquired using an Illumina HiSeq 2000 sequencer system and processed for RNA sequencing (RNA-Seq) expression analysis.

59 BASIC BIOLOGICAL SCIENCES

A Melanoma Brain Metastasis CTC Signature and CTC:B-cell Clusters Associate with Secondary Liver Metastasis: A Melanoma Brain–Liver Metastasis Axis

Melanoma brain metastasis is linked to dismal prognosis and low overall survival and is detected in up to 80% of patients at autopsy. Circulating tumor cells (CTC) are the smallest functional units of cancer and precursors of fatal metastasis. We previously used an unbiased multilevel approach to discover a unique ribosomal protein large/small subunit (RPL/RPS) CTC gene signature associated with melanoma brain metastasis. In this study, we hypothesized that CTC-driven melanoma brain metastasis secondary metastasis (“metastasis of metastasis” per clinical scenarios) has targeted organ specificity for the liver. We injected parallel cohorts of immunodeficient and newly developed humanized NBSGW (huNBSGW) mice with cells from CTC-derived melanoma brain metastasis to identify secondary metastatic patterns. We found the presence of a melanoma brain–liver metastasis axis in huNBSGW mice. Furthermore, RNA sequencing analysis of tissues showed a significant upregulation of the RPL/RPS CTC gene signature linked to metastatic spread to the liver. Additional RNA sequencing of CTCs from huNBSGW blood revealed extensive CTC clustering with human B cells in these mice. CTC:B-cell clusters were also upregulated in the blood of patients with primary melanoma and maintained either in CTC-driven melanoma brain metastasis or melanoma brain metastasis CTC–derived cells promoting liver metastasis. CTC-generated tumor tissues were interrogated at single-cell gene and protein expression levels (10x Genomics Xenium and HALO spatial biology platforms, respectively). Collectively, our findings suggest that heterotypic CTC:B-cell interactions can be critical at multiple stages of metastasis.

60 APPLIED LIFE SCIENCES

Human Primary Airway Epithelium +/- Macrophages Response to HCoV-229E Infection Transcriptomics (ACS-DP3)

The purpose of this experiment was to evaluate the human host cellular response to wild-type Human coronavirus strain 229E (HCoV-299E) infection. Sample data was obtained for mock and infected (MOI 3) primary human airway epithelial cells with and without macrophages and grown in air-liquid interface conditions. Sample data was acquired using an Illumina Hi-Seq 4000 sequencer system and further processed for RNA sequencing (RNA-Seq) expression analysis.

59 BASIC BIOLOGICAL SCIENCES

Microbial identification by immunohybridization assay of artificial RNA labels

Ribosomal RNA (rRNA) and engineered stable artificial RNAs (aRNAs) are frequently used to monitor bacteria in complex ecosystems. In this work, we describe a solid-phase immunocapture hybridization assay that can be used with low molecular weight RNA targets. A biotinylated DNA probe is efficiently hybridized in solution with the target RNA, and the DNA-RNA hybrids are captured on streptavidin-coated plates and quantified using a DNA-RNA heteroduplex-specific antibody conjugated to alkaline phosphatase. The assay was shown to be specific for both 5S rRNA and low molecular weight (LMW) artificial RNAs and highly sensitive, allowing detection of as little as 5.2 ng (0.15 pmol) in the case of 5S rRNA. Target RNAs were readily detected even in the presence of excess nontarget RNA. Detection using DNA probes as small as 17 bases targeting a repetitive artificial RNA sequence in an engineered RNA was more efficient than the detection of a unique sequence.

Non-NASA Center

A 5.8S nuclear ribosomal RNA gene sequence database: applications to ecology and evolution

We complied a 5.8S nuclear ribosomal gene sequence database for animals, plants, and fungi using both newly generated and GenBank sequences. We demonstrate the utility of this database as an internal check to determine whether the target organism and not a contaminant has been sequenced, as a diagnostic tool for ecologists and evolutionary biologists to determine the placement of asexual fungi within larger taxonomic groups, and as a tool to help identify fungi that form ectomycorrhizae.

Databases, Factual

Single‐Cell Nanodroplet Processing Proteomics Pipeline for Analysis of Human‐Derived Microglia

Single-cell omics tools provide unique insights into heterogeneous cell populations and their responses to stimuli. For example, single-cell RNA sequencing has identified several transcriptionally distinct populations of microglia, which are resident immune cells of the central nervous system (CNS) that are responsive to CNS injury, infection, and neurodegeneration. To date, single-cell studies of microglia have focused on RNA-sequencing or cytometry by time of flight (CyTOF), which provide indirect readouts of protein abundance or quantification of a limited number of targets. Herein, we present a workflow based on FACS-assisted isolation, cryopreservation, and nanodroplet-based processing for single-cell mass spectrometry proteomics analysis of the postmortem human brain cortex-derived microglia. From a single microglial cell, 1039 proteins could be identified on average. As a proof-of-principle, we applied single-cell proteomics for exploring the heterogeneity of brain microglia at the cellular level. This pilot proteomics data partially recapitulates the prior microglia subtypes. Specifically, we determined that mitochondrial proteins, in particular members of NADH dehydrogenase (Complex I), cytochrome b-c1 (Complex III), cytochrome c oxidase (Complex IV), F1-ATPase (Complex V), and Na+/K+-ATPase complex, drive variation across microglia. This pipeline offers the potential for identifying functionally and analytically relevant protein targets for microglia in Alzheimer's disease and other neurological disorders.

59 BASIC BIOLOGICAL SCIENCES

[Characterization of Black and Dichothrix Cyanobacteria Based on the 16S Ribosomal RNA Gene Sequence]

My project focuses on characterizing different cyanobacteria in thrombolitic mats found on the island of Highborn Cay, Bahamas. Thrombolites are interesting ecosystems because of the ability of bacteria in these mats to remove carbon dioxide from the atmosphere and mineralize it as calcium carbonate. In the future they may be used as models to develop carbon sequestration technologies, which could be used as part of regenerative life systems in space. These thrombolitic communities are also significant because of their similarities to early communities of life on Earth. I targeted two cyanobacteria in my research, Dichothrix spp. and whatever black is, since they are believed to be important to carbon sequestration in these thrombolitic mats. The goal of my summer research project was to molecularly identify these two cyanobacteria. DNA was isolated from each organism through mat dissections and DNA extractions. I ran Polymerase Chain Reactions (PCR) to amplify the 16S ribosomal RNA (rRNA) gene in each cyanobacteria. This specific gene is found in almost all bacteria and is highly conserved, meaning any changes in the sequence are most likely due to evolution. As a result, the 16S rRNA gene can be used for bacterial identification of different species based on the sequence of their 16S rRNA gene. Since the exact sequence of the Dichothrix gene was unknown, I designed different primers that flanked the gene based on the known sequences from other taxonomically similar cyanobacteria. Once the 16S rRNA gene was amplified, I cloned the gene into specialized Escherichia coli cells and sent the gene products for sequencing. Once the sequence is obtained, it will be added to a genetic database for future reference to and classification of other Dichothrix sp.

Ortega, Maya