Search NASASearch

SEARCH · Search NASA

Results for “Introns”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

An intron within the 16S ribosomal RNA gene of the archaeon Pyrobaculum aerophilum

The 16S rRNA genes of Pyrobaculum aerophilum and Pyrobaculum islandicum were amplified by the polymerase chain reaction, and the resulting products were sequenced directly. The two organisms are closely related by this measure (over 98% similar). However, they differ in that the (lone) 16S rRNA gene of Pyrobaculum aerophilum contains a 713-bp intron not seen in the corresponding gene of Pyrobaculum islandicum. To our knowledge, this is the only intron so far reported in the small subunit rRNA gene of a prokaryote. Upon excision the intron is circularized. A secondary structure model of the intron-containing rRNA suggests a splicing mechanism of the same type as that invoked for the tRNA introns of the Archaea and Eucarya and 23S rRNAs of the Archaea. The intron contains an open reading frame whose protein translation shows no certain homology with any known protein sequence.

NASA Discipline Exobiology

COL1A1 transgene expression in stably transfected osteoblastic cells. Relative contributions of first intron, 3'-flanking sequences, and sequences derived from the body of the human COL1A1 minigene

Collagen reporter gene constructs have be used to identify cell-specific sequences needed for transcriptional activation. The elements required for endogenous levels of COL1A1 expression, however, have not been elucidated. The human COL1A1 minigene is expressed at high levels and likely harbors sequence elements required for endogenous levels of activity. Using stably transfected osteoblastic Py1a cells, we studied a series of constructs (pOBColCAT) designed to characterize further the elements required for high level of expression. pOBColCAT, which contains the COL1A1 first intron, was expressed at 50-100-fold higher levels than ColCAT 3.6, which lacks the first intron. This difference is best explained by improved mRNA processing rather than a transcriptional effect. Furthermore, variation in activity observed with the intron deletion constructs is best explained by altered mRNA splicing. Two major regions of the human COL1A1 minigene, the 3'-flanking sequences and the minigene body, were introduced into pOBColCAT to assess both transcriptional enhancing activity and the effect on mRNA stability. Analysis of the minigene body, which includes the first five exons and introns fused with the terminal six introns and exons, revealed an orientation-independent 5-fold increase in CAT activity. In contrast the 3'-flanking sequences gave rise to a modest 61% increase in CAT activity. Neither region increased the mRNA half-life of the parent construct, suggesting that CAT-specific mRNA instability elements may serve as dominant negative regulators of stability. This study suggests that other sites within the body of the COL1A1 minigene are important for high expression, e.g. during periods of rapid extracellular matrix production.

NASA Discipline Cell Biology

The first intron and promoter of Arabidopsis DIACYLGLYCEROL ACYLTRANSFERASE 1 exert synergistic effects on pollen and embryo lipid accumulation

Summary Accumulation of triacylglycerols (TAGs) is crucial during various stages of plant development. In Arabidopsis , two enzymes share overlapping functions to produce TAGs, namely acyl‐CoA:diacylglycerol acyltransferase 1 (DGAT1) and phospholipid:diacylglycerol acyltransferase 1 (PDAT1). Loss of function of both genes in a dgat1‐1/pdat1‐2 double mutant is gametophyte lethal. However, the key regulatory elements controlling tissue‐specific expression of either gene has not yet been identified. We transformed a dgat1‐1/dgat1‐1//PDAT1/pdat1‐2 parent with transgenic constructs containing the Arabidopsis DGAT1 promoter fused to the AtDGAT1 open reading frame either with or without the first intron. Triple homozygous plants were obtained, however, in the absence of the DGAT1 first intron anthers fail to fill with pollen, seed yield is c . 10% of wild‐type, seed oil content remains reduced (similar to dgat1‐1/dgat1‐1 ), and non‐Mendelian segregation of the PDAT1/pdat1‐2 locus occurs. Whereas plants expressing the AtDGAT1pro:AtDGAT1 transgene containing the first intron mostly recover phenotypes to wild‐type. This study establishes that a combination of the promoter and first intron of AtDGAT1 provides the proper context for temporal and tissue‐specific expression of AtDGAT1 in pollen. Furthermore, we discuss possible mechanisms of intron mediated regulation and how regulatory elements can be used as genetic tools to functionally replace TAG biosynthetic enzymes in Arabidopsis .

McGuire, Sean T.

The role of chromatin state in intron retention: A case study in leveraging large scale deep learning models

Complex deep learning models trained on very large datasets have become key enabling tools for current research in natural language processing and computer vision. By providing pre-trained models that can be fine-tuned for specific applications, they enable researchers to create accurate models with minimal effort and computational resources. Large scale genomics deep learning models come in two flavors: the first are large language models of DNA sequences trained in a self-supervised fashion, similar to the corresponding natural language models; the second are supervised learning models that leverage large scale genomics datasets from ENCODE and other sources. We argue that these models are the equivalent of foundation models in natural language processing in their utility, as they encode within them chromatin state in its different aspects, providing useful representations that allow quick deployment of accurate models of gene regulation. We demonstrate this premise by leveraging the recently created Sei model to develop simple, interpretable models of intron retention, and demonstrate their advantage over models based on the DNA language model DNABERT-2. Our work also demonstrates the impact of chromatin state on the regulation of intron retention. Using representations learned by Sei, our model is able to discover the involvement of transcription factors and chromatin marks in regulating intron retention, providing better accuracy than a recently published custom model developed for this purpose.

Biochemistry & Molecular Biology

Modeling study on the cleavage step of the self-splicing reaction in group I introns

A three-dimensional model of the Tetrahymena thermophila group I intron is used to further explore the catalytic mechanism of the transphosphorylation reaction of the cleavage step. Based on the coordinates of the catalytic core model proposed by Michel and Westhof (Michel, F., Westhof, E. J. Mol. Biol. 216, 585-610 (1990)), we first converted their ligation step model into a model of the cleavage step by the substitution of several bases and the removal of helix P9. Next, an attempt to place a trigonal bipyramidal transition state model in the active site revealed that this modified model for the cleavage step could not accommodate the transition state due to insufficient space. A lowering of P1 helix relative to surrounding helices provided the additional space required. Simultaneously, it provided a better starting geometry to model the molecular contacts proposed by Pyle et al. (Pyle, A. M., Murphy, F. L., Cech, T. R. Nature 358, 123-128. (1992)), based on mutational studies involving the J8/7 segment. Two hydrated Mg2+ complexes were placed in the active site of the ribozyme model, using the crystal structure of the functionally similar Klenow fragment (Beese, L.S., Steitz, T.A. EMBO J. 10, 25-33 (1991)) as a guide. The presence of two metal ions in the active site of the intron differs from previous models, which incorporate one metal ion in the catalytic site to fulfill the postulated roles of Mg2+ in catalysis. The reaction profile is simulated based on a trigonal bipyramidal transition state, and the role of the hydrated Mg2+ complexes in catalysis is further explored using molecular orbital calculations.

NASA Discipline Exobiology

Regulation of sarcomere formation and function in the healthy heart requires a titin intronic enhancer

Heterozygous truncating variants in the sarcomere protein titin (TTN) are the most common genetic cause of heart failure. To understand mechanisms that regulate abundant cardiomyocyte (CM) TTN expression, we characterized highly conserved intron 1 sequences that exhibited dynamic changes in chromatin accessibility during differentiation of human CMs from induced pluripotent stem cells (hiPSC-CMs). Homozygous deletion of these sequences in mice caused embryonic lethality, whereas heterozygous mice showed an allele-specific reduction in Ttn expression. A 296 bp fragment of this element, denoted E1, was sufficient to drive expression of a reporter gene in hiPSC-CMs. Deletion of E1 downregulated TTN expression, impaired sarcomerogenesis, and decreased contractility in hiPSC-CMs. Site-directed mutagenesis of predicted binding sites of NK2 homeobox 5 (NKX2-5) and myocyte enhancer factor 2 (MEF2) within E1 abolished its transcriptional activity. In embryonic mice expressing E1 reporter gene constructs, we validated in vivo cardiac-specific activity of E1 and the requirement for NKX2-5- and MEF2-binding sequences. Moreover, isogenic hiPSC-CMs containing a rare E1 variant in the predicted MEF2-binding motif that was identified in a patient with unexplained dilated cardiomyopathy (DCM) showed reduced TTN expression. Together, these discoveries define an essential, functional enhancer that regulates TTN expression. Manipulation of this element may advance therapeutic strategies to treat DCM caused by TTN haploinsufficiency.

Kim, Yuri

T. thermophila group I introns that cleave amide bonds

The present invention relates to nucleic acid enzymes or enzymatic RNA molecules that are capable of cleaving a variety of bonds, including phosphodiester bonds and amide bonds, in a variety of substrates. Thus, the disclosed enzymatic RNA molecules are capable of functioning as nucleases and/or peptidases. The present invention also relates to compositions containing the disclosed enzymatic RNA molecule and to methods of making, selecting, and using such enzymes and compositions.

Joyce, Gerald F.

Conserved Function of RNA Binding Motif Protein48 (RBM48) and Armadillo Repeat Containing 7 (ARMC7) in Maize and Human U12 Splicing

Splicing of pre-mRNA is fundamental for genes containing introns. Emerging data points to a deeply conserved role of this process in eukaryotic cell differentiation and proliferation. The vast majority of introns termed U2-type introns are spliced by a major spliceosome; however, there also exist rare and more conserved U12-type introns that are spliced by a minor spliceosome. Mutations that disrupt U12 splicing inhibit cell differentiation in both maize endosperm and human blood cells. However, the mechanism underlying this process is not well understood. The maize RNA Binding Motif Protein 48 (RBM48) plays an essential role as a U12 splicing factor and is required for proper maize kernel development. Using human cell lines, CRISPR-Cas9 knockdown of RBM48 demonstrated a conserved function in human U12 intron splicing. RBM48 and Armadillo Repeat Containing 7 (ARMC7) protein interact in both maize and human as part of the activated minor spliceosome. Here we show RBM48 and ARMC7 co-localization in the nucleus of human cell lines. Vertebrate ARMC7 is normally localized in the cytosol, whereas RBM48 is found in the nucleus. Our data suggests that the interaction plays a role in regulating ARMC7 localization or the efficiency of U12 intron splicing. We also performed comprehensive transcriptome profiling and identified a common subset of conserved Minor Intron Containing Genes (MIGs) impacted in both human and maize RBM48 knockout mutants. Of importance, the vast majority of these MIGs are associated with developmental defects in both plants and animals. This suggests that aberrant splicing of these targets has a high likelihood of mediating abnormal cell phenotypes. These data support evolutionarily conserved U12 splicing mechanisms between maize and humans with both RBM48 and ARMC7 having roles in the activated spliceosome.

maize

Evolution of EF-hand calcium-modulated proteins. IV. Exon shuffling did not determine the domain compositions of EF-hand proteins

In the previous three reports in this series we demonstrated that the EF-hand family of proteins evolved by a complex pattern of gene duplication, transposition, and splicing. The dendrograms based on exon sequences are nearly identical to those based on protein sequences for troponin C, the essential light chain myosin, the regulatory light chain, and calpain. This validates both the computational methods and the dendrograms for these subfamilies. The proposal of congruence for calmodulin, troponin C, essential light chain, and regulatory light chain was confirmed. There are, however, significant differences in the calmodulin dendrograms computed from DNA and from protein sequences. In this study we find that introns are distributed throughout the EF-hand domain and the interdomain regions. Further, dendrograms based on intron type and distribution bear little resemblance to those based on protein or on DNA sequences. We conclude that introns are inserted, and probably deleted, with relatively high frequency. Further, in the EF-hand family exons do not correspond to structural domains and exon shuffling played little if any role in the evolution of this widely distributed homolog family. Calmodulin has had a turbulent evolution. Its dendrograms based on protein sequence, exon sequence, 3'-tail sequence, intron sequences, and intron positions all show significant differences.

NASA Discipline Exobiology

Fractal landscapes in biological systems: long-range correlations in DNA and interbeat heart intervals

Here we discuss recent advances in applying ideas of fractals and disordered systems to two topics of biological interest, both topics having common the appearance of scale-free phenomena, i.e., correlations that have no characteristic length scale, typically exhibited by physical systems near a critical point and dynamical systems far from equilibrium. (i) DNA nucleotide sequences have traditionally been analyzed using models which incorporate the possibility of short-range nucleotide correlations. We found, instead, a remarkably long-range power law correlation. We found such long-range correlations in intron-containing genes and in non-transcribed regulatory DNA sequences as well as intragenomic DNA, but not in cDNA sequences or intron-less genes. We also found that the myosin heavy chain family gene evolution increases the fractal complexity of the DNA landscapes, consistent with the intron-late hypothesis of gene evolution. (ii) The healthy heartbeat is traditionally thought to be regulated according to the classical principle of homeostasis, whereby physiologic systems operate to reduce variability and achieve an equilibrium-like state. We found, however, that under normal conditions, beat-to-beat fluctuations in heart rate display long-range power law correlations.

Non-NASA Center

Development of Plant Gene Vectors for Tissue-Specific Expression Using GFP as a Reporter Gene

Reporter genes are widely employed in plant molecular biology research to analyze gene expression and to identify promoters. Gus (UidA) is currently the most popular reporter gene but its detection requires a destructive assay. The use of jellyfish green fluorescent protein (GFP) gene from Aequorea Victoria holds promise for noninvasive detection of in vivo gene expression. To study how various plant promoters are expressed in sweet potato (Ipomoea batatas), we are transcriptionally fusing the intron-modified (mGFP) or synthetic (modified for codon-usage) GFP coding regions to these promoters: double cauliflower mosaic virus 35S (CaMV 35S) with AMV translational enhancer, ubiquitin7-intron-ubiquitin coding region (ubi7-intron-UQ) and sporaminA. A few of these vectors have been constructed and introduced into E. coli DH5a and Agrobacterium tumefaciens EHA105. Transient expression studies are underway using protoplast-electroporation and particle bombardment of leaf tissues.

Jackson, Jacquelyn

Fractal landscape analysis of DNA walks

By mapping nucleotide sequences onto a "DNA walk", we uncovered remarkably long-range power law correlations [Nature 356 (1992) 168] that imply a new scale invariant property of DNA. We found such long-range correlations in intron-containing genes and in non-transcribed regulatory DNA sequences, but not in cDNA sequences or intron-less genes. In this paper, we present more explicit evidences to support our findings.

NASA Program Space Physiology and Countermeasures

Long-range correlations in nucleotide sequences

DNA sequences have been analysed using models, such as an n-step Markov chain, that incorporate the possibility of short-range nucleotide correlations. We propose here a method for studying the stochastic properties of nucleotide sequences by constructing a 1:1 map of the nucleotide sequence onto a walk, which we term a 'DNA walk'. We then use the mapping to provide a quantitative measure of the correlation between nucleotides over long distances along the DNA chain. Thus we uncover in the nucleotide sequence a remarkably long-range power law correlation that implies a new scale-invariant property of DNA. We find such long-range correlations in intron-containing genes and in nontranscribed regulatory DNA sequences, but not in complementary DNA sequences or intron-less genes.

NASA Discipline Cardiopulmonary

Expansion of the tmRNA sequence database and new tools for search and visualization

Abstract Transfer–messenger RNA (tmRNA) contributes essential tRNA-like and mRNA-like functions during the process of trans-translation, a mechanism of quality control for the translating bacterial ribosome. Proper tmRNA identification benefits the study of trans-translation and also the study of genomic islands, which frequently use the tmRNA gene as an integration site. Automated tmRNA gene identification tools are available, but manual inspection is still important for eliminating false positives. We have increased our database of precisely mapped tmRNA sequences over 50-fold to 97 179 unique sequences. Group I introns had previously been found integrated within a single subsite within the TψC-loop; they have now been identified at four distinct subsites, suggesting multiple founding events of invasion of tmRNA genes by group I introns, all in the same vicinity. tmRNA genes were found in metagenomic archaeal genomes, perhaps a result of misbinning of bacterial sequences during genome assembly. With the expanded database, we have produced new covariance models for improved tmRNA sequence search and new secondary structure visualization tools.

59 BASIC BIOLOGICAL SCIENCES

Complete replacement of Arabidopsis oil-producing enzymes with heterologous diacylglycerol acyltransferases

Acyl-CoA:diacylglycerol acyltransferase 1 (DGAT1) and phospholipid:diacylglycerol acyltransferase 1 (PDAT1) share responsibility for triacylglycerol (TAG) biosynthesis, and their selectivities control TAG fatty acid (FA) compositions. For rational metabolic engineering of seed oils, replacing endogenous TAG biosynthesis with exogenous enzymes containing different substrate FA selectivities is desirable; however, the dgat1-1/pdat1-2 double mutant is pollen lethal. Here, we evaluated the ability of 3 DGAT1s, from phylogenetically diverse plants with distinct TAG assembly processes, to completely replace endogenous TAG biosynthesis in Arabidopsis ( Arabidopsis thaliana ). We transformed dgat1-1 mutant plants with expression constructs for DGAT1 s from Camelina sativa , Physaria fendleri , and castor ( Ricinus communis ). Transgene expression was properly “contextualized” by using a previously determined minimum necessary expression unit containing the promoter/5′ UTR and first intron of native AtDGAT1 ; both of these DNA elements are essential for pollen expression. Next, we crossed homozygous lines with a DGAT1/DGAT1/PDAT1/pdat1-2 parent. C. sativa and P. fendleri DGAT1s restored the FA compositions and transcriptional differences of dgat1-1 to near wild-type and rescued the dgat1-1/pdat1-2 pollen lethality. R. communis DGAT1 was active in dgat1-1 seeds but produced unique oil profiles and alterations in the expression of lipid metabolic genes; it also failed to rescue dgat1-1/pdat1-2 lethality. This study confirms that the promoter and first intron of AtDGAT1 can modulate the expression of foreign DGAT1 genes to fit the correct spatiotemporal profile necessary for completely replacing endogenous TAG biosynthesis. Furthermore, it demonstrates an additional layer of unexpected enzyme incompatibility between oilseed lineages, which may complicate bioengineering approaches that seek to replace essential genes with orthologs.

McGuire, Sean T. [Washington State Univ., Pullman,

RNA Splicing Events in Circulation Distinguish Individuals With and Without New-onset Type 1 Diabetes

Context: Alterations in RNA splicing may influence protein isoform diversity that contributes to or reflects the pathophysiology of certain diseases. Whereas specific RNA splicing events in pancreatic islets have been investigated in models of inflammation in vitro, how RNA splicing in the circulation correlates with or is reflective of type 1 diabetes (T1D) disease pathophysiology in humans remains unexplored. Objective: To use machine learning to investigate if alternative RNA splicing events differ between individuals with and without new-onset T1D and to determine if these splicing events provide insight into T1D pathophysiology. Methods: RNA deep sequencing was performed on whole blood samples from 2 independent cohorts: a training cohort consisting of 12 individuals with new-onset T1D and 12 age- and sex-matched nondiabetic controls and a validation cohort of the same size and demographics. Machine learning analysis was used to identify specific isoforms that could distinguish individuals with T1D from controls. Results: Distinct patterns of RNA splicing differentiated participants with T1D from unaffected controls. Notably, certain splicing events, particularly involving retained introns, showed significant association with T1D. Machine learning analysis using these splicing events as features from the training cohort demonstrated high accuracy in distinguishing between T1D subjects and controls in the validation cohort. Gene Ontology pathway enrichment analysis of the retained intron category showed evidence for a systemic viral response in T1D subjects. Conclusion: Alternative RNA splicing events in whole blood are significantly enriched in individuals with new-onset T1D and can effectively distinguish these individuals from unaffected controls. Further, our findings also suggest that RNA splicing profiles offer the potential to provide insights into disease pathogenesis.

60 APPLIED LIFE SCIENCES

Leiomodins: larger members of the tropomodulin (Tmod) gene family

The 64-kDa autoantigen D1 or 1D, first identified as a potential autoantigen in Graves' disease, is similar to the tropomodulin (Tmod) family of actin filament pointed end-capping proteins. A novel gene with significant similarity to the 64-kDa human autoantigen D1 has been cloned from both humans and mice, and the genomic sequences of both genes have been identified. These genes form a subfamily closely related to the Tmods and are here named the Leiomodins (Lmods). Both Lmod genes display a conserved intron-exon structure, as do three Tmod genes, but the intron-exon structure of the Lmods and the Tmods is divergent. mRNA expression analysis indicates that the gene formerly known as the 64-kDa autoantigen D1 is most highly expressed in a variety of human tissues that contain smooth muscle, earning it the name smooth muscle Leiomodin (SM-Lmod; HGMW-approved symbol LMOD1). Transcripts encoding the novel Lmod gene are present exclusively in fetal and adult heart and adult skeletal muscle, and it is here named cardiac Leiomodin (C-Lmod; HGMW-approved symbol LMOD2). Human C-Lmod is located near the hypertrophic cardiomyopathy locus CMH6 on human chromosome 7q3, potentially implicating it in this disease. Our data demonstrate that the Lmods are evolutionarily related and display tissue-specific patterns of expression distinct from, but overlapping with, the expression of Tmod isoforms. Copyright 2001 Academic Press.

Carrier Proteins/biosynthesis/genetics

Left-handed Z-DNA: structure and function

Z-DNA is a high energy conformer of B-DNA that forms in vivo during transcription as a result of torsional strain generated by a moving polymerase. An understanding of the biological role of Z-DNA has advanced with the discovery that the RNA editing enzyme double-stranded RNA adenosine deaminase type I (ADAR1) has motifs specific for the Z-DNA conformation. Editing by ADAR1 requires a double-stranded RNA substrate. In the cases known, the substrate is formed by folding an intron back onto the exon that is targeted for modification. The use of introns to direct processing of exons requires that editing occurs before splicing. Recognition of Z-DNA by ADAR1 may allow editing of nascent transcripts to be initiated immediately after transcription, ensuring that editing and splicing are performed in the correct sequence. Structural characterization of the Z-DNA binding domain indicates that it belongs to the winged helix-turn-helix class of proteins and is similar to the globular domain of histone-H5.

Review