Search NASASearch

SEARCH · Search NASA

Results for “Introns”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

The first intron and promoter of Arabidopsis DIACYLGLYCEROL ACYLTRANSFERASE 1 exert synergistic effects on pollen and embryo lipid accumulation

Summary Accumulation of triacylglycerols (TAGs) is crucial during various stages of plant development. In Arabidopsis , two enzymes share overlapping functions to produce TAGs, namely acyl‐CoA:diacylglycerol acyltransferase 1 (DGAT1) and phospholipid:diacylglycerol acyltransferase 1 (PDAT1). Loss of function of both genes in a dgat1‐1/pdat1‐2 double mutant is gametophyte lethal. However, the key regulatory elements controlling tissue‐specific expression of either gene has not yet been identified. We transformed a dgat1‐1/dgat1‐1//PDAT1/pdat1‐2 parent with transgenic constructs containing the Arabidopsis DGAT1 promoter fused to the AtDGAT1 open reading frame either with or without the first intron. Triple homozygous plants were obtained, however, in the absence of the DGAT1 first intron anthers fail to fill with pollen, seed yield is c . 10% of wild‐type, seed oil content remains reduced (similar to dgat1‐1/dgat1‐1 ), and non‐Mendelian segregation of the PDAT1/pdat1‐2 locus occurs. Whereas plants expressing the AtDGAT1pro:AtDGAT1 transgene containing the first intron mostly recover phenotypes to wild‐type. This study establishes that a combination of the promoter and first intron of AtDGAT1 provides the proper context for temporal and tissue‐specific expression of AtDGAT1 in pollen. Furthermore, we discuss possible mechanisms of intron mediated regulation and how regulatory elements can be used as genetic tools to functionally replace TAG biosynthetic enzymes in Arabidopsis .

McGuire, Sean T.

The role of chromatin state in intron retention: A case study in leveraging large scale deep learning models

Complex deep learning models trained on very large datasets have become key enabling tools for current research in natural language processing and computer vision. By providing pre-trained models that can be fine-tuned for specific applications, they enable researchers to create accurate models with minimal effort and computational resources. Large scale genomics deep learning models come in two flavors: the first are large language models of DNA sequences trained in a self-supervised fashion, similar to the corresponding natural language models; the second are supervised learning models that leverage large scale genomics datasets from ENCODE and other sources. We argue that these models are the equivalent of foundation models in natural language processing in their utility, as they encode within them chromatin state in its different aspects, providing useful representations that allow quick deployment of accurate models of gene regulation. We demonstrate this premise by leveraging the recently created Sei model to develop simple, interpretable models of intron retention, and demonstrate their advantage over models based on the DNA language model DNABERT-2. Our work also demonstrates the impact of chromatin state on the regulation of intron retention. Using representations learned by Sei, our model is able to discover the involvement of transcription factors and chromatin marks in regulating intron retention, providing better accuracy than a recently published custom model developed for this purpose.

Biochemistry & Molecular Biology

Regulation of sarcomere formation and function in the healthy heart requires a titin intronic enhancer

Heterozygous truncating variants in the sarcomere protein titin (TTN) are the most common genetic cause of heart failure. To understand mechanisms that regulate abundant cardiomyocyte (CM) TTN expression, we characterized highly conserved intron 1 sequences that exhibited dynamic changes in chromatin accessibility during differentiation of human CMs from induced pluripotent stem cells (hiPSC-CMs). Homozygous deletion of these sequences in mice caused embryonic lethality, whereas heterozygous mice showed an allele-specific reduction in Ttn expression. A 296 bp fragment of this element, denoted E1, was sufficient to drive expression of a reporter gene in hiPSC-CMs. Deletion of E1 downregulated TTN expression, impaired sarcomerogenesis, and decreased contractility in hiPSC-CMs. Site-directed mutagenesis of predicted binding sites of NK2 homeobox 5 (NKX2-5) and myocyte enhancer factor 2 (MEF2) within E1 abolished its transcriptional activity. In embryonic mice expressing E1 reporter gene constructs, we validated in vivo cardiac-specific activity of E1 and the requirement for NKX2-5- and MEF2-binding sequences. Moreover, isogenic hiPSC-CMs containing a rare E1 variant in the predicted MEF2-binding motif that was identified in a patient with unexplained dilated cardiomyopathy (DCM) showed reduced TTN expression. Together, these discoveries define an essential, functional enhancer that regulates TTN expression. Manipulation of this element may advance therapeutic strategies to treat DCM caused by TTN haploinsufficiency.

Kim, Yuri

Expansion of the tmRNA sequence database and new tools for search and visualization

Abstract Transfer–messenger RNA (tmRNA) contributes essential tRNA-like and mRNA-like functions during the process of trans-translation, a mechanism of quality control for the translating bacterial ribosome. Proper tmRNA identification benefits the study of trans-translation and also the study of genomic islands, which frequently use the tmRNA gene as an integration site. Automated tmRNA gene identification tools are available, but manual inspection is still important for eliminating false positives. We have increased our database of precisely mapped tmRNA sequences over 50-fold to 97 179 unique sequences. Group I introns had previously been found integrated within a single subsite within the TψC-loop; they have now been identified at four distinct subsites, suggesting multiple founding events of invasion of tmRNA genes by group I introns, all in the same vicinity. tmRNA genes were found in metagenomic archaeal genomes, perhaps a result of misbinning of bacterial sequences during genome assembly. With the expanded database, we have produced new covariance models for improved tmRNA sequence search and new secondary structure visualization tools.

59 BASIC BIOLOGICAL SCIENCES

Complete replacement of Arabidopsis oil-producing enzymes with heterologous diacylglycerol acyltransferases

Acyl-CoA:diacylglycerol acyltransferase 1 (DGAT1) and phospholipid:diacylglycerol acyltransferase 1 (PDAT1) share responsibility for triacylglycerol (TAG) biosynthesis, and their selectivities control TAG fatty acid (FA) compositions. For rational metabolic engineering of seed oils, replacing endogenous TAG biosynthesis with exogenous enzymes containing different substrate FA selectivities is desirable; however, the dgat1-1/pdat1-2 double mutant is pollen lethal. Here, we evaluated the ability of 3 DGAT1s, from phylogenetically diverse plants with distinct TAG assembly processes, to completely replace endogenous TAG biosynthesis in Arabidopsis ( Arabidopsis thaliana ). We transformed dgat1-1 mutant plants with expression constructs for DGAT1 s from Camelina sativa , Physaria fendleri , and castor ( Ricinus communis ). Transgene expression was properly “contextualized” by using a previously determined minimum necessary expression unit containing the promoter/5′ UTR and first intron of native AtDGAT1 ; both of these DNA elements are essential for pollen expression. Next, we crossed homozygous lines with a DGAT1/DGAT1/PDAT1/pdat1-2 parent. C. sativa and P. fendleri DGAT1s restored the FA compositions and transcriptional differences of dgat1-1 to near wild-type and rescued the dgat1-1/pdat1-2 pollen lethality. R. communis DGAT1 was active in dgat1-1 seeds but produced unique oil profiles and alterations in the expression of lipid metabolic genes; it also failed to rescue dgat1-1/pdat1-2 lethality. This study confirms that the promoter and first intron of AtDGAT1 can modulate the expression of foreign DGAT1 genes to fit the correct spatiotemporal profile necessary for completely replacing endogenous TAG biosynthesis. Furthermore, it demonstrates an additional layer of unexpected enzyme incompatibility between oilseed lineages, which may complicate bioengineering approaches that seek to replace essential genes with orthologs.

McGuire, Sean T. [Washington State Univ., Pullman,

RNA Splicing Events in Circulation Distinguish Individuals With and Without New-onset Type 1 Diabetes

Context: Alterations in RNA splicing may influence protein isoform diversity that contributes to or reflects the pathophysiology of certain diseases. Whereas specific RNA splicing events in pancreatic islets have been investigated in models of inflammation in vitro, how RNA splicing in the circulation correlates with or is reflective of type 1 diabetes (T1D) disease pathophysiology in humans remains unexplored. Objective: To use machine learning to investigate if alternative RNA splicing events differ between individuals with and without new-onset T1D and to determine if these splicing events provide insight into T1D pathophysiology. Methods: RNA deep sequencing was performed on whole blood samples from 2 independent cohorts: a training cohort consisting of 12 individuals with new-onset T1D and 12 age- and sex-matched nondiabetic controls and a validation cohort of the same size and demographics. Machine learning analysis was used to identify specific isoforms that could distinguish individuals with T1D from controls. Results: Distinct patterns of RNA splicing differentiated participants with T1D from unaffected controls. Notably, certain splicing events, particularly involving retained introns, showed significant association with T1D. Machine learning analysis using these splicing events as features from the training cohort demonstrated high accuracy in distinguishing between T1D subjects and controls in the validation cohort. Gene Ontology pathway enrichment analysis of the retained intron category showed evidence for a systemic viral response in T1D subjects. Conclusion: Alternative RNA splicing events in whole blood are significantly enriched in individuals with new-onset T1D and can effectively distinguish these individuals from unaffected controls. Further, our findings also suggest that RNA splicing profiles offer the potential to provide insights into disease pathogenesis.

60 APPLIED LIFE SCIENCES

Expanding the genetic toolkit: adenine and cytosine base editors for gene disruption in Aspergillus niger

Despite revolutionizing fungal genetic engineering, conventional CRISPR/Cas9-mediated knockouts rely on DNA double-strand breaks (DSBs), which can cause unwanted insertions and deletions, chromosomal abnormalities, and cytotoxicity. Base editors such as adenine base editors (ABEs), which convert A‧T to G‧C, and cytosine base editors (CBEs), which convert C‧G to T‧A, offer a safer alternative by enabling predictable, target-specific single-nucleotide changes without introducing DSBs. To overcome the limitations of traditional genome editing in filamentous fungi, we developed efficient base-editing systems in Aspergillus niger . For the first time, we constructed an ABE in A. niger , achieving up to 80% editing efficiency and inducing predictable A-to-G mutations at the intended intron sites, disrupting gene function through mRNA mis-splicing. We also developed a highly efficient CBE system, capable of introducing premature stop codons with 50–100% efficiency. To broaden the editing scope, we implemented a Cas9-NG variant recognizing a relaxed PAM sequence requiring only a single guanine (G), enabling editing at start codons and splice sites. Leveraging this expanded scope, we established gene disruption approaches by targeting start codons via ABE-mediated A-to-G conversions (ATG-to-GTG and ATG-to-ACG) and CBE-mediated C-to-T conversion (ATG-to-ATA). Additionally, our base-editing systems enable multiplex gRNA delivery and marker-free editing of multiple genes. Collectively, the scope-expanding strategies increase the number of genes targetable for disruption by base-editing in A. niger by 26.3% and enable near-complete coverage of 96% of the coding genes. Overall, this work demonstrates the potential of ABE and CBE systems as versatile, efficient, and safer alternatives to DSBs-based gene disruption in filamentous fungi.

Aspergillus

Unique trajectory of gene family evolution from genomic analysis of nearly all known species in an ancient yeast lineage

Gene gains and losses are a major driver of genome evolution; their precise characterization can provide insights into the origin and diversification of major lineages. Here, we examined gene family evolution of 1154 genomes from nearly all known species in the medically and technologically important yeast subphylum Saccharomycotina. We found that yeast gene family evolution differs from that of plants, animals, and filamentous ascomycetes, and is characterized by smaller overall gene numbers yet larger gene family sizes for a given gene number. Faster-evolving lineages (FELs) in yeasts experienced significantly higher rates of gene losses—commensurate with a narrowing of metabolic niche breadth—but higher speciation rates than their slower-evolving sister lineages (SELs). Gene families most often lost are those involved in mRNA splicing, carbohydrate metabolism, and cell division and are likely associated with intron loss, metabolic breadth, and non-canonical cell cycle processes. Our results highlight the significant role of gene family contractions in the evolution of yeast metabolism, genome function, and speciation, and suggest that gene family evolutionary trajectories have differed markedly across major eukaryotic lineages.

Comparative Genomics

A haplotype-resolved reference genome for Eucalyptus grandis

Eucalyptus grandis is a hardwood tree used worldwide as pure species or hybrid partner to breed fast-growing plantation forestry crops that serve as feedstocks of timber and lignocellulosic biomass for pulp, paper, biomaterials, and biorefinery products. The current v2.0 genome reference for the species served as the first reference for the genus and has helped drive the development of molecular breeding tools for eucalypts. Using PacBio HiFi long reads and Omni-C proximity ligation sequencing, we produced an improved, haplotype-phased assembly (v4.0) for TAG0014, an early-generation selection of E. grandis. The 2 haplotypes are 571 Mbp (HAP1) and 552 Mbp (HAP2) in size and consist of 37 and 46 contigs scaffolded onto 11 chromosomes (contig N50 of 28.9 and 16.7 Mbp), respectively. These haplotype assemblies are 70-90 Mbp smaller than the diploid v2.0 assembly but capture all except one of the 22 telomeres, suggesting that substantial redundant sequence was included in the previous assembly. A total of 35,929 (HAP1) and 35,583 (HAP2) gene models were annotated, of which 438 and 472 contain long introns (>10 kbp) in gene models previously (v2.0) identified as multiple smaller genes. These and other improvements have increased gene annotation completeness levels from 93.8 to 99.4% in the v4.0 assembly. We found that 6,493 and 6,346 genes are within tandem duplicate arrays (HAP1 and HAP2, respectively, 18.4 and 17.8% of the total) and >43.8% of the haplotype assemblies consists of repeat elements. Analysis of synteny between the haplotypes and the E. grandis v2.0 reference genome revealed extensive regions of collinearity, but also some major rearrangements, and provided a preview of population and pangenome variation in the species.

Lötter, Anneri

Mutation-driven RRE stem-loop II conformational change induces HIV-1 nuclear export dysfunction

Abstract The Rev response element (RRE) forms an oligomeric complex with the viral protein Rev to facilitate the nuclear export of intron-retaining viral RNAs during the late phase of HIV-1 (human immunodeficiency virus type 1) infection. However, the structures and mechanisms underlying this process remain largely unknown. Here, we determined the crystal structure of the HIV-1 RRE stem-loop II (SLII), revealing a unique three-way junction architecture in which the base stem (IIa) bifurcates into the stem-loops (IIb and IIc) to compose Rev binding sites. The crystal structures of various SLII mutants demonstrated that while some mutants retain the same “compact” fold as the wild type, other single-nucleotide mutants induce drastic conformational changes, forming an “extended” SLII structure. Through in vitro Rev binding assays and Rev activity measurements in HIV-1-infected cells using structure-guided SLII mutants designed to favor specific conformers, we showed that while the compact fold represents a functional SLII, the alternative extended conformation inhibits Rev binding and oligomerization and consequently stimulates HIV-1 RNA nuclear export dysfunction. The propensity of SLII to adopt multiple conformations as captured in crystal structures and their influence on Rev oligomerization illuminate emerging perspectives on RRE structural plasticity-based regulation of HIV-1 nuclear export and provide opportunities for developing anti-HIV drugs targeting specific RRE conformations.

Biochemistry & Molecular Biology

Characterisation and comparative analysis of mitochondrial genomes of false, yellow, black and blushing morels provide insights on their structure and evolution

Morchella species have considerable significance in terrestrial ecosystems, exhibiting a range of ecological lifestyles along the saprotrophism-to-symbiosis continuum. However, the mitochondrial genomes of these ascomycetous fungi have not been thoroughly studied, thereby impeding a comprehensive understanding of their genetic makeup and ecological role. In this study, we analysed the mitogenomes of 30 Morchellaceae species, including yellow, black, blushing and false morels. These mitogenomes are either circular or linear DNA molecules with lengths ranging from 217 to 565 kbp and GC content ranging from 38% to 48%. Fifteen core protein-coding genes, 28–37 tRNA genes and 3–8 rRNA genes were identified in these Morchellaceae mitogenomes. The gene order demonstrated a high level of conservation, with the cox1 gene consistently positioned adjacent to the rnS gene and cob gene flanked by apt genes. Some exceptions were observed, such as the rearrangement of atp6 and rps3 in Morchella importuna and the reversed order of atp6 and atp8 in certain morel mitogenomes. However, the arrangement of the tRNA genes remains conserved. We additionally investigated the distribution and phylogeny of homing endonuclease genes (HEGs) of the LAGLIDADG (LAGs) and GIY-YIG (GIYs) families. A total of 925 LAG and GIY sequences were detected, with individual species containing 19–48HEGs. These HEGs were primarily located in the cox1, cob, cox2 and nad5 introns and their presence and distribution displayed significant diversity amongst morel species. These elements significantly contribute to shaping their mitogenome diversity. Overall, this study provides novel insights into the phylogeny and evolution of the Morchellaceae.

59 BASIC BIOLOGICAL SCIENCES