Search NASA⌕ Search

SEARCH · Search NASA

Results for “functional genomics”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 199 records · Page 11

MicroFisher: Fungal taxonomic classification for metatranscriptomic and metagenomic data using multiple short hypervariable markers

AbstractProfiling the taxonomic and functional composition of microbes using metagenomic (MG) and metatranscriptomic (MT) sequencing is advancing our understanding of microbial functions. However, the sensitivity and accuracy of microbial classification using genome– or core protein-based approaches, especially the classification of eukaryotic organisms, is limited by the availability of genomes and the resolution of sequence databases. To address this, we propose the MicroFisher, a novel approach that applies multiple hypervariable marker genes to profile fungal communities from MGs and MTs. This approach utilizes the hypervariable regions of ITS and large subunit (LSU) rRNA genes for fungal identification with high sensitivity and resolution. Simultaneously, we propose a computational pipeline (MicroFisher) to optimize and integrate the results from classifications using multiple hypervariable markers. To test the performance of our method, we applied MicroFisher to the synthetic community profiling and found high performance in fungal prediction and abundance estimation. In addition, we also used MGs from forest soil and MTs of root eukaryotic microbes to test our method and the results showed that MicroFisher provided more accurate profiling of environmental microbiomes compared to other classification tools. Overall, MicroFisher serves as a novel pipeline for classification of fungal communities from MGs and MTs.

Wang, Haihua↗

Sequential membrane- and protein-bound organelles compartmentalize genomes during phage infection

Many eukaryotic viruses require membrane-bound compartments for replication, but no such organelles are known to be formed by prokaryotic viruses. Bacteriophages of the Chimalliviridae family sequester their genomes within a phage-generated organelle, the phage nucleus, which is enclosed by a lattice of the viral protein ChmA. We show that inhibiting phage nucleus formation arrests infections at an early stage in which the injected phage genome is enclosed within a membrane-bound early phage infection (EPI) vesicle. Early phage genes are expressed from the EPI vesicle, demonstrating its functionality as a prokaryotic, transcriptionally active, membrane-bound organelle. We also show that the phage nucleus is essential, with genome replication beginning after the injected DNA is transferred from the EPI vesicle to the phage nucleus. Our results show that Chimalliviridae require two sophisticated subcellular compartments of distinct compositions and functions that facilitate successive stages of the viral life cycle.

59 BASIC BIOLOGICAL SCIENCES↗

Development, optimization, and application of an episomal plasmid system for Rhodotorula toruloides

Rhodotorula toruloides is an emerging oleaginous yeast with strong potential as a microbial cell factory for the production of acetyl-CoA-derived bioproducts. However, engineering of this organism has been limited by the absence of a functional episomal plasmid system, a foundational genetic tool for rapid gene expression, pathway testing, and CRISPR-based genome engineering. Here, we report the first episomal plasmid system for R. toruloides . Through systematic screening of candidate autonomously replicating sequences (ARSs) from diverse sources, we identified multiple functional ARS elements and selected C63F4, a fragment derived from Contig 63 of R. toruloides CBS14, because of its stable performance. The resulting pC63F4 plasmid was maintained episomally, supported GFP reporter expression, exhibited a copy number of 2.39 ± 0.13, and showed good stability during long term cultivation. To overcome poor transformation efficiency, we developed a Cre- loxP -mediated in vivo re-circularization strategy that enabled reliable delivery of the episomal plasmid. Using this improved system, we demonstrated functional episomal expression of metabolic engineering genes and multi-gene pathways for the production of triacetic acid lactone, fatty alcohols, and limonene. Finally, we leveraged this platform to establish a redesigned CRISPR system that enables seamless genome editing in R. toruloides for the first time, while also simplifying marker recycling. Together, this work establishes a long-needed episomal plasmid platform and associated CRISPR toolkit that will accelerate metabolic engineering, synthetic biology, and fundamental studies in R. toruloides .

CRISPR-Cas9↗

Data for Development, Optimization, and Application of an Episomal Plasmid System for Rhodotorula toruloides

Rhodotorula toruloides is an emerging oleaginous yeast with strong potential as a microbial cell factory for the production of acetyl-CoA-derived bioproducts. However, engineering of this organism has been limited by the absence of a functional episomal plasmid system, a foundational genetic tool for rapid gene expression, pathway testing, and CRISPR-based genome engineering. Here, we report the first episomal plasmid system for R. toruloides . Through systematic screening of candidate autonomously replicating sequences (ARSs) from diverse sources, we identified multiple functional ARS elements and selected C63F4, a fragment derived from Contig 63 of R. toruloides CBS14, because of its stable performance. The resulting pC63F4 plasmid was maintained episomally, supported GFP reporter expression, exhibited a copy number of 2.39 ± 0.13, and showed good stability during long term cultivation. To overcome poor transformation efficiency, we developed a Cre-loxP-mediated in vivo re-circularization strategy that enabled reliable delivery of the episomal plasmid. Using this improved system, we demonstrated functional episomal expression of metabolic engineering genes and multi-gene pathways for the production of triacetic acid lactone, fatty alcohols, and limonene. Finally, we leveraged this platform to establish a redesigned CRISPR system that enables seamless genome editing in R. toruloides for the first time, while also simplifying marker recycling. Together, this work establishes a long-needed episomal plasmid platform and associated CRISPR toolkit that will accelerate metabolic engineering, synthetic biology, and fundamental studies in R. toruloides .

Gene Editing↗

Evolutionary Cell Computing: From Protocells to Self-Organized Computing

On the path from inanimate to animate matter, a key step was the self-organization of molecules into protocells - the earliest ancestors of contemporary cells. Studies of the properties of protocells and the mechanisms by which they maintained themselves and reproduced are an important part of astrobiology. These studies also have the potential to greatly impact research in nanotechnology and computer science. Previous studies of protocells have focussed on self-replication. In these systems, Darwinian evolution occurs through a series of small alterations to functional molecules whose identities are stored. Protocells, however, may have been incapable of such storage. We hypothesize that under such conditions, the replication of functions and their interrelationships, rather than the precise identities of the functional molecules, is sufficient for survival and evolution. This process is called non-genomic evolution. Recent breakthroughs in experimental protein chemistry have opened the gates for experimental tests of non-genomic evolution. On the basis of these achievements, we have developed a stochastic model for examining the evolutionary potential of non-genomic systems. In this model, the formation and destruction (hydrolysis) of bonds joining amino acids in proteins occur through catalyzed, albeit possibly inefficient, pathways. Each protein can act as a substrate for polymerization or hydrolysis, or as a catalyst of these chemical reactions. When a protein is hydrolyzed to form two new proteins, or two proteins are joined into a single protein, the catalytic abilities of the product proteins are related to the catalytic abilities of the reactants. We will demonstrate that the catalytic capabilities of such a system can increase. Its evolutionary potential is dependent upon the competition between the formation of bond-forming and bond-cutting catalysts. The degree to which hydrolysis preferentially affects bonds in less efficient, and therefore less well-ordered, peptides is also critical to evolution of a non-genomic system. Based on these results, a new computational object called a "molnet" is defined. Like a neural network, it is formed of interconnected units that send "signals" to each other. Like molecules, neural networks have a specific function once their structure is defined. The difference between a molnet and traditional neural networks, is that input to molnets is not simply passed along and processed from input to output units, but rather it is utilized to form and break connections(bonds), and thus to form new structures. Molnets represent a powerful tool that can be used to understand the conditions under which chemical systems can form large molecules, such as proteins, and display ever more complex functions. This has direct applications, for example to the design of smart,synthetic fabrics. Additional information is contained in the original.

Colombano, Silvano↗

Empirical evaluation of all unique Cas9 protospacers in E. coli reveal widespread functionality and rules for gRNA design

The Cas9 nuclease has become central to modern methods and technologies in synthetic biology, largely due to the ease with which it can be targeted to specific DNA loci via guide RNAs (gRNAs). Reports vary widely on the actual specificity of this targeting, with some studies observing 60% of gRNAs possessing no activity against the genome, yet an assumption persists within the E. coli community that inactive gRNAs are rare. To resolve these contradictions, we evaluated the activity of 463 000 unique gRNAs in the E. coli K12 MG1655 genome. We show that the overwhelming majority (at least 93%) of unique gRNAs are functional while only 0.3% are nonfunctional. These nonfunctional gRNAs exhibit strong spacer self-interaction, which can either be excluded using a simple design rule or “repaired” during library design. Finally, this work provides the greater microbial synthetic biology community both a set of nearly half a million empirically evaluated E. coli gRNAs as well as a thoroughly evaluated experimental procedure, complete with appropriate controls for Cas9 activity, for conducting Cas9 assays in E. coli specifically and bacteria more generally. Lastly, we have produced a webapp to allow users to easily browse and extract gRNA sequences from the E. coli genome, which can be accessed at https://grna.ornl.gov.

Kammerdiener, Elise K. [Oak Ridge National Laborat↗

Adaptive laboratory evolution and genetic engineering improved terephthalate utilization in Pseudomonas putida KT2440

Poly(ethylene terephthalate) (PET) is one of the most ubiquitous plastics and can be depolymerized through biological and chemo-catalytic routes to its constituent monomers, terephthalic acid (TPA) and ethylene glycol (EG). TPA and EG can be re-synthesized into PET for closed-loop recycling or microbially converted into higher-value products for open-loop recycling. Here, in this study, we expand on our previous efforts engineering and applying Pseudomonas putida KT2440 for PET conversion by employing adaptive laboratory evolution (ALE) to improve TPA catabolism. Three P. putida strains with varying degrees of metabolic engineering for EG catabolism underwent an automation-enabled ALE campaign on TPA, a TPA and EG mixture, and glucose as a control. ALE increased the growth rate on TPA and TPA-EG mixtures by 4.1- and 3.5-fold, respectively, in approximately 350 generations. Evolved isolates were collected at the midpoints and endpoints of 39 independent ALE experiments, and growth rates were increased by 0.15 and 0.20 h -1 on TPA and a TPA-EG, respectively, in the best performing isolates. Whole-genome re-sequencing identified multiple converged mutations, including loss-of-function mutations to global regulators gacS, gacA, and turA along with large duplication and intergenic deletion events that impacted the heterologously-expressed tphAB II catabolic genes. Reverse engineering of these targets confirmed causality, and a strain with all three regulators deleted and second copies of tphAB II and tpaK displayed improved TPA utilization compared to the base strain. Taken together, an iterative strain engineering process involving heterologous pathway engineering, ALE, whole genome sequencing, and genome editing identified five genetic interventions that improve P. putida growth on TPA, aimed at developing enhanced whole-cell biocatalysts for PET upcycling.

36 MATERIALS SCIENCE↗

Data from: A high-quality genome assembly of the tetraploid Teucrium chamaedrys unveils a recent whole genome duplication and a large biosynthetic gene cluster for diterpenoid metabolism

Teucrium is well known for making clerodane-type diterpenoids that are produced from the backbone kolavanyl diphosphate. In order to begin to elucidate some of the complex biosynthetic pathways of these medicinal compounds, we identified and functionally characterized several kolavanyl diphosphate synthases from T. chamaedrys . Along the way, we discovered the genome of this species to be one of the largest genomes published from the Lamiaceae family, to which it belongs. This tetraploid, 3 Gbp genome is especially rich in diterpene synthase genes, with 74 putative sequences identified.

biosynthetic gene cluster (BGC)↗

Identification and Classification of Fungal GPCR Gene Families

G protein-coupled receptors (GPCRs) are transmembrane proteins crucial for signal transduction in eukaryotes, responding to diverse extracellular signals. Researchers have found and systematically summarized 14 distinct types of GPCRs in fungi but their distribution among numerous fungal species remained largely unexamined. Additionally, three families of mammalian homologs (Rhodopsin, Glutamate, and Frizzled) have been found in previous studies, but they are not included in the systematic classification of fungal GPCRs. Our study establishes a unified classification of 17 GPCR classes in fungi, combining 14 fungal and 3 mammalian previously recognized groups, and classifies 28,294 GPCRs across 1357 fungal species, significantly expanding the scale of GPCRs in fungi and demonstrating their broader distribution. We found that mammalian homologs are notably more prevalent in Early Diverging Fungi (EDF), whereas the previous 14 classes are predominantly found in Ascomycota and Basidiomycota. The most abundant class detected in fungi was Pth11-like GPCRs, exclusively found in Pezizomycotina and involved in fungal pathogenicity. Our analysis suggested that Pezizomycotina ancestor possessed an extensive array of Pth11-like GPCRs, but over time, some species underwent considerable reductions in these GPCRs in conjunction with genome contractions. Utilizing a custom-built convolutional neural network (CNN) for the identification of fungal GPCRs, we identified several putative novel fungal GPCRs. Predicted interactions between these prospective new GPCRs and G-alpha proteins, as simulated by AlphaFold Multimer, provided additional support for their functional relevance. In conclusion, our work defines the first large-scale, unified classification of fungal GPCRs, reveals lineage-specific expansions and contractions, and uncovers previously unrecognized GPCR candidates with potential functional roles in fungal signaling.

G protein-coupled receptors↗

Gene Fusion: A Genome Wide Survey

As a well known fact, organisms form larger and complex multimodular (composite or chimeric) and mostly multi-functional proteins through gene fusion of two or more individual genes which have independent evolution histories and functions. We call each of these components a module. The existence of multimodular proteins may improves the efficiency in gene regulation and in cellular functions, and thus may give the host organism advantages in adaptation to environments. Analysis of all gene fusions in present-day organisms should allow us to examine the patterns of gene fusion in context with cellular functions, to trace back the evolution processes from the ancient smaller and uni-functional proteins to the present-day larger and complex multi-functional proteins, and to estimate the minimal number of ancestor proteins that existed in the last common ancestor for all life on earth. Although many multimodular proteins have been experimentally known, identification of gene fusion events systematically at genome scale had not been possible until recently when large number of completed genome sequences have been becoming available. In addition, technical difficulties for such analysis also exist due to the complexity of this biological and evolutionary process. We report from this study a new strategy to computationally identify multimodular proteins using completed genome sequences and the results surveyed from 22 organisms with the data from over 40 organisms to be presented during the meeting. Additional information is contained in the original extended abstract.

Liang, Ping↗

Modeling the Activity of Single Genes

The central dogma of molecular biology states that information is stored in DNA, transcribed to messenger RNA (mRNA) and then translated into proteins. This picture is significantly augmentated when we consider the action of certain proteins in regulating transcription. These transcription factors provide a feedback pathway by which genes can regulate one another's expression as mRNA and then as protein. To review: DNA, RNA and proteins have different functions. DNA is the molecular storehouse of genetic information. When cells divide, the DNA is replicated, so that each daughter cell maintains the same genetic information as the mother cell. RNA acts as a go-between from DNA to proteins. Only a single copy of DNA is present, but multiple copies of the same piece of RNA may be present, allowing cells to make huge amounts of protein. In eukaryotes (organisms with a nucleus), DNA is found in the nucleus only. RNA is copied in the nucleus then translocates(moves) outside the nucleus, where it is transcribed into proteins. Along the way, the RNA may be spliced, i.e., may have pieces cut out. RNA then attaches to ribosomes and is translated to proteins. Proteins are the machinery of the cell other than DNA and RNA, all the complex molecules of the cell are proteins. Proteins are specialized machines, each of which fulfills its own task, which may be transporting oxygen, catalyzing reactions, or responding to extracellular signals, just to name a few. One of the more interesting functions a protein may have is binding directly or indirectly to DNA to perform transcriptional regulation, thus forming a closed feedback loop of gene regulation. The structure of DNA and the central dogma were understood in the 50s; in the early 80s it became possible to make arbitrary modifications to DNA and use cellular machinery to transcribe and translate the resulting genes; more recently, genomes (i.e., the complete DNA sequence) of many organisms have been sequenced. This large-scale sequencing began with simple organisms, viruses and bacteria, progressed to eukaryotes such as yeast, and more recently (1998) progressed to a multi-cellular animal, the nematode Caenorhabditis elegans. Sequencers have now moved on to the fruit fly Drosophila melanogaster, whose sequence is slated for completion by the end of 1999. The human genome project is expected to determine the complete sequence of all 3 billion bases of human DNA within the next five years. In the wake of genome-scale sequencing, further instrumentation is being developed to assay gene expression and function on a comparably large scale. Much of the work in computational biology focuses on computational tools used in sequencing, finding genes that are related to a particular gene, finding which parts of the DNA code for proteins and which do not, understanding what proteins will be formed from a given length of DNA, predicting how the proteins will fold from a one-dimensional structure into a three dimensional structure, and so on. Much less computational work has been done regarding the function of proteins. One reason for this is that different proteins function very differently, and so work on protein function is very specific to certain classes of proteins. There are, for example, proteins such enzymes that catalyze various intracellular reactions, receptors that respond to extracellular signals and ion channels that regulate the flow of charged particles into and out of the cell. In this chapter, we will consider a particular class of proteins called transcription factors(TFs), which are responsible for regulating when a certain gene is expressed in a certain cell, which cells it is express in, and how much is expressed. Understanding these processes will involve developing a deeper understanding of transcription, translation, and the cellular processes that control those processes. All of these elements fall under the aegis of gene regulation or more narrowly transcriptional regulation. Some of the key questions in gene regulation are: What genes are expressed in a certain cell at a certain time? How does gene expression differ from cell to cell in a multicellular organism? Which proteins act as transcription factors, i.e., are important in regulating gene expression? From questions like these, we hope to understand which genes are important for various macroscopic processes. Nearly all of the cells of a multicellular organism contain the same DNA. Yet this same genetic information yields a large number of different cell types. The fundamental difference between a neuron and a liver cell, for example, is which genes are expressed. Thus understanding gene regulation is an important step in understanding development. Furthermore, understanding the usual genes that are expressed in cells may give important clues about various diseases. Some diseases, such as sickle cell anemia and cystic fibrosis, are caused by defects in single, non-regulatory genes; others, such as certain cancers, are caused when the cellular control circuitry malfunctions - an understanding of these diseases will involve pathways of multiple interacting gene products. There are numerous challenges in the area of understanding and modeling gene regulation. First and foremost, biologists would like to develop a deeper understanding of the processes involved, including which genes and families of genes are important, how they interact, etc. From a computation point of view, there has been embarrassingly little work done. In this chapter there are many areas in which we can phrase meaningful, non-trivial computational questions, but questions that have not been addressed. Some of these are purely computational (what is a good algorithm for dealing with a model of type X) and others are more mathematical (given a system with certain characteristics, what sort of model can one use? How does one find biochemical parameters from system-level behavior using as few experiments as possible?). In addition to biological and algorithmic problems, there is also the ever-present issue of theoretical biology - what general principles can be derived from these systems, what can one do with models other than just simulate time-courses, what can be deduced about a class of systems without knowing all the details? The fundamental challenge to computationalists and theorists is to add value to the biology - to use models, modeling techniques and algorithms to understand the biology in new ways.

Mjolsness, Eric↗

Ca X ML: Chemistry‐informed machine learning explains mutual changes between protein conformations and calcium ions in calcium‐binding proteins using structural and topological features

Proteins' flexibility is a feature in communicating changes in cell signaling instigated by binding with secondary messengers, such as calcium ions, associated with the coordination of muscle contraction, neurotransmitter release, and gene expression. When binding with the disordered parts of a protein, calcium ions must balance their charge states with the shape of calcium-binding proteins and their versatile pool of partners depending on the circumstances they transmit. Accurately determining the ionic charges of those ions is essential for understanding their role in such processes. However, it is unclear whether the limited experimental data available can be effectively used to train models to accurately predict the charges of calcium-binding protein variants. Here, we developed a chemistry-informed, machine-learning algorithm that implements a game theoretic approach to explain the output of a machine-learning model without the prerequisite of an excessively large database for high-performance prediction of atomic charges. We used the ab initio electronic structure data representing calcium ions and the structures of the disordered segments of calcium-binding peptides with surrounding water molecules to train several explainable models. Network theory was used to extract the topological features of atomic interactions in the structurally complex data dictated by the coordination chemistry of a calcium ion, a potent indicator of its charge state in protein. Our design created a computational tool of Ca X ML, which provided a framework of explainable machine learning model to annotate ionic charges of calcium ions in calcium-binding proteins in response to the chemical changes in an environment. Our framework will provide new insights into protein design for engineering functionality based on the limited size of scientific data in a genome space.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Exabiome: Advancing Microbial Science through Exascale Computing

The Exabiome project seeks to improve the understanding of microbiomes through the development of methods for accelerating metagenomic science using exascale computing. This article gives an overview of scientific impact of the three components of the project: metagenome assembly, protein family detection, and comparative analysis of metagenomes. Exabiome developed MetaHipMer, the only metagenome assembler capable of scaling to full exascale systems. MetaHipMer has enabled ground-breaking assemblies on the Frontier supercomputer, with many scientific benefits, such as the discovery of rare species and viral genomes. To investigate protein families, Exabiome developed two exascale tools, PASTIS and HipMCL. Together, these can utilize exascale resources to understand the functional diversity of billions of dark matter proteins and novel protein families. For comparative analysis, Exabiome developed kmerprof, a tool that can be used to compare huge metagenomes for many different scientific purposes, for example, grouping human microbiomes according to body location.

59 BASIC BIOLOGICAL SCIENCES↗

The reference genome for the northeastern Pacific bull kelp, Nereocystis luetkeana

Bull kelp, Nereocystis luetkeana, is a northeastern Pacific kelp with broad distribution from Alaska to central California. Its population declines have caused severe concerns in northern California, the Salish Sea in Washington, and recently in some populations in Oregon. Despite bull kelp's accumulated ecological and physiological studies, an assembled and annotated genomic reference was still unavailable. Here, we report the complete and annotated genome of Nereocystis luetkeana, produced by the California Conservation Genomics Project (CCGP), which aims to reveal genomic diversity patterns across California by sequencing the complete genomes of approximately 150 carefully selected species. The genome was assembled into 1562 scaffolds with 449.82 Mb, 80x of coverage and 22 952 gene models. BUSCO assembly showed a completeness score of 72% for the stramenopiles gene set. The mitochondria and chloroplast genome sequences have 37 Kb and 131 Mb, respectively. The orthology analysis between 10 Phaeophycean genomes showed 1065 expanded and 286 unique orthogroups for this species. Pairwise comparisons showed 542 orthogroups present only in N. luetkeana and M. pyrifera, another large-body kelp. The enrichment analysis of these orthogroups showed important functions related to central metabolism and signaling due to ATPases enrichment in these two species. This genome assembly will provide an essential resource for the ecology, evolution, conservation, and breeding of bull kelp.

California Conservation Genomics Project—CCGP↗

The protein structurome of Orthornavirae and its dark matter

Metatranscriptomics is uncovering more and more diverse families of viruses with RNA genomes comprising the viral kingdom Orthornavirae in the realm Riboviria. Thorough protein annotation and comparison are essential to get insights into the functions of viral proteins and virus evolution. In addition to sequence- and hmm profile-based methods, protein structure comparison adds a powerful tool to uncover protein functions and relationships. We constructed an Orthornavirae “structurome” consisting of already annotated as well as unannotated (“dark matter”) proteins and domains encoded in viral genomes. We used protein structure modeling and similarity searches to illuminate the remaining dark matter in hundreds of thousands of orthornavirus genomes. The vast majority of the dark matter domains showed either “generic” folds, such as single α-helices, or no high confidence structure predictions. Nevertheless, a variety of lineage-specific globular domains that were new either to orthornaviruses in general or to particular virus families were identified within the proteomic dark matter of orthornaviruses, including several predicted nucleic acid-binding domains and nucleases. In addition, we identified a case of exaptation of a cellular nucleoside monophosphate kinase as an RNA-binding protein in several virus families. Notwithstanding the continuing discovery of numerous orthornaviruses, it appears that all the protein domains conserved in large groups of viruses have already been identified. The rest of the viral proteome seems to be dominated by poorly structured domains including intrinsically disordered ones that likely mediate specific virus-host interactions.

59 BASIC BIOLOGICAL SCIENCES↗

Multi-layered heterochromatin interaction as a switch for DIM2-mediated DNA methylation

Functional crosstalk between DNA methylation, histone H3 lysine-9 trimethylation (H3K9me3) and heterochromatin protein 1 (HP1) is essential for proper heterochromatin assembly and genome stability. However, how repressive chromatin cues guide DNA methyltransferases for region-specific DNA methylation remains largely unknown. Here, we report structure-function characterizations of DNA methyltransferase Defective-In-Methylation-2 (DIM2) in Neurospora . The DNA methylation activity of DIM2 requires the presence of both H3K9me3 and HP1. Our structural study reveals a bipartite DIM2-HP1 interaction, leading to a disorder-to-order transition of the DIM2 target-recognition domain that is essential for substrate binding. Furthermore, the structure of DIM2-HP1-H3K9me3-DNA complex reveals a substrate-binding mechanism distinct from that for its mammalian orthologue DNMT1. In addition, the dual recognition of H3K9me3 peptide by the DIM2 RFTS and BAH1 domains allosterically impacts the DIM2-substrate binding, thereby controlling DIM2-mediated DNA methylation. Together, this study uncovers how multiple heterochromatin factors coordinately orchestrate an activity-switching mechanism for region-specific DNA methylation.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

A Fluorescence‐Based Transient Expression Assay for the Analysis of Upstream Open Reading Frames in Plants

Upstream open reading frames (uORFs) are regulatory elements present in the 5′ leaders of mRNA that can significantly impact downstream gene expression in eukaryotes. In crop engineering, editing of uORFs can provide an avenue to upregulate expression of native genes without the need to add persistent transgenic copies. Even with genome-wide methods to identify translated uORFs such as ribosome profiling, their functional characterization depends on validation through reporter gene assays and mutagenesis studies. Current screening methods for plants use luciferases or protoplasts to measure differential gene expression between wild-type and mutated transcript leaders, which requires tissue processing and/or substrate addition. Here, we present a time- and cost-efficient alternative to investigate transcript leaders by co-expression of two fluorescent proteins in Nicotiana benthamiana leaf tissue and test our assay on genes involved in photoprotection, editing of which could provide a pathway to increase CO 2 assimilation during sun–shade transitions.

Nicotiana benthamiana↗

Predictive models of the genetic bases underlying budding yeast fitness in multiple environments

Abstract The ability of organisms to adapt and survive depends on the effects of genes and the environment on fitness. However, the multigenic nature of fitness and genotype-by-environment interactions hinder our understanding of the genetic basis of fitness. Here, we established fitness prediction models for 35 environments using machine learning and existing fitness data and different genetic variant types for a Saccharomyces cerevisiae population. Models revealed that the predictive ability of genetic variants varied across environments, with copy number variants explaining the majority of fitness variation in most cases. Model interpretation showed that different variant types identified distinct gene sets associated with predictive variants. These gene sets were significantly enriched in experimentally validated genes affecting fitness in only a subset of environments, indicating that many genes influencing fitness remain unexplored. Notably, non-experimentally validated genes were more important than validated ones for fitness predictions. Gene contributions to predictions were both isolate- and environment-dependent, pointing to gene-by-gene and gene-by-environment interactions. Furthermore, models uncovered experimentally validated and novel candidate genetic interactions for a well-characterized stress, the fungicide benomyl. These findings highlight the feasibility of identifying the genetic basis of fitness by using different genetic variant types and offer novel targets for future functional analysis.

DNA copy number variations↗