Search NASASearch

SEARCH · Search NASA

Results for “Gene Duplication”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

Orthologs, paralogs and genome comparisons

During the past decade, ancient gene duplications were recognized as one of the main forces in the generation of diverse gene families and the creation of new functional capabilities. New tools developed to search data banks for homologous sequences, and an increased availability of reliable three-dimensional structural information led to the recognition that proteins with diverse functions can belong to the same superfamily. Analyses of the evolution of these superfamilies promises to provide insights into early evolution but are complicated by several important evolutionary processes. Horizontal transfer of genes can lead to a vertical spread of innovations among organisms, therefore finding a certain property in some descendants of an ancestor does not guarantee that it was present in that ancestor. Complete or partial gene conversion between duplicated genes can yield phylogenetic trees with several, apparently independent gene duplications, suggesting an often surprising parallelism in the evolution of independent lineages. Additionally, the breakup of domains within a protein and the fusion of domains into multifunctional proteins makes the delineation of superfamilies a task that remains difficult to automate.

Non-NASA Center

How long did it take for life to begin and evolve to cyanobacteria?

There is convincing paleontological evidence showing that stromatolite-building phototactic prokaryotes were already in existence 3.5 x 10(9) years ago. Late accretion impacts may have killed off life on our planet as late as 3.8 x 10(9) years ago. This leaves only 300 million years to go from the prebiotic soup to the RNA world and to cyanobacteria. However, 300 million years should be more than sufficient time. All known prebiotic reactions take place in geologically rapid time scales, and very slow prebiotic reactions are not feasible because the intermediate compounds would have been destroyed due to the passage of the entire ocean through deep-sea vents every 10(7) years or in even less time. Therefore, it is likely that self-replicating systems capable of undergoing Darwinian evolution emerged in a period shorter than the destruction rates of its components (<5 million years). The time for evolution from the first DNA/protein organisms to cyanobacteria is usually thought to be very long. However, the similarities of many enzymatic reactions, together with the analysis of the available sequence data, suggest that a significant number of the components involved in basic biological processes are the result of ancient gene duplication events. Assuming that the rate of gene duplication of ancient prokaryotes was comparable to today's present values, the development of a filamentous cyanobacterial-like genome would require approximately 7 x 10(6) years--or perhaps much less. Thus, in spite of the many uncertainties involved in the estimates of time for life to arise and evolve to cyanobacteria, we see no compelling reason to assume that this process, from the beginning of the primitive soup to cyanobacteria, took more than 10 million years.

Non-NASA Center

Classification and evolution of EF-hand proteins

Forty-five distinct subfamilies of EF-hand proteins have been identified. They contain from two to eight EF-hands that are recognizable by amino acid sequence as being statistically similar to other EF-hand domains. All proteins within one subfamily are congruent to one another, i.e. the dendrogram computed from one of the EF-hand domains is similar, within statistical error, to the dendrogram computed from another(s) domain. Thirteen subfamilies--including Calmodulin, Troponin C, Essential light chain, Regulatory light chain--referred to collectively as CTER, are congruent with one another. They appear to have evolved from a single ur-domain by two cycles of gene duplication and fusion. The subfamilies of CTER subsequently evolved by gene duplications and speciations. The remaining 32 subfamilies do not show such general patterns of congruence; however, some--such as S100, intestinal calcium binding protein (calbindin 9 kd), and trichohylin--do not form congruent clusters of subfamilies. Nearly all of the domains 1, 3, 5, and 7 are most similar to other ODD domains. Correspondingly the EVEN numbered domains of all 45 subfamilies most closely resemble EVEN domains of other subfamilies. Many sequence and chemical characteristics do not show systemic trends by subfamily or species of host organisms; such homoplasy is widespread. Eighteen of the subfamilies are heterochimeric; in addition to multiple EF-hands they contain domains of other evolutionary origins.

Non-NASA Center

Vicennial metagenomic time series unveils evolutionary dynamics of giant viruses in a freshwater ecosystem

Giant viruses play crucial ecological roles in aquatic ecosystems, yet their evolutionary dynamics in response to environmental changes, particularly in freshwater environments, are not well understood. We analyzed a 20-year time series (2000-2019) of 471 co-assembled metagenomes from Lake Mendota (USA) to reconstruct 1512 giant virus metagenome-assembled genomes, providing insights into viral genome evolution. Viruses in the order Imitervirales dominate the virome, remaining consistent across seasons and years. Our findings reveal gene duplication (23% of genes) and horizontal gene transfer (29% of genes) as key drivers of genomic innovation. A co-occurrence network analysis indicates increased virus-host interactions following the introduction of an invasive predatory zooplankton in 2009, highlighting potential hosts in Bigyra, Perkinsea, and Euglenozoa. While single nucleotide polymorphism analysis shows predominantly purifying selection in viral genes, there is a significant increase in positively selected genes post-invasion, particularly those related to infection. Comparative evolutionary analyses reveal that giant viruses exhibit genome-wide substitution rates similar to co-occurring bacteria but significantly slower than smaller dsDNA phages, suggesting both stability and adaptability. Our study demonstrates that freshwater giant viruses employ various evolutionary strategies to respond to environmental change. These results underscore their significant yet often underappreciated role in freshwater ecosystem dynamics.

Vasquez, Yumary M

Comparative modeling reveals the molecular determinants of aneuploidy fitness cost in a wild yeast model

Although implicated as deleterious in many organisms, aneuploidy can underlie rapid phenotypic evolution. However, aneuploidy will be maintained only if the benefit outweighs the cost, which remains incompletely understood. To quantify this cost and the molecular determinants behind it, we generated a panel of chromosome duplications in Saccharomyces cerevisiae and applied comparative modeling and molecular validation to understand aneuploidy toxicity. We show that 74%–94% of the variance in aneuploid strains’ growth rates is explained by the cumulative cost of genes on each chromosome, measured for single-gene duplications using a genomic library, along with the deleterious contribution of small nucleolar RNAs (snoRNAs) and beneficial effects of tRNAs. Machine learning to identify properties of detrimental gene duplicates provided no support for the balance hypothesis of aneuploidy toxicity and instead identified gene length as the best predictor of toxicity. Our results present a generalized framework for the cost of aneuploidy with implications for disease biology and evolution.

genic load

Lost in translation: What we have learned from attributes that do not translate from Arabidopsis to other plants

Abstract Research in Arabidopsis thaliana has a powerful influence on our understanding of gene functions and pathways. However, not everything translates from Arabidopsis to crops and other plants. Here, a group of experts consider instances where translation has been lost and why such translation is not possible or is challenging. First, despite great efforts, floral dip transformation has not succeeded in other species outside Brassicaceae. Second, due to gene duplications and losses throughout evolution, it can be complex to establish which genes are orthologs of Arabidopsis genes. Third, during evolution Arabidopsis has lost arbuscular mycorrhizal symbiosis. Fourth, other plants have evolved specialized cell types that are not present in Arabidopsis. Fifth, similarly, C4 photosynthesis cannot be studied in Arabidopsis, which is a C3 plant. Sixth, many other plant species have larger genomes, which has given rise to innovations in transcriptional regulation that are not present in Arabidopsis. Seventh, phenotypes such as acclimation to water stress can be challenging to translate due to different measurement strategies. And eighth, while the circadian oscillator is conserved, there are important nuances in the roles of circadian regulators in crop plants. A key theme emerging across these vignettes is that even when translation is lost, insights can still be gained through comparison with Arabidopsis.

Biochemistry & Molecular Biology

Gene and genome duplications have contrasting impacts on biosynthetic and flower developmental pathways in California poppy

Benzylisoquinoline alkaloids (BIAs) represent a vast group of specialized plant metabolites with diverse pharmaceutical applications, synthesized by a variety of gene families. Among the multiple plant lineages that produce BIAs, the most notable is the poppy family (Papaveraceae), with California poppy (Eschscholzia californica) emerging as a model organism. Here, we report a haplotype-resolved genome assembly, in combination with a high-density expression atlas, for California poppy. Genome analyses reveal recent diversification of BIA biosynthesis genes in poppy through localized duplications. Furthermore, we demonstrate that the degree of phylogenetic relatedness among paralogs within BIA biosynthesis-associated gene families correlates with similarities in gene expression. In contrast, gene families involved in carotenoid biosynthesis, which contributes to the intense orange petal pigmentation, are not phylogenetically clustered, and floral developmental regulators exhibit a high degree of retention of gene duplicates associated with ancient polyploidy events. These findings illustrate alternative roles for gene and genome duplications as drivers of trait evolution. Given the position of California poppy in the angiosperm phylogeny, the high-quality genomic resources generated for this work constitute a valuable resource for comparative genomic and transcriptomic analyses for poppies and flowering plants more generally.

Rössner, Le-Han [Justus-Liebig University, Giessen

Evolutionary conservation of the presumptive neural plate markers AmphiSox1/2/3 and AmphiNeurogenin in the invertebrate chordate amphioxus

Amphioxus, as the closest living invertebrate relative of the vertebrates, can give insights into the evolutionary origin of the vertebrate body plan. Therefore, to investigate the evolution of genetic mechanisms for establishing and patterning the neuroectoderm, we cloned and determined the embryonic expression of two amphioxus transcription factors, AmphiSox1/2/3 and AmphiNeurogenin. These genes are the earliest known markers for presumptive neuroectoderm in amphioxus. By the early neurula stage, AmphiNeurogenin expression becomes restricted to two bilateral columns of segmentally arranged neural plate cells, which probably include precursors of motor neurons. This is the earliest indication of segmentation in the amphioxus nerve cord. Later, expression extends to dorsal cells in the nerve cord, which may include precursors of sensory neurons. By the midneurula, AmphiSox1/2/3 expression becomes limited to the dorsal part of the forming neural tube. These patterns resemble those of their vertebrate and Drosophila homologs. Taken together with the evolutionarily conserved expression of the dorsoventral patterning genes, BMP2/4 and chordin, in nonneural and neural ectoderm, respectively, of chordates and Drosophila, our results are consistent with the evolution of the chordate dorsal nerve cord and the insect ventral nerve cord from a longitudinal nerve cord in a common bilaterian ancestor. However, AmphiSox1/2/3 differs from its vertebrate homologs in not being expressed outside the CNS, suggesting that additional roles for this gene have evolved in connection with gene duplication in the vertebrate lineage. In contrast, expression in the midgut of AmphiNeurogenin together with the gene encoding the insulin-like peptide suggests that amphioxus may have homologs of vertebrate pancreatic islet cells, which express neurogenin3. In addition, AmphiNeurogenin, like its vertebrate and Drosophila homologs, is expressed in apparent precursors of epidermal chemosensory and possibly mechanosensory cells, suggesting a common origin for protostome and deuterostome epidermal sensory cells in the ancestral bilaterian. Copyright 2000 Academic Press.

NASA Discipline Evolutionary Biology

Analysis of the myosins encoded in the recently completed Arabidopsis thaliana genome sequence

BACKGROUND: Three types of molecular motors play an important role in the organization, dynamics and transport processes associated with the cytoskeleton. The myosin family of molecular motors move cargo on actin filaments, whereas kinesin and dynein motors move cargo along microtubules. These motors have been highly characterized in non-plant systems and information is becoming available about plant motors. The actin cytoskeleton in plants has been shown to be involved in processes such as transportation, signaling, cell division, cytoplasmic streaming and morphogenesis. The role of myosin in these processes has been established in a few cases but many questions remain to be answered about the number, types and roles of myosins in plants. RESULTS: Using the motor domain of an Arabidopsis myosin we identified 17 myosin sequences in the Arabidopsis genome. Phylogenetic analysis of the Arabidopsis myosins with non-plant and plant myosins revealed that all the Arabidopsis myosins and other plant myosins fall into two groups - class VIII and class XI. These groups contain exclusively plant or algal myosins with no animal or fungal myosins. Exon/intron data suggest that the myosins are highly conserved and that some may be a result of gene duplication. CONCLUSIONS: Plant myosins are unlike myosins from any other organisms except algae. As a percentage of the total gene number, the number of myosins is small overall in Arabidopsis compared with the other sequenced eukaryotic genomes. There are, however, a large number of class XI myosins. The function of each myosin has yet to be determined.

NASA Discipline Plant Biology

Identification of candidate host-specificity genes in Exserohilum turcicum using comparative genomics and transcriptomics

Abstract Exserohilum turcicum causes northern corn leaf blight and sorghum leaf blight. While the same species cause disease in both crops, the strains are host-specific. Here, we report the sequence and de novo annotated assemblies of one sorghum- and one maize-specific E. turcicum strain. The strains were sequenced using the PacBio Sequel II system. The total genome length for both assemblies was between 44 and 45 Mb with N50 of ∼2.5 Mb. Ninety-eight percent of the Benchmarking Universal Single-Copy Orthologs (BUSCO) for both assemblies had complete status. The estimated number of genes was 11,762 and 12,029 in the sorghum- and maize-specific isolates, respectively. Funannotate, EffectorP, SignalP, and transcriptome data were used to create functional annotation of each genome. The whole-genome comparison identified ten large-scale inversions and three translocations between the maize- and sorghum-specific strains, along with homologous genes and gene duplications. RNA was sequenced from the maize- and sorghum-specific isolate 10 days post-inoculation in maize and sorghum and from axenic cultures. Gene expression data from planta and axenic growth experiments were compared for each strain. Candidate host-specificity genes were identified by combining results from whole-genome comparison, synteny analysis, gene annotations, and transcriptome data. Overall, this study identified several candidate host-specificity genes that provide insights into E. turcicum interaction with its hosts.

Krone, Mara J. (ORCID:0000000159006624)

Inferences from protein and nucleic acid sequences - Early molecular evolution, divergence of kingdoms and rates of change

Description of new sensitive, objective methods for establishing the probable common ancestry of very distantly related sequences and the quantitative evolutionary change which has taken place. These methods are applied to four families of proteins and nucleic acids and evolutionary trees will be derived where possible. Of the three families containing duplications of genetic material, two are nucleic acids: transfer RNA and 5S ribosomal RNA. Both of these structures are functional in the synthesis of coded proteins, and prototypes must have been present in the cell at the inception of the fundamental coding process that all living things share. There are many types of tRNA which recognize the various nucleotide triplets and the 20 amino acids. These types are thought to have arisen as a result of many gene duplications. Relationships among these types are discussed. The 5S ribosomal RNA, presently functional in both eukaryotes and prokaryotes, is very likely descended from an early form incorporating almost a complete duplication of genetic material. The amount of evolution in the various lines can again be compared. The other two families containing duplications are proteins; ferredoxin and cytochrome c.

Dayhoff, M. O.

Discovery of additional ancient genome duplications in yeasts

Whole-genome duplication (WGD) has had profound macroevolutionary impacts on diverse lineages, preceding adaptive radiations in vertebrates, teleost fish, and angiosperms. In contrast to the many known ancient WGDs in animals, and especially plants, we are aware of evidence for only four WGDs in fungi. The oldest of these occurred ∼100 million years ago (mya) and is shared by ∼60 extant Saccharomycetales species, including the baker’s yeast Saccharomyces cerevisiae. Notably, this is the only known ancient WGD event in the yeast subphylum Saccharomycotina. The dearth of ancient WGD events in fungi remains a mystery. Some studies have suggested that fungal lineages that experience chromosome and genome duplication quickly go extinct, leaving no trace in the genomic record, while others contend that the lack of known WGDs is due to an absence of data. Under the second hypothesis, additional sampling and deeper sequencing of fungal genomes should lead to the discovery of more WGD events. Coupling hundreds of recently published genomes from nearly every described Saccharomycotina species, with three additional long-read assemblies, we discovered three novel WGD events. Although the functions of retained duplicate genes originating from these events are broad, they bear similarities to the well-known WGD that occurred in the Saccharomycetales. In conclusion, our results suggest that WGD may be a more common evolutionary force in fungi than previously believed.

convergent evolution

Telomere-to-telomere assemblies of chromosome 10 reveal complex adaptive variation of 3-ketoacyl-CoA-synthases in Populus trichocarpa likely driven by Helitrons

The model woody plant Populus trichocarpa displays an atypical alkene-diverse wax cuticle likely driven by copy number variation (CNV) of 3-ketoacyl-CoA synthases ( KCS ), which has been difficult to confirm with short-read assemblies. Long-read sequencing enables the development of telomere-to-telomere resources to detect cryptic variation, including CNVs, which are currently missed. Integrating this information can improve genomic prediction for breeding and provide insights into the evolutionary basis of important traits. Our analysis of 78 long-read haplotypes from chromosome 10 identified more than twice as many KCS genes as previously reported, and numerous intragenic non-synonymous substitutions. Random Forest predictive models highlighted the importance of Potri.010G079500 in producing very long chain alkenes; however, its absence did not predict previously reported alkene-deficient phenotypes. Instead, alkene levels are best predicted by the combinations of KCS copies. Additionally, amino acid substitutions clustered around ligand and donor binding pockets, suggesting they contribute to differing wax cuticle composition. Finally, each KCS gene and copy was linked to a Helitron transposon. A phylogenetic analysis suggests Helitrons are the evolutionary mechanism for generating KCS tandem arrays. Long-read generated telomere-to-telomere assemblies of P. trichocarpa chromosome 10 revealed large-effect loci critical to genetic studies that are unattainable from short-reads. This new resource produced novel insights into genome structure and function, and a novel mechanism for generating tandem gene duplication. Our results highlight that, given current challenges in annotation and assembly, detailed and focused long-read sequences are key to interpreting complex genomic regions that contain tandem copy number variants.

09 BIOMASS FUELS

Functional diversification within the heme-binding split-barrel family

Due to neofunctionalization, a single fold can be identified in multiple proteins that have distinct molecular functions. Depending on the time that has passed since gene duplication and the number of mutations, the sequence similarity between functionally divergent proteins can be relatively high, eroding the value of sequence similarity as the sole tool for accurately annotating the function of uncharacterized homologs. Here, we combine bioinformatic approaches with targeted experimentation to reveal a large multifunctional family of putative enzymatic and nonenzymatic proteins involved in heme metabolism. This family (homolog of HugZ (HOZ)) is embedded in the “FMN-binding split barrel” superfamily and contains separate groups of proteins from prokaryotes, plants, and algae, which bind heme and either catalyze its degradation or function as nonenzymatic heme sensors. In prokaryotes these proteins are often involved in iron assimilation, whereas several plant and algal homologs are predicted to degrade heme in the plastid or regulate heme biosynthesis. In the plant Arabidopsis thaliana, which contains two HOZ subfamilies that can degrade heme in vitro (HOZ1 and HOZ2), disruption of AtHOZ1 (AT3G03890) or AtHOZ2A (AT1G51560) causes developmental delays, pointing to important biological roles in the plastid. In the tree Populus trichocarpa, a recent duplication event of a HOZ1 ancestor has resulted in localization of a paralog to the cytosol. Structural characterization of this cytosolic paralog and comparison to published homologous structures suggests conservation of heme-binding sites. This study unifies our understanding of the sequence-structure-function relationships within this multilineage family of heme-binding proteins and presents new molecular players in plant and bacterial heme metabolism.

59 BASIC BIOLOGICAL SCIENCES

Designing Peptide Fossils That Model the Evolution of the Bacterial Ferredoxin Fold

Electron transfer coupled to redox chemistry is at the heart of metabolism. The proteins responsible for moving electrons (protein electron carriers) must have emerged at the origin of life. The small iron–sulfur-binding bacterial ferredoxins were likely among these first proteins. Embedded within the ferredoxin sequence and structure is a symmetry that points to an ancient gene duplication event. Little is understood about the nature of ferredoxins prior to this duplication event or what environmental factors may have driven the selection for more complex forms. The deep-time molecular history of ferredoxins goes back billions of years and cannot be reconstructed by phylogenetic analyses based on amino acid sequences. Here, we use structure-guided protein design to model a fossil half-ferredoxin stage in the evolution of this fold, the semidoxins, and their symmetric full-length counterparts, the symdoxins. Semidoxin designs homodimerize, exhibiting structural, thermodynamic, and electrochemical behaviors in most cases identical to cognate symdoxins. However, the semi- and symdoxin fossil stages behave differently when incorporated into an in vivo electron transfer complementation assay. Both can support bacterial growth dependent on protein expression. Growth rates of bacteria expressing the semidoxins are much more sensitive to oxygen than those of bacteria expressing symdoxins. Motivated by the in vivo functionality of designed semidoxins, we identified putative naturally occurring semidoxins in extant anaerobic microorganisms. This is consistent with the observed in vivo oxygen sensitivity of the semidoxin designs. One natural semidoxin is shown to be folded and redox active. However, it exists as a mixture of monomers and dimers, suggesting a potential connection between semidoxins and even simpler single iron–sulfur cluster-binding peptides.

59 BASIC BIOLOGICAL SCIENCES

A single-cell atlas of the bobtail squid visual and nervous system highlights molecular principles of convergent evolution

Abstract The cephalopod and vertebrate visual systems are a textbook example of convergent evolution with unknown molecular underpinnings. Here we characterize 98,537 single-cell transcriptomes in the bobtail squidEuprymna berryito understand how the cephalopod retina and optic lobes relate to the vertebrate retina. We confirm the overall relative simplicity of the cephalopod retina but identify two related photoreceptor cell subtypes expressing distinct r-opsins. By contrast, the adult optic lobe contains a diverse repertoire of neuronal and glial cell types, with a predominance of dopaminergic neurons. We show that cephalopod-specific gene duplicates probably contributed to this cell type diversification. Comparing neuronal cell population in the optic lobes of hatchlings and adults, we reveal a switch towards dopaminergic neurotransmitter usage with age, indicative of a maturation process. We further identify an FMRF-amide-based retrograde signal from the optic lobe towards the retina that supports the functional analogy of the cephalopod optic lobe cortex and the vertebrate inner retina in visual signal processing from a molecular standpoint. Finally, comparative analyses with vertebrate and arthropod cells suggest a scenario in which two photoreceptor types and two neuronal populations may have already been present in the eye of the bilaterian ancestor.

Environmental Sciences & Ecology

Function and Evolution of the Plant MES Family of Methylesterases

Land plant evolution has been marked by numerous genetic innovations, including novel catalytic reactions. Plants produce various carboxyl methyl esters using carboxylic acids as substrates, both of which are involved in diverse biological processes. The biosynthesis of methyl esters is catalyzed by SABATH methyltransferases, and understanding of this family has broadened in recent years. Meanwhile, the enzymes catalyzing demethylation—known as methylesterases (MESs)—have received less attention. Here, we present a comprehensive review of the plant MES family, focusing on known biochemical and biological functions, and evolution in the plant kingdom. Thirty-two MES genes have been biochemically characterized, with substrates including methyl esters of plant hormones and several other specialized metabolites. One characterized member demonstrates non-esterase activity, indicating functional diversity in this family. MES genes regulate biological processes, including biotic and abiotic defense, as well as germination and root development. While MES genes are absent in green algae, they are ubiquitous among the land plants analyzed. Extant MES genes belong to three groups of deep origin, implying ancient gene duplication and functional divergence. Two of these groups have yet to have any characterized members. Much remains to be uncovered about the enzymatic functions, biological roles, and evolution of the MES family.

59 BASIC BIOLOGICAL SCIENCES