Search NASA⌕ Search

SEARCH · Search NASA

Results for “Comparative genomics”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

Section-level genome sequencing and comparative genomics of Aspergillus sections Cavernicolus and Usti.

The genus Aspergillus is diverse, including species of industrial importance, human pathogens, plant pests, and model organisms. Aspergillus includes species from sections Usti and Cavernicolus, which until recently were joined in section Usti, but have now been proposed to be non-monophyletic and were split by section Nidulantes, Aenei and Raperi. To learn more about these sections, we have sequenced the genomes of 13 Aspergillus species from section Cavernicolus (A. cavernicola, A. californicus, and A. egyptiacus), section Usti (A. carlsbadensis, A. germanicus, A. granulosus, A. heterothallicus, A. insuetus, A. keveii, A. lucknowensis, A. pseudodeflectus and A. pseudoustus), and section Nidulantes (A. quadrilineatus, previously A. tetrazonus). We compared these genomes with 16 additional species from Aspergillus to explore their genetic diversity, based on their genome content, repeat-induced point mutations (RIPs), transposable elements, carbohydrate-active enzyme (CAZyme) profile, growth on plant polysaccharides, and secondary metabolite gene clusters (SMGCs). All analyses support the split of section Usti and provide additional insights: Analyses of genes found only in single species show that these constitute genes which appear to be involved in adaptation to new carbon sources, regulation to fit new niches, and bioactive compounds for competitive advantages, suggesting that these support species differentiation in Aspergillus species. Sections Usti and Cavernicolus have mainly unique SMGCs. Section Usti contains very large and information-rich genomes, an expansion partially driven by CAZymes, as section Usti contains the most CAZyme-rich species seen in genus Aspergillus. Section Usti is clearly an underutilized source of plant biomass degraders and shows great potential as industrial enzyme producers. Citation: Nybo JL, Vesth TC, Theobald S, Frisvad JC, Larsen TO, Kjaerboelling I, Rothschild-Mancinelli K, Lyhne EK, Barry K, Clum A, Yoshinaga Y, Ledsgaard L, Daum C, Lipzen A, Kuo A, Riley R, Mondo S, LaButti K, Haridas S, Pangalinan J, Salamov AA, Simmons BA, Magnuson JK, Chen J, Drula E, Henrissat B, Wiebenga A, Lubbers RJM, Müller A, dos Santos Gomes AC, Mäkelä MR, Stajich JE, Grigoriev IV, Mortensen UH, de Vries RP, Baker SE, Andersen MR (2025). Section-level genome sequencing and comparative genomics of Aspergillus sections Cavernicolus and Usti. Studies in Mycology 111: 101-114. doi: 10.3114/sim.2025.111.03.

59 BASIC BIOLOGICAL SCIENCES↗

A practical approach to using the Genomic Standards Consortium MIxS reporting standard for comparative genomics and metagenomics

Comparative analysis of (meta)genomes necessitates aggregation, integration, and synthesis of well-annotated data using standards. The Genomic Standards Consortium (GSC) collaborates with the research community to develop and maintain the Minimal Information about any (x) Sequence (MIxS) reporting standard for genomic data. To facilitate use of the GSC’s MIxS reporting standard, we provide a description of the structure and terminology, how to navigate ontologies for required terms in MIxS, and demonstrate practical usage through a soil metagenome example.

standards, metadata, genome, metagenome, schema, v↗

Reclassification of Botryococcus braunii chemical races into separate species based on a comparative genomics analysis

The colonial green microalga Botryococcus braunii is well known for producing liquid hydrocarbons that can be utilized as biofuel feedstocks. B. braunii is taxonomically classified as a single species made up of three chemical races, A, B, and L, that are mainly distinguished by the hydrocarbons produced. We previously reported a B race draft nuclear genome, and here we report the draft nuclear genomes for the A and L races. A comparative genomic study of the three B. braunii races and 14 other algal species within Chlorophyta revealed significant differences in the genomes of each race of B. braunii. Phylogenomically, there was a clear divergence of the three races with the A race diverging earlier than both the B and L races, and the B and L races diverging from a later common ancestor not shared by the A race. DNA repeat content analysis suggested the B race had more repeat content than the A or L races. Orthogroup analysis revealed the B. braunii races displayed more gene orthogroup diversity than three closely related Chlamydomonas species, with nearly 24-36% of all genes in each B. braunii race being specific to each race. This analysis suggests the three races are distinct species based on sufficient differences in their respective genomes. We propose reclassification of the three chemical races to the following species names: Botryococcus alkenealis (A race), Botryococcus braunii (B race), and Botryococcus lycopadienor (L race).

59 BASIC BIOLOGICAL SCIENCES↗

Comparative genomics provides insights into the cold adaptation of endophytic fungi associated with Deschampsia antarctica

Endophytic fungi from Deschampsia antarctica , the southernmost flowering plant, provide insights into the cold adaptation mechanisms of plant-associated fungi in extreme environments. This study presents the genome sequences and comparative analysis of eight fungal isolates from D. antarctica leaves. These Antarctic fungal isolates were analyzed alongside 121 plant-associated fungal genomes to uncover signatures of adaptation and endophytic specialization. Antarctic endophytes show striking patterns, including reduced genome size (∼26.3 Mb on average), streamlined gene content (∼8844 genes), and notably small secretomes (∼288 proteins). Despite this reduced gene repertoire, they maintain a robust set of genes encoding carbohydrate-active enzymes (CAZymes) but lack those for lignin and bacterial cell wall degradation, indicating a symbiotic lifestyle that avoids host damage and predation. One isolate, Alternaria sp. UNIPAMPA017 stood out, with 26% of its genome occupied by transposable elements. Lifestyle, rather than phylogeny, was the main driver of CAZyme and secretome profiles, underscoring ecological convergence. Compared to endophytes from Arabidopsis and Populus, D. antarctica endophytes harbor fewer pectin-degrading enzymes, reflecting their adaptation to the cell wall structure of their monocot host. Together, these fungi reveal a pattern of genomic reduction and functional fine-tuning, hallmarks of life adapted to persist in cold, nutrient-scarce niches.

Ascomycota↗

GenomeDepot: data management system for microbial comparative genomics

Summary GenomeDepot is an open-source web-based platform for annotation, management, and comparative analysis of microbial genomic sequences and associated data including ortholog families, protein domains, operons, regulatory interactions, strain taxonomy, and sample metadata. GenomeDepot supports rapid creation of websites for user-defined genome collections that include bioinformatic tools for interactive genome browsing, Basic Local Alignment Search Tool (BLAST) search, annotation search, comparative genomic neighborhood visualization, and sequence download. Gene function annotations are generated by a customizable annotation pipeline. The pipeline runs annotation tools in Conda environments and can be easily extended with additional user-specified tools. Availability and implementation GenomeDepot is open source and distributed under the GNU General Public License via GitHub (https://github.com/aekazakov/genome-depot). GenomeDepot is implemented in Python and was tested in Ubuntu Linux. Full installation instructions and documentation are available at https://aekazakov.github.io/genome-depot/. GenomeDepot demo server is freely accessible at https://iseq.lbl.gov/demogd/.

Kazakov, Alexey [Lawrence Berkeley National Labora↗

Rethinking Suicide Thi4 Thiazole Synthases: Comparative Genomic Insights and Pilot Functional Evidence

Suicide thiazole synthases (Thi4) are mononuclear metal enzymes that form the thiazole moiety of thiamin from NAD + , glycine, and a sulfur atom that is stripped from an active-site cysteine residue, causing enzyme inactivation. Comparative genomic analysis shows that prokaryotic Thi4 genes often cluster on the chromosomal regions encoding ThiS, ThiF, and other proteins that can produce, relay, or use persulfide or thiocarboxylate sulfur. These recurring genomic associations raise the possibility that, in some microorganisms, suicide Thi4s may interact with sulfur-relay systems, i.e., they can possibly operate in a nonsuicide mode. This proof-of-concept study explores this possibility via complementation assays using Escherichia coli as a heterologous platform. A representative bacterial Thi4 that clustered with thiS and thiF complemented an E. coli ΔthiG (thiazole auxotroph) single mutant better than a ΔthiG ΔthiF ΔthiS triple mutant. Although (in)direct sulfur transfer could not be assessed in the scope of our investigation, the initial results suggest a dependence on host sulfur relay components, consistent with predicted interactions with the host sulfide transfer chain. Collectively, this new perspective provides a useful guide for future biochemical studies on alternative modes of action for “suicide Thi4s” and accessory proteins.

Bacteria↗

Cross-kingdom comparative genomics reveal the metabolic potential of fungi for lignin turnover in deadwood

Deadwood is a major carbon source in forests, and yet the fate of this carbon remains a gap in our understanding of global carbon cycling. Lignin, the most recalcitrant biopolymer in wood, is mainly decayed through extracellular enzymatic and chemical processes initiated by white-rot fungi. However, the intracellular conversion of lignin decay products has been overlooked in the fungal kingdom. Here we integrate comparative genomic and phylogenetic analyses to understand the distribution and evolution of enzymes responsible for modifying lignin-related aromatic compounds—such as decarboxylases, hydroxylases, dioxygenases and other downstream ring-cleavage enzymes—that funnel carbon to central metabolism across the bacterial and the fungal kingdoms. We demonstrate that specific fungal lineages conserve these enzyme families, and that the abilities to enzymatically depolymerize lignin and catabolize lignin-related aromatic compounds are not necessarily coupled. Our analyses also reveal an expanded substrate specificity of aromatic ring-cleavage enzymes during fungal evolution, as well as a clade of extracellular enzymes among them, broadening the spatial range of these biochemical capabilities. Together, our results highlight a large diversity of fungal enzymes and hosts that warrant further investigation for inclusion into carbon cycling models and biotechnological applications for the conversion of aromatic compounds.

59 BASIC BIOLOGICAL SCIENCES↗

Comparative genomics and stable isotope analysis reveal the saprotrophic-pathogenic lifestyle of a neotropical fungus

In terrestrial forested ecosystems, fungi may interact with trees in at least three distinct ways: (i) associated with roots as symbionts; (ii) as pathogens in roots, trunks, leaves, flowers, and fruits; or (iii) decomposing dead tree tissues on soil or even on dead tissues in living trees. Distinguishing the latter two nutrition modes is rather difficult in Hymenochaetaceae (Basidiomycota) species. Herein, we have used an integrative approach of comparative genomics, stable isotopes, host tree association, and bioclimatic data to investigate the lifestyle ecology of the scarcely known neotropical genus Phellinotus, focusing on the unique species Phellinotus piptadeniae. This species is strongly associated with living Piptadenia gonoacantha (Fabaceae) trees in the Atlantic Forest domain on a relatively high precipitation gradient. Phylogenomics resolved P. piptadeniae in a clade that also includes both plant pathogens and typical wood saprotrophs. Furthermore, both genome-predicted Carbohydrate-Active Enzymes (CAZy) and stable isotopes (δ 13 C and δ 15 N) revealed a rather flexible lifestyle for the species. Altogether, our findings suggest that P. piptadeniae has been undergoing a pathotrophic specialization in a particular tree species while maintaining all the metabolic repertoire of a wood saprothroph.

59 BASIC BIOLOGICAL SCIENCES↗

Comparative genomic analysis of thermophilic fungi reveals convergent evolutionary adaptations and gene losses

Thermophily is a trait scattered across the fungal tree of life, with its highest prevalence within three fungal families (Chaetomiaceae, Thermoascaceae, and Trichocomaceae), as well as some members of the phylum Mucoromycota. We examined 37 thermophilic and thermotolerant species and 42 mesophilic species for this study and identified thermophily as the ancestral state of all three prominent families of thermophilic fungi. Thermophilic fungal genomes were found to encode various thermostable enzymes, including carbohydrate-active enzymes such as endoxylanases, which are useful for many industrial applications. At the same time, the overall gene counts, especially in gene families responsible for microbial defense such as secondary metabolism, are reduced in thermophiles compared to mesophiles. We also found a reduction in the core genome size of thermophiles in both the Chaetomiaceae family and the Eurotiomycetes class. The Gene Ontology terms lost in thermophilic fungi include primary metabolism, transporters, UV response, and O-methyltransferases. Comparative genomics analysis also revealed higher GC content in the third base of codons (GC3) and a lower effective number of codons in fungal thermophiles than in both thermotolerant and mesophilic fungi. Furthermore, using the Support Vector Machine classifier, we identified several Pfam domains capable of discriminating between genomes of thermophiles and mesophiles with 94% accuracy. Using AlphaFold2 to predict protein structures of endoxylanases (GH10), we built a similarity network based on the structures. We found that the number of disulfide bonds appears important for protein structure, and the network clusters based on protein structures correlate with the optimal activity temperature. Thus, comparative genomics offers new insights into the biology, adaptation, and evolutionary history of thermophilic fungi while providing a parts list for bioengineering applications.

59 BASIC BIOLOGICAL SCIENCES↗

Comparative genomics of Aspergillus nidulans and section Nidulantes

Aspergillus nidulans is an important model organism for eukaryotic biology and the reference for the section Nidulantes in comparative studies. In this study, we de novo sequenced the genomes of 25 species of this section. Whole-genome phylogeny of 34 Aspergillus species and Penicillium chrysogenum clarifies the position of clades inside section Nidulantes. Comparative genomics reveals a high genetic diversity between species with 684 up to 2433 unique protein families. Furthermore, we categorized 2118 secondary metabolite gene clusters (SMGC) into 603 families across Aspergilli, with at least 40 % of the families shared between Nidulantes species. Genetic dereplication of SMGC and subsequent synteny analysis provides evidence for horizontal gene transfer of a SMGC. Proteins that have been investigated in A. nidulans as well as its SMGC families are generally present in the section Nidulantes, supporting its role as model organism. The set of genes encoding plant biomass-related CAZymes is highly conserved in section Nidulantes, while there is remarkable diversity of organization of MAT-loci both within and between the different clades. This study provides a deeper understanding of the genomic conservation and diversity of this section and supports the position of A. nidulans as a reference species for cell biology.

Theobald, Sebastian [Technical University of Denma↗

Comparative genomics reveals the high diversity and adaptation strategies of Polaromonas from polar environments

Abstract Background Bacteria from the genus Polaromonas are dominant phylotypes found in a variety of low-temperature environments in polar regions. The diversity and biogeographic distribution of Polaromonas have been largely expanded on the basis of 16 S rRNA gene amplicon sequencing. However, the evolution and cold adaptation mechanisms of Polaromonas from polar regions are poorly understood at the genomic level. Results A total of 202 genomes of the genus Polaromonas were analyzed, and 121 different species were delineated on the basis of average nucleotide identity (ANI) and phylogenomic placements. Remarkably, 8 genomes recovered from polar environments clustered into a separate clade (‘polar group’ hereafter). The genome size, coding density and coding sequences (CDSs) of the polar group were significantly different from those of other nonpolar Polaromonas . Furthermore, the enrichment of genes involved in carbohydrate and peptide metabolism was evident in the polar group. In addition, genes encoding proteins related to betaine synthesis and transport were increased in the genomes from the polar group. Phylogenomic analysis revealed that two different evolutionary scenarios may explain the adaptation of Polaromonas to cold environments in polar regions. Conclusions The global distribution of the genus Polaromonas highlights its strong adaptability in both polar and nonpolar environments. Species delineation significantly expands our understanding of the diversity of the Polaromonas genus on a global scale. In this study, a polar-specific clade was found, which may represent a specific ecotype well adapted to polar environments. Collectively, genomic insight into the metabolic diversity, evolution and adaptation of the genus Polaromonas at the genome level provides a genetic basis for understanding the potential response mechanisms of Polaromonas to global warming in polar regions.

54 ENVIRONMENTAL SCIENCES↗

A fast comparative genome browser for diverse bacteria and archaea

Genome sequencing has revealed an incredible diversity of bacteria and archaea, but there are no fast and convenient tools for browsing across these genomes. It is cumbersome to view the prevalence of homologs for a protein of interest, or the gene neighborhoods of those homologs, across the diversity of the prokaryotes. We developed a web-based tool, fast . genomics , that uses two strategies to support fast browsing across the diversity of prokaryotes. First, the database of genomes is split up. The main database contains one representative from each of the 6,377 genera that have a high-quality genome, and additional databases for each taxonomic order contain up to 10 representatives of each species. Second, homologs of proteins of interest are identified quickly by using accelerated searches, usually in a few seconds. Once homologs are identified, fast . genomics can quickly show their prevalence across taxa, view their neighboring genes, or compare the prevalence of two different proteins. Fast . genomics is available at https://fast.genomics.lbl.gov .

59 BASIC BIOLOGICAL SCIENCES↗

Identification of candidate host-specificity genes in Exserohilum turcicum using comparative genomics and transcriptomics

Abstract Exserohilum turcicum causes northern corn leaf blight and sorghum leaf blight. While the same species cause disease in both crops, the strains are host-specific. Here, we report the sequence and de novo annotated assemblies of one sorghum- and one maize-specific E. turcicum strain. The strains were sequenced using the PacBio Sequel II system. The total genome length for both assemblies was between 44 and 45 Mb with N50 of ∼2.5 Mb. Ninety-eight percent of the Benchmarking Universal Single-Copy Orthologs (BUSCO) for both assemblies had complete status. The estimated number of genes was 11,762 and 12,029 in the sorghum- and maize-specific isolates, respectively. Funannotate, EffectorP, SignalP, and transcriptome data were used to create functional annotation of each genome. The whole-genome comparison identified ten large-scale inversions and three translocations between the maize- and sorghum-specific strains, along with homologous genes and gene duplications. RNA was sequenced from the maize- and sorghum-specific isolate 10 days post-inoculation in maize and sorghum and from axenic cultures. Gene expression data from planta and axenic growth experiments were compared for each strain. Candidate host-specificity genes were identified by combining results from whole-genome comparison, synteny analysis, gene annotations, and transcriptome data. Overall, this study identified several candidate host-specificity genes that provide insights into E. turcicum interaction with its hosts.

Krone, Mara J. (ORCID:0000000159006624)↗

Viral niche-partitioning: comparative genomics of giant viruses across environmental gradients in a high Arctic freshwater-saltwater lake

Giant viruses (GVs; Nucleocytoviricota) impact the biology and ecology of a wide range of eukaryotic hosts, with implications for global biogeochemical cycles. Here, we investigated GV niche separation in highly stratified Lake A at the northern coast of Ellesmere Island, Nunavut, Canada. This lake is composed of a layer of ice-covered freshwater that overlies saltwater derived from the ancient Arctic Ocean, and it therefore provides a broad gradient of environmental conditions and ecological habitats, each with a distinct protist community and rich assemblages of associated GVs. The upper layer (mixolimnion) had measurable light and oxygen, and contained diverse GVs linked to photosynthetic protists, indicating adaptation to surface biotic and abiotic conditions. In contrast, the saline lower layer (monimolimnion), lacking oxygen and light, hosted GVs associated with predicted heterotrophic protists, some of which are known for a predatory lifestyle, and with several viral genes suggesting adaptation to deep-water anaerobic conditions. Our observations underscore the coupling between physical and chemical gradients, microeukaryotes and their associated GVs in Lake A, and provide insight into the potential for GVs to directly and indirectly impact host metabolism. There were similarities between the genetic composition of GVs and the metabolic processes of their potential hosts, implying co-evolution and niche-adaptation within the lake habitats. Notably, we found a greater presence of viral rhodopsins in deeper water layers, suggesting an evolutionary relationship with potential hosts capable of supplementing their energetic needs to thrive in low energy, anoxic conditions.

59 BASIC BIOLOGICAL SCIENCES↗

Unveiling the Arsenal of Apple Bitter Rot Fungi: Comparative Genomics Identifies Candidate Effectors, CAZymes, and Biosynthetic Gene Clusters in Colletotrichum Species

The bitter rot of apple is caused by Colletotrichum spp. and is a serious pre-harvest disease that can manifest in postharvest losses on harvested fruit. In this study, we obtained genome sequences from four different species, C. chrysophilum, C. noveboracense, C. nupharicola, and C. fioriniae, that infect apple and cause diseases on other fruits, vegetables, and flowers. Our genomic data were obtained from isolates/species that have not yet been sequenced and represent geographic-specific regions. Genome sequencing allowed for the construction of phylogenetic trees, which corroborated the overall concordance observed in prior MLST studies. Bioinformatic pipelines were used to discover CAZyme, effector, and secondary metabolic (SM) gene clusters in all nine Colletotrichum isolates. We found redundancy and a high level of similarity across species regarding CAZyme classes and predicted cytoplastic and apoplastic effectors. SM gene clusters displayed the most diversity in type and the most common cluster was one that encodes genes involved in the production of alternapyrone. Our study provides a solid platform to identify targets for functional studies that underpin pathogenicity, virulence, and/or quiescence that can be targeted for the development of new control strategies. With these new genomics resources, exploration via omics-based technologies using these isolates will help ascertain the biological underpinnings of their widespread success and observed geographic dominance in specific areas throughout the country.

59 BASIC BIOLOGICAL SCIENCES↗

Scaffolded and annotated nuclear and organelle genomes of the North American brown alga Saccharina latissima

Increasing the genomic resources of emerging aquaculture crop targets can expedite breeding processes as seen in molecular breeding advances in agriculture. High quality annotated reference genomes are essential to implement this relatively new molecular breeding scheme and benefit research areas such as population genetics, gene discovery, and gene mechanics by providing a tool for standard comparison. The brown macroalga Saccharina latissima (sugar kelp) is an ecologically and economically important kelp that is found in both the northern Pacific and Atlantic Oceans. Cultivation of Saccharina latissima for human consumption has increased significantly this century in both North America and Europe, and its single blade morphology allows for dense seeding practices used in the cultivation of its Asian sister species, Saccharina japonica. While Saccharina latissima has potential as a human food crop, insufficient information from genetic resources has limited molecular breeding in sugar kelp aquaculture. We present scaffolded and annotated Saccharina latissima nuclear and organelle genomes from a female gametophyte collected from Black Ledge, Groton, Connecticut. This Saccharina latissima genome compares well with other published kelp genomes and contains 218 scaffolds with a scaffold N50 of 1.35 Mb, a GC content of 49.84%, and 25,012 predicted genes. We also validated this genome by comparing the synteny and completeness of this Saccharina latissima genome to other kelp genomes. Our team has successfully performed initial genomic selection trials with sugar kelp using a draft version of this genome. This Saccharina latissima genome expands the genetic toolkit for the economically and ecologically important sugar kelp and will be a fundamental resource for future foundational science, breeding, and conservation efforts.

DeWeese, Kelly↗