Search NASA⌕ Search

Engineering topics

Paez-Espino, David

Publications and source records attributed to Paez-Espino, David.

At least 19 records

A global atlas of soil viruses reveals unexplored biodiversity and potential biogeochemical impacts

Historically neglected by microbial ecologists, soil viruses are now thought to be critical to global biogeochemical cycles. However, our understanding of their global distribution, activities and interactions with the soil microbiome remains limited. Here we present the Global Soil Virus Atlas, a comprehensive dataset compiled from 2,953 previously sequenced soil metagenomes and composed of 616,935 uncultivated viral genomes and 38,508 unique viral operational taxonomic units. Rarefaction curves from the Global Soil Virus Atlas indicate that most soil viral diversity remains unexplored, further underscored by high spatial turnover and low rates of shared viral operational taxonomic units across samples. By examining genes associated with biogeochemical functions, we also demonstrate the viral potential to impact soil carbon and nutrient cycling. This study represents an extensive characterization of soil viral diversity and provides a foundation for developing testable hypotheses regarding the role of the virosphere in the soil microbiome and global biogeochemistry.

59 BASIC BIOLOGICAL SCIENCES↗

Class 2 CRISPR/Cas compositions and methods of use

Provided are compositions and methods that include one or more of: (1) a Class 2 CRISPR/Cas effector protein, a nucleic acid encoding the effector protein, and/or a modified host cell comprising the effector protein (and/or a nucleic acid encoding the same); (2) a CRISPR/Cas guide RNA that binds to and provides sequence specificity to the Class 2 CRISPR/Cas effector protein, a nucleic acid encoding the CRISPR/Cas guide RNA, and/or a modified host cell comprising the CRISPR/Cas guide RNA (and/or a nucleic acid encoding the same); and (3) a CRISPR/Cas transactivating noncoding RNA (trancRNA), a nucleic acid encoding the CRISPR/Cas trancRNA, and/or a modified host cell comprising the CRISPR/Cas trancRNA (and/or a nucleic acid encoding the same).

Doudna, Jennifer A.↗

Identification of hidden N4-like viruses and their interactions with hosts

The N4-like viruses, which were recently assigned to the novel viral family Schitoviridae in 2021, belong to a podoviral-like viral lineage and possess conserved genomic characteristics and a unique replication mechanism. Despite their significance, our understanding of N4-like viruses is primarily based on viral isolates. To address this knowledge gap, this study has established a comprehensive N4-like viral data sets comprising 342 high-quality N4-like viruses/proviruses (144 viral isolates, 158 uncultured viruses, and 40 integrated N4-like proviruses). These viruses were classified into 97 subfamilies (89 of which are newly identified), 148 genera (100 of which are newly identified), and 253 species (177 of which are newly identified). The study reveals that N4-like viruses inhibit the polar region, oligotrophic open oceans, and the human gut, where they infect various bacterial lineages, such as Alpha/Beta/Gamma/Epsilon-proteobacteria in the Proteobacteria phylum. Although N4-like viral endogenization appears to be prevalent in Proteobacteria, it has also been observed in Firmicutes. Additionally, the phylogenetic analysis has identified evolutionary divergence within the hallmark genes of N4-like viruses, indicating a complex origin of the different conserved parts of viral genomes. Moreover, 1,101 putative auxiliary metabolic genes (AMGs) were identified in the N4-like viral pan-proteome, which mainly participate in nucleotide and cofactor/vitamin metabolisms. Of these AMGs, 27 were found to be associated with virulence, suggesting their potential involvement in the spread of bacterial pathogenicity. The findings of this study are significant, as N4-like viruses represent a unique viral lineage with a distinct replication mechanism and a conserved core genome. This work has resulted in a comprehensive global map of the entire N4-like viral lineage, including information on their distribution in different biomes, evolutionary divergence, genomic diversity, and the potential for viral-mediated host metabolic reprogramming. As such, this work significantly contributes to our understanding of the ecological function and viral-host interactions of bacteriophages.

60 APPLIED LIFE SCIENCES↗

Expansion of the global RNA virome reveals diverse clades of bacteriophages

High-throughput RNA sequencing offers broad opportunities to explore the Earth RNA virome. Mining 5,150 diverse metatranscriptomes uncovered >2.5 million RNA virus contigs. Analysis of >330,000 RNA-dependent RNA polymerases (RdRPs) shows that this expansion corresponds to a 5-fold increase of the known RNA virus diversity. Gene content analysis revealed multiple protein domains previously not found in RNA viruses and implicated in virus-host interactions. Extended RdRP phylogeny supports the monophyly of the five established phyla and reveals two putative additional bacteriophage phyla and numerous putative additional classes and orders. The dramatically expanded phylum Lenarviricota, consisting of bacterial and related eukaryotic viruses, now accounts for a third of the RNA virome. Identification of CRISPR spacer matches and bacteriolytic proteins suggests that subsets of picobirnaviruses and partitiviruses, previously associated with eukaryotes, infect prokaryotic hosts.

59 BASIC BIOLOGICAL SCIENCES↗

CASZ compositions and methods of use

Provided are compositions and methods that include one or more of: (1) a “CasZ” protein (also referred to as a CasZ polypeptide), a nucleic acid encoding the CasZ protein, and/or a modified host cell comprising the CasZ protein (and/or a nucleic acid encoding the same); (2) a CasZ guide RNA that binds to and provides sequence specificity to the CasZ protein, a nucleic acid encoding the CasZ guide RNA, and/or a modified host cell comprising the CasZ guide RNA (and/or a nucleic acid encoding the same); and (3) a CasZ transactivating noncoding RNA (trancRNA) (referred to herein as a “CasZ trancRNA”), a nucleic acid encoding the CasZ trancRNA, and/or a modified host cell comprising the CasZ trancRNA (and/or a nucleic acid encoding the same).

Doudna, Jennifer A.↗

Structural characterization of a soil viral auxiliary metabolic gene product – a functional chitosanase

Metagenomics is unearthing the previously hidden world of soil viruses. Many soil viral sequences in metagenomes contain putative auxiliary metabolic genes (AMGs) that are not associated with viral replication. Here, we establish that AMGs on soil viruses actually produce functional, active proteins. We focus on AMGs that potentially encode chitosanase enzymes that metabolize chitin – a common carbon polymer. We express and functionally screen several chitosanase genes identified from environmental metagenomes. One expressed protein showing endo-chitosanase activity (V-Csn) is crystalized and structurally characterized at ultra-high resolution, thus representing the structure of a soil viral AMG product. This structure provides details about the active site, and together with structure models determined using AlphaFold, facilitates understanding of substrate specificity and enzyme mechanism. Our findings support the hypothesis that soil viruses contribute auxiliary functions to their hosts.

59 BASIC BIOLOGICAL SCIENCES↗

CasZ compositions and methods of use

Provided are compositions and methods that include one or more of: (1) a “CasZ” protein (also referred to as a CasZ polypeptide), a nucleic acid encoding the CasZ protein, and/or a modified host cell comprising the CasZ protein (and/or a nucleic acid encoding the same); (2) a CasZ guide RNA that binds to and provides sequence specificity to the CasZ protein, a nucleic acid encoding the CasZ guide RNA, and/or a modified host cell comprising the CasZ guide RNA (and/or a nucleic acid encoding the same); and (3) a CasZ transactivating noncoding RNA (trancRNA) (referred to herein as a “CasZ trancRNA”), a nucleic acid encoding the CasZ trancRNA, and/or a modified host cell comprising the CasZ trancRNA (and/or a nucleic acid encoding the same).

59 BASIC BIOLOGICAL SCIENCES↗

CasZ compositions and methods of use

Provided are compositions and methods that include one or more of: (1) a “CasZ” protein (also referred to as a CasZ polypeptide), a nucleic acid encoding the CasZ protein, and/or a modified host cell comprising the CasZ protein (and/or a nucleic acid encoding the same); (2) a CasZ guide RNA that binds to and provides sequence specificity to the CasZ protein, a nucleic acid encoding the CasZ guide RNA, and/or a modified host cell comprising the CasZ guide RNA (and/or a nucleic acid encoding the same); and (3) a CasZ transactivating noncoding RNA (trancRNA) (referred to herein as a “CasZ trancRNA”), a nucleic acid encoding the CasZ trancRNA, and/or a modified host cell comprising the CasZ trancRNA (and/or a nucleic acid encoding the same).

Doudna, Jennifer A.↗

Virioplankton assemblages from challenger deep, the deepest place in the oceans

Hadal ocean biosphere, that is, the deepest part of the world’s oceans, harbors a unique microbial community, suggesting a potential uncovered co-occurring virioplankton assemblage. Herein, we reveal the unique virioplankton assemblages of the Challenger Deep, comprising 95,813 non-redundant viral contigs from the surface to the hadal zone. Almost all of the dominant viral contigs in the hadal zone were unclassified, potentially related to Alteromonadales and Oceanospirillales. 2,586 viral auxiliary metabolic genes from 132 different KEGG orthologous groups were mainly related to the carbon, nitrogen, sulfur, and arsenic metabolism. Lysogenic viral production and integrase genes were augmented in the hadal zone, suggesting the prevalence of viral lysogenic life strategy. Abundant rve genes in the hadal zone, which function as transposase in the caudoviruses, further suggest the prevalence of viral-mediated horizontal gene transfer. This study provides fundamental insights into the virioplankton assemblages of the hadal zone, reinforcing the necessity of incorporating virioplankton into the hadal biogeochemical cycles.

54 ENVIRONMENTAL SCIENCES↗

Patterns and ecological drivers of viral communities in acid mine drainage sediments across Southern China

Recent advances in environmental genomics have provided unprecedented opportunities for the investigation of viruses in natural settings. Yet, our knowledge of viral biogeographic patterns and the corresponding drivers is still limited. Here, we perform metagenomic deep sequencing on 90 acid mine drainage (AMD) sediments sampled across Southern China and examine the biogeography of viruses in this extreme environment. The results demonstrate that prokaryotic communities dictate viral taxonomic and functional diversity, abundance and structure, whereas other factors especially latitude and mean annual temperature also impact viral populations and functions. In silico predictions highlight lineage-specific virus-host abundance ratios and richness-dependent virus-host interaction structure. Further functional analyses reveal important roles of environmental conditions and horizontal gene transfers in shaping viral auxiliary metabolic genes potentially involved in phosphorus assimilation. Our findings underscore the importance of both abiotic and biotic factors in predicting the taxonomic and functional biogeographic dynamics of viruses in the AMD sediments.

59 BASIC BIOLOGICAL SCIENCES↗

Remarkably coherent population structure for a dominant Antarctic Chlorobium species

Background: In Antarctica, summer sunlight enables phototrophic microorganisms to drive primary production, thereby “feeding” ecosystems to enable their persistence through the long, dark winter months. In Ace Lake, a stratified marine-derived system in the Vestfold Hills of East Antarctica, a Chlorobium species of green sulphur bacteria (GSB) is the dominant phototroph, although its seasonal abundance changes more than 100-fold. Here, we analysed 413 Gb of Antarctic metagenome data including 59 Chlorobium metagenome-assembled genomes (MAGs) from Ace Lake and nearby stratified marine basins to determine how genome variation and population structure across a 7-year period impacted ecosystem function. Results: A single species, Candidatus Chlorobium antarcticum (most similar to Chlorobium phaeovibrioides DSM265) prevails in all three aquatic systems and harbours very little genomic variation (≥ 99% average nucleotide identity). A notable feature of variation that did exist related to the genomic capacity to biosynthesize cobalamin. The abundance of phylotypes with this capacity changed seasonally ~ 2-fold, consistent with the population balancing the value of a bolstered photosynthetic capacity in summer against an energetic cost in winter. The very high GSB concentration (> 10 8 cells ml –1 in Ace Lake) and seasonal cycle of cell lysis likely make Ca. Chlorobium antarcticum a major provider of cobalamin to the food web. Analysis of Ca. Chlorobium antarcticum viruses revealed the species to be infected by generalist (rather than specialist) viruses with a broad host range (e.g., infecting Gammaproteobacteria) that were present in diverse Antarctic lakes. The marked seasonal decrease in Ca. Chlorobium antarcticum abundance may restrict specialist viruses from establishing effective lifecycles, whereas generalist viruses may augment their proliferation using other hosts. Conclusion: The factors shaping Antarctic microbial communities are gradually being defined. In addition to the cold, the annual variation in sunlight hours dictates which phototrophic species can grow and the extent to which they contribute to ecosystem processes. The Chlorobium population studied was inferred to provide cobalamin, in addition to carbon, nitrogen, hydrogen, and sulphur cycling, as critical ecosystem services. The specific Antarctic environmental factors and major ecosystem benefits afforded by this GSB likely explain why such a coherent population structure has developed in this Chlorobium species.

59 BASIC BIOLOGICAL SCIENCES↗

CasZ compositions and methods of use

Provided are compositions and methods that include one or more of: (1) a “CasZ” protein (also referred to as a CasZ polypeptide), a nucleic acid encoding the CasZ protein, and/or a modified host cell comprising the CasZ protein (and/or a nucleic acid encoding the same); (2) a CasZ guide RNA that binds to and provides sequence specificity to the CasZ protein, a nucleic acid encoding the CasZ guide RNA, and/or a modified host cell comprising the CasZ guide RNA (and/or a nucleic acid encoding the same); and (3) a CasZ transactivating noncoding RNA (trancRNA) (referred to herein as a “CasZ trancRNA”), a nucleic acid encoding the CasZ trancRNA, and/or a modified host cell comprising the CasZ trancRNA (and/or a nucleic acid encoding the same).

Doudna, Jennifer A.↗

Prokaryotic viruses impact functional microorganisms in nutrient removal and carbon cycle in wastewater treatment plants

As one of the largest biotechnological applications, activated sludge (AS) systems in wastewater treatment plants (WWTPs) harbor enormous viruses, with 10-1,000-fold higher concentrations than in natural environments. However, the compositional variation and host-connections of AS viruses remain poorly explored. Here, we report a catalogue of ~50,000 prokaryotic viruses from six WWTPs, increasing the number of described viral species of AS by 23-fold, and showing the very high viral diversity which is largely unknown (98.4-99.6% of total viral contigs). Most viral genera are represented in more than one AS system with 53 identified across all. Viral infection widely spans 8 archaeal and 58 bacterial phyla, linking viruses with aerobic/anaerobic heterotrophs, and other functional microorganisms controlling nitrogen/phosphorous removal. Notably, Mycobacterium, notorious for causing AS foaming, is associated with 402 viral genera. Our findings expand the current AS virus catalogue and provide reference for the phage treatment to control undesired microorganisms in WWTPs.

54 ENVIRONMENTAL SCIENCES↗

A global metagenomic map of urban microbiomes and antimicrobial resistance

We present a global atlas of 4,728 metagenomic samples from mass-transit systems in 60 cities over 3 years, representing the first systematic, worldwide catalog of the urban microbial ecosystem. This atlas provides an annotated, geospatial profile of microbial strains, functional characteristics, antimicrobial resistance (AMR) markers, and genetic elements, including 10,928 viruses, 1,302 bacteria, 2 archaea, and 838,532 CRISPR arrays not found in reference databases. We identified 4,246 known species of urban microorganisms and a consistent set of 31 species found in 97% of samples that were distinct from human commensal organisms. Profiles of AMR genes varied widely in type and density across cities. Cities showed distinct microbial taxonomic signatures that were driven by climate and geographic differences. These results constitute a high-resolution global metagenomic atlas that enables discovery of organisms and genes, highlights potential public health and forensic applications, and provides a culture-independent view of AMR burden in cities.

59 BASIC BIOLOGICAL SCIENCES↗

Author Correction: A genomic catalog of Earth’s microbiomes

In the version of this article initially published, four people were missing from the alphabetical list of IMG/M Data Consortium members: Lauren V. Alteio of the Centre for Microbiology and Environmental Systems Science, University of Vienna, Vienna, Austria; Jeffrey L. Blanchard of the Biology Department, University of Massachusetts Amherst, Amherst, MA, USA; Kristen M. DeAngelis of the Department of Microbiology, University of Massachusetts Amherst, Amherst, MA, USA; and William Rodriguez-Reillo of the Research Computing Division, Harvard Medical School, Boston, MA, USA. The error has been corrected in the PDF and HTML versions of the article.

59 BASIC BIOLOGICAL SCIENCES↗

Searching for fat tails in CRISPR-Cas systems: Data analysis and mathematical modeling

Understanding CRISPR-Cas systems—the adaptive defence mechanism that about half of bacterial species and most of archaea use to neutralise viral attacks—is important for explaining the biodiversity observed in the microbial world as well as for editing animal and plant genomes effectively. The CRISPR-Cas system learns from previous viral infections and integrates small pieces from phage genomes called spacers into the microbial genome. The resulting library of spacers collected in CRISPR arrays is then compared with the DNA of potential invaders. One of the most intriguing and least well understood questions about CRISPR-Cas systems is the distribution of spacers across the microbial population. Here, using empirical data, we show that the global distribution of spacer numbers in CRISPR arrays across multiple biomes worldwide typically exhibits scale-invariant power law behaviour, and the standard deviation is greater than the sample mean. We develop a mathematical model of spacer loss and acquisition dynamics which fits observed data from almost four thousand metagenomes well. In analogy to the classical ‘rich-get-richer’ mechanism of power law emergence, the rate of spacer acquisition is proportional to the CRISPR array size, which allows a small proportion of CRISPRs within the population to possess a significant number of spacers. Our study provides an alternative explanation for the rarity of all-resistant super microbes in nature and why proliferation of phages can be highly successful despite the effectiveness of CRISPR-Cas systems.

59 BASIC BIOLOGICAL SCIENCES↗

VPF-Class: taxonomic assignment and host prediction of uncultivated viruses based on viral protein families

Abstract Motivation Two key steps in the analysis of uncultured viruses recovered from metagenomes are the taxonomic classification of the viral sequences and the identification of putative host(s). Both steps rely mainly on the assignment of viral proteins to orthologs in cultivated viruses. Viral Protein Families (VPFs) can be used for the robust identification of new viral sequences in large metagenomics datasets. Despite the importance of VPF information for viral discovery, VPFs have not yet been explored for determining viral taxonomy and host targets. Results In this work, we classified the set of VPFs from the IMG/VR database and developed VPF-Class. VPF-Class is a tool that automates the taxonomic classification and host prediction of viral contigs based on the assignment of their proteins to a set of classified VPFs. Applying VPF-Class on 731K uncultivated virus contigs from the IMG/VR database, we were able to classify 363K contigs at the genus level and predict the host of over 461K contigs. In the RefSeq database, VPF-class reported an accuracy of nearly 100% to classify dsDNA, ssDNA and retroviruses, at the genus level, considering a membership ratio and a confidence score of 0.2. The accuracy in host prediction was 86.4%, also at the genus level, considering a membership ratio of 0.3 and a confidence score of 0.5. And, in the prophages dataset, the accuracy in host prediction was 86% considering a membership ratio of 0.6 and a confidence score of 0.8. Moreover, from the Global Ocean Virome dataset, over 817K viral contigs out of 1 million were classified. Availability and implementation The implementation of VPF-Class can be downloaded from https://github.com/biocom-uib/vpf-tools. Supplementary information Supplementary data are available at Bioinformatics online.

59 BASIC BIOLOGICAL SCIENCES↗