Search NASA⌕ Search

Engineering topics

Huntemann, Marcel

Publications and source records attributed to Huntemann, Marcel.

Tetranucleotide frequencies differentiate genomic boundaries and metabolic strategies across environmental microbiomes

Microbiomes are constrained by physicochemical conditions, nutrient regimes, and community interactions across diverse environments, yet genomic signatures of this adaptation remain unclear. Metagenome sequencing is a powerful technique to analyze genomic content in the context of natural environments, establishing concepts of microbial ecological trends. Here, we developed a data discovery tool-a tetranucleotide-informed metagenome stability diagram-that is publicly available in the integrated microbial genomes and microbiomes (IMG/M) platform for metagenome ecosystem analyses. We analyzed the tetranucleotide frequencies from quality-filtered and unassembled sequence data of over 12,000 metagenomes to assess ecosystem-specific microbial community composition and function. We found that tetranucleotide frequencies can differentiate communities across various natural environments and that specific functional and metabolic trends can be observed in this structuring. Our tool places metagenomes sampled from diverse environments into clusters and along gradients of tetranucleotide frequency similarity, suggesting microbiome community compositions specific to gradient conditions. Within the resulting metagenome clusters, we identify protein-coding gene identifiers that are most differentiated between ecosystem classifications. We plan for annual updates to the metagenome stability diagram in IMG/M with new data, allowing for refinement of the ecosystem classifications delineated here. This framework has the potential to inform future studies on microbiome engineering, bioremediation, and the prediction of microbial community responses to environmental change. IMPORTANCE: Microbes adapt to diverse environments influenced by factors like temperature, acidity, and nutrient availability. We developed a new tool to analyze and visualize the genetic makeup of over 12,000 microbial communities, revealing patterns linked to specific functions and metabolic processes. This tool groups similar microbial communities and identifies characteristic genes within environments. By continually updating this tool, we aim to advance our understanding of microbial ecology, enabling applications like microbial engineering, bioremediation, and predicting responses to environmental change.

Kellom, Matthew↗

Microbial Metagenomes Across a Complete Phytoplankton Bloom Cycle: High-Resolution Sampling Every 4 Hours Over 22 Days

In May and June of 2021, marine microbial samples were collected for DNA sequencing in East Sound, WA, USA every 4 hours for 22 days. This high temporal resolution sampling effort captured the last 3 days of a Rhizosolenia sp. bloom, the initiation and complete bloom cycle of Chaetoceros socialis (8 days), and the following bacterial bloom (2 days). Metagenomes were completed on the time series, and the dataset includes 128 size-fractionated microbial samples (0.22–1.2 µm), providing gene abundances for the dominant members of bacteria, archaea, and viruses. This dataset also has time-matched nutrient analyses, flow cytometry data, and physical parameters of the environment at a single point of sampling within a coastal ecosystem that experiences regular bloom events, facilitating a range of modeling efforts that can be leveraged to understand microbial community structure and their influences on the growth, maintenance, and senescence of phytoplankton blooms.

59 BASIC BIOLOGICAL SCIENCES↗

Metatranscriptomes of California grassland soil microbial communities in response to rewetting

When very dry soil is rewet, rapid stimulation of microbial activity has important implications for ecosystem biogeochemistry, yet associated changes in microbial transcription are poorly known. Here, we present metatranscriptomes of California annual grassland soil microbial communities, collected over 1 week from soils rewet after a summer drought—providing a time series of short-term transcriptional response during rewetting.

59 BASIC BIOLOGICAL SCIENCES↗

IMG/PR: a database of plasmids from genomes and metagenomes with rich annotations and metadata

Plasmids are mobile genetic elements found in many clades of Archaea and Bacteria. They drive horizontal gene transfer, impacting ecological and evolutionary processes within microbial communities, and hold substantial importance in human health and biotechnology. To support plasmid research and provide scientists with data of an unprecedented diversity of plasmid sequences, we introduce the IMG/PR database, a new resource encompassing 699 973 plasmid sequences derived from genomes, metagenomes and metatranscriptomes. IMG/PR is the first database to provide data of plasmid that were systematically identified from diverse microbiome samples. IMG/PR plasmids are associated with rich metadata that includes geographical and ecosystem information, host taxonomy, similarity to other plasmids, functional annotation, presence of genes involved in conjugation and antibiotic resistance. The database offers diverse methods for exploring its extensive plasmid collection, enabling users to navigate plasmids through metadata-centric queries, plasmid comparisons and BLAST searches. The web interface for IMG/PR is accessible at https://img.jgi.doe.gov/pr. Plasmid metadata and sequences can be downloaded from https://genome.jgi.doe.gov/portal/IMG_PR.

59 BASIC BIOLOGICAL SCIENCES↗

Unraveling the functional dark matter through global metagenomics

Metagenomes encode an enormous diversity of proteins, reflecting a multiplicity of functions and activities1,2. Exploration of this vast sequence space has been limited to a comparative analysis against reference microbial genomes and protein families derived from those genomes. Here, to examine the scale of yet untapped functional diversity beyond what is currently possible through the lens of reference genomes, we develop a computational approach to generate reference-free protein families from the sequence space in metagenomes. We analyse 26,931 metagenomes and identify 1.17 billion protein sequences longer than 35 amino acids with no similarity to any sequences from 102,491 reference genomes or the Pfam database3. Using massively parallel graph-based clustering, we group these proteins into 106,198 novel sequence clusters with more than 100 members, doubling the number of protein families obtained from the reference genomes clustered using the same approach. We annotate these families on the basis of their taxonomic, habitat, geographical and gene neighbourhood distributions and, where sufficient sequence diversity is available, predict protein three-dimensional models, revealing novel structures. Overall, our results uncover an enormously diverse functional space, highlighting the importance of further exploring the microbial functional dark matter.

54 ENVIRONMENTAL SCIENCES↗

Terabase-Scale Coassembly of a Tropical Soil Microbiome

Petabases of reads are being produced by environmental metagenome sequencing. An essential step in analyzing these data is metagenome assembly, the computational reconstruction of genome sequences from microbial communities. “Coassembly” of metagenomic sequence data, in which multiple samples are assembled together, enables more complete detection of microbial genomes in an environment than “multiassembly,” in which samples are assembled individually.

54 ENVIRONMENTAL SCIENCES↗

IMG Annotation Pipeline (IMGAP) v5.1.13

The IMG Annotation Pipeline is a collection of Bash and Python scripts to control a workflow for structural and functional annotation of prokaryotic genomes, metagenomes, and metatranscriptomes. The bash scripts in general control the overall workflow and are wrappers around 3rd party executables (not included in repo) that predict features or functions. Whereas the Python scripts do post-processing of raw output in terms of filtering or format transformation and in some cases contain some logic for picking the correct predictions or resolving overlaps. The pipeline is tailored to produce results required by IMG (https://img.jgi.doe.gov/) and is executed on every dataset submitted to IMG via https://img.jgi.doe.gov/submit. These consist of internal genomes, metagenomes and metatranscriptomes sequenced and assembled at the JGI, as well as datasets submitted by external users (non-lab/JGI affiliates).

Huntemann, Marcel↗

Discovery of a novel filamentous prophage in the genome of the Mimosa pudica microsymbiont Cupriavidus taiwanensis STM 6018

Integrated virus genomes (prophages) are commonly found in sequenced bacterial genomes but have rarely been described in detail for rhizobial genomes. Cupriavidus taiwanensis STM 6018 is a rhizobial Betaproteobacteria strain that was isolated in 2006 from a root nodule of a Mimosa pudica host in French Guiana, South America. Here we describe features of the genome of STM 6018, focusing on the characterization of two different types of prophages that have been identified in its genome. The draft genome of STM 6018 is 6,553,639 bp, and consists of 80 scaffolds, containing 5,864 protein-coding genes and 61 RNA genes. STM 6018 contains all the nodulation and nitrogen fixation gene clusters common to symbiotic Cupriavidus species; sharing >99.97% bp identity homology to the nod / nif / noeM gene clusters from C. taiwanensis LMG19424 T and “ Cupriavidus neocalidonicus” STM 6070. The STM 6018 genome contains the genomes of two prophages: one complete Mu-like capsular phage and one filamentous phage, which integrates into a putative dif site. This is the first characterization of a filamentous phage found within the genome of a rhizobial strain. Further examination of sequenced rhizobial genomes identified filamentous prophage sequences in several Beta-rhizobial strains but not in any Alphaproteobacterial rhizobia.

59 BASIC BIOLOGICAL SCIENCES↗