Search NASA⌕ Search

SEARCH · Search NASA

Results for “biological databases”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 343 records · Page 19

AstroAmpSeq: Microbial Bioinformatics Education with NASA GeneLab’s Amplicon Pipeline

The prevalence and importance of large sequencing datasets in microbiology has led to a movement to share microbial ecology experimental data through open-access databases. This is particularly true of experiments that are difficult to replicate, such as those conducted in the spaceflight environment and shared via NASA GeneLab. It is now possible and indeed valuable for students to access and re-analyze these shared datasets for educational and research purposes. To provide students with experience utilizing microbial bioinformatics tools, GeneLab for Colleges and Universities (GL4U) has designed AstroAmpSeq, a week-long, virtually implemented project-based learning (PBL) minicourse to instruct undergraduate students on 16S amplicon sequencing. AstroAmpSeq was created to be accessible to students without prior bioinformatics or microbial ecology experience. During the minicourse students work in teams to process, analyze, and visualize a subsample of GeneLab dataset GLDS-280 using GeneLab’s standard amplicon processing pipeline, which is based in R. Students develop a hypothesis related to the dataset then generate and analyze figures to evaluate their hypothesis. Formative assessment of student learning is determined via pre- and post-evaluations, peer feedback, and self-reflection. Project and presentation rubrics serve as a summative assessment of student learning. GL4U AstroAmpSeq not only meets American Society for Microbiology Curriculum Guidelines, but also incites student interest in research by an inquiry-based approach and can be made part of a larger semester-long curriculum. GL4U AstroAmpSeq raises awareness of space microbiology and bioinformatics as a field and career path among undergraduates. Further, by using a GeneLab dataset and nesting microbiology techniques into the real-world application of space biology, AstroAmpSeq enforces deeper and longer-lasting student learning.

microbiology↗

U.S. Pacific Coast Workshop Report on Preconstruction Research Recommendations (U.S. Offshore Wind Synthesis of Environmental Effects Research (SEER) Project)

In May 2022, the U.S. Offshore Wind Synthesis of Environmental Effects Research (SEER) project team hosted a stakeholder workshop focused on preconstruction (baseline) research needs for potential floating offshore wind (OSW) energy development on the U.S. Pacific Coast, including California, Oregon, and Washington. Prior to the workshop, the SEER team developed a set of initial synthesized research recommendations that were identified based on a review of relevant, publicly available resources and with advisory group input. The workshop covered three marine life breakout groups on subsequent days to discuss research recommendations related to 1) marine mammals and sea turtles, 2) fish and invertebrates, and 3) birds and bats. As part of the workshop, over a hundred participants from the public and private sectors provided feedback on various aspects of the initial research recommendations, including associated data and knowledge gaps, benefits/limitations of available methods and technologies, and technological advancements or infrastructure needed to address the recommendation. Approximately 1,000 total comments were received on the workshop MURAL boards and were synthesized in this report. Based on workshop feedback, SEER developed a final database of over 500 specific research recommendations based on more than 40 resources. In Fall 2022, the full database and a tool with updated synthesized research recommendations were disseminated on Tethys (https://tethys.pnnl.gov) to assist with informing future funding opportunities and research programming. There is a continued need to improve awareness of the potential environmental effects, monitoring technologies, and management strategies for floating OSW energy development on the U.S. Pacific Coast. Coordination of these activities will require the sustained involvement of multiple stakeholders from across sectors. Beyond the baseline considerations discussed in this workshop, future state-of-the-science activities should be planned to consider research needs across wind energy life cycle phases for all relevant wildlife taxa and associated habitat and ecosystem processes.

17 WIND ENERGY↗

Nuclear model calculations and their role in space radiation research

Proper assessments of spacecraft shielding requirements and concomitant estimates of risk to spacecraft crews from energetic space radiation requires accurate, quantitative methods of characterizing the compositional changes in these radiation fields as they pass through thick absorbers. These quantitative methods are also needed for characterizing accelerator beams used in space radiobiology studies. Because of the impracticality/impossibility of measuring these altered radiation fields inside critical internal body organs of biological test specimens and humans, computational methods rather than direct measurements must be used. Since composition changes in the fields arise from nuclear interaction processes (elastic, inelastic and breakup), knowledge of the appropriate cross sections and spectra must be available. Experiments alone cannot provide the necessary cross section and secondary particle (neutron and charged particle) spectral data because of the large number of nuclear species and wide range of energies involved in space radiation research. Hence, nuclear models are needed. In this paper current methods of predicting total and absorption cross sections and secondary particle (neutrons and ions) yields and spectra for space radiation protection analyses are reviewed. Model shortcomings are discussed and future needs presented. c2002 COSPAR. Published by Elsevier Science Ltd. All right reserved.

NASA Center JSC↗

Warming is Associated With More Encoded Antimicrobial Resistance Genes and Transcriptions Within Five Drug Classes in Soil Bacteria: A Case Study and Synthesis

ABSTRACT The effect of warming on anti‐microbial resistance (AMR) genes in the environment has critical implications for public health but is little studied. We collected published soil bacterial genomes from the BV‐BRC database and tested the correlation between reported optimal growth temperature and the number of encoded AMR genes. Furthermore, we tested the relationship between temperature and AMR gene transcription in a natural ecosystem by analysing soil transcriptomes from a warming manipulation experiment in an Alaskan boreal forest. We hypothesised that there is a positive relationship between warming and AMR prevalence in gene content in bacterial genomes and transcriptomic sequences, and that this effect would vary by drug class. Regarding the bacterial genomes, we found a positive relationship between the fraction of encoded AMR genes and the reported optimal temperature of soil bacteria. The drug classes tetracycline and lincosamide/macrolide/streptogramin had the strongest positive relationship with reported optimal temperature. For the case study in a natural ecosystem, we found 61 significantly upregulated AMR gene‐associated transcripts spanning eight drug classes in warmed plots. In the Alaskan soil samples, we found that warming elicited the strongest positive effect on transcripts targeting lincosamide/streptogramin, beta‐lactam and phenicol/quinolone antibiotics. Overall, higher temperatures were linked to AMR gene prevalence.

Hacopian, Melanie T. [Department of Ecology and Ev↗

Rapid recovery from the Late Ordovician mass extinction

Understanding the evolutionary role of mass extinctions requires detailed knowledge of postextinction recoveries. However, most models of recovery hinge on a direct reading of the fossil record, and several recent studies have suggested that the fossil record is especially incomplete for recovery intervals immediately after mass extinctions. Here, we analyze a database of genus occurrences for the paleocontinent of Laurentia to determine the effects of regional processes on recovery and the effects of variations in preservation and sampling intensity on perceived diversity trends and taxonomic rates during the Late Ordovician mass extinction and Early Silurian recovery. After accounting for variation in sampling intensity, we find that marine benthic diversity in Laurentia recovered to preextinction levels within 5 million years, which is nearly 15 million years sooner than suggested by global compilations. The rapid turnover in Laurentia suggests that processes such as immigration may have been particularly important in the recovery of regional ecosystems from environmental perturbations. However, additional regional studies and a global analysis of the Late Ordovician mass extinction that accounts for variations in sampling intensity are necessary to confirm this pattern. Because the record of Phanerozoic mass extinctions and postextinction recoveries may be compromised by variations in preservation and sampling intensity, all should be reevaluated with sampling-standardized analyses if the evolutionary role of mass extinctions is to be fully understood.

Evolution↗

An in silico assessment of gene function and organization of the phenylpropanoid pathway metabolic networks in Arabidopsis thaliana and limitations thereof

The Arabidopsis genome sequencing in 2000 gave to science the first blueprint of a vascular plant. Its successful completion also prompted the US National Science Foundation to launch the Arabidopsis 2010 initiative, the goal of which is to identify the function of each gene by 2010. In this study, an exhaustive analysis of The Institute for Genomic Research (TIGR) and The Arabidopsis Information Resource (TAIR) databases, together with all currently compiled EST sequence data, was carried out in order to determine to what extent the various metabolic networks from phenylalanine ammonia lyase (PAL) to the monolignols were organized and/or could be predicted. In these databases, there are some 65 genes which have been annotated as encoding putative enzymatic steps in monolignol biosynthesis, although many of them have only very low homology to monolignol pathway genes of known function in other plant systems. Our detailed analysis revealed that presently only 13 genes (two PALs, a cinnamate-4-hydroxylase, a p-coumarate-3-hydroxylase, a ferulate-5-hydroxylase, three 4-coumarate-CoA ligases, a cinnamic acid O-methyl transferase, two cinnamoyl-CoA reductases) and two cinnamyl alcohol dehydrogenases can be classified as having a bona fide (definitive) function; the remaining 52 genes currently have undetermined physiological roles. The EST database entries for this particular set of genes also provided little new insight into how the monolignol pathway was organized in the different tissues and organs, this being perhaps a consequence of both limitations in how tissue samples were collected and in the incomplete nature of the EST collections. This analysis thus underscores the fact that even with genomic sequencing, presumed to provide the entire suite of putative genes in the monolignol-forming pathway, a very large effort needs to be conducted to establish actual catalytic roles (including enzyme versatility), as well as the physiological function(s) for each member of the (multi)gene families present and the metabolic networks that are operative. Additionally, one key to identifying physiological functions for many of these (and other) unknown genes, and their corresponding metabolic networks, awaits the development of technologies to comprehensively study molecular processes at the single cell level in particular tissues and organs, in order to establish the actual metabolic context.

NASA Program Fundamental Space Biology↗

The secondary metabolism collaboratory: a database and web discussion portal for secondary metabolite biosynthetic gene clusters

Secondary metabolites are small molecules produced by all corners of life, often with specialized bioactive functions with clinical and environmental relevance. Secondary metabolite biosynthetic gene clusters (BGCs) can often be identified within DNA sequences by various sequence similarity tools, but determining the exact functions of genes in the pathway and predicting their chemical products can often only be done by careful, manual comparative analysis. To facilitate this, we report the first release of the secondary metabolism collaboratory (SMC), which aims to provide a comprehensive, tool-agnostic repository of BGC sequence data drawn from all publicly available and user-submitted bacterial and archaeal genome and contig sources. On the website, users are provided a searchable catalog of putative BGCs identified from each source, along with visualizations of gene and domain annotations derived from multiple sequence analysis tools. SMC’s data is also available through publicly-accessible application programming interface (API) endpoints to facilitate programmatic access. Users are encouraged to share their findings (and search for others’) through comment posts on BGC and source pages. At the time of writing, SMC is the largest repository of BGC information, holding 13.1M BGC regions from 1.3M source sequences and growing, and can be found at https://smc.jgi.doe.gov.

59 BASIC BIOLOGICAL SCIENCES↗

Metabolic interactions underpinning high methane fluxes across terrestrial freshwater wetlands

Current estimates of wetland contributions to the global methane budget carry high uncertainty, particularly in accurately predicting emissions from high methane-emitting wetlands. Microorganisms drive methane cycling, but little is known about their conservation across wetlands. To address this, we integrate 16S rRNA amplicon datasets, metagenomes, metatranscriptomes, and annual methane flux data across 9 wetlands, creating the Multi-Omics for Understanding Climate Change (MUCC) v2.0.0 database. This resource is used to link microbiome composition to function and methane emissions, focusing on methane-cycling microbes and the networks driving carbon decomposition. We identify eight methane-cycling genera shared across wetlands and show wetland-specific metabolic interactions in marshes, revealing low connections between methanogens and methanotrophs in high-emitting wetlands. Methanoregula emerged as a hub methanogen across networks and is a strong predictor of methane flux. In these wetlands it also displays the functional potential for methylotrophic methanogenesis, highlighting the importance of this pathway in these ecosystems. Collectively, our findings illuminate trends between microbial decomposition networks and methane flux while providing an extensive publicly available database to advance future wetland research.

54 ENVIRONMENTAL SCIENCES↗

AlgaeOrtho, a bioinformatics tool for processing ortholog inference results in algae

Introduction: Microalgae constitute a prominent feedstock for producing biofuels and biochemicals by virtue of their prolific reproduction, high bioproduct accumulation, and the ability to grow in brackish and saline water. However, naturally occurring wild type algal strains are rarely optimal for industrial use; therefore, bioengineering of algae is necessary to generate superior performing strains that can address production challenges in industrial settings, particularly the bioenergy and bioproduct sectors. One of the crucial steps in this process is deciding on a bioengineering target: namely, which gene/protein to differentially express. These targets are often orthologs which are defined as genes/proteins originating from a common ancestor in divergent species. Although bioinformatics tools for the identification of protein orthologs already exist, processing the output from such tools is nontrivial, especially for a researcher with little or no bioinformatics experience. Methods: The present study introduces AlgaeOrtho, a user-friendly tool that builds upon the SonicParanoid orthology inference tool (based on an algorithm that identifies potential protein orthologs based on amino acid sequences) and the PhycoCosm database from JGI (Joint Genome Institute) to help researchers identify orthologs of their proteins of interest in multiple diverse algal species. Results: The output of this application includes a table of the putative orthologs of their protein of interest, a heatmap showing sequence similarity (%), and an unrooted tree of the putative protein orthologs. Notably, the tool would be instrumental in identifying novel bioengineering targets in different algal strains, including targets in not-fully annotated algal species, since it does not depend on existing protein annotations. We tested AlgaeOrtho using three case studies, for which orthologs of proteins relevant to bioengineering targets, were identified from diverse algal species, demonstrating its ease of use and utility for bioengineering researchers. Discussion: This tool is unique in the protein ortholog identification space as it can visualize putative orthologs, as desired by the user, across several algal species.

09 BIOMASS FUELS↗

PAVC: The foundation for a Pan-Arctic Vegetation Cover database

Field-measured Arctic vegetation cover data is essential for creating accurate, high-quality vegetation structure and composition maps. Extrapolating field data into high-resolution cover maps provides detailed, function-specific information for use in Earth System Models, vegetation classifications, and monitoring vegetation change over time and space. However, field campaigns that collect plant cover vary substantially in scope, method, and purpose, which makes them difficult to unify across data stores, and they are often not designed to meet remote sensing needs. In this work, we synthesized and harmonized field-based fractional cover data from various data stores to create a high-quality, consistent repository schema for remote sensing-based vegetation cover mapping applications. We developed a reproducible workflow for synthesizing visual estimate and point-intercept fractional cover data. The resultant Pan-Arctic Vegetation Cover (PAVC) database contains synthesized fractional cover at both the species and plant functional type levels. The latter includes absolute foliar cover for deciduous shrubs and trees, evergreen shrubs and trees, forbs, graminoids, lichen, bryophytes, and “other” vegetation, as well as absolute cover for litter and top cover for water and bare ground.

Steckler, Morgan R. [Oak Ridge National Laboratory↗

Characterization of Two Microbial Isolates from Andean Lakes in Bolivia

We are currently investigating the biological population present in the highest and least explored perennial lakes on earth in the Bolivian and Chilean Andes, including several volcanic crater lakes of more than 6000 m elevation, in combination of microbiological and molecular biological methods. Our samples were collected in saline lakes of the Laguna Blanca Laguna Verde area in the Bolivian Altiplano and in the Licancabur volcano crater (27 deg. 47 min S/67 deg. 47 min. W) in the ongoing project studying high altitude lakes. The main goal of the project is to look for analogies with Martian paleolakes. These Bolivian lakes can be described as Andean lakes following the classification of Chong. We have attempted to isolate pure cultures and phylogenetically characterize prokaryotes that grew under laboratory conditions. Sediment samples taken from the Licancabur crater lake (LC), Laguna Verde (LV), and Laguna Blanca (LB) were analyzed and cultured using enriched liquid media under both aerobic and anaerobic conditions. All cultures were incubated at room temperature (15 to 20 C) and under light exposure. For the reported isolates, 36 hours incubation were necessary for reaching optimal optical densities to consider them viable cultures. Ten serial dilutions starting from 1% inoculum were required to obtain a suitable enriched cell culture to transfer into solid media. Cultures on solid medium were necessary to verify the formation of colonies in order to isolate pure cultures. Different solid media were prepared using several combinations of both trace minerals and carbohydrates sources in order to fit their nutrient requirements. The microorganisms formed individual colonies on solid media enriched with tryptone, yeast extract and sodium chloride. Cells morphology was studied by optical and electronic microscopy. Rodshape morphologies were observed in most cases. Total bacterial genomic DNA was isolated from 50 ml late-exponential phase culture by using the CTAB miniprep protocol. The 16S rRNA genes were amplified by PCR using both Bacteria- and Archaeauniversal primer sets: 27f and 1492r, 21f and 1492r respectively. Sequences of 16S rRNA gene were determined and initially compared with reference sequences contained in the EMBL nucleotide sequence database by using the BLAST program and were subsequently aligned with 16S rRNA reference sequences in the ARB package (http://www.mikro.biologie.tu-muenchen.de). Aligned sequences were inserted within a stable phylogenetic tree by using the ARB parsimony tool. In this work we report the morphology and phylogenetic characterization of two isolates belonged to Laguna Blanca sediments.

Demergasso, C.↗

The impact of curation errors in the PDBBind Database on machine learning predictions of protein–protein binding affinity

The PDBBind database has been widely utilized for the computational prediction of protein–protein binding affinities. While the accuracy of the PDBBind-curated equilibrium dissociation constants (K D ) has been reported for the protein–ligand subset of the PDBBind database, the curation accuracy has not been reported for the protein–protein subset. Here, we present a detailed manual analysis for the subset of PDBBind records with PubMed Central Open Access primary publications and find that ~19% of these records had K D values that were not supported by their primary publications. The impact of these putative curation errors on the machine learning-based prediction of K D from experimental protein–protein 3D structures was evaluated and correcting the curation errors improved the Pearson correlation coefficient between measured and random forest-predicted log 10 (K D ) values by ~8 percentage points. This finding underscores the importance of dataset accuracy for computational modelling and highlights the need for more stringent curation processes when extracting information from the scientific literature.

59 BASIC BIOLOGICAL SCIENCES↗

Produced Water DNA Database (PW-DNA): Utilizing KBase to generate an environmental specific curated molecular database

The deep subsurface is estimated to host the majority of Earth’s microbial biomass yet remains one of the most challenging environments to access and study. One common approach to investigate these microbial communities is through the analysis of produced water from subsurface reservoirs, where researchers can assess water and gas chemistry along with molecular (DNA/RNA) sequence data. Advances in high-throughput sequencing have greatly expanded our understanding of these environments and their biotechnological potential. However, further progress requires large-scale, integrative meta-analyses across diverse datasets. To address this need, we developed the Produced Water-DNA (PW-DNA) Database, a curated, publicly available resource that consolidates microbial DNA/RNA sequences, geochemical data, and relevant metadata from in situ hydrocarbon environments such as coal beds, oil reservoirs, and natural gas systems. The PW-DNA database delivers three core benefits to the research community: (1) it improves data sharing by linking environmental microbial datasets with corresponding geochemical parameters, enabling more robust filtering and analysis; (2) it connects with complementary research databases to promote broader dissemination and interoperability; and (3) it supports technological innovation by serving as a resource for identifying microbial trends and exploring genetic potential. While individual studies have highlighted basin-specific microbial communities and functional redundancy in biogeochemical cycling, a comprehensive, system-wide perspective is needed to better understand connectivity and novelty across subsurface ecosystems. By designing the PW-DNA in the KBase platform, we provide a reproducible, visual framework for integrating large-scale genomic and geochemical data, enabling researchers to perform more informed analyses and experimental design. Ultimately, this resource enhances the ability to identify, characterize, and interpret microbial functions across diverse subsurface environments, thereby accelerating discovery in subsurface microbiology and biotechnology.

59 BASIC BIOLOGICAL SCIENCES↗

A new dynamical atmospheric ionizing radiation (AIR) model for epidemiological studies

A new Atmospheric Ionizing Radiation (AIR) model is currently being developed for use in radiation dose evaluation in epidemiological studies targeted to atmospheric flight personnel such as civilian airlines crewmembers. The model will allow computing values for biologically relevant parameters, e.g. dose equivalent and effective dose, for individual flights from 1945. Each flight is described by its actual three dimensional flight profile, i.e. geographic coordinates and altitudes varying with time. Solar modulated primary particles are filtered with a new analytical fully angular dependent geomagnetic cut off rigidity model, as a function of latitude, longitude, arrival direction, altitude and time. The particle transport results have been obtained with a technique based on the three-dimensional Monte Carlo transport code FLUKA, with a special procedure to deal with HZE particles. Particle fluxes are transformed into dose-related quantities and then integrated all along the flight path to obtain the overall flight dose. Preliminary validations of the particle transport technique using data from the AIR Project ER-2 flight campaign of measurements are encouraging. Future efforts will deal with modeling of the effects of the aircraft structure as well as inclusion of solar particle events. Published by Elsevier Ltd on behalf of COSPAR.

Aviation↗

A functional microbiome catalogue crowdsourced from North American rivers

Predicting elemental cycles and maintaining water quality under increasing anthropogenic influence requires knowledge of the spatial drivers of river microbiomes. However, understanding of the core microbial processes governing river biogeochemistry is hindered by a lack of genome-resolved functional insights and sampling across multiple rivers. Here we used a community science effort to accelerate the sampling, sequencing and genome-resolved analyses of river microbiomes to create the Genome Resolved Open Watersheds database (GROWdb). GROWdb profiles the identity, distribution, function and expression of microbial genomes across river surface waters covering 90% of United States watersheds. Specifically, GROWdb encompasses microbial lineages from 27 phyla, including novel members from 10 families and 128 genera, and defines the core river microbiome at the genome level. GROWdb analyses coupled to extensive geospatial information reveals local and regional drivers of microbial community structuring, while also presenting foundational hypotheses about ecosystem function. Building on the previously conceived River Continuum Concept, we layer on microbial functional trait expression, which suggests that the structure and function of river microbiomes is predictable. We make GROWdb available through various collaborative cyberinfrastructures, so that it can be widely accessed across disciplines for watershed predictive modelling and microbiome-based management practices.

59 BASIC BIOLOGICAL SCIENCES↗

Genomic fingerprints of the world’s soil ecosystems

Despite the explosion of soil metagenomic data, we lack a synthesized understanding of patterns in the distribution and functions of soil microorganisms. These patterns are critical to predictions of soil microbiome responses to climate change and resulting feedbacks that regulate greenhouse gas release from soils. To address this gap, we assay 1,512 manually curated soil metagenomes using complementary annotation databases, read-based taxonomy, and machine learning to extract multidimensional genomic fingerprints of global soil microbiomes. Our objective is to uncover novel biogeographical patterns of soil microbiomes across environmental factors and ecological biomes with high molecular resolution. We reveal shifts in the potential for (i) microbial nutrient acquisition across pH gradients; (ii) stress-, transport-, and redox-based processes across changes in soil bulk density; and (iii) greenhouse gas emissions across biomes. We also use an unsupervised approach to reveal a collection of soils with distinct genomic signatures, characterized by coordinated changes in soil organic carbon, nitrogen, and cation exchange capacity and in bulk density and clay content that may ultimately reflect soil environments with high microbial activity. Genomic fingerprints for these soils highlight the importance of resource scavenging, plant-microbe interactions, fungi, and heterotrophic metabolisms. Across all analyses, we observed phylogenetic coherence in soil microbiomes—more closely related microorganisms tended to move congruently in response to soil factors. Collectively, the genomic fingerprints uncovered here present a basis for global patterns in the microbial mechanisms underlying soil biogeochemistry and help beget tractable microbial reaction networks for incorporation into process-based models of soil carbon and nutrient cycling.

59 BASIC BIOLOGICAL SCIENCES↗

Speeding genomic island discovery through systematic design of reference database composition

Background Genomic islands (GIs) are mobile genetic elements that integrate site-specifically into bacterial chromosomes, bearing genes that affect phenotypes such as pathogenicity and metabolism. GIs typically occur sporadically among related bacterial strains, enabling comparative genomic approaches to GI identification. For a candidate GI in a query genome, the number of reference genomes with a precise deletion of the GI serves as a support value for the GI. Our comparative software for GI identification was slowed by our original use of large reference genome databases (DBs). Here we explore smaller species-focused DBs. Results With increasing DB size, recovery of our reliable prophage GI calls reached a plateau, while recovery of less reliable GI calls (FPs) increased rapidly as DB sizes exceeded ~500 genomes; i.e., overlarge DBs can increase FP rates. Paradoxically, relative to prophages, FPs were both more frequently supported only by genomes outside the species and more frequently supported only by genomes inside the species; this may be due to their generally lower support values. Setting a DB size limit for our SMA ll R anked T ailored (SMART) DB design speeded runtime ~65-fold. Strictly intra-species DBs would tend to lower yields of prophages for small species (with few genomes available); simulations with large species showed that this could be partially overcome by reaching outside the species to closely related taxa, without an FP burden. Employing such taxonomic outreach in DB design generated redundancy in the DB set; as few as 2984 DBs were needed to cover all 47894 prokaryotic species. Conclusions Runtime decreased dramatically with SMART DB design, with only minor losses of prophages. We also describe potential utility in other comparative genomics projects.

59 BASIC BIOLOGICAL SCIENCES↗

A metagenomic perspective on the microbial prokaryotic genome census

Following 30 years of sequencing, we assessed the phylogenetic diversity (PD) of >1.5 million microbial genomes in public databases, including metagenome-assembled genomes (MAGs) of uncultivated microbes. As compared to the vast diversity uncovered by metagenomic sequences, cultivated taxa account for a modest portion of the overall diversity, 9.73% in bacteria and 6.55% in archaea, while MAGs contribute 48.54% and 57.05%, respectively. Therefore, a substantial fraction of bacterial (41.73%) and archaeal PD (36.39%) still lacks any genomic representation. This unrepresented diversity manifests primarily at lower taxonomic ranks, exemplified by 134,966 species identified in 18,087 metagenomic samples. Our study exposes diversity hotspots in freshwater, marine subsurface, sediment, soil, and other environments, whereas human samples yielded minimal novelty within the context of existing datasets. These results offer a roadmap for future genome recovery efforts, delineating uncaptured taxa in underexplored environments and underscoring the necessity for renewed isolation and sequencing.

59 BASIC BIOLOGICAL SCIENCES↗