Search NASASearch

SEARCH · Search NASA

Results for “Biological databases”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 73 records · Page 4

Metagenome-assembled genomes from East River floodplain sediments near Crested Butte, CO, USA (June to September 2017)

Microorganisms play a key role in cycling nutrients and contaminants in the terrestrial environment depending on their genetic potential. Here, we present metagenome-assembled genomes (MAGs) for the bacterial and archaeal community in floodplain sediment samples taken in 2017 in June (flooded conditions) and September (drained conditions) at two locations (MCB1 and MCB3) in an active meander (Meander C) of the East River. Sediment cores were collected from 2 depths, a near-surface, generally unsaturated depth (15-40 centimeter (cm) depth below surface) and a deeper depth influenced by flooding with redoximorphic features (50-88 cm depth below surface). Sediments were homogenized from the ~10 cm cores for microbial analyses. A total of 24 metagenomes were sequenced through the Joint genome institute (JGI) corresponding to 8 samples sequenced in triplicate. These metagenomes can be found under Genomes Online Database (GOLD) sequencing project: Gs0151851. Metagenomes were assembled, binned, and refined using metawrap to generate MAGs (>50% complete and < 10% contamination based on checkM scores). This dataset includes a zip file of 405 MAG fasta files and a csv file with quality, taxonomic classification (Genome Taxonomy Database Release RS220), and metagenome accessions for MAGs. This dataset also includes a file-level metadata (flmd.csv) file that lists each file contained in the dataset with associated metadata and a data dictionary (dd.csv) file that contains column/row headers used throughout the files along with a definition, units, and data type.This work was supported by the Watershed Function Science Focus Area at Lawrence Berkeley National Laboratory funded by the US Department of Energy, Office of Science, Biological and Environmental Research under Contract No. DE-AC02-05CH11231. Part of this work was performed at SLAC Accelerator Laboratory funded by the US Department of Energy, Office of Science, Biological and Environmental Research under Contract No. DE-AC02-76SF00515.

54 ENVIRONMENTAL SCIENCES

Metagenome-assembled genomes from East River floodplain sediments near Crested Butte, CO, USA (May to September 2018)

Microorganisms play a key role in cycling nutrients and contaminants in the terrestrial environment depending on their genetic potential. Here, we present metagenome-assembled genomes (MAGs) for the bacterial and archaeal community in floodplain sediment samples taken in 2018 in May (flooded conditions) and September (drained conditions) at two locations (MCB1 and MCB3) near the Meander C/Pumphouse floodplain sites of the East River. Sediment cores were collected from 2 depths, a near-surface, generally unsaturated depth (30-40 centimeter (cm) depth below surface) and a deeper depth influenced by flooding with redoximorphic features (70-80 cm depth below surface). Sediments were homogenized from the 10 cm core for microbial analyses. A total of 24 metagenomes were sequenced through the Joint genome institute (JGI) corresponding to 8 samples sequenced in triplicate. These metagenomes can be found under Genomes Online Database (GOLD) sequencing project: Gs0141020. Metagenomes were assembled, binned, and refined using metawrap to generate MAGs (>50% complete and < 10% contamination based on checkM scores). This dataset includes a zip file of 478 MAG fasta files and a csv file with quality, taxonomic classification (Genome Taxonomy Database Release RS220), and metagenome accessions for MAGs. This dataset also includes a file-level metadata (flmd.csv) file that lists each file contained in the dataset with associated metadata and a data dictionary (dd.csv) file that contains column/row headers used throughout the files along with a definition, units, and data type.This work was supported by the Watershed Function Science Focus Area at Lawrence Berkeley National Laboratory funded by the US Department of Energy, Office of Science, Biological and Environmental Research under Contract No. DE-AC02-05CH11231. Part of this work was performed at SLAC Accelerator Laboratory funded by the US Department of Energy, Office of Science, Biological and Environmental Research under Contract No. DE-AC02-76SF00515.

54 ENVIRONMENTAL SCIENCES

Mondo: integrating disease terminology across communities

Precision medicine aims to enhance diagnosis, treatment, and prognosis by integrating multimodal data at the point of care. However, challenges arise due to the vast number of diseases, differing methods of classification, and conflicting terminological coding systems and practices used to represent molecular definitions of disease. This lack of interoperability artificially constrains the potential for diagnosis, clinical decision support, care outcome analysis, as well as data linkage across research domains to support the development or repurposing of therapeutics. There is a clear and pressing need for a unified system for managing disease entities⁠—including identifiers, synonyms, and definitions. To address these issues, we created the Mondo disease ontology—a community-driven, open-source, unified disease classification system that harmonizes diverse terminologies into a consistent, computable framework. Mondo integrates key medical and biomedical terminologies, including Online Mendelian Inheritance in Man (OMIM), Orphanet, Medical Subject Headings (MeSH), National Cancer Institute Thesaurus (NCIt), and more, to provide a comprehensive and accurate representation of disease concepts with fully provenanced and attributed links back to the sources. Mondo can be used as the handle for curation of gene–disease associations utilized in diagnostic applications, research applications such as computational phenotyping, and in clinical coding systems in clinical decision support by pointing the clinician to the numerous knowledge resources linked to the Mondo identifier. Mondo's community-centric approach, stewarded by the Monarch Initiative's expertise in ontologies, ensures that the ontology remains adaptable to the evolving needs of biomedical research and clinical communities, as well as the knowledge providers.

biomedical informatics

The Pan-Arctic Vegetation Cover (PAVC) database v1.1

The Pan-Arctic Vegetation Cover (PAVC) database contains synthesized field-data observations of vegetation cover from 978 Arctic Alaska plots with observations from 2010 to 2021. The cover datasets contain plot data at both the plant functional type (PFT) and species-level resolution, with standardized PFT definitions and species names. We synthesized publicly available point-intercept and visual estimate plots from the Arctic Vegetation Archive of Alaska, the Alaska Vegetation Plots Database, the North Slope Science Catalog, and the National Ecological Observatory Network; as well as previously unpublished data from the Next-Generation Ecosystem Experiments: Arctic (NGEE Arctic).Users will find four synthesized datasets, 4 associated data descriptor (dd) files, and 1 metadata file in the PAVC database:synthesized_species_fcover.csv contains fractional cover (fcover) for unique accepted species names, where names include vegetation identified at the family, genus, species, subspecies, and variety levels, as well as general functional types across all 5 data sources. The synthesized_species_fcover_dd.csv accompanies this dataset with header information.synthesized_pft_fcover.csv contains fcover for the following PFTs: non-vascular plants with lichen and bryophyte subcategories, trees with deciduous and evergreen subcategories, shrubs with deciduous and evergreen subcategories, graminoids (grasses), and forbs (herbaceous flowering plants) measured as total cover. Litter and “other” cover are also included as total cover. Additional “types” include water and bare ground, which were measured as top cover. The synthesized_pft_fcover_dd.csv accompanies this dataset with header information.species_pft_checklist.csv is a lookup table containing the translation from a dataset species name to an accepted species name and to a PFT. This table can be used to clarify our species to PFT adjudications, and to aid users in assigning their own PFTs. Any issues found in this checklist should be reported in the Issues tab of our github.survey_unit_information.csv contains auxiliary information about the plots synthesized in this database. It contains useful information for filtering plots of interest based on temporal, geospatial, and contextual information about the plot surveys.flmd.csv contains metadata information about each file in the database.This research was performed as a part of the NGEE Arctic project. The NGEE Arctic project was a research effort to reduce uncertainty in Earth System Models by developing a predictive understanding of carbon-rich Arctic ecosystems and feedbacks to climate. NGEE Arctic was supported by the Department of Energy's Office of Biological and Environmental Research.The NGEE Arctic project had two field research sites: 1) located within the Arctic polygonal tundra coastal region on the Barrow Environmental Observatory (BEO) and the North Slope near Utqiagvik (Barrow), Alaska and 2) multiple areas on the discontinuous permafrost region of the Seward Peninsula north of Nome, Alaska.Through observations, experiments, and synthesis with existing datasets, NGEE Arctic provided an enhanced knowledge base for multi-scale modeling and contributed to improved process representation at global pan-Arctic scales within the Department of Energy's Earth system Model (the Energy Exascale Earth System Model, or E3SM), and specifically within the E3SM Land Model component (ELM).

54 ENVIRONMENTAL SCIENCES

Gender in Mineral Names

Minerals are the fundamental constituents of Earth, and mineral names appear in scientific literature for disciplines including geology, chemistry, materials science, biology, and medicine, among others. Choosing a name is the full responsibility of the authors of new mineral proposals submitted to the International Mineralogical Association (IMA). Scientific nomenclature and its traditions have evolved over time and, consequently, mineral names track changes in the landscape of mineralogy with respect to language, technology, and culture. To evaluate these changes, the namesake information for all 5896 minerals approved by the IMA or ‘grandfathered’ into use as of December 2022 was recorded and categorized within a workable database. The compiled information yields diverse insights into the intersection of science and culture and could also be used to project future trends. In this study, we used the name database to investigate gender diversity among mineral eponyms. More than half (c. 54%) of all mineral species are named after people, the identities of whom are largely a reflection of the people that have historically been involved, in one way or another, in the geosciences and in the mining industry. Of the 2738 people with minerals named for them, approximately 6.1% are (interpreted to be) women. Nearly all minerals named for women were named during the last sixty years, although the rate of growth in the year-on-year percentage of women among new mineral namesakes has slowed since about 1985. If current and historical trends hold, our model predicts that women will not comprise more than about 10.35% of newly established mineral namesakes in future years. The representation of women among mineral namesakes also differs starkly among countries. For example, Russians comprise 43.11% of women with minerals named for them, but account for only 15.12% of all eponyms. However, there are additional disparities beyond the proportions of namesakes. For scientists who were alive when a mineral was named for them, women were an average of 3.74 years older than men when evaluated over the same timespan (1954–2022). These results demonstrate that gender-based disparities are imprinted into current mineral nomenclature and indicate that gender parity among new mineral namesakes is impossible without unprecedented changes in the upstream demographics that are most likely to affect naming trends.

58 GEOSCIENCES

FatPlants: a comprehensive information system for lipid-related genes and metabolic pathways in plants

Abstract FatPlants, an open-access, web-based database, consolidates data, annotations, analysis results, and visualizations of lipid-related genes, proteins, and metabolic pathways in plants. Serving as a minable resource, FatPlants offers a user-friendly interface for facilitating studies into the regulation of plant lipid metabolism and supporting breeding efforts aimed at increasing crop oil content. This web resource, developed using data derived from our own research, curated from public resources, and gleaned from academic literature, comprises information on known fatty-acid-related proteins, genes, and pathways in multiple plants, with an emphasis on Glycine max, Arabidopsis thaliana, and Camelina sativa. Furthermore, the platform includes machine-learning based methods and navigation tools designed to aid in characterizing metabolic pathways and protein interactions. Comprehensive gene and protein information cards, a Basic Local Alignment Search Tool search function, similar structure search capacities from AphaFold, and ChatGPT-based query for protein information are additional features. Database URL: https://www.fatplants.net/

59 BASIC BIOLOGICAL SCIENCES

Drought shifts dissolved organic matter sources from above- to belowground and stress-induced processes in Amazon white-sand forests

White-sand forests contribute significantly to dissolved organic matter (DOM) production in the central Amazon, forming blackwater rivers that dominate organic matter export from the Amazon basin to the ocean. Despite their importance in controlling DOM export, white-sand forests are understudied, and it remains unclear whether systematic changes in the formation of blackwater DOM occur and how seasonal variations and extremes like El Niño-associated droughts impact them. We collected soil porewater from two central Amazon white-sand forests for two years, spanning a wet La Niña year followed by an El Niño drought year. The molecular composition of DOM was analyzed using high-resolution mass spectrometry, and correlation network analysis was employed to identify ecologically meaningful DOM subsets. Using additional chemical characterization, database annotations, correlation with 14C-age of DOM and climatic variables, and ecological null modeling, we propose five distinct DOM sources: plant litter and throughfall, soil organic matter (SOM) decomposition, root exudation, and two drought response subsets of likely microbial and plant origin. During drought conditions, aboveground plant-derived compounds decreased, while SOM products, root exudates, and drought response compounds increased. These drought responses were qualitatively similar in both years but notably amplified in the drier El Niño year. Drought amplified deterministic control over DOM composition, indicating that DOM reflected directed biological responses and that future droughts are likely to generate similar shifts. Overall, drought substantially altered belowground carbon cycling by shifting DOM sources and inducing stress responses, effects expected to recur and potentially intensify under future climate scenarios.

Lange, Dan F.

Montane Conifer, Aspen, Meadow, and Sagebrush Metagenome Resolved Genomes and Traits in East River Watershed, Colorado, USA

Climate change is driving vegetation shifts in mountain watersheds, with unknown impacts on biogeochemical cycles. We hypothesize that these shifts will reshape soil microbiomes and associated biogeochemical processes. As a part of Lawrence Berkeley National Laboratory (LBNL) Watershed Science Focus Area (SFA), we assessed microbiome and microbial functional trait differences between soils under conifer, aspen, forby meadows, and sagebrush across the East River Watershed, CO, controlling for elevation and aspect.Here we present metagenome assembled genomes (MAGs) for the bacterial and archaeal communities from soils 0-20cm in depth across three locations in the watershed—Headwaters, Upper Reaches, and Lower Reaches from August 3-11th 2016. Each location was further subdivided into two blocks, with one block on a west facing aspect, and two on the east aspect of the valley. Within blocks, two samples per vegetation type were taken (one at each depth). This resulted in 66 samples, which were sequenced at JGI and can be found under the Joint Genome Institute (JGI) Genomes Online Database (GOLD) sequencing project Gs0118068. Metagenomes were assembled through an inhouse pipeline (see methods), binned using four autobinners (concoct, maxbin2, metabat2, and vamb) and consolidated using dastool. The consolidated bins from all metagenomes were pooled, filtered by completeness (>75%) and contamination (<25%), and dereplicated at 95% ANI using drep. The dataset includes a zip file of 687 genomes (Vegtype_MAGS.zip), the accession numbers for the underlying metagenomes, a csv file with MAG quality metrics and taxonomy from Genome Taxonomy Database (GTDB) and National Center for Biotechnology Information (NCBI) taxonomic representative genome proteins (EastRiver_Vegtype_drep_genome_info.csv), and a file containing MAG quality metrics and taxonomy (gtdb_drep_bin_taxonomy.csv). The dataset additionally includes a sample metadata file (EastRiver_Vegtype_sample_metadata.csv), a metadata file used to register associated samples with IGSNs (International Generic Sample Numbers) (samples.csv), a Google KML file for the sampled locations (sample_collection_sites.kml), a location metadata file (locations.csv), a file-level metadata file (flmd.csv), and a data dictionary (dd.csv) file.This work was supported by the Watershed Function Science Focus Area at Lawrence Berkeley National Laboratory funded by the US Department of Energy, Office of Science, Biological and Environmental Research under Contract No. DE-AC02-05CH11231.

54 ENVIRONMENTAL SCIENCES

Outcomes of PAX sapiens-Supported Global Wildlife Data Sharing Conferences for Enhanced One Health Security (GWDSC)

Across two consecutive Global Wildlife Data Sharing Conferences supported by PAX sapiens—Year 1 (May 2024) at Pacific Northwest National Laboratory and Year 2 (2025) in Ciudad Real, Spain—the initiative converted wildlife data sharing from aspiration into operational reality, producing measurable impacts in platform development, data mobilization, standards harmonization, and international partnership formation. The conferences addressed a critical gap in global health security: while 75% of emerging infectious diseases affect both humans and animals and over 60% originate in wildlife, wildlife health surveillance has historically lagged behind human and agricultural sectors due to fragmented databases, inconsistent terminology, uneven capacity, and limited cross-border coordination. By convening practitioners, government agencies, international organizations, academic institutions, and NGOs, the GWDSC catalyzed trust-based relationships and practical workflows that enable earlier detection, better risk assessment, and more effective prevention of threats at the wildlife–domestic animal–human–environment interface.

54 ENVIRONMENTAL SCIENCES

Gaia: An AI-enabled genomic context–aware platform for protein sequence annotation

Protein sequence similarity search is fundamental to biology research, but current methods are typically not able to consider crucial genomic context information indicative of protein function, especially in microbial systems. Here, we present Gaia (Genomic AI Annotator), a sequence annotation platform that enables rapid, context-aware protein sequence search across genomic datasets. Gaia leverages gLM2, a mixed-modality genomic language model trained on both amino acid sequences and their genomic neighborhoods to generate embeddings that integrate sequence-structure-context information. This approach allows for the identification of functionally and/or evolutionarily related genes that are found in conserved genomic contexts, which may be missed by traditional sequence- or structure-based search alone. Gaia enables real-time search of a curated database comprising more than 85 million protein clusters from 131,744 microbial genomes. We compare the homolog retrieval performance of Gaia search against other embedding and alignment-based approaches. We provide Gaia as a web-based, freely available tool.

Jha, Nishant

Missing microbial eukaryotes and misleading meta-omic conclusions

Meta-omics is commonly used for large-scale analyses of microbial eukaryotes, including species or taxonomic group distribution mapping, gene catalog construction, and inference on the functional roles and activities of microbial eukaryotes in situ. Here, we explore the potential pitfalls of common approaches to taxonomic annotation of protistan meta-omic datasets. We re-analyze three environmental datasets at three levels of taxonomic hierarchy in order to illustrate the crucial importance of database completeness and curation in enabling accurate environmental interpretation. We show that taxonomic membership of sequence clusters estimates community composition more accurately than returning exact sequence labels, and overlap between clusters can address database shortcomings. Clustering approaches can be applied to diverse environments while continuing to exploit the wealth of annotation data collated in databases, and selecting and evaluating these databases is a critical part of correctly annotating protistan taxonomy in environmental datasets. We argue that ongoing curation of genetic resources is crucial in accurately annotating protists in in situ meta-omic datasets. Moreover, we propose that precise taxonomic annotation of meta-omic data is a clustering problem rather than a feasible alignment problem.

59 BASIC BIOLOGICAL SCIENCES

Poplar: a phylogenomics pipeline

Motivation Generating phylogenomic trees from the genomic data is essential in understanding biological systems. Each step of this complex process has received extensive attention and has been significantly streamlined over the years. Given the public availability of data, obtaining genomes for a wide selection of species is straightforward. However, analyzing that data to generate a phylogenomic tree is a multistep process with legitimate scientific and technical challenges, often requiring a significant input from a domain-area scientist. Results We present Poplar, a new, streamlined computational pipeline, to address the computational logistical issues that arise when constructing the phylogenomic trees. It provides a framework that runs state-of-the-art software for essential steps in the phylogenomic pipeline, beginning from a genome with or without an annotation, and resulting in a species tree. Running Poplar requires no external databases. In the execution, it enables parallelism for execution for clusters and cloud computing. The trees generated by Poplar match closely with state-of-the-art published trees. The usage and performance of Poplar is far simpler and quicker than manually running a phylogenomic pipeline. Availability and implementation Freely available on GitHub at https://github.com/sandialabs/poplar. Implemented using Python and supported on Linux.

Koning, Elizabeth [Sandia National Laboratories (S

SLAB: simultaneous labeling and binding affinity prediction for protein–ligand structures

Machine learning models are often used as scoring functions to predict the binding affinity of a protein–ligand complex. These models are trained with limited amounts of data with experimentally measured binding affinity values. A large number of compounds are labeled inactive through single-concentration screens without measuring binding affinities. These inactive compounds, along with the active ones, can be used to train binary classification models, while regression models are trained using compounds with binding affinities only. However, the classification and regression tasks are often handled separately, without sharing the learned feature representations. In this paper, we propose a novel model architecture that jointly performs regression and classification objectives, aiming to maximize data utilization and improve predictive performance by leveraging two complementary tasks. In our setup, the regression yields the binding affinity, whereas the classification task yields the label as active or inactive. We demonstrate our method using PDBbind, the standard 3D structure database, as well as a dataset of flavivirus protease compounds with binding affinity data. Our experiments show that the new joint training strategy improves the accuracy of the model, increasing applicability in various practical drug screening scenarios.

Biological and medical sciences

From soil to sequence: filling the critical gap in genome-resolved metagenomics is essential to the future of soil microbial ecology

Abstract Soil microbiomes are heterogeneous, complex microbial communities. Metagenomic analysis is generating vast amounts of data, creating immense challenges in sequence assembly and analysis. Although advances in technology have resulted in the ability to easily collect large amounts of sequence data, soil samples containing thousands of unique taxa are often poorly characterized. These challenges reduce the usefulness of genome-resolved metagenomic (GRM) analysis seen in other fields of microbiology, such as the creation of high quality metagenomic assembled genomes and the adoption of genome scale modeling approaches. The absence of these resources restricts the scale of future research, limiting hypothesis generation and the predictive modeling of microbial communities. Creating publicly available databases of soil MAGs, similar to databases produced for other microbiomes, has the potential to transform scientific insights about soil microbiomes without requiring the computational resources and domain expertise for assembly and binning.

59 BASIC BIOLOGICAL SCIENCES

Metagenome-assembled genomes from topsoils along a hillslope water gradient across early snowmelt to late summer in East River, CO

Drought is changing the American Mountain West at unprecedented rates with unknown consequences to soil microbiome composition and function. As a part of LBNL Watershed Science Focus Area (SFA), we investigated shifts in microbial community and transcriptional activity on a subalpine conifer-meadow transition zone throughout the summer of 2023 as soil dried down. This work took place in Crested Butte, CO on Snodgrass mountain, using a proxy for drought conditions.Here we present metagenome assembled genomes (MAGs) for the bacterial and archaeal community at 0-10cm from three sites along a hillslope water gradient across five timepoints from early snowmelt to late summer. 42 metagenomes were sequenced at Joint Genome Institute (JGI) and can be found under the JGI GOLD (Genomes Online Database) sequencing project Gs0166660. Metagenomes were assembled through an inhouse pipeline (see methods), binned using four autobinners (concoct, maxbin2, metabat2, and vamb) and consolidated using dastool. The consolidated bins from all metagenomes were pooled, filtered by completeness (>70%) and contamination (<10%), and dereplicated at 95% ANI using drep. This dataset (1) a zip file of 157 MAGs (as fasta files, Gs0166660_bins_tar.gz), (2) sample metadata file with sample IGSNs (International Generic Sample Numbers) (samples.csv), (3) bounding box coordinates for the sampled locations (Gs0166660.kml), (4) metagenome assembly and coassembly metadata file listing IMG/M (Integrated Microbial Genomes/Metagenomes) metagenome accessions linking samples to metagenomes (EastRiver_Drought_ESSDive_Metadata.csv), (5) location metadata file (locations.csv), (6) file-level metadata file (flmd.csv) and (7) data dictionary (dd.csv) file.This work was supported by the Watershed Function Science Focus Area at Lawrence Berkeley National Laboratory funded by the US Department of Energy, Office of Science, Biological and Environmental Research under Contract No. DE-AC02-05CH11231.

54 ENVIRONMENTAL SCIENCES

Interactive tools for functional annotation of bacterial genomes

Automated annotations of protein functions are error-prone because of our lack of knowledge of protein functions. For example, it is often impossible to predict the correct substrate for an enzyme or a transporter. Furthermore, much of the knowledge that we do have about the functions of proteins is missing from the underlying databases. We discuss how to use interactive tools to quickly find different kinds of information relevant to a protein’s function. Many of these tools are available via PaperBLAST (http://papers.genomics.lbl.gov). Combining these tools often allows us to infer a protein’s function. Ideally, accurate annotations would allow us to predict a bacterium’s capabilities from its genome sequence, but in practice, this remains challenging. We describe interactive tools that infer potential capabilities from a genome sequence or that search a genome to find proteins that might perform a specific function of interest.

59 BASIC BIOLOGICAL SCIENCES

Bactericidal effectors of the Stenotrophomonas maltophilia type IV secretion system: functional definition of the nuclease TfdA and structural determination of TfcB

ABSTRACT Stenotrophomonas maltophilia expresses a type IV protein secretion system (T4SS) that promotes contact-dependent killing of other bacteria and does so partly by secreting the effector TfcB. Here, we report the structure of TfcB, comprising an N-terminal domain similar to the catalytic domain of glycosyl hydrolase (GH-19) chitinases and a C-terminal domain for recognition and translocation by the T4SS. Utilizing a two-hybrid assay to measure effector interactions with the T4SS coupling protein VirD4, we documented the existence of five more T4SS substrates. One of these was protein 20845, an annotated nuclease. A S. maltophilia mutant lacking the gene for 20845 was impaired for killing Escherichia coli , Klebsiella pneumoniae , and Pseudomonas aeruginosa . Moreover, the cloned 20845 gene conferred robust toxicity, with the recombinant E. coli being rescued when 20845 was co-expressed with its cognate immunity protein. The 20845 effector was an 899 amino-acid protein, comprised of a GHH-nuclease domain in its N-terminus, a large central region of indeterminant function, and a C-terminus for secretion. Engineered variants of the 20845 gene that had mutations in the predicted catalytic site did not impede E. coli , indicating that the antibacterial effect of 20845 involves its nuclease activity. Using flow cytometry with DNA staining, we determined that 20845, but not its mutant variants, confers a loss in DNA content of target bacteria. Database searches revealed that uncharacterized homologs of 20845 occur within a range of bacteria. These data indicate that the S. maltophilia T4SS promotes interbacterial competition through the action of multiple toxic effectors, including a potent, novel DNase. IMPORTANCE Stenotrophomonas maltophilia is a multi-drug-resistant, Gram-negative bacterium that is an emerging pathogen of humans. Patients with cystic fibrosis are particularly susceptible to S. maltophilia infection. In hospital water systems and various types of infections, S. maltophilia co-exists with other bacteria, including other pathogens such as Pseudomonas aeruginosa . We previously demonstrated that S. maltophilia has a functional VirB/D4 type VI protein secretion system (T4SS) that promotes contact-dependent killing of other bacteria. Since most work on antibacterial systems involves the type VI secretion system, this observation remains noteworthy. Moreover, S. maltophilia currently stands alone as a model for a human pathogen expressing an antibacterial T4SS. Using biochemical, genetic, and cell biological approaches, we now report both the discovery of a novel antibacterial nuclease (TfdA) and the first structural determination of a bactericidal T4SS effector (TfcB).

59 BASIC BIOLOGICAL SCIENCES

Unveiling the microbial realm with VEBA 2.0: a modular bioinformatics suite for end-to-end genome-resolved prokaryotic, (micro)eukaryotic and viral multi-omics from either short- or long-read sequencing

Abstract The microbiome is a complex community of microorganisms, encompassing prokaryotic (bacterial and archaeal), eukaryotic, and viral entities. This microbial ensemble plays a pivotal role in influencing the health and productivity of diverse ecosystems while shaping the web of life. However, many software suites developed to study microbiomes analyze only the prokaryotic community and provide limited to no support for viruses and microeukaryotes. Previously, we introduced the Viral Eukaryotic Bacterial Archaeal (VEBA) open-source software suite to address this critical gap in microbiome research by extending genome-resolved analysis beyond prokaryotes to encompass the understudied realms of eukaryotes and viruses. Here we present VEBA 2.0 with key updates including a comprehensive clustered microeukaryotic protein database, rapid genome/protein-level clustering, bioprospecting, non-coding/organelle gene modeling, genome-resolved taxonomic/pathway profiling, long-read support, and containerization. We demonstrate VEBA’s versatile application through the analysis of diverse case studies including marine water, Siberian permafrost, and white-tailed deer lung tissues with the latter showcasing how to identify integrated viruses. VEBA represents a crucial advancement in microbiome research, offering a powerful and accessible software suite that bridges the gap between genomics and biotechnological solutions.

59 BASIC BIOLOGICAL SCIENCES