Search NASASearch

SEARCH · Search NASA

Results for “Sequence Function Data”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 73 records · Page 4

Discovery of additional ancient genome duplications in yeasts

Whole-genome duplication (WGD) has had profound macroevolutionary impacts on diverse lineages, preceding adaptive radiations in vertebrates, teleost fish, and angiosperms. In contrast to the many known ancient WGDs in animals, and especially plants, we are aware of evidence for only four WGDs in fungi. The oldest of these occurred ∼100 million years ago (mya) and is shared by ∼60 extant Saccharomycetales species, including the baker’s yeast Saccharomyces cerevisiae. Notably, this is the only known ancient WGD event in the yeast subphylum Saccharomycotina. The dearth of ancient WGD events in fungi remains a mystery. Some studies have suggested that fungal lineages that experience chromosome and genome duplication quickly go extinct, leaving no trace in the genomic record, while others contend that the lack of known WGDs is due to an absence of data. Under the second hypothesis, additional sampling and deeper sequencing of fungal genomes should lead to the discovery of more WGD events. Coupling hundreds of recently published genomes from nearly every described Saccharomycotina species, with three additional long-read assemblies, we discovered three novel WGD events. Although the functions of retained duplicate genes originating from these events are broad, they bear similarities to the well-known WGD that occurred in the Saccharomycetales. In conclusion, our results suggest that WGD may be a more common evolutionary force in fungi than previously believed.

convergent evolution

GenomeDepot: data management system for microbial comparative genomics

Summary GenomeDepot is an open-source web-based platform for annotation, management, and comparative analysis of microbial genomic sequences and associated data including ortholog families, protein domains, operons, regulatory interactions, strain taxonomy, and sample metadata. GenomeDepot supports rapid creation of websites for user-defined genome collections that include bioinformatic tools for interactive genome browsing, Basic Local Alignment Search Tool (BLAST) search, annotation search, comparative genomic neighborhood visualization, and sequence download. Gene function annotations are generated by a customizable annotation pipeline. The pipeline runs annotation tools in Conda environments and can be easily extended with additional user-specified tools. Availability and implementation GenomeDepot is open source and distributed under the GNU General Public License via GitHub (https://github.com/aekazakov/genome-depot). GenomeDepot is implemented in Python and was tested in Ubuntu Linux. Full installation instructions and documentation are available at https://aekazakov.github.io/genome-depot/. GenomeDepot demo server is freely accessible at https://iseq.lbl.gov/demogd/.

Kazakov, Alexey [Lawrence Berkeley National Labora

Depth-resolved sagebrush root metabolomics, rhizosphere microbial communities, and geochemistry at the East River Watershed

This data set consists of results from soil nutrient profile, untargeted metabolomics, mass spec imaging, and amplicon sequencing. Data for soil nutrient profile includes common cations (Ca, Mg, Na, and K etc.) extracted from 3 digesting steps – ammonia acetate (for exchangeable cations), nitric acid (for acid dissolved fraction), and hydrofluoric acid/perchloric acid (HF/HClO4) for whole soil digestion. It also includes concentration of organic carbon, inorganic nitrogen (ammonia and nitrate) and phosphorus (Bray-1 P and nitric acid extract), and total nitrogen and phosphorus. Data for untargeted metabolomics includes metabolomic profile for root exudate/tissues and soil extracts from depths at surface soil to saprolite, that were measured using gas chromatography – mass spectrometry (GC-MS), and liquid chromatography – tandem mass spectrometry (LC-MS/MS). Data for mass spec imaging includes spatial distribution of metabolites that were detected and annotated with Fourier transformation ion cyclotron resonance mass spectrometer (FTICR-MS). Data for amplicon sequencing includes the base paired 16S and ITS ribosomal RNA sequences from Miseq Illumina sequencing. All samples were collected from 2 sampling campaign October 2022 and June 2023. Collectively, these datasets enable a mechanistic evaluation of how nutrient acquisition, especially nitrogen and phosphorus, differs between shallow roots operating in soil and deep roots functioning within the fractured bedrock zone. All files are provided as comma-separated values (CSV) fies (.csv) and (GZIP) file (.gz). The compressed .gz FASTQ files can be read directly in R using the dada2 package as part of the amplicon sequence analysis workflow. This work was supported by the Watershed Function Science Focus Area at Lawrence Berkeley National Laboratory funded by the US Department of Energy, Office of Science, Biological and Environmental Research under Contract No. DE-AC02-05CH11231. This research was performed on a project award 60563 (https://dx.doi.org/10.46936/expl.proj.2022.60563/60008727) from the Environmental Molecular Sciences Laboratory, a DOE Office of Science User Facility sponsored by the Biological and Environmental Research program under Contract No. DE-AC05-76RL01830.

EARTH SCIENCE > AGRICULTURE > SOILS > CARBON

A map of the rubisco biochemical landscape

Rubisco is the primary CO 2 -fixing enzyme of the biosphere, yet it has slow kinetics. The roles of evolution and chemical mechanism in constraining its biochemical function remain debated. Engineering efforts aimed at adjusting the biochemical parameters of rubisco have largely failed, although recent results indicate that the functional potential of rubisco has a wider scope than previously known. Here we developed a massively parallel assay, using an engineered Escherichia coli in which enzyme activity is coupled to growth, to systematically map the sequence–function landscape of rubisco. Composite assay of more than 99% of single-amino acid mutants versus CO 2 concentration enabled inference of enzyme velocity and apparent CO 2 affinity parameters for thousands of substitutions. This approach identified many highly conserved positions that tolerate mutation and rare mutations that improve CO 2 affinity. These data indicate that non-trivial biochemical changes are readily accessible and that the functional distance between rubiscos from diverse organisms can be traversed, laying the groundwork for further enzyme engineering efforts.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH

Measurements of soil protist richness and community composition are influenced by primer pair, annealing temperature, and bioinformatics choices

ABSTRACT Protists are a diverse and understudied group of microbial eukaryotic organisms especially in terrestrial environments. Advances in molecular methods are increasing our understanding of the distribution and functions of these creatures; however, there is a vast array of choices researchers make including barcoding genes, primer pairs, PCR settings, and bioinformatic options that can impact the outcome of protist community surveys. Here, we tested four commonly used primer pairs targeting the V4 and V9 regions of the 18S rRNA gene using different PCR annealing temperatures and processed the sequences with different bioinformatic parameters in 10 diverse soils to evaluate how primer pair, amplification parameters, and bioinformatic choices influence the composition and richness of protist and non-protist taxa using Illumina sequencing. Our results showed that annealing temperature influenced sequencing depth and protist taxon richness for most primer pairs, and that merging forward and reverse sequencing reads for the V4 primer pairs dramatically reduced the number of sequences and taxon richness of protists. The data sets of primers that targeted the same 18S rRNA gene region (e.g., V4 or V9) had similar protist community compositions; however, data sets from primers targeting the V4 18S rRNA gene region detected a greater number of protist taxa compared to those prepared with primers targeting the V9 18S rRNA region. There was limited overlap of protist taxa between data sets targeting the two different gene regions (80/549 taxa). Together, we show that laboratory and bioinformatic choices can substantially affect the results and conclusions about protist diversity and community composition using metabarcoding. IMPORTANCE Ecosystem functioning is driven by the activity and interactions of the microbial community, in both aquatic and terrestrial environments. Protists are a group of highly diverse, mostly unicellular microbes whose identity and roles in terrestrial ecosystem ecology have been largely ignored until recently. This study highlights the importance of choices researchers make, such as primer pair, on the results and conclusions about protist diversity and community composition in soils. In order to better understand the roles protist taxa play in terrestrial ecosystems, biases in methodological and analytical choices should be understood and acknowledged.

Biotechnology & Applied Microbiology

Development of high throughput and in vitro assays for analyzing RNA modifications

Modifications on RNAs play major roles in their stability, translation, and enzymatic activity. Despite its importance, the current techniques are insufficient to study the structure and function of RNA modifications. Indeed, the National Academies of Science, Engineering and Medicine indicate that developing new tools and further study the function of RNA modifications is strategically a high priority for advancing science in the coming years (https://www.nationalacademies.org/our-work/toward-sequencing-and-mapping-of-rna-modifications). RNA modifications occur in all domains of life controlling processes such as RNA turnover, translation regulation, cellular defenses and bioproduction. Our preliminary data indicated that the insulin mRNA might get ADP-ribosylated by the ADP-ribosyltransferase PARP12. RNA ADP-ribosylation has been described in Escherichia coli. Combined to the fact that ADP-ribosyltransferase (PARP) genes are conserved throughout evolution we hypothesize that this modification might play essential roles in cells. Therefore, we proposed to develop sequencing techniques and in vitro enzymatic assays to identify and validate ADP-ribosylation motifs and sites. Here we report the development of RNA-seq and qPCR assays to identify ADP-ribosylated RNAs, in addition to a nicotinamide adenosine dinucleotide (NAD – ADP-ribosylation donor) consumption assay and an enzyme-linked immunosorbent assay (ELISA) to measure ADP-ribosyltransferase activity. Testing these assays with the insulin mRNA confirmed that this transcript is ADP-ribosylated. These assays will not only enable studying the function of ADP-ribosylation but can be easily adapted for studying other RNA modifications. This will open opportunities to study RNA modifications in different model systems from bacteria to viruses to plants, bringing insights into their cellular functions and the possibility of targeting them for biotechnological applications.

59 BASIC BIOLOGICAL SCIENCES

Time-series metagenomics reveals changing protistan ecology of a temperate dimictic lake

Abstract Background Protists, single-celled eukaryotic organisms, are critical to food web ecology, contributing to primary productivity and connecting small bacteria and archaea to higher trophic levels. Lake Mendota is a large, eutrophic natural lake that is a Long-Term Ecological Research site and among the world’s best-studied freshwater systems. Metagenomic samples have been collected and shotgun sequenced from Lake Mendota for the last 20 years. Here, we analyze this comprehensive time series to infer changes to the structure and function of the protistan community and to hypothesize about their interactions with bacteria. Results Based on small subunit rRNA genes extracted from the metagenomes and metagenome-assembled genomes of microeukaryotes, we identify shifts in the eukaryotic phytoplankton community over time, which we predict to be a consequence of reduced zooplankton grazing pressures after the invasion of a invasive predator (the spiny water flea) to the lake. The metagenomic data also reveal the presence of the spiny water flea and the zebra mussel, a second invasive species to Lake Mendota, prior to their visual identification during routine monitoring. Furthermore, we use species co-occurrence and co-abundance analysis to connect the protistan community with bacterial taxa. Correlation analysis suggests that protists and bacteria may interact or respond similarly to environmental conditions. Cryptophytes declined in the second decade of the timeseries, while many alveolate groups (e.g., ciliates and dinoflagellates) and diatoms increased in abundance, changes that have implications for food web efficiency in Lake Mendota. Conclusions We demonstrate that metagenomic sequence-based community analysis can complement existing efforts to monitor protists in Lake Mendota based on microscopy-based count surveys. We observed patterns of seasonal abundance in microeukaryotes in Lake Mendota that corroborated expectations from other systems, including high abundance of cryptophytes in winter and diatoms in fall and spring, but with much higher resolution than previous surveys. Our study identified long-term changes in the abundance of eukaryotic microbes and provided context for the known establishment of an invasive species that catalyzes a trophic cascade involving protists. Our findings are important for decoding potential long-term consequences of human interventions, including invasive species introduction.

59 BASIC BIOLOGICAL SCIENCES

Machine learning prediction of enzyme optimum pH

The relationship between pH and enzyme catalytic activity, especially the optimal pH (pH opt ) at which enzymes function, is critical for biotechnological applications. Hence, computational methods to predict pH opt will enhance enzyme discovery and design by facilitating accurate identification of enzymes that function optimally at specific pH levels, and by elucidating sequence-function relationships. Here, in this study, we proposed and evaluated various machine learning methods for predicting pH opt , conducting extensive hyperparameter optimization and training over 11,000 model instances. Our results demonstrate that models utilizing language model embeddings markedly outperform other methods in predicting pHopt. We present EpHod, the best-performing model, to predict pHopt, making it publicly available to researchers. From sequence data, EpHod directly learns structural and biophysical features that relate to pH opt , including proximity of residues to the catalytic centre and the accessibility of solvent molecules. Overall, EpHod presents a promising advancement in pH opt prediction and will potentially speed up the development of enzyme technologies.

97 MATHEMATICS AND COMPUTING

PRIME: An evaluation framework for protein representation inference and generalization in viral mutation space

Background Protein language models (PLMs) have revolutionized protein fitness prediction, yet their application to rapidly evolving viral pathogens is often confounded by extreme sequence homology. This homology leads to “data leakage” in standard random validation splits, yielding inflated performance metrics that fail to translate into real-world biosurveillance utility. Results We present Protein Representation Inference for Mutation Evaluation (PRIME), a framework that integrates domain-specific fine-tuning with a rigorous position-stratified validation protocol to evaluate viral threats. Using a dataset of 347,432 SARS-CoV-2 receptor binding domain (RBD) sequences, we demonstrate that while random training data split yields deceptive R 2 values (> 0.90), they fail to generalize to novel mutational sites. By benchmarking models up to 650 M parameters, we show that domain-specific fine-tuning of the ESM-C 600 M model with correctly stratified data provides an initial demonstration of predictive signal for binding affinity and expression at unseen mutational sites of binding affinity and expression on unseen sites (R 2 ~0.23), a significant advancement over base foundation models which exhibit no predictive power (R 2 <0). PRIME’s embedding-based clustering identified 3.03% of bat coronavirus sequences as candidates for further experimental prioritization based on their functional similarity to human-infective strains in embedding space, offering a perspective complementary to traditional phylogenetic methods. Conclusion PRIME establishes a new benchmark for the application of PLMs in pathogen surveillance. Our findings demonstrate that state-of-the-art models and fine-tuning, when paired with stratified validation, provide biologically meaningful insights into pathogen evolution and zoonotic risk.

59 BASIC BIOLOGICAL SCIENCES

Multimodal Approaches for Leveraging Domain Knowledge with State-of-the-Art Machine Learning to Engineer Biocatalysts

This grant aimed to accelerate the development of specialized enzymes—biological catalysts essential for sustainable manufacturing and medicine—by integrating traditional laboratory evolution with cutting-edge artificial intelligence. To achieve this, we developed a suite of high-throughput sequencing tools and a centralized database to bridge the gap between a protein’s genetic "code" and its physical function. By training machine learning models on large datasets, we also demonstrated the ability to move beyond slow, trial-and-error testing to a "generative" approach, where AI can independently design new, versatile enzymes like tryptophan synthases. Ultimately, these findings demonstrate that combining laboratory data with computer-guided design enables the engineering of highly efficient biological tools with unprecedented speed and precision.

59 BASIC BIOLOGICAL SCIENCES

Learning genetic perturbation effects with variational causal inference

Advances in sequencing technologies have enhanced the understanding of gene regulation in cells. In particular, Perturb-seq has enabled high-resolution profiling of the transcriptomic response to genetic perturbations at the single-cell level. This understanding has implications in functional genomics and potentially for identifying therapeutic targets. Various computational models have been developed to predict perturbational effects. While deep learning models excel at interpolating observed perturbational data, they tend to overfit in the lack of enough data and may not generalize well to unseen perturbations. In contrast, mechanistic models, such as linear causal models based on gene regulatory networks, hold greater potential for extrapolation, as they encapsulate regulatory information that can predict responses to unseen perturbations. However, their application has been limited to small studies due to overly simplistic assumptions, making them less effective in handling noisy, large-scale single-cell data. We propose a hybrid approach that combines a mechanistic causal model with variational deep learning, termed Single Cell Causal Variational Autoencoder (SCCVAE). The mechanistic model employs a learned regulatory network to represent perturbational changes as shift interventions that propagate through the learned network. SCCVAE integrates this mechanistic causal model into a variational autoencoder, generating rich, comprehensive transcriptomic responses. Our results indicate that SCCVAE exhibits superior performance over current state-of-the-art baselines for extrapolating to predict unseen perturbational responses. Additionally, for the observed perturbations, the latent space learned by SCCVAE allows for the identification of functional perturbation modules and simulation of single-gene knockdown experiments of varying penetrance, presenting a robust tool for interpreting and interpolating perturbational responses at the single-cell level.

59 BASIC BIOLOGICAL SCIENCES

Data for Development, Optimization, and Application of an Episomal Plasmid System for Rhodotorula toruloides

Rhodotorula toruloides is an emerging oleaginous yeast with strong potential as a microbial cell factory for the production of acetyl-CoA-derived bioproducts. However, engineering of this organism has been limited by the absence of a functional episomal plasmid system, a foundational genetic tool for rapid gene expression, pathway testing, and CRISPR-based genome engineering. Here, we report the first episomal plasmid system for R. toruloides . Through systematic screening of candidate autonomously replicating sequences (ARSs) from diverse sources, we identified multiple functional ARS elements and selected C63F4, a fragment derived from Contig 63 of R. toruloides CBS14, because of its stable performance. The resulting pC63F4 plasmid was maintained episomally, supported GFP reporter expression, exhibited a copy number of 2.39 ± 0.13, and showed good stability during long term cultivation. To overcome poor transformation efficiency, we developed a Cre-loxP-mediated in vivo re-circularization strategy that enabled reliable delivery of the episomal plasmid. Using this improved system, we demonstrated functional episomal expression of metabolic engineering genes and multi-gene pathways for the production of triacetic acid lactone, fatty alcohols, and limonene. Finally, we leveraged this platform to establish a redesigned CRISPR system that enables seamless genome editing in R. toruloides for the first time, while also simplifying marker recycling. Together, this work establishes a long-needed episomal plasmid platform and associated CRISPR toolkit that will accelerate metabolic engineering, synthetic biology, and fundamental studies in R. toruloides .

Gene Editing

RCSB protein data Bank: Next‐generation advanced search for exploration of experimental structures and computed structure models

Abstract The Protein Data Bank (PDB), established in 1971, is the primary global, open‐access archive for experimentally determined 3D macromolecular structures (proteins, RNA, DNA). The research‐focused RCSB.org web‐portal provides access to these data alongside more than one million machine‐learning‐predicted structure models, greatly expanding the available structural landscape. Rapid growth of both experimental and computational structures has increased the need for powerful yet accessible search tools that serve a broad and diverse scientific community. Herein, we describe a redesigned RCSB Protein Data Bank RCSB.org Advanced Search capability that supports intuitive discovery of 3D structures through a unified interface. This interface integrates annotation‐, sequence‐, and 3D structure‐based searches, embeds an interactive 3D viewer, and incorporates curated biological knowledge, such as catalytic site definitions from Mechanism and Catalytic Site Atlas and ligand‐guided structural motifs, for constructing geometry‐driven queries. A new Chemical Search tool allows definition of chemical queries via an integrated drawing tool or standard identifiers, seamlessly combining them with annotation filters. By allowing query definition directly within spatial and chemical contexts, these search interfaces reduce the need for detailed knowledge of residue numbering, chain identifiers, or external cheminformatics software. This capability enables efficient exploration of structures, chemical diversity, and structure–function relationships across all life domains. The redesigned interfaces can be accessed directly at rcsb.org/search/advanced for Advanced Search and rcsb.org/search/chemical for Chemical Search.

Rose, Yana [Research Collaboratory for Structural

Structural basis for intermodular communication in assembly-line polyketide biosynthesis

Assembly-line polyketide synthases (PKSs) are modular multi-enzyme systems with considerable potential for genetic reprogramming. Understanding how they selectively transport biosynthetic intermediates along a defined sequence of active sites could be harnessed to rationally alter PKS product structures. Here, to investigate functional interactions between PKS catalytic and substrate acyl carrier protein (ACP) domains, we employed a bifunctional reagent to crosslink transient domain–domain interfaces of a prototypical assembly line, the 6-deoxyerythronolide B synthase, and resolved their structures by single-particle cryogenic electron microscopy (cryo-EM). Together with statistical per-particle image analysis of cryo-EM data, we uncovered interactions between ketosynthase (KS) and ACP domains that discriminate between intra-modular and inter-modular communication while reinforcing the relevance of conformational asymmetry during the catalytic cycle. Our findings provide a foundation for the structure-based design of hybrid PKSs comprising biosynthetic modules from different naturally occurring assembly lines.

59 BASIC BIOLOGICAL SCIENCES

Decoding substrate specificity determining factors in glycosyltransferase-B enzymes – insights from machine learning models

Substrate specificity is an essential characteristic of any enzyme's function and an understanding of the factors that determine this specificity is crucial for enzyme engineering. Unlike the structure of an enzyme which is directly impacted by its sequence, substrate specificity as an enzyme attribute involves a rather indirect relationship with sequence as it also depends on structural aspects that dictate substrate accessibility and active site dynamics. In this study, we explore the performance of classifier-based machine learning models trained on curated sequence and structural data for a class of glycosyltransferases (GTs), namely GT-Bs, to understand their substrate specificity determining factors. GTs enable the transfer of sugar moieties to other biomolecules such as oligosaccharides or proteins and are found in all kingdoms of life. In plants, GTs participate in the biosynthesis of plant cell wall biopolymers (e.g.: hemicelluloses and pectins) and are an integral part of the enzymatic machinery that enables the storage of carbon and energy as plant biomass. To elucidate the substrate specificity of uncharacterized GT-Bs, we constructed multi-label machine learning models (Support Vector Classifier, K-Nearest Neighbors, Gaussian Naïve-Bayes, Random Forest) that incorporate both sequence and structural features. These models achieve good predictive accuracies on test datasets. However, despite our use of structural information, we highlight that there is further scope for improvement in training these models to draw interpretable relationships between sequence, structure and substrate specificity determining motifs in GT-Bs.

97 MATHEMATICS AND COMPUTING

Investigating biological nitrogen fixation via single-cell transcriptomics

The extensive use of nitrogen fertilizers has detrimental environmental consequences, and it is essential for society to explore sustainable alternatives. One promising avenue is engineering root nodule symbiosis, a naturally occurring process in certain plant species within the nitrogen-fixing clade, into non-leguminous crops. Advancements in single-cell transcriptomics provide unprecedented opportunities to dissect the molecular mechanisms underlying root nodule symbiosis at the cellular level. This review summarizes key findings from single-cell studies in Medicago truncatula, Lotus japonicus, and Glycine max. We highlight how these studies address fundamental questions about the development of root nodule symbiosis, including the following findings: (i) single-cell transcriptomics has revealed a conserved transcriptional program in root hair and cortical cells during rhizobial infection, suggesting a common infection pathway across legume species; (ii) characterization of determinate and indeterminate nodules using single-cell technologies supports the compartmentalization of nitrogen fixation, assimilation, and transport into distinct cell populations; (iii) single-cell transcriptomics data have enabled the identification of novel root nodule symbiosis genes and provided new approaches for prioritizing candidate genes for functional characterization; and (iv) trajectory inference and RNA velocity analyses of single-cell transcriptomics data have allowed the reconstruction of cellular lineages and dynamic transcriptional states during root nodule symbiosis.

Lotus japonicus

Metagenome-assembled genomes from topsoils collected during NEON campaign in East River, CO (06/14/2018-06/28/2018)

The Watershed Function Science Focus Area (WF SFA) at Lawrence Berkeley National Lab is working to build a mechanistic understanding of the distribution and dynamics of biogeochemical processes in mountainous watersheds and their response to perturbation. In June 2018, the NEON (National Ecological Observatory Network) Airborne Observatory Platform (AOP) performed a taskable airborne imaging campaign to collect visible to shortwave infrared (VSWIR) imaging spectroscopy and LiDAR data across 330 km2 in the Upper East River at Crested Butte, CO. We conducted a parallel ground sampling campaign to sample vegetation traits, as well as soil physical, chemical, and microbiological characteristics. We collected these samples from 438 sites across 12 locations spanning much of the elevation, topographic, and geologic variability across the study area. A subset of 250 samples were used for soil metagenomics which is presented here. In addition, at each site, vegetation samples were collected to measure species-specific leaf water content and leaf mass area, foliar elemental composition and foliar CN stable isotope ratios. Soil samples were collected to measure soil physical properties which include bulk density and soil texture analysis. A suite of soil chemical properties was measured from the samples collected at each site, including pH, organic matter, concentrations exchangeable cations, total elemental composition, and the concentrations of extractable N pools (e.g. total free amino acids, ammonium, nitrate, dissolved organic N, and total dissolved N). Additionally, we have measured soil microbial biomass CN stoichiometry. Here, we present 1982 metagenome-assembled genomes (MAGs) for the bacterial and archaeal community from topsoil collected from during NEON 2018 campaign. All metagenomes were sequenced at JGI (Joint Genome Institute) (GOLD Study ID: Gs0149986). Metagenomes were assembled using JGI Metagenome Workflow (10.1128/mSystems.00804-20). The dataset includes (1) zip files for 1982 MAG fasta files (neon_genomes1-5.tar.gz, split into 5 tarballs to keep tarballs under 0.5 GB), (2) neon_Gs0149986_samples_soilproperties_metagenomes.csv: the sample information together with the accession numbers for the underlying metagenomes and the associated soil physical and chemical measurements in NMDC (National Microbiome Data Collaborative) compliant format, (3) neon_Gs0149986.kml: location bounding box file for the sampled locations, (4) samples.csv: sample metadata file used to register Internationall Generic Sample Numbers (IGSNs), (5) flmd.csv: file level metadata file, and (6) dd.csv: data dictionary file. This work was supported by the Watershed Function Science Focus Area at Lawrence Berkeley National Laboratory funded by the US Department of Energy, Office of Science, Biological and Environmental Research under Contract No. DE-AC02-05CH11231.

2018 NEON and 2025 CHESS Campaigns

Prevalence and diversity of TAL effector-like proteins in fungal endosymbiotic Mycetohabitans spp.

EndofungalMycetohabitans(formerlyBurkholderia) spp. rely on a type III secretion system to deliver mostly unidentified effector proteins when colonizing their host fungus,Rhizopus microsporus. The one known secreted effector family fromMycetohabitansconsists of homologues of transcription activator-like (TAL) effectors, which are used by plant pathogenicXanthomonasandRalstoniaspp. to activate host genes that promote disease. These ‘BurkholderiaTAL-like (Btl)’ proteins bind corresponding specific DNA sequences in a predictable manner, but their genomic target(s) and impact on transcription in the fungus are unknown. Recent phenotyping of Btl mutants of twoMycetohabitansstrains revealed that the single Btl in oneMycetohabitans endofungorumstrain enhances fungal membrane stress tolerance, while others in aMycetohabitans rhizoxinicastrain promote bacterial colonization of the fungus. The phenotypic diversity underscores the need to assess the sequence diversity and, given that sequence diversity translates to DNA targeting specificity, the functional diversity of Btl proteins. Using a dual approach to maximize capture of Btl protein sequences for our analysis, we sequenced and assembled nineMycetohabitansspp. genomes using long-read PacBio technology and also mined available short-read Illumina fungal–bacterial metagenomes. We show thatbtlgenes are present across diverseMycetohabitansstrains from Mucoromycota fungal hosts yet vary in sequences and predicted DNA binding specificity. Phylogenetic analysis revealed distinct clades of Btl proteins and suggested thatMycetohabitansmight contain more species than previously recognized. Within our data set, Btl proteins were more conserved acrossM. rhizoxinicastrains than acrossM. endofungorum, but there was also evidence of greater overall strain diversity within the latter clade. Overall, the results suggest that Btl proteins contribute to bacterial–fungal symbioses in myriad ways.

Genetics & Heredity