Search NASA⌕ Search

SEARCH · Search NASA

Results for “functional genomics”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 127 records · Page 7

Genome-resolved analysis of Serratia marcescens strain SMTT infers niche specialization as a hydrocarbon-degrader

Abstract Bacteria that are chronically exposed to high levels of pollutants demonstrate genomic and corresponding metabolic diversity that complement their strategies for adaptation to hydrocarbon-rich environments. Whole genome sequencing was carried out to infer functional traits of Serratia marcescens strain SMTT recovered from soil contaminated with crude oil. The genome size (Mb) was 5,013,981 with a total gene count of 4,842. Comparative analyses with carefully selected S. marcescens strains, 2 of which are associated with contaminated soil, show conservation of central metabolic pathways in addition to intra-specific genetic diversity and metabolic flexibility. Genome comparisons also indicated an enrichment of genes associated with multidrug resistance and efflux pumps for SMTT. The SMTT genome contained genes that enable the catabolism of aromatic compounds via the protocatechuate para-degradation pathway, in addition to meta-cleavage of catechol (meta-cleavage pathway II); gene enrichment for aromatic compound degradation was markedly higher for SMTT compared to the other S. marcescens strains analysed. Our data presents a valuable genetic inventory for future studies on strains of S. marcescens and provides insights into those genomic features of SMTT with industrial potential.

Genetics & Heredity↗

A compendium of human gene functions derived from evolutionary modelling

A comprehensive, computable representation of the functional repertoire of all macromolecules encoded within the human genome is a foundational resource for biology and biomedical research. The Gene Ontology Consortium has been working towards this goal by generating a structured body of information about gene functions, which now includes experimental findings reported in more than 175,000 publications for human genes and genes in experimentally tractable model organisms 1,2 . Here, we describe the results of a large, international effort to integrate all of these findings to create a representation of human gene functions that is as complete and accurate as possible. Specifically, we apply an expert-curated, explicit evolutionary modelling approach to all human protein-coding genes. This approach integrates available experimental information across families of related genes into models that reconstruct the gain and loss of functional characteristics over evolutionary time. The models and the resulting set of 68,667 integrated gene functions cover approximately 82% of human protein-coding genes. The functional repertoire reveals a marked preponderance of molecular regulatory functions, and the models provide insights into the evolutionary origins of human gene functions. We show that our set of descriptions of functions can improve the widely used genomic technique of Gene Ontology enrichment analysis. The experimental evidence for each functional characteristic is recorded, thereby enabling the scientific community to help review and improve the resource, which we have made publicly available.

59 BASIC BIOLOGICAL SCIENCES↗

Unraveling the Dynamics of Nucleosome Arrays

The organization of genomic DNA into chromatin is a fundamental determinant of genome stability, regulation, and cellular function. Nucleosomes, the basic repeating units of chromatin, assemble into higher-order structures whose organization and heterogeneity remain difficult to characterize using conventional ensemble-averaged techniques. A key need in the field is the development of experimental approaches capable of directly visualizing nucleosome assemblies and their structural variability at the single-molecule level. This LDRD Lab-Wide project focused on establishing and evaluating atomic force microscopy (AFM)–based approaches for the characterization of nucleosome assemblies. The work emphasized experimental workflows for preparing, imaging, and assessing multi-nucleosome systems, rather than isolated single nucleosomes. Through method development and exploratory measurements, the project demonstrated the feasibility of applying scanning probe microscopy to investigate chromatin-relevant assemblies and provided preliminary insight into the strengths and limitations of this approach for future quantitative studies. Results and lessons learned from this effort were disseminated to the broader scientific community through multiple national conference presentations, helping to position LLNL for continued work in chromatin and genome organization research.

59 BASIC BIOLOGICAL SCIENCES↗

Myco-Ed: Mycological curriculum for education and discovery

Fungi are important and hyperdiverse organisms, yet chronically understudied. Most fungal clades have no reference genomes, impeding our understanding of their ecosystem functions and use as solutions in health and biotechnology. Also, opportunities for training in fungal biology and genomics are lacking, creating a bottleneck that hinders the recruitment and cultivation of a talented future mycological workforce. To address these issues, we developed Myco-Ed, an educational program offering training and scientific contributions through genome sequencing and analysis. Myco-Ed empowers students to pursue careers in fungal biology while improving fungal resources. Myco-Ed has been piloted at 12 institutions (15 classrooms) ranging from online e-Campuses to R1 universities, resulting in hundreds of fungal observations and many new high-quality reference genomes.

Branco, Sara↗

Microbial Community Analysis & Functional Evaluation in Soils

The overall objective of this proposal was to develop technologies to alter the composition and function of important members of microbial communities. In particular, the overall objective of the microbial community editing portion of the proposal focuses on developing foundational tools and understanding required to predict, alter and design grass rhizosphere communities impacting DOE missions. Specifically, the project is centered on the Microbial Community Analysis & Functional Evaluation in Soils (m-CAFES) to manipulate microbial consortia associated with plants of interest for the bioenergy sector, under the presumption that bacterial communities can be manipulated to enhance plant health. For tasks of specific interest to us, we are focusing on developing novel Clustered Regularly Interspaced Short Palindromic Repeats (CRISPR) based technologies (primarily focusing on Aim 1) and their delivery modalities (notably subaim 1.2) to edit specific bacterial genomes of interest to enhance their functionalities, and programmably ablate specific undesirable members of bacterial communities for plant health. We are focusing on engineering bacteriophages (bacterial viruses, for subaim 1.2) to carry programmable CRISPR-Cas systems (subaim 1.1) to target (ablate) or alter (edit) genomes of interest. This will enable us to carry out microbial perturbations that will impact community composition and function and ultimately plant growth and health, to enable the next phase of the project by deploying them in situ (subaims 1.3 and 1.4).

59 BASIC BIOLOGICAL SCIENCES↗

Community‐Level Metabolic Shifts Following Land Use Change in the Amazon Rainforest Identified by a Supervised Machine Leaning Approach

ABSTRACT The Amazon rainforest has been subjected to high rates of deforestation, mostly for pasturelands, over the last few decades. This change in plant cover is known to alter the soil microbiome and the functions it mediates, but the genomic changes underlying this response are still unresolved. In this study, we used a combination of deep shotgun metagenomics complemented by a supervised machine learning approach to compare the metabolic strategies of tropical soil microbial communities in pristine forests and long‐term established pastures in the Amazon. Machine learning‐derived metagenome analysis indicated that microbial community structures (bacteria, archaea and viruses) and the composition of protein‐coding genes were distinct in each plant cover type environment. Forest and pasture soils had different genomic diversities for the above three taxonomic groups, characterised by their protein‐coding genes. These differences in metagenome profiles in soils under forests and pastures suggest that metabolic strategies related to carbohydrate and energy metabolisms were altered at community level. Changes were also consistent with known modifications to the C and N cycles caused by long‐term shifts in aboveground vegetation and were also associated with several soil physicochemical properties known to change with land use, such as the C/N ratio, soil temperature and exchangeable acidity. In addition, our analysis reveals that these alterations in land use can also result in changes to the composition and diversity of the soil DNA virome. Collectively, our study indicates that soil microbial communities shift their overall metabolic strategies, driven by genomic alterations observed in pristine forests and long‐term established pastures with implications for the C and N cycles.

carbon and nitrogen cycles↗

DNA parts and gene constructs for plant biodesign

Plant biodesign requires the knowledge of DNA parts (e.g., genes, promoters, terminators), along with their combinations (as gene constructs) linked to engineered traits. DNA parts with validated or predicted functions in plants have been deposited in various online databases. However, these existing databases focus on basic biological functions of individual DNA parts, leaving a gap between basic knowledge and bioengineering applications. To fill this knowledge gap, we have created a user-friendly, open-ended database as a knowledge graph linking DNA parts to gene constructs to traits. This database contains experimentally validated DNA parts and gene constructs documented in peer-reviewed publications. The DNA parts include 1) molecular components with biological functions, such as genes involved in various biological processes (e.g., metabolic and signal transduction pathways) and 2) molecular components with technical functions, such as gene expression, genome engineering and sequence splicing. The gene constructs deposited in this database include both single-gene and multi-gene constructs. This database allows users to submit DNA parts and gene construct compositions linked to engineered traits described in peer-reviewed publications, providing a public digital repository for sharing the biodesign information among the researchers in the fields of plant biotechnology and plant synthetic biology.

plant biodesign synthetic biology gene constructs ↗

MAP kinase pathways in the yeast Saccharomyces cerevisiae

A cascade of three protein kinases known as a mitogen-activated protein kinase (MAPK) cascade is commonly found as part of the signaling pathways in eukaryotic cells. Almost two decades of genetic and biochemical experimentation plus the recently completed DNA sequence of the Saccharomyces cerevisiae genome have revealed just five functionally distinct MAPK cascades in this yeast. Sexual conjugation, cell growth, and adaptation to stress, for example, all require MAPK-mediated cellular responses. A primary function of these cascades appears to be the regulation of gene expression in response to extracellular signals or as part of specific developmental processes. In addition, the MAPK cascades often appear to regulate the cell cycle and vice versa. Despite the success of the gene hunter era in revealing these pathways, there are still many significant gaps in our knowledge of the molecular mechanisms for activation of these cascades and how the cascades regulate cell function. For example, comparison of different yeast signaling pathways reveals a surprising variety of different types of upstream signaling proteins that function to activate a MAPK cascade, yet how the upstream proteins actually activate the cascade remains unclear. We also know that the yeast MAPK pathways regulate each other and interact with other signaling pathways to produce a coordinated pattern of gene expression, but the molecular mechanisms of this cross talk are poorly understood. This review is therefore an attempt to present the current knowledge of MAPK pathways in yeast and some directions for future research in this area.

Non-NASA Center↗

Structural and functional analyses of SARS-CoV-2 Nsp3 and its specific interactions with the 5’ UTR of the viral genome

ABSTRACT Non-structural protein 3 (Nsp3) is the largest open reading frame encoded in the SARS-CoV-2 genome, essential for the formation of double-membrane vesicles (DMV) wherein viral RNA replication occurs. We conducted an extensive structure-function analysis of Nsp3 and determined the crystal structures of the ubiquitin-like 1 (Ubl1), nucleic acid binding (NAB), β-coronavirus-specific marker (βSM) domains, and a sub-region of the Y domain of this protein. We show that the Ubl1, ADP-ribose phosphatase (ADRP), human SARS Unique (HSUD), NAB, and Y domains of Nsp3 bind the 5’ UTR of the viral genome and that the Ubl1 and Y domains possess affinity for recognition of this region, suggesting high specificity. The Ubl1-Nucleocapsid (N) protein complex binds the 5’ UTR with greater affinity than the individual proteins alone. Our results suggest that multiple domains of Nsp3, particularly Ubl1 and Y, shepherd the 5’ UTR of the viral genome during translocation through the DMV membrane, priming the Ubl1 domain to load the genome onto N protein. IMPORTANCE The largest protein encoded by the SARS-CoV-2 genome is Nsp3. In infected cells, this multi-domain protein forms a pore structure in the virus-induced double-membrane vesicles (DMV). We have incomplete data on Nsp3 molecular structure, and here, we describe crystal structures for multiple domains of Nsp3. It is thought that newly replicated viral RNA transits through the DMV pore; however, we possess incomplete data on which regions of Nsp3 actually interact with RNA. Here, we present data showing that five domains of Nsp3 interact with the 5’ UTR of the SARS-CoV-2 RNA, including the Y domain for which no function has ever been discovered. These data suggest that the pore structure plays an active role in recognizing the terminal end of the genome, transiting and loading the viral RNA onto the cytoplasmic nucleocapsid protein. These data help expand our knowledge of Nsp3 structure and function and the SARS-CoV-2 replication cycle.

Microbiology↗

Structure and sequence evolution in the pennycress ( Thlaspi arvense ) pangenome

Eukaryotic genomes harbor many forms of variation, including nucleotide diversity and structural polymorphisms, which experience natural selection and contribute to genome evolution and biodiversity. Harnessing this variation for agriculture hinges on our ability to detect, quantify, catalog, and deploy genetic diversity. Here, we explore seven complete genomes of the emerging biofuel crop pennycress ( Thlaspi arvense ) drawn from across the species' current genetic diversity to catalog variation in genome structure and content. Across this new pangenome resource, we find contrasting evolutionary modes in different genomic zones. Gene-poor, repeat-rich pericentromeric regions experience frequent rearrangements, including repeated centromere repositioning. By contrast, conserved gene-dense chromosome arms maintain large-scale synteny across accessions even in fast-evolving NOD-like receptor immune genes, where microsynteny breaks down across species, but gene cluster positioning macrosynteny is maintained. Our findings highlight that multiple elements of the genome experience dynamic evolution that conserves functional content on the chromosome scale but allows repositioning and presence–absence variation on a local scale. This diversity is invisible to classical reference-based strategies and highlights the strength and utility of pangenomic resources. These results provide a valuable case study of rapid genomic structural evolution within a species and powerful resources for crop development in an emerging biofuel crop.

Thlaspi arvense↗

Systematic identification of transcriptional activation domains from non-transcription factor proteins in plants and yeast

Transcription factors can promote gene expression through activation domains. Whole-genome screens have systematically mapped activation domains in transcription factors but not in non-transcription factor proteins (e.g., chromatin regulators and coactivators). To fill this knowledge gap, we employed the activation domain predictor PADDLE to analyze the proteomes of Arabidopsis thaliana and Saccharomyces cerevisiae. We screened 18,000 predicted activation domains from >800 non-transcription factor genes in both species, confirming that 89% of candidate proteins contain active fragments. Our work enables the annotation of hundreds of nuclear proteins as putative coactivators, many of which have never been ascribed any function in plants. Analysis of peptide sequence compositions reveals how the distribution of key amino acids dictates activity. Finally, we validated short, "universal" activation domains with comparable performance to state-of-the-art activation domains used for genome engineering. Our approach enables the genome-wide discovery and annotation of activation domains that can function across diverse eukaryotes.

59 BASIC BIOLOGICAL SCIENCES↗

Exploring Connectivity in Sequence Space of Functional RNA

Emergence of replicable genetic molecules was one of the marking points in the origin of life, evolution of which can be conceptualized as a walk through the space of all possible sequences. A theoretical concept of fitness landscape helps to understand evolutionary processes through assigning a value of fitness to each genotype. Then, evolution of a phenotype is viewed as a series of consecutive, single-point mutations. Natural selection biases evolution toward peaks of high fitness and away from valleys of low fitness. whereas neutral drift occurs in the sequence space without direction as mutations are introduced at random. Large networks of neutral or near-neutral mutations on a fitness landscape, especially for sufficiently long genomes, are possible or even inevitable. Their detection in experiments, however, has been elusive. Although a few near-neutral evolutionary pathways have been found, recent experimental evidence indicates landscapes consist of largely isolated islands. The generality of these results, however, is not clear, as the genome length or the fraction of functional molecules in the genotypic space might have been insufficient for the emergence of large, neutral networks. Thorough investigation on the structure of the fitness landscape is essential to understand the mechanisms of evolution of early genomes. RNA molecules are commonly assumed to play the pivotal role in the origin of genetic systems. They are widely believed to be early, if not the earliest, genetic and catalytic molecules, with abundant biochemical activities as aptamers and ribozymes, i.e. RNA molecules capable, respectively, to bind small molecules or catalyze chemical reactions. Here, we present results of our recent studies on the structure of the sequence space of RNA ligase ribozymes selected through in vitro evolution. Several hundred thousands of sequences active to a different degree were obtained by way of deep sequencing. Analysis of these sequences revealed several large clusters defined such that every sequence in a cluster can be reached from any other sequence in the same cluster through a series of single point mutations. Sequences in a single cluster appear to adopt more than one secondary structure. The mechanism of refolding within a single cluster was examined. To shed light on possible evolutionary paths in the space of ribozymes, the connectivity between clusters was investigated. The effect of length of RNA molecules on the structure of the fitness landscape and possible evolutionary paths was examined by way of comparing functional sequences of 20 and 80 nucleobases in length. It was found that sequences of different lengths shared secondary structure motifs that were presumed responsible for catalytic activity, with increasing complexity and global structural rearrangements emerging in longer molecules.

Wei, Chenyu↗

MjCyc: Rediscovering the pathway-genome landscape of the first sequenced archaeon, Methanocaldococcus (Methanococcus) jannaschii

The genome of Methanocaldococcus (Methanococcus) jannaschii DSM 2661 was the first Archaeal genome to be sequenced in 1996. Subsequent sequence-based annotation cycles led to its first metabolic reconstruction in 2005. Leveraging new experimental results and function assignments, we have now re-annotated M. jannaschii, creating an updated resource with novel information and testable predictions in a pathway-genome database available at BioCyc.org. This reannotation effort has resulted in 652 function assignments with enzyme roles, accounting for a third of the total protein-coding entries for this genome. The updated resource includes 883 reactions, 540 enzymes, and 142 individual pathways. Despite notable progress in computational genomics, more than a third of the genome remains functionally uncharacterized. The publicly available MjCyc pathway-genome database holds great potential for the wider community to conduct research on the biology of methanogenic Archaea.

59 BASIC BIOLOGICAL SCIENCES↗

Transcriptional response of Methanosarcina acetivorans to repression of the energy-conserving methanophenazine: CoM-CoB heterodisulfide reductase enzyme HdrED

ABSTRACT Methane-producing archaea are key organisms in the anaerobic carbon cycle. These organisms, also called methanogens, grow by converting substrate to methane gas in a process called methanogenesis. Previous research showed that the reduction of the terminal electron acceptor is the rate-limiting step in methanogenesis by Methanosarcina acetivorans . In order to gain insight into how the cells sense and respond to the availability of the terminal electron acceptor, we designed an experiment to deplete cells of the essential terminal oxidase enzyme, HdrED. We found that the depletion of HdrED in vivo results in a higher abundance of transcripts for methyltransferases ( mtaC2, mtaB3, mtaC3 ), coenzyme B biosynthesis, C1 metabolism, and pyrimidine compounds. In most cases, these changes were distinct from transcript abundance changes observed during the transition from exponential growth to stationary phase cultures. These data implicate the methylotrophic methanogenesis regulator MsrC (MA4383) in CoM-S-S-CoB heterodisulfide sensing and indicate cells have a specific mechanism to sense intracellular ratio of CoM-S-S-CoB, coenzyme M, and coenzyme B thiols and further suggest transcripts encoding translation and methanogenesis functions are controlled by feed-forward regulation depending on substrate availability. IMPORTANCE Methanosarcina is an emerging model archaeon and synthetic biology platform for the production of renewable energy and sustainable chemicals to reduce dependence on petroleum. Research into metabolic networks and gene regulation in this organism and other methanogens will inform genome-scale metabolic modeling and microbial function prediction in uncultured or non-model anaerobes and archaea. This study suggests methanogens use unknown mechanisms to efficiently couple methanogenesis to gene regulation via CoM-S-S-CoB and ATP availability.

Buan, Nicole R. (ORCID:000000027560973X)↗

Proteomic Assessment of Fluid Shifts and Association with Visual Impairment and Intracranial Pressure in Twin Astronauts

BACKGROUND: Astronauts participating in long duration space missions are at an increased risk of physiological disruptions. The development of visual impairment and intracranial pressure (VIIP) syndrome is one of the leading health concerns for crew members on long-duration space missions; microgravity-induced fluid shifts and chronic elevated cabin CO2 may be contributing factors. By studying physiological and molecular changes in one identical twin during his 1-year ISS mission and his ground-based co-twin, this work extends a current NASA-funded investigation to assess space flight induced "Fluid Shifts" in association with the development of VIIP. This twin study uniquely integrates physiological and -omic signatures to further our understanding of the molecular mechanisms underlying space flight-induced VIIP. We are: (i) conducting longitudinal proteomic assessments of plasma to identify fluid regulation-related molecular pathways altered by long-term space flight; and (ii) integrating physiological and proteomic data with genomic data to understand the genomic mechanism by which these proteomic signatures are regulated. PURPOSE: We are exploring proteomic signatures and genomic mechanisms underlying space flight-induced VIIP symptoms with the future goal of developing early biomarkers to detect and monitor the progression of VIIP. This study is first to employ a male monozygous twin pair to systematically determine the impact of fluid distribution in microgravity, integrating a comprehensive set of structural and functional measures with proteomic, metabolomic and genomic data. This project has a broader impact on Earth-based clinical areas, such as traumatic brain injury-induced elevations of intracranial pressure, hydrocephalus, and glaucoma. HYPOTHESIS: We predict that the space-flown twin will experience a space flight-induced alteration in proteins and peptides related to fluid balance, fluid control and brain injury as compared to his pre-flight protein/peptide signatures. Conversely, the trajectory of these protein signatures will remain relatively constant in his ground based co-twin. METHODS: We are using proteomic and standard immunoelectrophoresis techniques to delineate the change in protein signatures throughout the course of a long duration space flight in relation to the development of VIIP. We are also applying a novel cell-based metaboloic organ system assay ("Organs on a Plate") to address how these circulating biomarkers affect physiological processes at the cellular and organ level which could result in VIIP symptoms. These molecular data will be correlated with physiological measures (eg. extra and intracellular fluid volume, vascular filling/flow patterns, MRI, and Optic Coherence Tomography. DISCUSSION: Pre- and in-flight data collection is in progress for the space-flown twin, and similar data have been obtained from the ground-based twin. Biosamples will be batch processed when received from ISS after the conclusion of the 1-year mission. Omic and Physiological measures from the twin astronauts will be compared to similar data being collected on twin subjects who participated in simulated microgravity study. bed rest study.

Rana, Brinda K.↗

The Molecular Ecology of Guerrero Negro: Justifying the Need for Environmental Genomics

The record of life on the only planet where it is known to exist is contained in the biogeochemical processes that organisms catalyze for their survival, in the compounds that they produce, and in their phylogenetic (evolutionary) relationships to each other. We manipulated sulfate and nutrient concentrations in intact microbial mats over periods of time up to a year. The objectives of the manipulations were: 1) characterize the diversity of process-associated functional genes; 2) understand environmental conditions leading to shifts in microbial guilds; 3) monitor/identify competitive responses of organisms sharing a metabolic niche. Characterization of functional genes associated with carbon (mcrA), nitrogen (nifH, nirK) and sulfur (dsrkB) cycling performed to date provided insight into the diversity and metabolic potential of the system; however, we only identified broad scale correlations between gene abundances and changes in mat physiology. For instance, increases in methane production by mats subjected to lowered sulfate and salinity concentrations were correlated with an observed increase in abundance of hydrogenotroph-like mcrA genes. However, due to low sequence similarity to any cultured isolates, phylogenetic associations only allow order level taxonomic commentary, preventing any associations being made on the cellular level. In each of the genes characterized from these experiments, a significant portion of sequences recovered show minimal phylogenetic affiliation to cultured organisms, preventing any understanding of inter-community dynamics and the functional capacities of these unknown organisms. Environmental genomics may provide insight into mat systems by allowing the correlation of functional genes with phylogenetic markers.

Smith, Jason M.↗

Differential Expression of Core Metabolic Functions in Candidatus Altiarchaeum Inhabiting Distinct Subsurface Ecosystems

Candidatus Altiarchaea are widespread across aquatic subsurface ecosystems and possess a highly conserved core genome, yet adaptations of this core genome to different biotic and abiotic factors based on gene expression remain unknown. Here, we investigated the metatranscriptome of two Ca. Altiarchaeum populations that thrive in two substantially different subsurface ecosystems. In Crystal Geyser, a high-CO2 groundwater system in the USA, Ca. Altiarchaeum crystalense co-occurs with the symbiont Ca. Huberiarchaeum crystalense, while in the Muehlbacher sulfidic spring in Germany, an artesian spring high in sulfide concentration, Ca. A. hamiconexum is heavily infected with viruses. We here mapped metatranscriptome reads against their genomes to analyse the in situ expression profile of their core genomes. Out of 537 shared gene clusters, 331 were functionally annotated and 130 differed significantly in expression between the two sites. Main differences were related to genes involved in cell defence like CRISPR-Cas, virus defence, replication, transcription and energy and carbon metabolism. Our results demonstrate that altiarchaeal populations in the subsurface are likely adapted to their environment while influenced by other biological entities that tamper with their core metabolism. We consequently posit that viruses and symbiotic interactions can be major energy sinks for organisms in the deep biosphere.

archaea↗

High-throughput protein characterization by complementation using DNA barcoded fragment libraries

Abstract Our ability to predict, control, or design biological function is fundamentally limited by poorly annotated gene function. This can be particularly challenging in non-model systems. Accordingly, there is motivation for new high-throughput methods for accurate functional annotation. Here, we used co mplementation of aux otrophs and DNA barcode seq uencing (Coaux-Seq) to enable high-throughput characterization of protein function. Fragment libraries from eleven genetically diverse bacteria were tested in twenty different auxotrophic strains of Escherichia coli to identify genes that complement missing biochemical activity. We recovered 41% of expected hits, with effectiveness ranging per source genome, and observed success even with distant E. coli relatives like Bacillus subtilis and Bacteroides thetaiotaomicron . Coaux-Seq provided the first experimental validation for 53 proteins, of which 11 are less than 40% identical to an experimentally characterized protein. Among the unexpected function identified was a sulfate uptake transporter, an O-succinylhomoserine sulfhydrylase for methionine synthesis, and an aminotransferase. We also identified instances of cross-feeding wherein protein overexpression and nearby non-auxotrophic strains enabled growth. Altogether, Coaux-Seq’s utility is demonstrated, with future applications in ecology, health, and engineering.

59 BASIC BIOLOGICAL SCIENCES↗