Search NASA⌕ Search

SEARCH · Search NASA

Results for “Protein Family”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

nf-core/proteinfamilies: a scalable pipeline for the generation of protein families

The growth of metagenomics-derived amino acid sequence data has transformed our understanding of protein function, microbial diversity, and evolutionary relationships. However, the vast majority of these proteins remain functionally uncharacterized. Grouping the millions of such uncharacterized sequences with the few experimentally characterized ones allows the transfer of annotations, while the inspection of conserved residues with multiple sequence alignments can provide clues to function, even in the absence of existing functional information. To address the challenges associated with this data surge and the need to group sequences, we present a scalable, open-source, parametrizable Nextflow pipeline (nf-core/proteinfamilies) that generates nascent protein families or assigns new proteins to existing families. The computational benchmarks demonstrated that resource usage scales approximately linearly with input size, and the biological benchmarks showed that the generated protein families closely resemble manually curated families in widely used databases.

Nextflow↗

metagRoot: a comprehensive database of protein families associated with plant root microbiomes

The plant root microbiome is vital in plant health, nutrient uptake, and environmental resilience. To explore and harness this diversity, we present metagRoot, a specialized and enriched database focused on the protein families of the plant root microbiome. MetagRoot integrates metagenomic, metatranscriptomic, and reference genome-derived protein data to characterize 71 091 enriched protein families, each containing at least 100 sequences. These families are annotated with multiple sequence alignments, CRISPR elements, hidden Markov models, taxonomic and functional classifications, ecosystem and geolocation metadata, and predicted 3D structures using AlphaFold2. MetagRoot is a powerful tool for decoding the molecular landscape of root-associated microbial communities and advancing microbiome-informed agricultural practices by enriching protein family information with ecological and structural context. The database is available at https://pavlopoulos-lab.org/metagroot/ or https://www.metagroot.org.

Chasapi, Maria N↗

Sequence, structure prediction, and epitope analysis of the polymorphic membrane protein family in Chlamydia trachomatis

The polymorphic membrane proteins (Pmps) are a family of autotransporters that play an important role in infection, adhesion and immunity in Chlamydia trachomatis. Here we show that the characteristic GGA(I,L,V) and FxxN tetrapeptide repeats fit into a larger repeat sequence, which correspond to the coils of a large beta-helical domain in high quality structure predictions. Analysis of the protein using structure prediction algorithms provided novel insight to the chlamydial Pmp family of proteins. While the tetrapeptide motifs themselves are predicted to play a structural role in folding and close stacking of the beta-helical backbone of the passenger domain, we found many of the interesting features of Pmps are localized to the side loops jutting out from the beta helix including protease cleavage, host cell adhesion, and B-cell epitopes; while T-cell epitopes are predominantly found in the beta-helix itself. This analysis more accurately defines the Pmp family of Chlamydia and may better inform rational vaccine design and functional studies.

59 BASIC BIOLOGICAL SCIENCES↗

A novel bacterial protein family that catalyses nitrous oxide reduction

Nitrous oxide (N 2 O), a driver of global warming and climate change, has reached unprecedented concentrations in Earth’s atmosphere. Current N 2 O sources outpace N 2 O sinks, emphasizing the need for comprehensive understanding of processes that consume N 2 O. Microbes that express the enzyme N 2 O reductase (N 2 OR) convert N 2 O to climate change-neutral dinitrogen (N 2 ). Known N 2 ORs belong to the canonical clade I and clade II NosZ reductases and are considered key enzymes for N 2 O reduction. Here we report a previously unrecognized protein family with a role in N 2 O reduction, clade III lactonase-type N 2 OR (L-N 2 OR), which diverges in sequence from canonical NosZ but conserves three-dimensional protein structural features. Integrated physiological, metagenomic, proteomic and structural modelling studies demonstrate that L-N 2 ORs catalyse N 2 O reduction. L-N 2 OR genes occur in several phyla, predominantly in uncultured taxa with broad geographic distribution. Our findings expand the known diversity of N 2 ORs and implicate previously unrecognized taxa (for example, Nitrospinota) in N 2 O consumption. In conclusion, the expansion of N 2 OR diversity and the identification of a novel type of catalyst for N 2 O reduction advances the understanding of N 2 O sinks, has implications for greenhouse gas emission and climate change modelling, and expands opportunities for innovative biotechnologies aimed at curbing N 2 O emissions.

He, Guang 何广 [Univ. of Tennessee, Knoxville, TN (U↗

An evolutionarily conserved tryptophan cage promotes folding of the extended RNA recognition motif in the hnRNPR ‐like protein family

Abstract The heterogeneous nuclear ribonucleoprotein (hnRNP) R‐like family is a class of RNA binding proteins in the hnRNP superfamily with diverse functions in RNA processing. Here, we present the 1.90 Å X‐ray crystal structure and solution NMR studies of the first RNA recognition motif (RRM) of human hnRNPR. We find that this domain adopts an extended RRM (eRRM1) featuring a canonical RRM with a structured N‐terminal extension (N ext ) motif that docks against the RRM and extends the β‐sheet surface. The adjoining loop is structured and forms a tryptophan cage motif to position the N ext motif for docking to the RRM. Combining mutagenesis, solution NMR spectroscopy, and thermal denaturation studies, we evaluate the importance of residues in the N ext –RRM interface and adjoining loop on eRRM folding and conformational dynamics. We find that these sites are essential for protein solubility, conformational ordering, and thermal stability. Consistent with their importance, mutations in the N ext –RRM interface and loop are associated with several cancers in a survey of somatic mutations in cancer studies. Sequence and structure comparison of the human hnRNPR eRRM1 to experimentally verified and predicted hnRNPR‐like proteins reveals conserved features in the eRRM.

Biochemistry & Molecular Biology↗

Exabiome: Advancing Microbial Science through Exascale Computing

The Exabiome project seeks to improve the understanding of microbiomes through the development of methods for accelerating metagenomic science using exascale computing. This article gives an overview of scientific impact of the three components of the project: metagenome assembly, protein family detection, and comparative analysis of metagenomes. Exabiome developed MetaHipMer, the only metagenome assembler capable of scaling to full exascale systems. MetaHipMer has enabled ground-breaking assemblies on the Frontier supercomputer, with many scientific benefits, such as the discovery of rare species and viral genomes. To investigate protein families, Exabiome developed two exascale tools, PASTIS and HipMCL. Together, these can utilize exascale resources to understand the functional diversity of billions of dark matter proteins and novel protein families. For comparative analysis, Exabiome developed kmerprof, a tool that can be used to compare huge metagenomes for many different scientific purposes, for example, grouping human microbiomes according to body location.

59 BASIC BIOLOGICAL SCIENCES↗

A conserved chaperone protein is required for the formation of a noncanonical type VI secretion system spike tip complex

Type VI secretion systems (T6SSs) are dynamic protein nanomachines found in Gram-negative bacteria that deliver toxic effector proteins into target cells in a contact-dependent manner. Prior to secretion, many T6SS effector proteins require chaperones and/or accessory proteins for proper loading onto the structural components of the T6SS apparatus. However, despite their established importance, the precise molecular function of several T6SS accessory protein families remains unclear. In this study, we set out to characterize the DUF2169 family of T6SS accessory proteins. Using gene co-occurrence analyses, we find that DUF2169-encoding genes strictly co-occur with genes encoding T6SS spike complexes formed by valine-glycine repeat protein G (VgrG) and DUF4150 domains. Although structurally similar to Pro-Ala-Ala-Arg (PAAR) domains, “PAAR-like” DUF4150 domains lack PAAR motifs and instead contain a conserved PIPY motif, leading us to designate them PIPY domains. Next, we present both genetic and biochemical evidence that PIPY domains require a cognate DUF2169 protein to form a functional T6SS spike complex with VgrG. This contrasts with canonical PAAR proteins, which bind VgrG on their own to form functional spike complexes. By solving the first crystal structure of a DUF2169 protein, we show that this T6SS accessory protein adopts a novel protein fold. Furthermore, biophysical and structural modeling data suggest that DUF2169 contains a dynamic loop that physically interacts with a hydrophobic patch on the surface of its cognate PIPY domain. Based on these findings, we propose a model whereby DUF2169 proteins function as molecular chaperones that maintain VgrG–PIPY spike complexes in a secretion-competent state prior to their export by the T6SS apparatus.

DUF2169↗

A periplasmic zinc capture protein enhances the resistance of Neisseria gonorrhoeae to nutritional immunity

During microbial infection, mammalian hosts reduce the availability of free metals such as zinc in a process known as nutritional immunity. Pathogens counteract nutritional immunity by expressing gene products that enhance growth in metal-limited conditions. One of the most transcriptionally induced genes in zinc-limited Neisseria gonorrhoeae , ngo1049, encodes a DUF4198 family protein we have named Zcp. This family of proteins is widely distributed in Gram-negative bacteria. Here, we provide the first structural, biochemical, and functional characterization of a DUF4198 protein. Zcp is a periplasmic, homodimeric substrate-binding protein (SBP), which binds one zinc ion per subunit with submicromolar affinity. We identified a zinc binding pocket in each subunit, composed of three histidine residues. Zcp enables maximal growth of N. gonorrhoeae in zinc-limited conditions but is dispensable for zinc uptake, in contrast to the cluster A-I SBP ZnuA, which is required for zinc import. The growth defect of zcp mutant N. gonorrhoeae is rescued by zinc supplementation. Zcp associates with proteins with roles in maintaining cell envelope integrity, and N. gonorrhoeae lacking zcp is more sensitive to envelope-targeting antimicrobials. Zcp enables infectivity of human epithelial cells and neutrophils by zinc-limited N. gonorrhoeae . We conclude that N. gonorrhoeae produces Zcp to buffer periplasmic zinc, which enables ZnuA to balance import of different metals and ensures the bioavailability of zinc for extracytoplasmic zinc-requiring proteins, as part of the coordinated response to host-imposed nutritional immunity.

Liyayi, Ian K. [Department of Microbiology, Immuno↗

Specialization Restricts the Evolutionary Paths Available to Yeast Sugar Transporters

Functional innovation at the protein level is a key source of evolutionary novelties. The constraints on functional innovations are likely to be highly specific in different proteins, which are shaped by their unique histories and the extent of global epistasis that arises from their structures and biochemistries. These contextual nuances in the sequence–function relationship have implications both for a basic understanding of the evolutionary process and for engineering proteins with desirable properties. Here, we have investigated the molecular basis of novel function in a model member of an ancient, conserved, and biotechnologically relevant protein family. These Major Facilitator Superfamily sugar porters are a functionally diverse group of proteins that are thought to be highly plastic and evolvable. By dissecting a recent evolutionary innovation in an α-glucoside transporter from the yeast Saccharomyces eubayanus, we show that the ability to transport a novel substrate requires high-order interactions between many protein regions and numerous specific residues proximal to the transport channel. To reconcile the functional diversity of this family with the constrained evolution of this model protein, we generated new, state-of-the-art genome annotations for 332 Saccharomycotina yeast species spanning ~400 My of evolution. By integrating phylogenetic and phenotypic analyses across these species, we show that the model yeast α-glucoside transporters likely evolved from a multifunctional ancestor and became subfunctionalized. The accumulation of additive and epistatic substitutions likely entrenched this subfunction, which made the simultaneous acquisition of multiple interacting substitutions the only reasonably accessible path to novelty.

59 BASIC BIOLOGICAL SCIENCES↗

A widespread family of molecular chaperones promotes the intracellular stability of type VIIb secretion system– exported toxins

To survive in highly competitive environments, bacteria use specialized secretion systems to deliver antibacterial toxins into neighboring cells, thereby inhibiting their growth. In many Gram-positive bacteria, the export of such toxins requires a membrane-bound molecular apparatus known as the type VIIb secretion system (T7SSb). Recently, it was shown that toxin recruitment to the T7SSb requires a physical interaction between a toxin and two or more so-called targeting factors, which harbor key residues required for T7SS-dependent protein export. However, in addition to these targeting factors, some toxins additionally require a protein belonging to the DUF4176 protein family. Here, by examining two toxin–DUF4176 protein pairs, we demonstrate that DUF4176 constitutes a family of toxin-specific molecular chaperones. In addition to being required for toxin stability in producing cells, we find that DUF4176 proteins facilitate toxin export by specifically interacting with a previously uncharacterized intrinsically disordered region found in many T7SS toxins. Using X-ray crystallography, we determine structures of several DUF4176 chaperones in their unbound state, and of a DUF4176 chaperone in complex with the binding site of its cognate toxin. These structures reveal that this binding site consists of a disordered amphipathic α-helix that requires interaction with its cognate chaperone for proper folding. Overall, we have identified a family of secretion system associated molecular chaperones found throughout T7SSb-containing Gram-positive bacteria.

Gkragkopoulou, Polyniki↗

Comparative genomics of Aspergillus nidulans and section Nidulantes

Aspergillus nidulans is an important model organism for eukaryotic biology and the reference for the section Nidulantes in comparative studies. In this study, we de novo sequenced the genomes of 25 species of this section. Whole-genome phylogeny of 34 Aspergillus species and Penicillium chrysogenum clarifies the position of clades inside section Nidulantes. Comparative genomics reveals a high genetic diversity between species with 684 up to 2433 unique protein families. Furthermore, we categorized 2118 secondary metabolite gene clusters (SMGC) into 603 families across Aspergilli, with at least 40 % of the families shared between Nidulantes species. Genetic dereplication of SMGC and subsequent synteny analysis provides evidence for horizontal gene transfer of a SMGC. Proteins that have been investigated in A. nidulans as well as its SMGC families are generally present in the section Nidulantes, supporting its role as model organism. The set of genes encoding plant biomass-related CAZymes is highly conserved in section Nidulantes, while there is remarkable diversity of organization of MAT-loci both within and between the different clades. This study provides a deeper understanding of the genomic conservation and diversity of this section and supports the position of A. nidulans as a reference species for cell biology.

Theobald, Sebastian [Technical University of Denma↗

CAHS: Context-Aware Homology Search

Protein homology search is foundational to bioinformatics: it supports annotation transfer, structure/function inference, and evolutionary analysis over rapidly expanding sequence repositories (e.g., UniProtKB). Profile hidden Markov models (pHMMs), as implemented in HMMER, remain the most widely trusted approach because they provide statistically calibrated E-values; however, their gap behavior is fixed once a profile is trained, despite biological evidence that insertion/deletion tolerance varies across flexible loops and intrinsically disordered regions. We present CAHS (Context-Aware Homology Search), a lightweight query-time adapter for pHMM search that incorporates learned and biologically motivated signals without changing HMMER's downstream search pipeline or its calibrated E-value reporting. Given a query sequence, CAHS computes per-residue representations from a protein language model and a disorder predictor, maps these to profile coordinates, and modulates only match-state transition rows (gap-open and gap-extension probabilities) while preserving Plan7 constraints. We comprehensively evaluate CAHS across six structurally diverse protein families and multi-domain architectures against a 570k-sequence target corpus. CAHS expands detection capability, retrieving thousands of additional remote homologs at relaxed thresholds by maintaining alignment quality through flexible regions. For multi-domain proteins, context-aware modulation resolves 94% of fragmented alignments. Crucially, CAHS preserves hit-set invariance at stringent operating points (E<10-10), demonstrating increased statistical confidence without inflating false positives. Furthermore, sharper statistical distinction between homologs and background noise during early filter stages yields up to a 3.87× acceleration in end-to-end wall-clock time on high-performance computing clusters. Overall, CAHS illustrates a practical AI-for-science design pattern: augmenting a trusted probabilistic model with query-specific learned signals to improve interpretable, reproducible inference in data-rich biology.

Bhattaram, Swethasree [Georgia Institute of Techno↗

Copper-dependent halogenase catalyses unactivated C−H bond functionalization

Carbon–hydrogen (C–H) bonds are the foundation of essentially every organic molecule, making them an ideal place to do chemical synthesis. The key challenge is achieving selectivity for one particular C(sp 3 )−H bond. In recent years, metalloenzymes have been found to perform C(sp 3 )−H bond functionalization. Despite substantial progresses in the past two decades, enzymatic halogenation and pseudohalogenation of unactivated C(sp 3 )−H—providing a functional handle for further modification—have been achieved with only non-haem iron/α-ketoglutarate-dependent halogenases, and are therefore limited by the chemistry possible with these enzymes. Here, in this work, we report the discovery and characterization of a previously unknown halogenase ApnU, part of a protein family containing domain of unknown function 3328 (DUF3328). ApnU uses copper in its active site to catalyse iterative chlorinations on multiple unactivated C(sp 3 )−H bonds. By taking advantage of the softer copper centre, we demonstrate that ApnU can catalyse unprecedented enzymatic C(sp 3 )−H bond functionalization such as iodination and thiocyanation. Using biochemical characterization and proteomics analysis, we identified the functional oligomeric state of ApnU as a covalently linked homodimer, which contains three essential pairs—one interchain and two intrachain—of disulfide bonds. The metal-coordination active site in ApnU consists of binuclear type II copper centres, as revealed by electron paramagnetic resonance spectroscopy. This discovery expands the enzymatic capability of C(sp 3 )−H halogenases and provides a foundational understanding of this family of binuclear copper-dependent oxidative enzymes.

biocatalysis↗

GenomeDepot: data management system for microbial comparative genomics

Summary GenomeDepot is an open-source web-based platform for annotation, management, and comparative analysis of microbial genomic sequences and associated data including ortholog families, protein domains, operons, regulatory interactions, strain taxonomy, and sample metadata. GenomeDepot supports rapid creation of websites for user-defined genome collections that include bioinformatic tools for interactive genome browsing, Basic Local Alignment Search Tool (BLAST) search, annotation search, comparative genomic neighborhood visualization, and sequence download. Gene function annotations are generated by a customizable annotation pipeline. The pipeline runs annotation tools in Conda environments and can be easily extended with additional user-specified tools. Availability and implementation GenomeDepot is open source and distributed under the GNU General Public License via GitHub (https://github.com/aekazakov/genome-depot). GenomeDepot is implemented in Python and was tested in Ubuntu Linux. Full installation instructions and documentation are available at https://aekazakov.github.io/genome-depot/. GenomeDepot demo server is freely accessible at https://iseq.lbl.gov/demogd/.

Kazakov, Alexey [Lawrence Berkeley National Labora↗

GenomeDepot v1.0

GenomeDepot is a web-based platform for annotation, management, and comparative analysis of microbial genomic sequences and associated data including ortholog families, protein domains, operons, regulatory interactions, strain taxonomy, and sample metadata. GenomeDepot supports rapid creation of web-sites for user-defined genome collections that include bioinformatic tools for interactive genome browsing, BLAST search, annotation search, comparative genomic neighborhood visualization, and sequence download. Gene function annotations are generated by a customizable annotation pipeline. The pipeline runs annotation tools in Conda environments and can be easily extended with additional user-specified tools.

Kazakov, Alexey [Lawrence Berkeley National Labora↗

Seed coat transcriptomic profiling of 5-593, a genotype important for genetic studies of seed coat color and patterning in common bean ( Phaseolus vulgaris L.)

Common bean (Phaseolus vulgaris L.) market classes have distinct seed coat colors, which are directly related to the diverse flavonoids found in the mature seed coat. To understand and elucidate the molecular mechanisms underlying the regulation of seed coat color, RNA-Seq data was collected from the black bean 5-593 and used for a differential gene expression and enrichment analysis from four different seed coat color development stages. 5-593 carries dominant alleles for 10 of the 11 major genes that control seed coat color and expression and has historically been used to develop introgression lines used for seed coat genetic analysis. Pairwise comparison among the four stages identified 6,294 differentially expressed genes (DEGs) varying from 508 to 5,780 DEGs depending on the compared stages. Kyoto Encyclopedia of Genes and Genomes (KEGG) enrichment analysis revealed that phenylpropanoid biosynthesis, flavonoid biosynthesis, and plant hormone signal transduction comprised the principal pathways expressed during bean seed coat pigment development. Transcriptome analysis suggested that most structural genes for flavonoid biosynthesis and some potential regulatory genes were significantly differentially expressed. Further studies detected 29 DEGs as important candidate genes governing the key enzymatic flavonoid biosynthetic pathways for common bean seed coat color development. Additionally, four gene models, Pv5-593.02G016100, 593.02G078700, Pv5-593.02G090900, and Pv5-593.06G121300, encode MYB-like transcription factor family protein were identified as strong candidate regulatory genes in anthocyanin biosynthesis which could regulate the expression levels of some important structural genes in flavonoid biosynthesis pathway. These findings provide a framework to draw new insights into the molecular networks underlying common bean seed coat pigment development.

60 APPLIED LIFE SCIENCES↗

Integrative Modeling and Analysis of Fungal Central Carbon Metabolism

Over a thousand fungal genomes have been sequenced, yet manually curated genome-scale metabolic models (GEMs) are available for only a limited number of species. Moreover, these models have often been developed independently, leading to inconsistencies in namespaces, compartment definitions, and pathway representations that hinder comparative analysis, the systematic reuse of prior curation efforts, and the integration of consolidated metabolic knowledge. Here, we present the Consolidated Fungal Core Metabolism Model (CFCMM), constructed by integrating thirteen published fungal models spanning Ascomycota, Mucoromycota, and both Crabtree-positive and Crabtree-negative yeasts. We harmonized metabolites and reactions into a non-redundant shared ModelSEED ontological space, standardized compartmentalization, and refined gene–protein–reaction (GPR) rules. Using pathway-level visualization and systematic gap detection, we further improved the integrated network through literature-guided curation to correct stoichiometry, stereospecificity, and pathway architecture. Orthologous protein family reconstruction and functional annotation workflows were used to validate and inform GPR associations, with particular emphasis on ambiguous enzyme superfamilies and membrane-associated components. Using the resulting CFCMM, we built high-quality central carbon core models for each fungus and performed flux balance analysis to quantify ATP-yield variation under aerobic and anaerobic conditions, explicitly evaluating scenarios driven by differences in electron transport chain (ETC) composition. Simulations reproduced the expected fermentative yield of approximately 2 mmol ATP per mmol glucose under anaerobic conditions and separated the thirteen fungi into two bioenergetic groups under aerobic respiration based on Complex I status, with predicted yields of approximately 30 versus 22 mmol ATP per mmol glucose. Forcing flux through the alternative oxidase bypass further reduced ATP yields to approximately 12 and 4 mmol ATP per mmol glucose in Complex I-containing and Complex I-lacking fungi, respectively. Collectively, this work provides a manually curated, ModelSEED-consistent, and extensible fungal core metabolic template, deployed in DOE KBase as a resource for automated reconstruction of central carbon core models from any sequenced fungal genome. In addition, the CFCMM provides modular components for developing GEMs with more accurate energy predictions and enables robust comparative analyses of fungal bioenergetics and core metabolic diversity

59 BASIC BIOLOGICAL SCIENCES↗