Search NASA⌕ Search

SEARCH · Search NASA

Results for “protein evolution”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

Substitution Models of Protein Evolution with Selection on Enzymatic Activity

Abstract Substitution models of evolution are necessary for diverse evolutionary analyses including phylogenetic tree and ancestral sequence reconstructions. At the protein level, empirical substitution models are traditionally used due to their simplicity, but they ignore the variability of substitution patterns among protein sites. Next, in order to improve the realism of the modeling of protein evolution, a series of structurally constrained substitution models were presented, but still they usually ignore constraints on the protein activity. Here, we present a substitution model of protein evolution with selection on both protein structure and enzymatic activity, and that can be applied to phylogenetics. In particular, the model considers the binding affinity of the enzyme–substrate complex as well as structural constraints that include the flexibility of structural flaps, hydrogen bonds, amino acids backbone radius of gyration, and solvent-accessible surface area that are quantified through molecular dynamics simulations. We applied the model to the HIV-1 protease and evaluated it by phylogenetic likelihood in comparison with the best-fitting empirical substitution model and a structurally constrained substitution model that ignores the enzymatic activity. We found that accounting for selection on the protein activity improves the fitting of the modeled functional regions with the real observations, especially in data with high molecular identity, which recommends considering constraints on the protein activity in the development of substitution models of evolution.

Ferreiro, David↗

Simultaneous enhancement of multiple functional properties using evolution-informed protein design

Abstract A major challenge in protein design is to augment existing functional proteins with multiple property enhancements. Altering several properties likely necessitates numerous primary sequence changes, and novel methods are needed to accurately predict combinations of mutations that maintain or enhance function. Models of sequence co-variation (e.g., EVcouplings), which leverage extensive information about various protein properties and activities from homologous protein sequences, have proven effective for many applications including structure determination and mutation effect prediction. We apply EVcouplings to computationally design variants of the model protein TEM-1 β -lactamase. Nearly all the 14 experimentally characterized designs were functional, including one with 84 mutations from the nearest natural homolog. The designs also had large increases in thermostability, increased activity on multiple substrates, and nearly identical structure to the wild type enzyme. This study highlights the efficacy of evolutionary models in guiding large sequence alterations to generate functional diversity for protein design applications.

59 BASIC BIOLOGICAL SCIENCES↗

Exploring the binding properties and activities of ancestral expansins

Bacterial expansins are non-lytic proteins capable of loosening cellulose networks, offering promising applications in agriculture, biotechnology, and material science. Their ability to disrupt noncovalent interactions in biopolymer matrices such as cellulose and chitin positions them as valuable tools for upgrading abundant natural materials. However, their industrial use remains limited due to their relatively low wall-loosening activity compared to plant expansins. To address this limitation, we applied Ancestral Sequence Resurrection (ASR) to reconstruct and characterize ancient variants of the Bacillus subtilis expansin BsEXLX1. ASR is a powerful evolutionary tool that enables the inference and synthesis of ancestral proteins, allowing researchers to explore functional traits that may have been lost over time. This approach not only provides insights into protein evolution but also facilitates the design of proteins with enhanced properties, such as improved substrate affinity or structural stability. In this study, we combined biochemical and biophysical assays to evaluate the activity and binding behavior of ancestral expansins. Our results reveal that ancestral variants exhibit increased cellulose affinity, reduced binding to acidic polysaccharides, and greater salt resistance. Furthermore, these traits enhance their wall-loosening activity and demonstrate the utility of ASR in engineering surface-active proteins for industrial applications, particularly in biomass processing and cellulose modification.

09 BIOMASS FUELS↗

Microbial sensor variation across biogeochemical conditions in the terrestrial deep subsurface

ABSTRACT Microbes can be found in abundance many kilometers underground. While microbial metabolic capabilities have been examined across different geochemical settings, it remains unclear how changes in subsurface niches affect microbial needs to sense and respond to their environment. To address this question, we examined how microbial extracellular sensor systems vary with environmental conditions across metagenomes at different Deep Mine Microbial Observatory (DeMMO) subsurface sites. Because two-component systems (TCSs) directly sense extracellular conditions and convert this information into intracellular biochemical responses, we expected that this sensor family would vary across isolated oligotrophic subterranean environments that differ in abiotic and biotic conditions. TCSs were found at all six subsurface sites, the service water control, and the surface site, with an average of 0.88 sensor histidine kinases (HKs) per 100 genes across all sites. Abundance was greater in subsurface fracture fluids compared with surface-derived fluids, and candidate phyla radiation (CPR) bacteria presented the lowest HK frequencies. Measures of microbial diversity, such as the Shannon diversity index, revealed that HK abundance is inversely correlated with microbial diversity ( r 2 = 0.81). Among the geochemical parameters measured, HK frequency correlated most strongly with variance in dissolved organic carbon ( r 2 = 0.82). Taken together, these results implicate the abiotic and biotic properties of an ecological niche as drivers of sensor needs, and they suggest that microbes in environments with large fluctuations in organic nutrients (e.g., lacustrine, terrestrial, and coastal ecosystems) may require greater TCS diversity than ecosystems with low nutrients (e.g., open ocean). IMPORTANCE The ability to detect extracellular environmental conditions is a fundamental property of all life forms. Because microbial two-component sensor systems convert information about extracellular conditions into biochemical information that controls their behaviors, we evaluated how two-component sensor systems evolved within the deep Earth across multiple sites where abiotic and biotic properties vary. We show that these sensor systems remain abundant in microbial consortia at all subterranean sampling sites and observe correlations between sensor system abundances and abiotic (dissolved organic carbon variation) and biotic (consortia diversity) properties. These results suggest that multiple environmental properties may drive sensor protein evolution and highlight the need for further studies of metagenomic and geochemical data in parallel to understand the drivers of microbial sensor evolution.

response regulator↗

Orange carotenoid proteins: structural understanding of evolution and function

Cyanobacteria uniquely contain a primitive water-soluble carotenoprotein, the orange carotenoid protein (OCP). Nearly all extant cyanobacterial genomes contain genes for the OCP or its homologs, implying an evolutionary constraint for cyanobacteria to conserve its function. Genes encoding the OCP and its two constituent structural domains, the N-terminal domain, helical carotenoid proteins (HCPs), and its C-terminal domain, are found in the most basal lineages of extant cyanobacteria. These three carotenoproteins exemplify the importance of the protein for carotenoid properties, including protein dynamics, in response to environmental changes in facilitating a photoresponse and energy quenching. Furthermore, we review new structural insights for these carotenoproteins and situate the role of the protein in what is currently understood about their functions.

59 BASIC BIOLOGICAL SCIENCES↗

Three rate-determining protein roles in photosynthetic O 2 -evolution addressed by time-resolved experiments on genetically modified photosystems

Light-driven water splitting by plants, algae and cyanobacteria is pivotal for global bioenergetics and biomass formation. A manganese cluster bound to the photosystem II proteins catalyzes the complex reaction at high rate, but the rate-determining factors are insufficiently understood. Here we trace the oxygen-evolution transition by time-resolved polarography and infrared spectroscopy for cyanobacterial photosystems genetically modified at two strategic sites, complemented by computational chemistry. Our results highlight three rate-determining roles of the protein environment of the metal cluster: acceleration of proton-coupled electron transfer, acceleration of substrate-water insertion after O 2 -formation, and balancing of rate-determining enthalpic and entropic contributions. Whereas in general the substrate-water insertion step may be unresolvable in time-resolved experiments, here it likely becomes traceable because of deceleration by genetic modification. Our results may stimulate new time-resolved experiments on substrate-water insertion in photosynthesis, clarification of enthalpy-entropy compensation in enzyme catalysis, and knowledge-guided development of inorganic catalyst materials.

Bioenergetics↗

The protein structurome of Orthornavirae and its dark matter

Metatranscriptomics is uncovering more and more diverse families of viruses with RNA genomes comprising the viral kingdom Orthornavirae in the realm Riboviria. Thorough protein annotation and comparison are essential to get insights into the functions of viral proteins and virus evolution. In addition to sequence- and hmm profile-based methods, protein structure comparison adds a powerful tool to uncover protein functions and relationships. We constructed an Orthornavirae “structurome” consisting of already annotated as well as unannotated (“dark matter”) proteins and domains encoded in viral genomes. We used protein structure modeling and similarity searches to illuminate the remaining dark matter in hundreds of thousands of orthornavirus genomes. The vast majority of the dark matter domains showed either “generic” folds, such as single α-helices, or no high confidence structure predictions. Nevertheless, a variety of lineage-specific globular domains that were new either to orthornaviruses in general or to particular virus families were identified within the proteomic dark matter of orthornaviruses, including several predicted nucleic acid-binding domains and nucleases. In addition, we identified a case of exaptation of a cellular nucleoside monophosphate kinase as an RNA-binding protein in several virus families. Notwithstanding the continuing discovery of numerous orthornaviruses, it appears that all the protein domains conserved in large groups of viruses have already been identified. The rest of the viral proteome seems to be dominated by poorly structured domains including intrinsically disordered ones that likely mediate specific virus-host interactions.

59 BASIC BIOLOGICAL SCIENCES↗

SARS-CoV-2 evolution balances conflicting roles of N protein phosphorylation

All lineages of SARS-CoV-2, the coronavirus responsible for the COVID-19 pandemic, contain mutations between amino acids 199 and 205 in the nucleocapsid (N) protein that are associated with increased infectivity. The effects of these mutations have been difficult to determine because N protein contributes to both viral replication and viral particle assembly during infection. Here, we used single-cycle infection and virus-like particle assays to show that N protein phosphorylation has opposing effects on viral assembly and genome replication. Ancestral SARS-CoV-2 N protein is densely phosphorylated, leading to higher levels of genome replication but 10-fold lower particle assembly compared to evolved variants with low N protein phosphorylation, such as Delta (N:R203M), Iota (N:S202R), and B.1.2 (N:P199L). A new open reading frame encoding a truncated N protein called N*, which occurs in the B.1.1 lineage and subsequent lineages of the Alpha, Gamma, and Omicron variants, supports high levels of both assembly and replication. Our findings help explain the enhanced fitness of viral variants of concern and a potential avenue for continued viral selection.

Microbiology↗

Birth of protein folds and functions in the virome

The rapid evolution of viruses generates proteins that are essential for infectivity and replication but with unknown functions, due to extreme sequence divergence. Here, using a database of 67,715 newly predicted protein structures from 4,463 eukaryotic viral species, we found that 62% of viral proteins are structurally distinct and lack homologues in the AlphaFold database. Among the remaining 38% of viral proteins, many have non-viral structural analogues that revealed surprising similarities between human pathogens and their eukaryotic hosts. Structural comparisons suggested putative functions for up to 25% of unannotated viral proteins, including those with roles in the evasion of innate immunity. In particular, RNA ligase T-like phosphodiesterases were found to resemble phage-encoded proteins that hydrolyse the host immune-activating cyclic dinucleotides 3',3'- and 2',3'-cyclic GMP-AMP (cGAMP). Experimental analysis showed that RNA ligase T homologues encoded by avian poxviruses similarly hydrolyse cGAMP, showing that RNA ligase T-mediated targeting of cGAMP is an evolutionarily conserved mechanism of immune evasion that is present in both bacteriophage and eukaryotic viruses. Together, the viral protein structural database and analyses presented here afford new opportunities to identify mechanisms of virus–host interactions that are common across the virome.

59 BASIC BIOLOGICAL SCIENCES↗

De novo design of pH-responsive self-assembling helical protein filaments

Abstract Biological evolution has led to precise and dynamic nanostructures that reconfigure in response to pH and other environmental conditions. However, designing micrometre-scale protein nanostructures that are environmentally responsive remains a challenge. Here we describe the de novo design of pH-responsive protein filaments built from subunits containing six or nine buried histidine residues that assemble into micrometre-scale, well-ordered fibres at neutral pH. The cryogenic electron microscopy structure of an optimized design is nearly identical to the computational design model for both the subunit internal geometry and the subunit packing into the fibre. Electron, fluorescent and atomic force microscopy characterization reveal a sharp and reversible transition from assembled to disassembled fibres over 0.3 pH units, and rapid fibre disassembly in less than 1 s following a drop in pH. The midpoint of the transition can be tuned by modulating buried histidine-containing hydrogen bond networks. Computational protein design thus provides a route to creating unbound nanomaterials that rapidly respond to small pH changes.

77 NANOSCIENCE AND NANOTECHNOLOGY↗

Bioelectrochemical crossbar architecture screening platform for extracellular electron transfer

Electroactive microbes can serve as living components in bioelectronic devices, where their unique ability to transfer electrons enables applications in sensing, energy conversion, and synthesis, but they remain challenging to engineer because the bioelectrochemical systems (BESs) used for characterization are low throughput. Here, we present a bioelectrochemical crossbar architecture screening platform (BiCASP) that uses stacked and orthogonally arrayed electrodes to enable individual sample selection for characterization in arrayed formats. This device reports on the current generated by electroactive bacteria on the minute timescale, decreasing the time for data acquisition by several orders of magnitude compared to conventional BESs. This device increases the throughput of screening engineered biological components in cells, identifying mutants of the membrane protein wire MtrA in Shewanella oneidensis that retain the ability to support extracellular electron transfer (EET). BiCASP may be integrated with bioelectronics that need directed evolution of electroactive proteins.

Shewanella↗

A potential role for RNA aminoacylation prior to its role in peptide synthesis

Coded ribosomal peptide synthesis could not have evolved unless its sequence and amino acid–specific aminoacylated tRNA substrates already existed. We therefore wondered whether aminoacylated RNAs might have served some primordial function prior to their role in protein synthesis. Here, we show that specific RNA sequences can be nonenzymatically aminoacylated and ligated to produce amino acid–bridged stem-loop RNAs. We used deep sequencing to identify RNAs that undergo highly efficient glycine aminoacylation followed by loop-closing ligation. The crystal structure of one such glycine-bridged RNA hairpin reveals a compact internally stabilized structure with the same eponymous T-loop architecture that is found in many noncoding RNAs, including the modern tRNA. We demonstrate that the T-loop-assisted amino acid bridging of RNA oligonucleotides enables the rapid template-free assembly of a chimeric version of an aminoacyl-RNA synthetase ribozyme. We suggest that the primordial assembly of amino acid–bridged chimeric ribozymes provides a direct and facile route for the covalent incorporation of amino acids into RNA. A greater functionality of covalently incorporated amino acids could contribute to enhanced ribozyme catalysis, providing a driving force for the evolution of sequence and amino acid–specific aminoacyl-RNA synthetase ribozymes in the RNA World. The synthesis of specifically aminoacylated RNAs, an unlikely prospect for nonenzymatic reactions but a likely one for ribozymes, could have set the stage for the subsequent evolution of coded protein synthesis.

Science & Technology - Other Topics↗

Conformational Dynamics and Catalytic Backups in a Hyper-thermostable Engineered Archaeal Protein Tyrosine Phosphatase

Protein tyrosine phosphatases (PTPs) are a family of enzymes that play important roles in regulating cellular signaling pathways. The activity of these enzymes is regulated by the motion of a catalytic loop that places a critical conserved aspartic acid side chain into the active site for acid–base catalysis upon loop closure. These enzymes also have a conserved phosphate-binding loop that is typically highly rigid and forms a well-defined anion-binding nest. The intimate links between loop dynamics and chemistry in these enzymes make PTPs an excellent model system for understanding the role of loop dynamics in protein function and evolution. In this context, archaeal PTPs, which have often evolved in extremophilic organisms, are highly understudied, despite their unusual biophysical properties. We present here an engineered chimeric PTP (ShufPTP) generated by shuffling the amino acid sequence of five extant hyperthermophilic archaeal PTPs. Despite ShufPTP’s high sequence similarity to its natural counterparts, it presents a suite of unique properties, including high flexibility of the phosphate binding P-loop, facile oxidation of the active-site cysteine, mechanistic promiscuity, and, most notably, hyperthermostability, with a denaturation temperature likely >130 °C (>8 °C higher than the highest recorded growth temperature of any archaeal strain). Our combined structural, biochemical, biophysical, and computational analysis provides insight both into how small steps in evolutionary space can radically modulate the biophysical properties of an enzyme and showcases the tremendous potential of archaeal enzymes for biotechnology, to generate novel enzymes capable of operating under extreme conditions.

archaea↗

A [FeFe] Hydrogenase–Rubrerythrin Chimeric Enzyme Functions to Couple H 2 Oxidation to Reduction of H 2 O 2 in the Foodborne Pathogen Clostridium perfringens

[FeFe] hydrogenases are a diverse class of H 2 -activating enzymes with a wide range of utilities in nature. As H 2 is a promising renewable energy carrier, exploration of the increasingly realized functional diversity of [FeFe] hydrogenases is instrumental for understanding how these remarkable enzymes can benefit society and inspire new technologies. In this work, we uncover the properties of a highly unusual natural chimera composed of a [FeFe] hydrogenase and rubrerythrin as a single polypeptide. The unique combination of [FeFe] hydrogenase with rubrerythrin, an enzyme that functions in H 2 O 2 detoxification, raises the question of whether catalytic reactions, such as H 2 oxidation and H 2 O 2 reduction, are functionally linked. Herein, we express and purify a representative chimera from Clostridium perfringens (termed Cper HydR) and apply various electrochemical and spectroscopic approaches to determine its activity and confirm the presence of each of the proposed metallocofactors. The cumulative data demonstrate that the enzyme contains a surprising array of metallocofactors: the catalytic site of [FeFe] hydrogenase termed the H-cluster, two [4Fe-4S] clusters, two rubredoxin Fe(Cys) 4 centers, and a hemerythrin-like diiron site. The absence of an H 2 -evolution current in protein film voltammetry highlights an exceptional bias of this enzyme toward H 2 oxidation to the greatest extent that has been observed for a [FeFe] hydrogenase. Here, we demonstrate that Cper HydR uses H 2 , catalytically split by the hydrogenase domain, to reduce H 2 O 2 by the diiron site. Structural modeling suggests a homodimeric nature of the protein. Overall, this study demonstrates that Cper HydR is an H 2 -dependent H 2 O 2 reductase. Equipped with this information, we discuss the possible role of this enzyme as a part of the oxygen-stress response system, proposing that Cper HydR constitutes a new pathway for H 2 O 2 mitigation.

08 HYDROGEN↗

The evolution of analytical techniques for multiplex analysis of protein biomarkers

Introduction: The landscape of biomarker development has evolved with advanced analytical technologies, particularly affinity- and mass spectrometry-based techniques. These advancements have deepened our understanding of disease mechanisms, enabling the development of precise diagnostic tools and personalized medicine. Protein biomarkers, which play pivotal roles in biological processes, have become invaluable in diagnosing and monitoring diseases, aided by their presence in various biological samples and the availability of established detection methods. Areas covered: This review covers the role of protein biomarkers in clinical practice, the development and dimensionality of protein biomarkers, advancements in detection technologies, a comparison of these technologies, and future directions in biomarker discovery and disease mechanism elucidation. Expert opinion: Advances in biomarker technologies have the potential to transform diagnostics and personalized treatment but face challenges such as high costs and technical complexity. Enhancing reproducibility and integrating multi-omics approaches may offer better insights. In conclusion, the field should evolve toward high-throughput, automated methods, continuously adapting research, and clinical practices.

59 BASIC BIOLOGICAL SCIENCES↗

Specialization Restricts the Evolutionary Paths Available to Yeast Sugar Transporters

Functional innovation at the protein level is a key source of evolutionary novelties. The constraints on functional innovations are likely to be highly specific in different proteins, which are shaped by their unique histories and the extent of global epistasis that arises from their structures and biochemistries. These contextual nuances in the sequence–function relationship have implications both for a basic understanding of the evolutionary process and for engineering proteins with desirable properties. Here, we have investigated the molecular basis of novel function in a model member of an ancient, conserved, and biotechnologically relevant protein family. These Major Facilitator Superfamily sugar porters are a functionally diverse group of proteins that are thought to be highly plastic and evolvable. By dissecting a recent evolutionary innovation in an α-glucoside transporter from the yeast Saccharomyces eubayanus, we show that the ability to transport a novel substrate requires high-order interactions between many protein regions and numerous specific residues proximal to the transport channel. To reconcile the functional diversity of this family with the constrained evolution of this model protein, we generated new, state-of-the-art genome annotations for 332 Saccharomycotina yeast species spanning ~400 My of evolution. By integrating phylogenetic and phenotypic analyses across these species, we show that the model yeast α-glucoside transporters likely evolved from a multifunctional ancestor and became subfunctionalized. The accumulation of additive and epistatic substitutions likely entrenched this subfunction, which made the simultaneous acquisition of multiple interacting substitutions the only reasonably accessible path to novelty.

59 BASIC BIOLOGICAL SCIENCES↗

Designing Peptide Fossils That Model the Evolution of the Bacterial Ferredoxin Fold

Electron transfer coupled to redox chemistry is at the heart of metabolism. The proteins responsible for moving electrons (protein electron carriers) must have emerged at the origin of life. The small iron–sulfur-binding bacterial ferredoxins were likely among these first proteins. Embedded within the ferredoxin sequence and structure is a symmetry that points to an ancient gene duplication event. Little is understood about the nature of ferredoxins prior to this duplication event or what environmental factors may have driven the selection for more complex forms. The deep-time molecular history of ferredoxins goes back billions of years and cannot be reconstructed by phylogenetic analyses based on amino acid sequences. Here, we use structure-guided protein design to model a fossil half-ferredoxin stage in the evolution of this fold, the semidoxins, and their symmetric full-length counterparts, the symdoxins. Semidoxin designs homodimerize, exhibiting structural, thermodynamic, and electrochemical behaviors in most cases identical to cognate symdoxins. However, the semi- and symdoxin fossil stages behave differently when incorporated into an in vivo electron transfer complementation assay. Both can support bacterial growth dependent on protein expression. Growth rates of bacteria expressing the semidoxins are much more sensitive to oxygen than those of bacteria expressing symdoxins. Motivated by the in vivo functionality of designed semidoxins, we identified putative naturally occurring semidoxins in extant anaerobic microorganisms. This is consistent with the observed in vivo oxygen sensitivity of the semidoxin designs. One natural semidoxin is shown to be folded and redox active. However, it exists as a mixture of monomers and dimers, suggesting a potential connection between semidoxins and even simpler single iron–sulfur cluster-binding peptides.

59 BASIC BIOLOGICAL SCIENCES↗

LevSeq: Rapid Generation of Sequence-Function Data for Directed Evolution and Machine Learning

Sequence-function data provides valuable information about the protein functional landscape but is rarely obtained during directed evolution campaigns. Here, we present Long-read every variant Sequencing (LevSeq), a pipeline that combines a dual barcoding strategy with nanopore sequencing to rapidly generate sequence-function data for entire protein-coding genes. LevSeq integrates into existing protein engineering workflows and comes with open-source software for data analysis and visualization. The pipeline facilitates data-driven protein engineering by consolidating sequence-function data to inform directed evolution and provide the requisite data for machine learning-guided protein engineering (MLPE). LevSeq enables quality control of mutagenesis libraries prior to screening, which reduces time and resource costs. Simulation studies demonstrate LevSeq’s ability to accurately detect variants under various experimental conditions. Lastly, we show LevSeq’s utility in engineering protoglobins for new-to-nature chemistry. Widespread adoption of LevSeq and sharing of the data will enhance our understanding of protein sequence-function landscapes and empower data-driven directed evolution.

59 BASIC BIOLOGICAL SCIENCES↗