Search NASA⌕ Search

SEARCH · Search NASA

Results for “fitness landscape”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

A combinatorially complete epistatic fitness landscape in an enzyme active site

Protein engineering often targets binding pockets or active sites which are enriched in epistasis—nonadditive interactions between amino acid substitutions—and where the combined effects of multiple single substitutions are difficult to predict. Few existing sequence-fitness datasets capture epistasis at large scale, especially for enzyme catalysis, limiting the development and assessment of model-guided enzyme engineering approaches. We present here a combinatorially complete, 160,000-variant fitness landscape across four residues in the active site of an enzyme. Assaying the native reaction of a thermostable β-subunit of tryptophan synthase (TrpB) in a nonnative environment yielded a landscape characterized by significant epistasis and many local optima. These effects prevent simulated directed evolution approaches from efficiently reaching the global optimum. There is nonetheless wide variability in the effectiveness of different directed evolution approaches, which together provide experimental benchmarks for computational and machine learning workflows. The most-fit TrpB variants contain a substitution that is nearly absent in natural TrpB sequences—a result that conservation-based predictions would not capture. Thus, although fitness prediction using evolutionary data can enrich in more-active variants, these approaches struggle to identify and differentiate among the most-active variants, even for this near-native function. Overall, this work presents a large-scale testing ground for model-guided enzyme engineering and suggests that efficient navigation of epistatic fitness landscapes can be improved by advances in both machine learning and physical modeling.

biocatalysis↗

Self-driving laboratories to autonomously navigate the protein fitness landscape

Abstract Protein engineering has nearly limitless applications across chemistry, energy and medicine, but creating new proteins with improved or novel functions remains slow, labor-intensive and inefficient. Here we present the Self-driving Autonomous Machines for Protein Landscape Exploration (SAMPLE) platform for fully autonomous protein engineering. SAMPLE is driven by an intelligent agent that learns protein sequence–function relationships, designs new proteins and sends designs to a fully automated robotic system that experimentally tests the designed proteins and provides feedback to improve the agent’s understanding of the system. We deploy four SAMPLE agents with the goal of engineering glycoside hydrolase enzymes with enhanced thermal tolerance. Despite showing individual differences in their search behavior, all four agents quickly converge on thermostable enzymes. Self-driving laboratories automate and accelerate the scientific discovery process and hold great potential for the fields of protein engineering and synthetic biology.

Rapp, Jacob T. (ORCID:0000000226835204)↗

Challenges and Solutions for Leave-One-Out Biosensor Design in the Context of a Rugged Fitness Landscape

The leave-one-out (LOO) green fluorescent protein (GFP) approach to biosensor design combines computational protein design with split protein reconstitution. LOO-GFPs reversibly fold and gain fluorescence upon encountering the target peptide, which can be redefined by computational design of the LOO site. Such an approach can be used to create reusable biosensors for the early detection of emerging biological threats. Enlightening biophysical inferences for nine LOO-GFP biosensor libraries are presented, with target sequences from dengue, influenza, or HIV, replacing beta strands 7, 8, or 11. An initially low hit rate was traced to components of the energy function, manifesting in the over-rewarding of over-tight side chain packing. Also, screening by colony picking required a low library complexity, but designing a biosensor against a peptide of at least 12 residues requires a high-complexity library. This double-bind was solved using a “piecemeal” iterative design strategy. Also, designed LOO-GFPs fluoresced in the unbound state due to unwanted dimerization, but this was solved by fusing a fully functional prototype LOO-GFP to a fiber-forming protein, Drosophila ultrabithorax, creating a biosensor fiber. One influenza hemagglutinin biosensor is characterized here in detail, showing a shifted excitation/emission spectrum, a micromolar affinity for the target peptide, and an unexpected photo-switching ability.

Chemistry↗

Fitness Landscape-Guided Engineering of Locally Supercharged Virus-like Particles with Enhanced Cell Uptake Properties

Protein-based nanoparticles are useful models for the study of self-assembly and attractive candidates for drug delivery. Virus-like particles (VLPs) are especially promising platforms for expanding the repertoire of therapeutics that can be delivered effectively as they can deliver many copies of a molecule per particle for each delivery event. However, their use is often limited due to poor uptake of VLPs into mammalian cells. In this study, we use the fitness landscape of the bacteriophage MS2 VLP as a guide to engineer capsid variants with positively charged surface residues to enhance their uptake into mammalian cells. By combining mutations with positive fitness scores that were likely to produce assembled capsids, we identified two key double mutants with internalization efficiencies as much as 67-fold higher than that of wtMS2. Internalization of these variants with positively charged surface residues depends on interactions with cell surface sulfated proteoglycans, and yet, they are biophysically similar to wtMS2 with low cytotoxicity and an overall negative charge. Additionally, the best-performing engineered MS2 capsids can deliver a potent anticancer small-molecule therapeutic with efficacy levels similar to antibody-drug conjugates. Through this work, we were able to establish fitness landscape-based engineering as a successful method for designing VLPs with improved cell penetration. These findings suggest that VLPs with positive surface charge could be useful in improving the delivery of small-molecule- and nucleic acid-based therapeutics.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Exploring Connectivity in Sequence Space of Functional RNA

Emergence of replicable genetic molecules was one of the marking points in the origin of life, evolution of which can be conceptualized as a walk through the space of all possible sequences. A theoretical concept of fitness landscape helps to understand evolutionary processes through assigning a value of fitness to each genotype. Then, evolution of a phenotype is viewed as a series of consecutive, single-point mutations. Natural selection biases evolution toward peaks of high fitness and away from valleys of low fitness. whereas neutral drift occurs in the sequence space without direction as mutations are introduced at random. Large networks of neutral or near-neutral mutations on a fitness landscape, especially for sufficiently long genomes, are possible or even inevitable. Their detection in experiments, however, has been elusive. Although a few near-neutral evolutionary pathways have been found, recent experimental evidence indicates landscapes consist of largely isolated islands. The generality of these results, however, is not clear, as the genome length or the fraction of functional molecules in the genotypic space might have been insufficient for the emergence of large, neutral networks. Thorough investigation on the structure of the fitness landscape is essential to understand the mechanisms of evolution of early genomes. RNA molecules are commonly assumed to play the pivotal role in the origin of genetic systems. They are widely believed to be early, if not the earliest, genetic and catalytic molecules, with abundant biochemical activities as aptamers and ribozymes, i.e. RNA molecules capable, respectively, to bind small molecules or catalyze chemical reactions. Here, we present results of our recent studies on the structure of the sequence space of RNA ligase ribozymes selected through in vitro evolution. Several hundred thousands of sequences active to a different degree were obtained by way of deep sequencing. Analysis of these sequences revealed several large clusters defined such that every sequence in a cluster can be reached from any other sequence in the same cluster through a series of single point mutations. Sequences in a single cluster appear to adopt more than one secondary structure. The mechanism of refolding within a single cluster was examined. To shed light on possible evolutionary paths in the space of ribozymes, the connectivity between clusters was investigated. The effect of length of RNA molecules on the structure of the fitness landscape and possible evolutionary paths was examined by way of comparing functional sequences of 20 and 80 nucleobases in length. It was found that sequences of different lengths shared secondary structure motifs that were presumed responsible for catalytic activity, with increasing complexity and global structural rearrangements emerging in longer molecules.

Wei, Chenyu↗

Opportunities and Challenges for Machine Learning-Assisted Enzyme Engineering

Enzymes can be engineered at the level of their amino acid sequences to optimize key properties such as expression, stability, substrate range, and catalytic efficiency or even to unlock new catalytic activities not found in nature. Because the search space of possible proteins is vast, enzyme engineering usually involves discovering an enzyme starting point that has some level of the desired activity followed by directed evolution to improve its “fitness” for a desired application. Recently, machine learning (ML) has emerged as a powerful tool to complement this empirical process. ML models can contribute to (1) starting point discovery by functional annotation of known protein sequences or generating novel protein sequences with desired functions and (2) navigating protein fitness landscapes for fitness optimization by learning mappings between protein sequences and their associated fitness values. In this Outlook, we explain how ML complements enzyme engineering and discuss its future potential to unlock improved engineering outcomes.

60 APPLIED LIFE SCIENCES↗

Dosage optimization for reducing tumor burden using a phenotype-structured population model with a drug-resistance continuum

Abstract Drug resistance is a significant obstacle to effective cancer treatment. To gain insights into how drug resistance develops, we adopted a concept called fitness landscape and employed a phenotype-structured population model by fitting to a set of experimental data on a drug used for ovarian cancer, olaparib. Our modeling approach allowed us to understand how a drug affects the fitness landscape and track the evolution of a population of cancer cells structured with a spectrum of drug resistance. We also incorporated pharmacokinetic (PK) modeling to identify the optimal dosages of the drug that could lead to long-term tumor reduction. We derived a formula that indicates that maximizing variation in plasma drug concentration over a dosing interval could be important in reducing drug resistance. Our findings suggest that it may be possible to achieve better treatment outcomes with a drug dose lower than the levels recommended by the drug label. Acknowledging the current limitations of our work, we believe that our approach, which combines modeling of both PK and drug resistance evolution, could contribute to a new direction for better designing drug treatment regimens to improve cancer treatment.

Life Sciences & Biomedicine - Other Topics↗

“Multiagent” Screening Improves Directed Enzyme Evolution by Identifying Epistatic Mutations

Enzyme evolution has enabled numerous advances in biotechnology and synthetic biology, yet still requires many iterative rounds of screening to identify optimal mutant sequences. This is due to the sparsity of the fitness landscape, which is caused by epistatic mutations that only offer improvements when combined with other mutations. We report an approach that incorporates diverse substrate analogues in the screening process, where multiple substrates act like multiple agents navigating the fitness landscape, identifying epistatic mutant residues without a need for testing the entire combinatorial search space. We initially validate this approach by engineering a malonyl-CoA synthetase and identify numerous epistatic mutations improving activity for several diverse substrates. The majority of these mutations would have been missed upon screening for a single substrate alone. We expect that this approach can accelerate a wide array of enzyme engineering programs.

60 APPLIED LIFE SCIENCES↗

Multi-objective optimization of root phenotypes for nutrient capture using evolutionary algorithms

Root phenotypes are avenues to the development of crop cultivars with improved nutrient capture, which is an important goal for global agriculture. The fitness landscape of root phenotypes is highly complex and multidimensional. It is difficult to predict which combinations of traits (phene states) will create the best performing integrated phenotypes in various environments. Brute force methods to map the fitness landscape by simulating millions of phenotypes in multiple environments are computationally challenging. Evolutionary optimization algorithms may provide more efficient avenues to explore high dimensional domains such as the root phenotypic space. We coupled the three-dimensional functional–structural plant model, SimRoot, to the Borg Multi-Objective Evolutionary Algorithm (MOEA) and the evolutionary search over several generations facilitated the identification of optimal root phenotypes balancing trade-offs across nutrient uptake, biomass accumulation, and root carbon costs in environments varying in nutrient availability. Our results show that several combinations of root phenes generate optimal integrated phenotypes where performance in one objective comes at the cost of reduced performance in one or more of the remaining objectives, and such combinations differed for mobile and non-mobile nutrients and for maize (a monocot) and bean (a dicot). Functional–structural plant models can be used with multi-objective optimization to identify optimal root phenotypes under various environments, including future climate scenarios, which will be useful in developing the more resilient, efficient crops urgently needed in global agriculture.

59 BASIC BIOLOGICAL SCIENCES↗

How microscopic epistasis and clonal interference shape the fitness trajectory in a spin glass model of microbial long-term evolution

The adaptive dynamics of evolving microbial populations takes place on a complex fitness landscape generated by epistatic interactions. The population generically consists of multiple competing strains, a phenomenon known as clonal interference. Microscopic epistasis and clonal interference are central aspects of evolution in microbes, but their combined effects on the functional form of the population’s mean fitness are poorly understood. Here, we develop a computational method that resolves the full microscopic complexity of a simulated evolving population subject to a standard serial dilution protocol. Through extensive numerical experimentation, we find that stronger microscopic epistasis gives rise to fitness trajectories with slower growth independent of the number of competing strains, which we quantify with power-law fits and understand mechanistically via a random walk model that neglects dynamical correlations between genes. We show that increasing the level of clonal interference leads to fitness trajectories with faster growth (in functional form) without microscopic epistasis, but leaves the rate of growth invariant when epistasis is sufficiently strong, indicating that the role of clonal interference depends intimately on the underlying fitness landscape. The simulation package for this work may be found at https://github.com/nmboffi/spin_glass_evodyn .

59 BASIC BIOLOGICAL SCIENCES↗

How microscopic epistasis and clonal interference shape the fitness trajectory in a spin glass model of microbial long-term evolution

The adaptive dynamics of evolving microbial populations takes place on a complex fitness landscape generated by epistatic interactions. The population generically consists of multiple competing strains, a phenomenon known as clonal interference. Microscopic epistasis and clonal interference are central aspects of evolution in microbes, but their combined effects on the functional form of the population’s mean fitness are poorly understood. Here, we develop a computational method that resolves the full microscopic complexity of a simulated evolving population subject to a standard serial dilution protocol. Through extensive numerical experimentation, we find that stronger microscopic epistasis gives rise to fitness trajectories with slower growth independent of the number of competing strains, which we quantify with power-law fits and understand mechanistically via a random walk model that neglects dynamical correlations between genes. We show that increasing the level of clonal interference leads to fitness trajectories with faster growth (in functional form) without microscopic epistasis, but leaves the rate of growth invariant when epistasis is sufficiently strong, indicating that the role of clonal interference depends intimately on the underlying fitness landscape. The simulation package for this work may be found at https://github.com/nmboffi/spin_glass_evodyn .

59 BASIC BIOLOGICAL SCIENCES↗

ALPS: The Age-Layered Population Structure for Reducing the Problem of Premature Convergence

To reduce the problem of premature convergence we define a new attribute of an individual, its age, and propose the Age-Layered Population Structure (ALPS), in which age is used to restrict competition and breeding between members of the population. ALPS differs from a typical EA by segregating individuals into different age-layers by their age - a measure of how long the genetic material has been in the population - and by regularly replacing all individuals in the bottom layer with randomly generated ones. The introduction of new, randomly generated individuals at regular intervals results in an EA that is never completely converged and is always looking at new parts of the fitness landscape. By using age to restrict competition and breeding search is able to develop promising young individuals without them being dominated by older ones. We demonstrate the effectiveness of the ALPS algorithm on an antenna design problem in which evolution with ALPS produces antennas more than twice as good as does evolution with two other types of EAs. Further analysis shows that the ALPS model does allow the offspring of newly generated individuals to move the population out of mediocre local-optima to better parts of the fitness landscape.

Hornby, Gregory S.↗

Ensemble Monte Carlo calculations with five novel moves

We introduce five novel types of Monte Carlo (MC) moves that brings the number of moves of ensemble MC calculations from three to eight. So far such calculations have relied on affine invariant stretch moves that were originally introduced by Christen (2007), walk moves by Goodman and Weare (2010) and quadratic moves by Militzer (2023). Ensemble MC methods have been very popular because they harness information about the fitness landscape from a population of walkers rather than relying on expert knowledge. Here we modified the affine method and employed a simplex of points to set the stretch direction. We adopt the simplex concept to quadratic moves. We also generalize quadratic moves to arbitrary order. Finally, we introduce directed moves that employ the values of the probability density while all other types of moves rely solely on the location of the walkers. We apply all algorithms to the Rosenbrock density in 2 and 20 dimensions and to the ring potential in 12 and 24 dimensions. We evaluate their efficiency by comparing error bars, autocorrelation time, travel time, and the level of cohesion that measures whether any walkers were left behind. Our code is open source.

97 MATHEMATICS AND COMPUTING↗

Energetic Basis and Design of Enzyme Function Demonstrated Using GFP, an Excited-State Enzyme

We report the past decades have witnessed an explosion of de novo protein designs with a remarkable range of scaffolds. It remains challenging, however, to design catalytic functions that are competitive with naturally occurring counterparts as well as biomimetic or nonbiological catalysts. Although directed evolution often offers efficient solutions, the fitness landscape remains opaque. Green fluorescent protein (GFP), which has revolutionized biological imaging and assays, is one of the most redesigned proteins. While not an enzyme in the conventional sense, GFPs feature competing excited-state decay pathways with the same steric and electrostatic origins as conventional ground-state catalysts, and they exert exquisite control over multiple reaction outcomes through the same principles. Thus, GFP is an “excited-state enzyme”. Herein we show that rationally designed mutants and hybrids that contain environmental mutations and substituted chromophores provide the basis for a quantitative model and prediction that describes the influence of sterics and electrostatics on excited-state catalysis of GFPs. As both perturbations can selectively bias photoisomerization pathways, GFPs with fluorescence quantum yields (FQYs) and photoswitching characteristics tailored for specific applications could be predicted and then demonstrated. The underlying energetic landscape, readily accessible via spectroscopy for GFPs, offers an important missing link in the design of protein function that is generalizable to catalyst design.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Nickase fidelity drives EvolvR-mediated diversification in mammalian cells

Abstract In vivo genetic diversifiers have previously enabled efficient searches of genetic variant fitness landscapes for continuous directed evolution. However, existing genomic diversification modalities for mammalian genomic loci exclusively rely on deaminases to generate transition mutations within target loci, forfeiting access to most missense mutations. Here, we engineer CRISPR-guided error-prone DNA polymerases (EvolvR) to diversify all four nucleotides within genomic loci in mammalian cells. We demonstrate that EvolvR generates both transition and transversion mutations throughout a mutation window of at least 40 bp and implement EvolvR to evolve previously unreported drug-resistantMAP2K1variants via substitutions not achievable with deaminases. Moreover, we discover that the nickase’s mismatch tolerance limits EvolvR’s mutation window and substitution biases in a gRNA-specific fashion. To compensate for gRNA-to-gRNA variability in mutagenesis, we maximize the number of gRNA target sequences by incorporating a PAM-flexible nickase into EvolvR. Finally, we find a strong correlation between predicted free energy changes underlying R-loop formation and EvolvR’s performance using a given gRNA. The EvolvR system diversifies all four nucleotides to enable the evolution of mammalian cells, while nuclease and gRNA-specific properties underlying nickase fidelity can be engineered to further enhance EvolvR’s mutation rates.

Science & Technology - Other Topics↗