Search NASASearch

SEARCH · Search NASA

Results for “Protein modeling”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 127 records · Page 7

Enzyme property prediction using artificial intelligence

Artificial intelligence (AI)-driven enzyme property prediction enables rapid discovery and engineering of enzymes for a wide range of biotechnological and therapeutic applications. Here, we first introduce the key components in AI model development, including enzyme datasets, protein representation methods, and model architectures. We then highlight a variety of AI tools developed for the prediction of enzyme properties and functional annotations, including enzyme structure, kinetic parameters, substrate specificity, thermostability, solubility, Enzyme Commission number, and Gene Ontology term. Moreover, we describe representative downstream applications enabled by these AI tools. Finally, we discuss some challenges and opportunities as well as future prospects.

Yuan, Le [University of Illinois at Urbana-Champai

Enzymatic carbon–fluorine bond cleavage by human gut microbes

Fluorinated compounds are used for agrochemical, pharmaceutical, and numerous industrial applications, resulting in global contamination. In many molecules, fluorine is incorporated to enhance the half-life and improve bioavailability. Fluorinated compounds enter the human body through food, water, and xenobiotics including pharmaceuticals, exposing gut microbes to these substances. The human gut microbiota is known for its xenobiotic biotransformation capabilities, but it was not previously known whether gut microbial enzymes could break carbon-fluorine bonds, potentially altering the toxicity of these compounds. Here, through the development of a rapid, miniaturized fluoride detection assay for whole-cell screening, we identified active gut microbial defluorinases. We biochemically characterized enzymes from diverse human gut microbial classes including Clostridia, Bacilli, and Coriobacteriia, with the capacity to hydrolyze (di)fluorinated organic acids and a fluorinated amino acid. Whole-protein alanine scanning, molecular dynamics simulations, and chimeric protein design enabled the identification of a disordered C-terminal protein segment involved in defluorination activity. Domain swapping exclusively of the C-terminus conferred defluorination activity to a nondefluorinating dehalogenase. To advance our understanding of the structural and sequence differences between defluorinating and nondefluorinating dehalogenases, we trained machine learning models which identified protein termini as important features. Models trained on 41-amino acid segments from protein C termini alone predicted defluorination activity with 83% accuracy (compared to 95% accuracy based on full-length protein features). This work is relevant for therapeutic interventions and environmental and human health by uncovering specificity-determining signatures of fluorine biochemistry from the gut microbiome.

Probst, Silke I

Specialization Restricts the Evolutionary Paths Available to Yeast Sugar Transporters

Functional innovation at the protein level is a key source of evolutionary novelties. The constraints on functional innovations are likely to be highly specific in different proteins, which are shaped by their unique histories and the extent of global epistasis that arises from their structures and biochemistries. These contextual nuances in the sequence–function relationship have implications both for a basic understanding of the evolutionary process and for engineering proteins with desirable properties. Here, we have investigated the molecular basis of novel function in a model member of an ancient, conserved, and biotechnologically relevant protein family. These Major Facilitator Superfamily sugar porters are a functionally diverse group of proteins that are thought to be highly plastic and evolvable. By dissecting a recent evolutionary innovation in an α-glucoside transporter from the yeast Saccharomyces eubayanus, we show that the ability to transport a novel substrate requires high-order interactions between many protein regions and numerous specific residues proximal to the transport channel. To reconcile the functional diversity of this family with the constrained evolution of this model protein, we generated new, state-of-the-art genome annotations for 332 Saccharomycotina yeast species spanning ~400 My of evolution. By integrating phylogenetic and phenotypic analyses across these species, we show that the model yeast α-glucoside transporters likely evolved from a multifunctional ancestor and became subfunctionalized. The accumulation of additive and epistatic substitutions likely entrenched this subfunction, which made the simultaneous acquisition of multiple interacting substitutions the only reasonably accessible path to novelty.

59 BASIC BIOLOGICAL SCIENCES

SLAB: simultaneous labeling and binding affinity prediction for protein–ligand structures

Machine learning models are often used as scoring functions to predict the binding affinity of a protein–ligand complex. These models are trained with limited amounts of data with experimentally measured binding affinity values. A large number of compounds are labeled inactive through single-concentration screens without measuring binding affinities. These inactive compounds, along with the active ones, can be used to train binary classification models, while regression models are trained using compounds with binding affinities only. However, the classification and regression tasks are often handled separately, without sharing the learned feature representations. In this paper, we propose a novel model architecture that jointly performs regression and classification objectives, aiming to maximize data utilization and improve predictive performance by leveraging two complementary tasks. In our setup, the regression yields the binding affinity, whereas the classification task yields the label as active or inactive. We demonstrate our method using PDBbind, the standard 3D structure database, as well as a dataset of flavivirus protease compounds with binding affinity data. Our experiments show that the new joint training strategy improves the accuracy of the model, increasing applicability in various practical drug screening scenarios.

Biological and medical sciences

Energetic and structural control of polyspecificity in a multidrug transporter

Multidrug efflux pumps are dynamic molecular machines that drive antibiotic resistance by harnessing ion gradients to export chemically diverse substrates. Despite their clinical importance, the molecular principles underlying multidrug promiscuity and energy efficiency remain poorly understood. Using multiparametric deep mutational scanning across eight substrates and two energy conditions, we deconvolute the contributions of substrate recognition, energetic coupling, and protein stability, providing an integrated, high-resolution view of multidrug transport. We find that substrate specificity arises from a distributed network of residues extending beyond the binding site, with mutations that reshape binding, coupling, conformational flexibility, and membrane interactions. Further, we apply a pH-based selection scheme to measure the effect of mutation on pH-dependent transport efficiency. By integrating these data, we reveal a fundamental relationship between efficiency and promiscuity: Highly efficient variants exhibit broad substrate profiles, while inefficient variants are narrower. In conclusion, these findings establish a direct link between energy coupling and polyspecificity, uncovering the biochemical logic underlying multidrug transport.

Biological Sciences

Genes encoding calmodulin-binding proteins in the Arabidopsis genome

Analysis of the recently completed Arabidopsis genome sequence indicates that approximately 31% of the predicted genes could not be assigned to functional categories, as they do not show any sequence similarity with proteins of known function from other organisms. Calmodulin (CaM), a ubiquitous and multifunctional Ca(2+) sensor, interacts with a wide variety of cellular proteins and modulates their activity/function in regulating diverse cellular processes. However, the primary amino acid sequence of the CaM-binding domain in different CaM-binding proteins (CBPs) is not conserved. One way to identify most of the CBPs in the Arabidopsis genome is by protein-protein interaction-based screening of expression libraries with CaM. Here, using a mixture of radiolabeled CaM isoforms from Arabidopsis, we screened several expression libraries prepared from flower meristem, seedlings, or tissues treated with hormones, an elicitor, or a pathogen. Sequence analysis of 77 positive clones that interact with CaM in a Ca(2+)-dependent manner revealed 20 CBPs, including 14 previously unknown CBPs. In addition, by searching the Arabidopsis genome sequence with the newly identified and known plant or animal CBPs, we identified a total of 27 CBPs. Among these, 16 CBPs are represented by families with 2-20 members in each family. Gene expression analysis revealed that CBPs and CBP paralogs are expressed differentially. Our data suggest that Arabidopsis has a large number of CBPs including several plant-specific ones. Although CaM is highly conserved between plants and animals, only a few CBPs are common to both plants and animals. Analysis of Arabidopsis CBPs revealed the presence of a variety of interesting domains. Our analyses identified several hypothetical proteins in the Arabidopsis genome as CaM targets, suggesting their involvement in Ca(2+)-mediated signaling networks.

NASA Discipline Plant Biology

Predicting the Functional State of Protein Kinases Using Interpretable Graph Neural Networks

Kinases are a family of proteins that function as molecular switches, regulating several essential cellular activities such as cell proliferation. Dysfunctional kinases are implicated in several types of cancers and hence they are actively pursued as drug targets. Given the vast number of complex kinase structures that are available in the protein data bank (PDB), there is a necessity to develop methodologies that can identify structurally important moieties of the kinases in an automated fashion, for such techniques can be instrumental in identifying novel drug targets. In this work, we develop a graph neural network (GNN) based deep learning framework for classifying the functionally active and inactive states of a large set of eukaryotic protein kinases, making use of their 3D structure from the PDB. We show that GNN based machine learning models can classify protein states with an accuracy greater than 97%. We further use the GNN models to automatically identify regions of the kinases that are important for its function. For this purpose, Gradient-weighted Class Activation Mapping (Grad-CAM) was implemented on the protein graphs. Remarkably, Grad-CAM consistently identifies the highly conserved DFG motif as the most important part of the protein across the entire kinome, without any prior input. Other regions of the hydrophobic core such as the HRD motif were also identified by the interpretable GNN framework, consistent with the literature. We discuss the significance of each of these regions in detail.

Ashwin Ravichandran

Correlating Protein Dynamics and Catalytic Activity of a Model Hydrogenase Using Paramagnetic and Biological Nuclear Magnetic Resonance Spectroscopy

Rational catalyst design remains a significant challenge, with electronic structure, steric, and electrostatic effects known to contribute to activity. Recently, dynamics has been recognized as another factor that impacts catalysis, though identifying and predicting these effects has remained out of reach. Nickel-substituted rubredoxin (NiRd), a protein-based mimic of a hydrogenase enzyme, serves as a model catalytic system in which dynamics can be systematically investigated with respect to activity. While over 30 secondary-sphere mutants of NiRd have been shown to be catalytically active, no significant correlation was observed between the rates and catalytic overpotential or electronic structure, prompting questions about the protein-derived factors that modulate activity. Here, in this work, NMR spectroscopy was used to investigate the roles of substrate accessibility, protein dynamics, and protein stability in controlling catalysis. Significant paramagnetic effects from the nickel center (S = 1) isolate the methylene proton resonances of the metal-coordinating cysteine residues. The sensitivity of resonance positions and linewidths to local environment offers an opportunity to study dynamical molecular changes around the metal center with high resolution. Machine learning algorithms were employed to identify correlations between the catalytic activity and the paramagnetic NMR spectra. These analyses revealed spectroscopic features of specific cysteine protons that report on catalytic overpotential and increased turnover rates, which are further supported by the results obtained using high-field NMR techniques. Collectively, these studies indicate the potential for multifrequency NMR techniques to resolve key contributors to catalytic activity and highlight the importance of local and outer-sphere dynamics.

Protein Engineering

Predicting metal-binding proteins and structures through integration of evolutionary-scale and physics-based modeling

Metals are essential elements in all living organisms, binding to approximately 50% of proteins. They serve to stabilize proteins, catalyze reactions, regulate activities, and fulfill various physiological and pathological functions. While there have been many advancements in determining the structures of protein-metal complexes, numerous metal-binding proteins still need to be identified through computational methods and validated through experiments. Here, to address this need, we have developed the ESMBind workflow, which combines evolutionary scale modeling (ESM) for metal-binding prediction and physics-based protein-metal modeling. Our approach utilizes the ESM-2 and ESM-IF models to predict metal-binding probability at the residue level. In addition, we have designed a metal-placement method and energy minimization technique to generate detailed 3D structures of protein-metal complexes. Our workflow outperforms other models in terms of residue and 3D-level predictions. To demonstrate its effectiveness, we applied the workflow to 142 uncharacterized fungal pathogen proteins and predicted metal-binding proteins involved in fungal infection and virulence.

59 BASIC BIOLOGICAL SCIENCES

Designing Peptide Fossils That Model the Evolution of the Bacterial Ferredoxin Fold

Electron transfer coupled to redox chemistry is at the heart of metabolism. The proteins responsible for moving electrons (protein electron carriers) must have emerged at the origin of life. The small iron–sulfur-binding bacterial ferredoxins were likely among these first proteins. Embedded within the ferredoxin sequence and structure is a symmetry that points to an ancient gene duplication event. Little is understood about the nature of ferredoxins prior to this duplication event or what environmental factors may have driven the selection for more complex forms. The deep-time molecular history of ferredoxins goes back billions of years and cannot be reconstructed by phylogenetic analyses based on amino acid sequences. Here, we use structure-guided protein design to model a fossil half-ferredoxin stage in the evolution of this fold, the semidoxins, and their symmetric full-length counterparts, the symdoxins. Semidoxin designs homodimerize, exhibiting structural, thermodynamic, and electrochemical behaviors in most cases identical to cognate symdoxins. However, the semi- and symdoxin fossil stages behave differently when incorporated into an in vivo electron transfer complementation assay. Both can support bacterial growth dependent on protein expression. Growth rates of bacteria expressing the semidoxins are much more sensitive to oxygen than those of bacteria expressing symdoxins. Motivated by the in vivo functionality of designed semidoxins, we identified putative naturally occurring semidoxins in extant anaerobic microorganisms. This is consistent with the observed in vivo oxygen sensitivity of the semidoxin designs. One natural semidoxin is shown to be folded and redox active. However, it exists as a mixture of monomers and dimers, suggesting a potential connection between semidoxins and even simpler single iron–sulfur cluster-binding peptides.

59 BASIC BIOLOGICAL SCIENCES

Q -score as a reliability measure for protein, nucleic acid and small-molecule atomic coordinate models derived from 3DEM maps

Atomic coordinate models are important for the interpretation of 3D maps produced with cryoEM and cryoET (3D electron microscopy; 3DEM). In addition to visual inspection of such maps and models, quantitative metrics can inform about the reliability of the atomic coordinates, in particular how well the model is supported by the experimentally determined 3DEM map. A recently introduced metric, Q-score, was shown to correlate well with the reported resolution of the map for well fitted models. Here, we present new statistical analyses of Q-score based on its application to ∼10 000 maps and models archived in the EMDB (Electron Microscopy Data Bank) and PDB (Protein Data Bank). Further, we introduce two new metrics based on Q-score to represent each map and model relative to all entries in the EMDB and those with similar resolution. We explore through illustrative examples of proteins, nucleic acids and small molecules how Q-scores can indicate whether the atomic coordinates are well fitted to 3DEM maps and also whether some parts of a map may be poorly resolved due to factors such as molecular flexibility, radiation damage and/or conformational heterogeneity. These examples and statistical analyses provide a basis for how Q-scores can be interpreted effectively in order to evaluate 3DEM maps and atomic coordinate models prior to publication and archiving.

B factors

Tetragonal Lysozyme Nucleation and Crystal Growth: The Role of the Solution Phase

Lysozyme, and most particularly the tetragonal form of the protein, has become the default standard protein for use in macromolecule crystal nucleation and growth studies. There is a substantial body of experimental evidence, from this and other laboratories, that strongly suggests this proteins crystal nucleation and growth is by addition of associated species that are preformed by standard reversible concentration-driven self association processes in the bulk solution. The evidence includes high resolution AFM studies of the surface packing and of growth unit size at incorporation, fluorescence resonance energy transfer measurements of intermolecular distances in dilute solution, dialysis kinetics, and modeling of the growth rate data. We have developed a selfassociation model for the proteins crystal nucleation and growth. The model accounts for the obtained crystal symmetry, explains the observed surface structures, and shows the importance of the symmetry obtained by self-association in solution to the process as a whole. Further, it indicates that nucleation and crystal growth are not distinct mechanistically, but identical, with the primary difference being the probability that the particle will continue to grow or dissolve. This model also offers a possible mechanism for fluid flow effects on the growth process and how microgravity may affect it. While a single lysozyme molecule is relatively small (M.W. = 14,400), a structured octamer in the 4(sub 3) helix configuration (the proposed average sized growth unit) would have a M.W. = 115,000 and dimensions of 5.6 x 5.6 x 7.6 nm. Direct AFM measurements of growth unit incorporation indicate that units as wide as 11.2 nm and as long as 11.4 nm commonly attach to the crystal. These measurements were made at approximately saturation conditions, and they reflect the sizes of species that both added or desorbed from the crystal surface. The larger and less isotropic the associated species the more likely that it will be oriented to some degree in a flowing boundary layer, even at the low flow velocities measured about macromolecule crystals. Flow-driven effects resulting in misorientation upon addition to and incorporation into the crystal need only be a small fraction of a percentage to significantly affect the resulting crystal. One Earth, concentration gradient driven flow will maintain a high interfacial concentration, i.e., a high level (essentially that of the bulk solution) of solute association at the interface and higher growth rate. Higher growth rates mean an increased probability that misaligned growth units are trapped by subsequent growth layers before they can be desorbed and try again, or that the desorbing species will be smaller than the adsorbing species. In microgravity the extended diffusive boundary layer will lower the interfacial concentration. This results in a net dissociation of aggregated species that diffuse in from the bulk solution, i.e., smaller associated species, which are more likely able to make multiple attempts to correctly bind, yielding higher quality crystals.

Pusey, Marc L.

Automated AI-driven Molecular Design for Therapeutic Discovery

In recent years, artificial intelligence and machine learning (AI/ML) approaches have revolutionized the process of designing new therapeutics, enabling scientists to rapidly respond to emerging threats from various pathogens. A prime example is the SARS-CoV-2 main protease, a key target for the development of antiviral inhibitors. In this study, we employed a novel, integrated approach that combines AI-driven iterative design of inhibitor candidates, screening based on physio-chemical properties and toxicity, physics-based computational modeling of protein-inhibitor interactions, and AI-assisted analysis of Native MS biophysical assay and characterization of designed candidates. Our deep learning 3D-scaffold model, which uses an input scaffold as a starting point, generated tens of thousands of compounds while preserving the key scaffold. To optimize these candidates, we calculated a comprehensive set of 136 descriptors, including both 2D and 3D molecular features, for compounds targeting the SARS-CoV-2 Main protease (Mpro) and a neurodegenerative disease-associated protein, cyclophilin (Cyp). The generated compounds were initially filtered based on their properties and then ranked according to their predicted binding affinity using our automated modeling and ML methods. Experimental validation of the Mpro candidates showing inhibitory activity demonstrates that our workflow can expedite the therapeutic discovery.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH

Simultaneous enhancement of multiple functional properties using evolution-informed protein design

Abstract A major challenge in protein design is to augment existing functional proteins with multiple property enhancements. Altering several properties likely necessitates numerous primary sequence changes, and novel methods are needed to accurately predict combinations of mutations that maintain or enhance function. Models of sequence co-variation (e.g., EVcouplings), which leverage extensive information about various protein properties and activities from homologous protein sequences, have proven effective for many applications including structure determination and mutation effect prediction. We apply EVcouplings to computationally design variants of the model protein TEM-1 β -lactamase. Nearly all the 14 experimentally characterized designs were functional, including one with 84 mutations from the nearest natural homolog. The designs also had large increases in thermostability, increased activity on multiple substrates, and nearly identical structure to the wild type enzyme. This study highlights the efficacy of evolutionary models in guiding large sequence alterations to generate functional diversity for protein design applications.

59 BASIC BIOLOGICAL SCIENCES

Protein Kinase Classification with 2866 Hidden Markov Models and One Support Vector Machine

The main application considered in this paper is predicting true kinases from randomly permuted kinases that share the same length and amino acid distributions as the true kinases. Numerous methods already exist for this classification task, such as HMMs, motif-matchers, and sequence comparison algorithms. We build on some of these efforts by creating a vector from the output of thousands of structurally based HMMs, created offline with Pfam-A seed alignments using SAM-T99, which then must be combined into an overall classification for the protein. Then we use a Support Vector Machine for classifying this large ensemble Pfam-Vector, with a polynomial and chisquared kernel. In particular, the chi-squared kernel SVM performs better than the HMMs and better than the BLAST pairwise comparisons, when predicting true from false kinases in some respects, but no one algorithm is best for all purposes or in all instances so we consider the particular strengths and weaknesses of each.

Weber, Ryan

On The Development of Biophysical Models for Space Radiation Risk Assessment

Experimental techniques in molecular biology are being applied to study biological risks from space radiation. The use of molecular assays presents a challenge to biophysical models which in the past have relied on descriptions of energy deposition and phenomenological treatments of repair. We describe a biochemical kinetics model of cell cycle control and DNA damage response proteins in order to model cellular responses to radiation exposures. Using models of cyclin-cdk, pRB, E2F's, p53, and GI inhibitors we show that simulations of cell cycle populations and GI arrest can be described by our biochemical approach. We consider radiation damaged DNA as a substrate for signal transduction processes and consider a dose and dose-rate reduction effectiveness factor (DDREF) for protein expression.

Cucinotta, F. A.

Linking secretion and cytoskeleton in immunity– a case for Arabidopsis TGNap1

In plants, robust defense depends on the efficient and resilient trafficking supply chains to the site of pathogen attack. Though the importance of intracellular trafficking in plant immunity has been well established, a lack of clarity remains regarding the contribution of the various trafficking pathways in transporting immune-related proteins. We have recently identified a trans-Golgi network protein, TGN-ASSOCIATED PROTEIN 1 (TGNap1), which functionally links post-Golgi vesicles with the cytoskeleton to transport immunity-related proteins in the model plant species Arabidopsis thaliana. We propose new hypotheses on the various functional implications of TGNap1 and then elaborate on the surprising heterogeneity of TGN vesicles during immunity revealed by the discovery of TGNap1 and other TGN-associated proteins in recent years.

59 BASIC BIOLOGICAL SCIENCES

SEC ‐ SAXS / MC Ensemble Structural Studies of the Microtubule Binding Protein Cdt1 Show Monomeric, Folded‐Over Conformations

ABSTRACT Cdt1 is a mixed folded protein critical for DNA replication licensing and it also has a “moonlighting” role at the kinetochore via direct binding to microtubules and the Ndc80 complex. However, it is unknown how the structure and conformations of Cdt1 could allow it to participate in these multiple, unique sets of protein complexes. While robust methods exist to study entirely folded or unfolded proteins, structure–function studies of combined, mixed folded/disordered proteins remain challenging. In this work, we employ orthogonal biophysical and computational techniques to provide structural characterization of mitosis‐competent human Cdt1. Thermal stability analyses shows that both folded winged helix domains1 are unstable. CD and NMR show that the N‐terminal and linker regions are intrinsically disordered. DLS shows that Cdt1 is monomeric and polydisperse, while SEC‐MALS confirms that it is monomeric at high concentrations, but without any apparent inter‐molecular self‐association. SEC‐SAXS enabled computational modeling of the protein structures. Using the program SASSIE, we performed rigid body Monte Carlo simulations to generate a conformational ensemble of structures. We observe that neither fully extended nor extremely compact Cdt1 conformations are consistent with SAXS. The best‐fit models have the N‐terminal and linker disordered regions extended into the solution and the two folded domains close to each other in apparent “folded over” conformations. We hypothesize the best‐fit Cdt1 conformations could be consistent with a function as a scaffold protein that may be sterically blocked without binding partners. Our study also provides a template for combining experimental and computational techniques to study mixed‐folded proteins.

Cell Biology