Search NASA⌕ Search

SEARCH · Search NASA

Results for “Protein structure predictions”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 91 records · Page 5

Free Energy Landscapes for Elucidating the Structural Consequences of Exon-20 mutations on the ErbB Family of Protein Kinases

The ErbB family of protein kinases plays an important role in major cellular functions and consequently mutations in the functional regions of these proteins are implicated in several types of cancer growths. To envision rational design of small molecule drugs that target the diseased proteins it is important to quantify the structural effects of the mutations, as some of these mutants render the protein resistant to tyrosine kinase inhibitors (TKIs). Herein we use advanced sampling techniques and long-timescale molecular dynamics simulations to predict the effect of major exon 20 mutations on the ErbB family, specifically EGFR and HER2 proteins. Exon 20 mutations have been clinically known to induce TKI resistance, though the mechanisms of such an effect is poorly understood. By mapping out the free energy landscape of the mutants and comparing them against the wild-type, we elucidate the structural differences in the binding pocket region that alter the nature of drug-protein interactions. We believe that these insights will play a pivotal role in developing small molecule drugs that overcome the TKI resistance.

Ashwin Ravichandran↗

Structure-Aware Unsupervised, Transformational Machine Learning for Drug Discovery (DTRA Basic Research Final Report)

The major goal of this project is to develop machine learning (ML) methods to enable improved predictive power on real drug discovery for novel targets. More specifically, we planned to demonstrate the capability and effectiveness of ML tools utilizing unlabeled large-volume protein-ligand datasets. We investigated multiple pre-training approaches for 3D protein-ligand structure-based foundation models, without relying on experimental binding data. We also addressed scenarios in which crystal structures are unavailable or binding data are limited. We also planned to develop a complete pipeline to screen novel compounds as well as to demonstrate the capability and effectiveness of the developed methods by testing on a realistic drug discovery task such as SARS-CoV-2. While the major goals and milestones remain consistent with the original proposal, certain technical details have been adjusted, based on the experimental results and related outcomes.

97 MATHEMATICS AND COMPUTING↗

Ca X ML: Chemistry‐informed machine learning explains mutual changes between protein conformations and calcium ions in calcium‐binding proteins using structural and topological features

Proteins' flexibility is a feature in communicating changes in cell signaling instigated by binding with secondary messengers, such as calcium ions, associated with the coordination of muscle contraction, neurotransmitter release, and gene expression. When binding with the disordered parts of a protein, calcium ions must balance their charge states with the shape of calcium-binding proteins and their versatile pool of partners depending on the circumstances they transmit. Accurately determining the ionic charges of those ions is essential for understanding their role in such processes. However, it is unclear whether the limited experimental data available can be effectively used to train models to accurately predict the charges of calcium-binding protein variants. Here, we developed a chemistry-informed, machine-learning algorithm that implements a game theoretic approach to explain the output of a machine-learning model without the prerequisite of an excessively large database for high-performance prediction of atomic charges. We used the ab initio electronic structure data representing calcium ions and the structures of the disordered segments of calcium-binding peptides with surrounding water molecules to train several explainable models. Network theory was used to extract the topological features of atomic interactions in the structurally complex data dictated by the coordination chemistry of a calcium ion, a potent indicator of its charge state in protein. Our design created a computational tool of Ca X ML, which provided a framework of explainable machine learning model to annotate ionic charges of calcium ions in calcium-binding proteins in response to the chemical changes in an environment. Our framework will provide new insights into protein design for engineering functionality based on the limited size of scientific data in a genome space.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Protein Kinase Classification with 2866 Hidden Markov Models and One Support Vector Machine

The main application considered in this paper is predicting true kinases from randomly permuted kinases that share the same length and amino acid distributions as the true kinases. Numerous methods already exist for this classification task, such as HMMs, motif-matchers, and sequence comparison algorithms. We build on some of these efforts by creating a vector from the output of thousands of structurally based HMMs, created offline with Pfam-A seed alignments using SAM-T99, which then must be combined into an overall classification for the protein. Then we use a Support Vector Machine for classifying this large ensemble Pfam-Vector, with a polynomial and chisquared kernel. In particular, the chi-squared kernel SVM performs better than the HMMs and better than the BLAST pairwise comparisons, when predicting true from false kinases in some respects, but no one algorithm is best for all purposes or in all instances so we consider the particular strengths and weaknesses of each.

Weber, Ryan↗

Visualizing and analyzing 3D biomolecular structures using Mol* at RCSB.org: Influenza A H5N1 virus proteome case study

The easiest and often most useful way to work with experimentally determined or computationally predicted structures of biomolecules is by viewing their three-dimensional (3D) shapes using a molecular visualization tool. Mol* was collaboratively developed by RCSB Protein Data Bank (RCSB PDB, RCSB.org) and Protein Data Bank in Europe (PDBe, PDBe.org) as an open-source, web-based, 3D visualization software suite for examination and analyses of biostructures. It is capable of displaying atomic coordinates and related experimental data of biomolecular structures together with a variety of annotations, facilitating basic and applied research, training, education, and information dissemination. Across RCSB.org, the RCSB PDB research-focused web portal, Mol* has been implemented to support single-mouse-click atomic-level visualization of biomolecules (e.g., proteins, nucleic acids, carbohydrates) with bound cofactors, small-molecule ligands, ions, water molecules, or other macromolecules. RCSB.org Mol* can seamlessly display 3D structures from various sources, allowing structure interrogation, superimposition, and comparison. Using influenza A H5N1 virus as a topical case study of an important pathogen, we exemplify how Mol* has been embedded within various RCSB.org tools—allowing users to view polymer sequence and structure-based annotations integrated from trusted bioinformatics data resources, assess patterns and trends in groups of structures, and view structures of any size and compositional complexity. In addition to being linked to every experimentally determined biostructure and Computed Structure Model made available at RCSB.org, Standalone Mol* is freely available for visualizing any atomic-level or multi-scale biostructure at rcsb.org/3d-view.

3D biostructure↗

Increasing thermostability of the key photorespiratory enzyme glycerate 3‐kinase by structure‐based recombination

As global temperatures rise, improving crop yields will require enhancing the thermotolerance of crops. One approach for improving thermotolerance is using bioengineering to increase the thermostability of enzymes catalysing essential biological processes. Photorespiration is an essential recycling process in plants that is integral to photosynthesis and crop growth. The enzymes of photorespiration are targets for enhancing plant thermotolerance as this pathway limits carbon fixation at elevated temperatures. We explored the effects of temperature on the activity of the photorespiratory enzyme glycerate kinase (GLYK) from various organisms and the homologue from the thermophilic alga Cyanidioschyzon merolae was more thermotolerant than those from mesophilic plants, including Arabidopsis thaliana. To understand enzyme features underlying the thermotolerance of C. merolae GLYK (CmGLYK), we performed molecular dynamics simulations using AlphaFold-predicted structures, which revealed greater movement of loop regions of mesophilic plant GLYKs at higher temperatures compared to CmGLYK. Based on these simulations, hybrid proteins were produced and analysed. These hybrid enzymes contained loop regions from CmGLYK replacing the most mobile corresponding loops of AtGLYK. Two of these hybrid enzymes had enhanced thermostability, with melting temperatures increased by 6 °C. One hybrid with three grafted loops maintained higher activity at elevated temperatures. Whilst this hybrid enzyme exhibited enhanced thermostability and a similar Km for ATP compared to AtGLYK, its Km for glycerate increased threefold. This study demonstrates that molecular dynamics simulation-guided structure-based recombination offers a promising strategy for enhancing the thermostability of other plant enzymes with possible application to increasing the thermotolerance of plants under warming climates.

59 BASIC BIOLOGICAL SCIENCES↗

Finding the global minimum: a fuzzy end elimination implementation

The 'fuzzy end elimination theorem' (FEE) is a mathematically proven theorem that identifies rotameric states in proteins which are incompatible with the global minimum energy conformation. While implementing the FEE we noticed two different aspects that directly affected the final results at convergence. First, the identification of a single dead-ending rotameric state can trigger a 'domino effect' that initiates the identification of additional rotameric states which become dead-ending. A recursive check for dead-ending rotameric states is therefore necessary every time a dead-ending rotameric state is identified. It is shown that, if the recursive check is omitted, it is possible to miss the identification of some dead-ending rotameric states causing a premature termination of the elimination process. Second, we examined the effects of removing dead-ending rotameric states from further considerations at different moments of time. Two different methods of rotameric state removal were examined for an order dependence. In one case, each rotamer found to be incompatible with the global minimum energy conformation was removed immediately following its identification. In the other, dead-ending rotamers were marked for deletion but retained during the search, so that they influenced the evaluation of other rotameric states. When the search was completed, all marked rotamers were removed simultaneously. In addition, to expand further the usefulness of the FEE, a novel method is presented that allows for further reduction in the remaining set of conformations at the FEE convergence. In this method, called a tree-based search, each dead-ending pair of rotamers which does not lead to the direct removal of either rotameric state is used to reduce significantly the number of remaining conformations. In the future this method can also be expanded to triplet and quadruplet sets of rotameric states. We tested our implementation of the FEE by exhaustively searching ten protein segments and found that the FEE identified the global minimum every time. For each segment, the global minimum was exhaustively searched in two different environments: (i) the segments were extracted from the protein and exhaustively searched in the absence of the surrounding residues; (ii) the segments were exhaustively searched in the presence of the remaining residues fixed at crystal structure conformations. We also evaluated the performance of the method for accurately predicting side chain conformations. We examined the influence of factors such as type and accuracy of backbone template used, and the restrictions imposed by the choice of potential function, parameterization and rotamer database. Conclusions are drawn on these results and future prospects are given.

NASA Program Exobiology↗

Enzyme property prediction using artificial intelligence

Artificial intelligence (AI)-driven enzyme property prediction enables rapid discovery and engineering of enzymes for a wide range of biotechnological and therapeutic applications. Here, we first introduce the key components in AI model development, including enzyme datasets, protein representation methods, and model architectures. We then highlight a variety of AI tools developed for the prediction of enzyme properties and functional annotations, including enzyme structure, kinetic parameters, substrate specificity, thermostability, solubility, Enzyme Commission number, and Gene Ontology term. Moreover, we describe representative downstream applications enabled by these AI tools. Finally, we discuss some challenges and opportunities as well as future prospects.

Yuan, Le [University of Illinois at Urbana-Champai↗

The Effect of Solution Conditions on the Nucleation Kinetics of Tetragonal Lysozyme Crystals

An understanding of protein crystal nucleation rates and the effect of solution conditions upon them, is fundamental to the preparation of protein crystals of the desired size and shape for X-ray diffraction analysis. The ability to predict the effect of supersaturation, temperature, pH and precipitant concentration on the number and size of crystals formed is of great benefit in the pursuit of protein structure analysis. In this study we experimentally examine the effect of supersaturation, temperature, pH and sodium chloride concentration on the nucleation rate of tetragonal chicken egg white lysozyme crystals. In order to do this batch crystallization plates were prepared at given solution concentrations and incubated at three different temperatures over the period of one week. The number of crystals per well with their size and dimensions were recorded and correlated against solution conditions. Duplicate experiments indicate the reproducibility of the technique. Although it is well known that crystal numbers increase with increasing supersaturation, large changes in crystal number were also correlated against solution conditions of temperature, pH and salt concentration over the same supersaturation ranges. Analysis of these results enhance our understanding of the effect of solution conditions such as the dramatic effect that small changes in charge and ionic strength can have on the number of tetragonal lysozyme crystals that form and grow in solution.

Judge, Russell A.↗

Characterization of alternate encounter assemblies of SARS-CoV-2 main protease

The assembly of two monomeric constructs spanning segments 1-199 (MPro 1-199 ) and 10-306 (MPro 10-306 ) of SARS-CoV-2 main protease (MPro) was examined to assess the existence of a transient heterodimer intermediate in the N-terminal autoprocessing pathway of MPro model precursor. Together, they form a heterodimer population accompanied by a 13-fold increase in catalytic activity. Addition of inhibitor GC373 to the proteins increases the activity further by ~7-fold with a 1:1 complex and higher order assemblies approaching 1:2 and 2:2 molecules of MPro 1-199 and MPro 10-306 detectable by analytical ultracentrifugation and native mass estimation by light scattering. Assemblies larger than a heterodimer (1:1) are discussed in terms of alternate pathways of domain III association, either through switching the location of helix 201 to 214 onto a second helical domain of MPro 10-306 and vice versa or direct interdomain III contacts like that of the native dimer, based on known structures and AlphaFold 3 prediction, respectively. At a constant concentration of MPro 1-199 with molar excess of GC373, the rate of substrate hydrolysis displays first order dependency on the MPro 10-306 concentration and vice versa. An equimolar composition of the two proteins with excess GC373 exhibits half-maximal activity at ~6 μM MPro 1-199 . Catalytic activity arises primarily from MPro 1-199 and is dependent on the interface interactions involving the N-finger residues 1 to 9 of MPro 1-199 and E290 of MPro 10-306 . Importantly, our results confirm that a single N-finger region with its associated intersubunit contacts is sufficient to form a heterodimeric MPro intermediate with enhanced catalytic activity.

60 APPLIED LIFE SCIENCES↗

A Trichomonas vaginalis C2-XYPPX-repeat protein with a structured C2 domain displaying dampened flexibility upon binding calcium

C2 domains are ubiquitous membrane-binding modules of ∼130 residues in eukaryotes that are often associated with proteins involved in membrane trafficking and lipid modification. The genome of Trichomonas vaginalis, the most common, non-viral, sexually transmitted human pathogen, encodes eight genes that contain a N-terminal C2 module linked to a XYPPX-repeat domain of more than four XYPPX repeats (C2-XYPPX). While the function of the XYPPX-repeat domain remains unknown, its multiple association with C2 domains in T. vaginalis suggests it is important. Here, the C2 domain from one of these C2-XYPPX-repeat proteins, Tv-C2-1, was structurally and physically characterized using X-ray crystallography and NMR spectroscopy. The crystal structure for Tv-C2-1 shows that this domain shares a fold common to all C2 domains, a compact Greek-key motif composed of eight anti-parallel β-strands in the type-2 topology. An NMR chemical shift perturbation study with Ca 2+ showed that Tv-C2-1 bound two Ca 2+ atoms primarily via two loops (loop-1 and loop-3) on the predicted calcium binding face of the protein with K d s of 58.0 ± 0.1 μM and 232 ± 6 μM. Estimations of the overall rotational correlation time, τ c , in the apo (11.1 ns) and Ca 2+ -bound (9.2 ns) state suggests the protein becomes more compact upon Ca 2+ binding, consistent with a decrease in dynamics in loop-3 and marginally in loop-1 suggested by amide 15 N heteronuclear steady-state { 1 H}- 15 N NOEs. Showing Tv-C2-1 binds calcium and adopts a compact Greek-key motif structure, two primary features of C2 domains, suggests understanding the function of the XYPPX-repeat domain may be warranted.

NMR spectroscopy↗

Thermophilic Chassis-Enabled High-Throughput Selection of a Thermostable Fluorogenic Reporter

Thermostable proteins show increased shelf life and performance at elevated temperatures and under harsh conditions, resulting in lower costs for various industrial and biotechnological applications. However, due to a limited understanding of the relationship between stability and function, protein stabilization remains primarily a trial-and-error approach. Therefore, building a combinatorial library of mutations predicted to improve stability, followed by experimental testing, represents a markedly improved methodology. However, the lack of high-throughput approaches to screen even a moderately sized library presents a major bottleneck in the field. Here, in this study, we use a thermophile, Parageobacillus thermoglucosidasius (Ptherm) to rapidly screen combinatorial libraries consisting of rationally designed thermostabilizing mutations (∼10 3 –10 4 ) of a mesophilic fluorescent reporter, Y-FAST. On a Petri dish, microbial growth at an elevated temperature and exposure to fluorogen yielded several colonies of Ptherm that showed distinct fluorescence at 55 and 68 °C in our two sequentially generated libraries using Rosetta and ProteinMPNN, respectively. The Y-FAST variants isolated from fluorescent colonies were brighter than Y-FAST and showed higher resistance to thermal and chemical denaturation. AlphaFold-predicted structures and MD simulations revealed stability-enhancing salt bridges and hydrogen bond networks in the isolated FAST variants. The moderately thermostable FAST (tsFAST) and hyperstable FAST (hsFAST) were then demonstrated as translation reporters for protein expression and folding at elevated temperatures, such as 55 and 68 °C. Our approach of combinatorial library generation and high-throughput screening in a thermophilic chassis could, in principle, be extended to other proteins fused to these translation reporters. Furthermore, the hsFAST protein is small─half the size of the green fluorescent protein─and does not require oxygen for maturation, making it ideal for engineering extremophilic anaerobes for biosensing and bioconversion.

59 BASIC BIOLOGICAL SCIENCES↗

Cryo-EM Structure of the Mnx Protein Complex Reveals a Tunnel Framework for the Mechanism of Manganese Biomineralization

The global manganese cycle relies on microbes to oxidize soluble Mn(II) to insoluble Mn(IV) oxides. Some microbes require peroxide or superoxide as oxidants, but others can use O 2 directly, via multicopper oxidase (MCO) enzymes. One of these, MnxG from Bacillus sp. strain PL-12, was isolated in tight association with small accessory proteins, MnxE and MnxF. The protein complex, called Mnx, has eluded crystallization efforts, but we now report the 3D structure of a point mutant using cryo-EM single particle analysis, cross-linking mass spectrometry, and AlphaFold Multimer prediction. The ß-sheet–rich complex features MnxG enzyme, capped by a heterohexameric ring of alternating MnxE and MnxF subunits, and a tunnel that runs through MnxG and its MnxE 3 F 3 cap. The tunnel dimensions and charges can accommodate the mechanistically inferred binuclear manganese intermediates. Furthermore, comparison with the Fe(II)-oxidizing MCO, ceruloplasmin, identifies likely coordinating groups for the Mn(II) substrate, at the entrance to the tunnel. Thus, the 3D structure provides a rationale for the established manganese oxidase mechanism, and a platform for further experiments to elucidate mechanistic details of manganese biomineralization.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Investigating Electron Conductivity Regimes in the Bacterial Cytochrome Wire OmcS

The anaerobic bacterium Geobacter sulfurreducens produces extracellular, electronically conductive cytochrome polymer wires that are conductive over micron length scales. Structure models from cryo-electron microscopy data show OmcS wires form a linear chain of hemes along the protein wire axis, which is proposed as the structural basis supporting their electronic properties. However, the mechanism by which this heme arrangement supports long-range electronic conduction remains unknown. Structure models from cryo-electron microscopy data show these wires form a linear chain of hemes along the protein wire axis, which is proposed as the structural basis supporting their electronic properties. Existing computational models using static heme redox potentials and coupling energies fail to explain experimental observations, predicting conductances 10,000 to 100,000 times lower than measured values. Here, we investigate how dynamic disorder affects site energies, interheme coupling, and long-range electronic conductivity within these cytochrome wires. We introduce an approach to extract charge carrier site information directly from Kohn–Sham density functional theory, without employing projector schemes, and show that site and coupling energies are highly sensitive to changes in interheme geometry and the surrounding electrostatic environment. Unlike models that incorporate dynamic disorder as a thermally averaged quantity, our quantum charge carrier model incorporates proxies for dynamic disorder through decoherence corrections, yielding predicted diffusion coefficient closer to what is expected from experiment and comparable with other organic-based electronic materials. Based on these simulations, we propose that the instantaneous fluctuations of the local electrostatic environment can transiently lift energy degeneracies and delocalize charge carriers. Furthermore, these studies reveal how incorporating dynamic fluctuations associated with the environment resolves the discrepancy between theory and experiment in microbial cytochrome wires and highlight design principles for bioinspired, heme-based conductive materials.

Bioinorganic chemistry↗

Stability of Magnetically-Suppressed Solutal Convection In Protein Crystal Growth

The effect of convection during the crystallization of proteins is not very well understood. In a gravitational field, convection is caused by crystal sedimentation and by solutal buoyancy induced flow and these can lead to crystal imperfections. While crystallization in microgravity can approach diffusion limited growth conditions (no convection), terrestrially strong magnetic fields can be used to control fluid flow and sedimentation effects. In this work, a theory is presented on the stability of solutal convection of a magnetized fluid in the presence of a magnetic field. The requirements for stability are developed and compared to experiments performed within the bore of a superconducting magnet. The theoretical predictions are in good agreement with the experiments and show solutal convection can be stabilized if the surrounding fluid has larger magnetic susceptibility and the magnetic field has a specific structure. Discussion on the application of the technique to protein crystallization is also provided.

Leslie, F. W.↗

Advancing molecular machine learning representations with stereoelectronics-infused molecular graphs

Molecular representation is a critical element in our understanding of the physical world and the foundation for modern molecular machine learning. Previous molecular machine learning models have used strings, fingerprints, global features and simple molecular graphs that are inherently information-sparse representations. However, as the complexity of prediction tasks increases, the molecular representation needs to encode higher fidelity information. This work introduces a new approach to infusing quantum-chemical-rich information into molecular graphs via stereoelectronic effects, enhancing expressivity and interpretability. Learning to predict the stereoelectronics-infused representation with a tailored double graph neural network workflow enables its application to any downstream molecular machine learning task without expensive quantum-chemical calculations. We show that the explicit addition of stereoelectronic information substantially improves the performance of message-passing two-dimensional machine learning models for molecular property prediction. We show that the learned representations trained on small molecules can accurately extrapolate to much larger molecular structures, yielding chemical insight into orbital interactions for previously intractable systems, such as entire proteins, opening new avenues of molecular design. Finally, we have developed a web application (simg.cheme.cmu.edu) where users can rapidly explore stereoelectronic information for their own molecular systems.

Boiko, Daniil A↗

Enzymatic carbon–fluorine bond cleavage by human gut microbes

Fluorinated compounds are used for agrochemical, pharmaceutical, and numerous industrial applications, resulting in global contamination. In many molecules, fluorine is incorporated to enhance the half-life and improve bioavailability. Fluorinated compounds enter the human body through food, water, and xenobiotics including pharmaceuticals, exposing gut microbes to these substances. The human gut microbiota is known for its xenobiotic biotransformation capabilities, but it was not previously known whether gut microbial enzymes could break carbon-fluorine bonds, potentially altering the toxicity of these compounds. Here, through the development of a rapid, miniaturized fluoride detection assay for whole-cell screening, we identified active gut microbial defluorinases. We biochemically characterized enzymes from diverse human gut microbial classes including Clostridia, Bacilli, and Coriobacteriia, with the capacity to hydrolyze (di)fluorinated organic acids and a fluorinated amino acid. Whole-protein alanine scanning, molecular dynamics simulations, and chimeric protein design enabled the identification of a disordered C-terminal protein segment involved in defluorination activity. Domain swapping exclusively of the C-terminus conferred defluorination activity to a nondefluorinating dehalogenase. To advance our understanding of the structural and sequence differences between defluorinating and nondefluorinating dehalogenases, we trained machine learning models which identified protein termini as important features. Models trained on 41-amino acid segments from protein C termini alone predicted defluorination activity with 83% accuracy (compared to 95% accuracy based on full-length protein features). This work is relevant for therapeutic interventions and environmental and human health by uncovering specificity-determining signatures of fluorine biochemistry from the gut microbiome.

Probst, Silke I↗