Search NASASearch

SEARCH · Search NASA

Results for “Protein modeling”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

(abstract) Modeling Protein Families and Human Genes: Hidden Markov Models and a Little Beyond

We will first give a brief overview of Hidden Markov Models (HMMs) and their use in Computational Molecular Biology. In particular, we will describe a detailed application of HMMs to the G-Protein-Coupled-Receptor Superfamily. We will also describe a number of analytical results on HMMs that can be used in discrimination tests and database mining. We will then discuss the limitations of HMMs and some new directions of research. We will conclude with some recent results on the application of HMMs to human gene modeling and parsing.

Hidden Markov Models HMMs proteins computational m

Modelling protein functional domains in signal transduction using Maude

Modelling of protein-protein interactions in signal transduction is receiving increased attention in computational biology. This paper describes recent research in the application of Maude, a symbolic language founded on rewriting logic, to the modelling of functional domains within signalling proteins. Protein functional domains (PFDs) are a critical focus of modern signal transduction research. In general, Maude models can simulate biological signalling networks and produce specific testable hypotheses at various levels of abstraction. Developing symbolic models of signalling proteins containing functional domains is important because of the potential to generate analyses of complex signalling networks based on structure-function relationships.

Signal Transduction

Integrating Ultra-Coarse-Grained Protein Models into Accessible Workflows for Multiscale Molecular Dynamics

To capture protein conformational transitions using molecular dynamics (MD), several simulation resolutions covering different spatial and temporal scales are typically needed. All-atom (AA) simulations provide fine resolution, but are computationally infeasible for large systems over longer durations. Coarse-grained (CG) and ultra-coarse-grained (UCG) models have a lower resolution and computational cost while still being able to conserve essential protein features. Prior work on a Multiscale Machinelearned Modeling Infrastructure (MuMMI) combined both AA and CG simulations to study RAS-RAF protein interactions, leveraging CG models for longer time scales and using AA to investigate unusual conformations in greater detail. However, MuMMI is still resource-intensive, and this study aims to maximize exploration of the protein conformational space while reducing computational cost. In this paper, we build on prior work that integrates UCG models based on heterogeneous elastic network modeling (hENM) into the MuMMI workflow. We demonstrate that UCG models enable accurate sampling of protein conformations, focusing on simulating RAS-RAF protein interactions. Using higher-resolution CG Martini simulation data, we can automatically refine intramolecular interactions in UCG models. We present a scalable Python package that uses fluctuations observed in higher-resolution CG Martini simulations to estimate bond coefficients of the UCG model. We built novel machine learning-based backmapping methods to recover more detailed CG Martini structures from UCG structures, using diffusion models to learn the mapping between scales. Finally, we present UCG-mini-MuMMI, an accessible and less compute-intensive version of MuMMI as a resource for the scientific community. Incorporating UCG models into MD studies is applicable to a broad range of systems and proteins, and our study offers insights into the advantages and limitations of these methods.

Chemical structure

Modeling Protein–Protein and Protein–Ligand Interactions by the ClusPro Team in CASP16

ABSTRACT In the CASP16 experiment, our team employed hybrid computational strategies to predict both protein–protein and protein–ligand complex structures. For protein–protein docking, we combined physics‐based sampling—using ClusPro FFT docking and molecular dynamics—with AlphaFold (AF)‐based sampling, followed by AF‐based refinement. Our method produced numerous high‐accuracy complex models, including cases where AF alone failed, underscoring the critical role of physics‐based sampling alongside deep learning‐based refinement. For protein–ligand docking, we integrated the ClusPro LigTBM template‐based approach with a machine learning‐based confidence model for rescoring. The method preserves conserved interaction fragments derived from homologous complexes, followed by local resampling using physics‐based sampling and a diffusion model. Our template‐based strategy achieved a mean lDDT‐PLI of 0.69 across 233 targets, which was highly competitive. These results demonstrate that combining physics‐based modeling with AI‐driven refinement can significantly enhance the accuracy of both protein–protein and protein–ligand structure predictions.

Ashizawa, Ryota [Department of Applied Mathematics

Predicting compatibility between ferredoxins and the Fe protein of nitrogenase using in silico protein modeling

Biological nitrogen fixation is the process by which certain bacteria and archaea use the enzyme nitrogenase to reduce atmospheric nitrogen into bioavailable ammonium. Engineering non‐nitrogen‐fixing organisms, like plants, to use nitrogenase could reduce dependency on synthetic fertilizer and mitigate the environmental impacts of industrial fertilizer production. However, nitrogenase activity requires delivery of reducing power by small electron carrying proteins known as ferredoxins and flavodoxins, and successfully engineering nitrogenase into new systems will require a mechanistic understanding of electron delivery by these proteins. Most organisms often have multiple ferredoxins, raising the question of which ferredoxin can support nitrogenase activity. The purpose of this study is to gain insight into how we can predict which ferredoxin is compatible with the Fe protein, the component of nitrogenase that interacts with ferredoxin or flavodoxin. Our in silico protein–protein docking simulations reveal that most ferredoxins and flavodoxins involved in nitrogen fixation have the shortest distance (≤10 Å) between their redox cofactor and the [4Fe‐4S] cluster of the Fe protein. We found shorter cofactor distance contributes to faster intermolecular electron tunneling rates. Bacterial ferredoxins that play a role in nitrogen fixation also exhibit more complementary interactions with the Fe protein than bacterial and plant ferredoxins not involved in this process. Heterologous expression of a set of ferredoxins from both nitrogen‐fixing and non‐nitrogen‐fixing bacteria in the diazotroph Rhodopseudomonas palustris supports our model‐derived prediction that shorter distances between the electron‐carrying cofactors favor nitrogenase compatibility. These findings offer a framework to predict and potentially enhance ferredoxin–nitrogenase compatibility, which will help to improve our ability to engineer nitrogen fixation into non‐nitrogen‐fixing organisms like plants.

59 BASIC BIOLOGICAL SCIENCES

Green Fluorescent Protein as a Model for Protein Crystal Growth Studies

Green fluorescent protein (GFP) from jellyfish Aequorea Victoria has become a popular marker for e.g. mutagenesis work. Its fluorescent property, which originates from a chromophore located in the center of the molecule, makes it widely applicable as a research too]. GFP clones have been produced with a variety of spectral properties, such as blue and yellow emitting species. The protein is a single chain of molecular weight 27 kDa and its structure has been determined at 1.9 Angstrom resolution. The combination of GFP's fluorescent property, the knowledge of its several crystallization conditions, and its increasing use in biophysical and biochemical studies, all led us to consider it as a model material for macromolecular crystal growth studies. Initial preparations of GFP were from E.coli with yields of approximately 5 mg/L of culture media. Current yields are now in the 50 - 120 mg/L range, and we hope to further increase this by expression of the GFP gene in the Pichia system. The results of these efforts and of preliminary crystal growth studies will be presented.

Agena, Sabine

Marginal protein stability drives subcellular proteome isoelectric point

There exists a positive correlation between the pH of subcellular compartments and the median isoelectric point (pI) for the associated proteomes. Proteins in the human lysosome—a highly acidic compartment in the cell—have a median pI of ∼6.5, whereas proteins in the more basic mitochondria have a median pI of ∼8.0. Proposed mechanisms reflect potential adaptations to pH. For example, enzyme active site general acid/base residue pKs are likely evolved to match environmental pH. However, such effects would be limited to a few residues on specific proteins, and might not affect the proteome at large. A protein model that considers residue burial upon folding recapitulates the correlation between proteome pI and environmental pH. This correlation can be fully described by a neutral evolution process; no functional selection is included in the model. Proteins in acidic environments incur a lower energetic penalty for burying acidic residues than basic residues, resulting in a net accumulation of acidic residues in the protein core. The inverse is true under alkaline conditions. The pI distributions of subcellular proteomes are likely not a direct result of functional adaptations to pH, but a molecular spandrel stemming from marginal stability.

Kaiser Loell

Reweighting configurations generated by transferable, machine learned models for protein sidechain backmapping

Multiscale modeling requires the linking of models at different levels of detail, with the goal of gaining accelerations from lower fidelity models while recovering fine details from higher resolution models. Communication across resolutions is particularly important in modeling soft matter, where tight couplings exist between molecular-level details and mesoscale structures. While multiscale modeling of biomolecules has become a critical component in exploring their structure and self-assembly, backmapping from coarse-grained to fine-grained, or atomistic, representations presents a challenge, despite recent advances through machine learning. A major hurdle, especially for strategies utilizing machine learning, is that backmappings can only approximately recover the atomistic ensemble of interest. We demonstrate conditions for which backmapped configurations may be reweighted to exactly recover the desired atomistic ensemble. By training separate decoding models for each sidechain type, we develop an algorithm based on normalizing flows and geometric algebra attention to autoregressively propose backmapped configurations for any protein sequence. Critical for reweighting with modern protein force fields, our trained models include all hydrogen atoms in the backmapping and make probabilities associated with atomistic configurations directly accessible. We also demonstrate, however, that reweighting is extremely challenging despite state-of-the-art performance on recently developed metrics and generation of configurations with low energies in atomistic protein force fields. Through detailed analysis of configurational weights, we show that machine-learned backmappings must not only generate configurations with reasonable energies, but also correctly assign relative probabilities under the generative model. These are broadly important considerations in generative modeling of atomistic molecular configurations.

Monroe, Jacob I. [Univ. of Arkansas, Fayetteville,

Protein solubility modeling

A thermodynamic framework (UNIQUAC model with temperature dependent parameters) is applied to model the salt-induced protein crystallization equilibrium, i.e., protein solubility. The framework introduces a term for the solubility product describing protein transfer between the liquid and solid phase and a term for the solution behavior describing deviation from ideal solution. Protein solubility is modeled as a function of salt concentration and temperature for a four-component system consisting of a protein, pseudo solvent (water and buffer), cation, and anion (salt). Two different systems, lysozyme with sodium chloride and concanavalin A with ammonium sulfate, are investigated. Comparison of the modeled and experimental protein solubility data results in an average root mean square deviation of 5.8%, demonstrating that the model closely follows the experimental behavior. Model calculations and model parameters are reviewed to examine the model and protein crystallization process. Copyright 1999 John Wiley & Sons, Inc.

Proteins/chemistry

Multistep modeling of protein structure: application to bungarotoxin

Modelling of bungarotoxin in atomic details is presented in this article. The model-building procedure utilizes the low-resolution crystal coordinates of the c-alpha atoms of bungarotoxin, sequence homology within the neurotoxin family, as well as high-resolution x-ray diffraction data of cobratoxin and erabutoxin. Our model-building procedure involves: (a) principles of comparative modelling, (b) embedding procedures of distance geometry, and (c) use of molecular mechanics for optimizing packing. The model is not only consistent with the c-alpha coordinates of crystal structure, but also agrees with solution conformational features of the triple-stranded beta sheet as observed by NOE measurements.

NASA Discipline Exobiology

Anisotropic interactions for continuum modeling of protein–membrane systems

In this work, a model for anisotropic interactions between proteins and cellular membranes is proposed for large-scale continuum simulations. The framework of the model is based on dynamic density functional theory, which provides a formalism to describe the lipid densities within the membrane as continuum fields while still maintaining the fidelity of the underlying molecular interactions. Within this framework, we extend recent results to include the anisotropic effects of protein–lipid interactions. As applications, we consider two membrane proteins of biological interest: a RAS–RAF complex tethered to the membrane and a membrane embedded G protein-coupled receptor. A strong qualitative and quantitative agreement is found between the numerical results and the corresponding molecular dynamics simulations. Combining the scope of continuum level simulations with the details from molecular level particle simulations enables research into protein–membrane behaviors at a more biologically relevant scale, which crucially can also be accessed via experiment.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH

Molecular modeling of calmodulin: a comparison with crystallographic data

Two methods of side-chain placement on a modeled protein have been examined. Two molecular models of calmodulin were constructed that differ in the treatment of side chains prior to optimization of the molecule. A virtual bond analysis program developed by Purisima and Scheraga was used to determine the backbone conformation based on 2.2 angstroms resolution C alpha coordinates for the molecules. In the first model, side chains were initially constructed in an extended conformation. In the second model, a conformational grid search technique was employed. Calcium ions were treated explicitly during energy optimization using CHARMM. The models are compared to a recently published refined crystal structure of calmodulin. The results indicate that the initial choices for side-chains, but also significant effects on the main-chain conformation and supersecondary structure. The conformational differences are discussed. Analysis of these and other methods makes possible the formulation of a methodology for more appropriate side-chain placement in modeled proteins.

NASA Discipline Exobiology

Multistep modeling of protein structure: application towards refinement of tyr-tRNA synthetase

The scope of multistep modeling (MSM) is expanding by adding a least-squares minimization step in the procedure to fit backbone reconstruction consistent with a set of C-alpha coordinates. The analytical solution of Phi and Psi angles, that fits a C-alpha x-ray coordinate is used for tyr-tRNA synthetase. Phi and Psi angles for the region where the above mentioned method fails, are obtained by minimizing the difference in C-alpha distances between the computed model and the crystal structure in a least-squares sense. We present a stepwise application of this part of MSM to the determination of the complete backbone geometry of the 321 N terminal residues of tyrosine tRNA synthetase to a root mean square deviation of 0.47 angstroms from the crystallographic C-alpha coordinates.

NASA Discipline Exobiology

A Proposed Model for Protein Crystal Nucleation and Growth

How does one take a molecule, strongly asymmetric in both shape and charge distribution, and assemble it into a crystal? We propose a model for the nucleation and crystal growth process for tetragonal lysozyme, based upon fluorescence, light, neutron, and X-ray scattering data, size exclusion chromatography experiments, dialysis kinetics, AFM, and modeling of growth rate data, from this and other laboratories. The first species formed is postulated to be a 'head to side' dimer. Through repeating associations involving the same intermolecular interactions this grows to a 4(sub 3) helix structure, that in turn serves as the basic unit for nucleation and subsequent crystal growth. High salt attenuates surface charges while promoting hydrophobic interactions. Symmetry facilitates subsequent helix-helix self-association. Assembly stability is enhanced when a four helix structure is obtained, with each bound to two neighbors. Only two unique interactions are required. The first are those for helix formation, where the dominant interaction is the intermolecular bridging anion. The second is the anti-parallel side-by-side helix-helix interaction, guided by alternating pairs of symmetry related salt bridges along each side. At this stage all eight unique positions of the P4(sub3)2(sub 1),2(sub 1) unit cell are filled. The process is one of a) attenuating the most strongly interacting groups, such that b) the molecules begin to self-associate in defined patterns, so that c) symmetry is obtained, which d) propagates as a growing crystal. Simple and conceptually obvious in hindsight, this tells much about what we are empirically doing when we crystallize macromolecules. By adjusting the growth parameters we are empirically balancing the intermolecular interactions, preferentially attenuating the dominant strong (for lysozyme the charged groups) while strengthening the lesser strong (hydrophobic) interactions. In the general case for proteins the lack of a singularly defined association pathway may lead to formation of multiple species, i.e., amorphous precipitation. Weak interactions, such as hydrogen bonds, are promiscuous, serving to strengthen rather than define specific interactions. Participation in an interaction sequesters that surface from subsequent interactions, and we expect the strongest bonds to form first. This model, its basis, how it fits into the currently understood osmotic second virial coefficient approach to crystallization, and what it suggests will be discussed.

Pusey, Marc

Fluorescent Applications to Crystallization

By covalently modifying a subpopulation, less than or equal to 1%, of a macromolecule with a fluorescent probe, the labeled material will add to a growing crystal as a microheterogeneous growth unit. Labeling procedures can be readily incorporated into the final stages of purification, and tests with model proteins have shown that labeling u to 5 percent of the protein molecules does not affect the X-ray data quality obtained . The presence of the trace fluorescent label gives a number of advantages. Since the label is covalently attached to the protein molecules, it "tracks" the protein s response to the crystallization conditions. The covalently attached probe will concentrate in the crystal relative to the solution, and under fluorescent illumination crystals show up as bright objects against a darker background. Non-protein structures, such as salt crystals, do not show up under fluorescent illumination. Crystals have the highest protein concentration and are readily observed against less bright precipitated phases, which under white light illumination may obscure the crystals. Automated image analysis to find crystals should be greatly facilitated, without having to first define crystallization drop boundaries as the protein or protein structures is all that shows up. Fluorescence intensity is a faster search parameter, whether visually or by automated methods, than looking for crystalline features. Preliminary tests, using model proteins, indicates that we can use high fluorescence intensity regions, in the absence of clear crystalline features or "hits", as a means for determining potential lead conditions. A working hypothesis is that more rapid amorphous precipitation kinetics may overwhelm and trap more slowly formed ordered assemblies, which subsequently show up as regions of brighter fluorescence intensity. Experiments are now being carried out to test this approach using a wider range, of proteins. The trace fluorescently labeled crystals will also emit with sufficient intensity to aid in the automation of crystal alignment using relatively low cost optics, further increasing throughput at synchrotrons.

Pusey, Marc L.

The evolution of the protein synthesis system. I - A model of a primitive protein synthesis system

A model is developed to describe the evolution of the protein synthesis system. The model is comprised of two independent autocatalytic systems, one including one gene (A-gene) and two activated amino acid polymerases (O and A-polymerases), and the other including the addition of another gene (N-gene) and a nucleotide polymerase. Simulation results have suggested that even a small enzymic activity and polymerase specificity could lead the system to the most accurate protein synthesis, as far as permitted by transitions to systems with higher accuracy.

Mizutani, H.

Quick-and-Easy Validation of Protein–Ligand Binding Models Using Fragment-Based Semiempirical Quantum Chemistry

Electronic structure calculations in enzymes converge very slowly with respect to the size of the model region that is described using quantum mechanics (QM), requiring hundreds of atoms to obtain converged results and exhibiting substantial sensitivity (at least in smaller models) to which amino acids are included in the QM region. As such, there is considerable interest in developing automated procedures to construct a QM model region based on well-defined criteria. However, testing such procedures is burdensome due to the cost of large-scale electronic structure calculations. Here, we show that semiempirical methods can be used as alternatives to density functional theory (DFT) to assess convergence in sequences of models generated by various automated protocols. The cost of these convergence tests is reduced even further by means of a many-body expansion. We use this approach to examine convergence (with respect to model size) of protein–ligand binding energies. Fragment-based semiempirical calculations afford well-converged interaction energies in a tiny fraction of the cost required for DFT calculations. Two-body interactions between the ligand and single-residue amino acid fragments afford a low-cost way to construct a “QM-informed” enzyme model of reduced size, furnishing an automatable active-site model-building procedure. This provides a streamlined, user-friendly approach for constructing ligand binding-site models that needs neither a priori information nor manual adjustments. Extension to model-building for thermochemical calculations should be straightforward.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH

Smart culture medium optimization for recombinant protein production: Experimental, modeling, and AI/ML-driven strategies

Recombinant protein production (RPP) is central to biotechnology, where recombinant proteins are used as either end products or catalysts in the synthesis of chemicals, fuels, and materials. Among the major cost drivers, culture medium plays a pivotal role in determining protein yield and quality. This review presents a comprehensive perspective on the critical stages of “smart” culture medium optimization: planning, screening, modeling, optimization, and validation. In the planning stage, we examine the nutritional and energetic roles of medium components, including carbon, nitrogen, amino acids, salts, and trace metals, and their impacts on culture parameters such as pH, oxidative state, and osmolality. We highlight the variability in trace metal content due to water sources, culture vessels, and raw materials, which can substantially influence RPP. The screening stage covers Design of Experiments (DoE) approaches, assessing their theoretical basis, implementation, and limitations. For modeling, we describe methods that integrate experimental data to develop predictive models for smart medium formulation. Model-based optimization strategies can then be employed to select optimal media compositions for a given application. The validation stage aims to evaluate model predictions and provide feedback for model training and refinement. Finally, we survey mechanistic and artificial intelligence/machine learning (AI/ML)-driven models as integrated, transformational tools for predictive modeling of bioprocess conditions, nutrient availability, cellular metabolism, and protein quality, with the goal of optimizing culture media to enhance protein yields while reducing costs and environmental impact. We conclude by addressing the challenges of translating laboratory-scale medium optimization to industrial-scale settings and exploring future AI/ML-driven approaches that may overcome current bottlenecks and accelerate medium design for RPP. Overall, this review provides a unified framework for advancing smart medium design in RPP.

Artificial Intelligence/Machine Learning (AI/ML)