Search NASASearch

SEARCH · Search NASA

Results for “Protein modeling”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

Integrating Ultra-Coarse-Grained Protein Models into Accessible Workflows for Multiscale Molecular Dynamics

To capture protein conformational transitions using molecular dynamics (MD), several simulation resolutions covering different spatial and temporal scales are typically needed. All-atom (AA) simulations provide fine resolution, but are computationally infeasible for large systems over longer durations. Coarse-grained (CG) and ultra-coarse-grained (UCG) models have a lower resolution and computational cost while still being able to conserve essential protein features. Prior work on a Multiscale Machinelearned Modeling Infrastructure (MuMMI) combined both AA and CG simulations to study RAS-RAF protein interactions, leveraging CG models for longer time scales and using AA to investigate unusual conformations in greater detail. However, MuMMI is still resource-intensive, and this study aims to maximize exploration of the protein conformational space while reducing computational cost. In this paper, we build on prior work that integrates UCG models based on heterogeneous elastic network modeling (hENM) into the MuMMI workflow. We demonstrate that UCG models enable accurate sampling of protein conformations, focusing on simulating RAS-RAF protein interactions. Using higher-resolution CG Martini simulation data, we can automatically refine intramolecular interactions in UCG models. We present a scalable Python package that uses fluctuations observed in higher-resolution CG Martini simulations to estimate bond coefficients of the UCG model. We built novel machine learning-based backmapping methods to recover more detailed CG Martini structures from UCG structures, using diffusion models to learn the mapping between scales. Finally, we present UCG-mini-MuMMI, an accessible and less compute-intensive version of MuMMI as a resource for the scientific community. Incorporating UCG models into MD studies is applicable to a broad range of systems and proteins, and our study offers insights into the advantages and limitations of these methods.

Chemical structure

Modeling Protein–Protein and Protein–Ligand Interactions by the ClusPro Team in CASP16

ABSTRACT In the CASP16 experiment, our team employed hybrid computational strategies to predict both protein–protein and protein–ligand complex structures. For protein–protein docking, we combined physics‐based sampling—using ClusPro FFT docking and molecular dynamics—with AlphaFold (AF)‐based sampling, followed by AF‐based refinement. Our method produced numerous high‐accuracy complex models, including cases where AF alone failed, underscoring the critical role of physics‐based sampling alongside deep learning‐based refinement. For protein–ligand docking, we integrated the ClusPro LigTBM template‐based approach with a machine learning‐based confidence model for rescoring. The method preserves conserved interaction fragments derived from homologous complexes, followed by local resampling using physics‐based sampling and a diffusion model. Our template‐based strategy achieved a mean lDDT‐PLI of 0.69 across 233 targets, which was highly competitive. These results demonstrate that combining physics‐based modeling with AI‐driven refinement can significantly enhance the accuracy of both protein–protein and protein–ligand structure predictions.

Ashizawa, Ryota [Department of Applied Mathematics

Predicting compatibility between ferredoxins and the Fe protein of nitrogenase using in silico protein modeling

Biological nitrogen fixation is the process by which certain bacteria and archaea use the enzyme nitrogenase to reduce atmospheric nitrogen into bioavailable ammonium. Engineering non‐nitrogen‐fixing organisms, like plants, to use nitrogenase could reduce dependency on synthetic fertilizer and mitigate the environmental impacts of industrial fertilizer production. However, nitrogenase activity requires delivery of reducing power by small electron carrying proteins known as ferredoxins and flavodoxins, and successfully engineering nitrogenase into new systems will require a mechanistic understanding of electron delivery by these proteins. Most organisms often have multiple ferredoxins, raising the question of which ferredoxin can support nitrogenase activity. The purpose of this study is to gain insight into how we can predict which ferredoxin is compatible with the Fe protein, the component of nitrogenase that interacts with ferredoxin or flavodoxin. Our in silico protein–protein docking simulations reveal that most ferredoxins and flavodoxins involved in nitrogen fixation have the shortest distance (≤10 Å) between their redox cofactor and the [4Fe‐4S] cluster of the Fe protein. We found shorter cofactor distance contributes to faster intermolecular electron tunneling rates. Bacterial ferredoxins that play a role in nitrogen fixation also exhibit more complementary interactions with the Fe protein than bacterial and plant ferredoxins not involved in this process. Heterologous expression of a set of ferredoxins from both nitrogen‐fixing and non‐nitrogen‐fixing bacteria in the diazotroph Rhodopseudomonas palustris supports our model‐derived prediction that shorter distances between the electron‐carrying cofactors favor nitrogenase compatibility. These findings offer a framework to predict and potentially enhance ferredoxin–nitrogenase compatibility, which will help to improve our ability to engineer nitrogen fixation into non‐nitrogen‐fixing organisms like plants.

59 BASIC BIOLOGICAL SCIENCES

Reweighting configurations generated by transferable, machine learned models for protein sidechain backmapping

Multiscale modeling requires the linking of models at different levels of detail, with the goal of gaining accelerations from lower fidelity models while recovering fine details from higher resolution models. Communication across resolutions is particularly important in modeling soft matter, where tight couplings exist between molecular-level details and mesoscale structures. While multiscale modeling of biomolecules has become a critical component in exploring their structure and self-assembly, backmapping from coarse-grained to fine-grained, or atomistic, representations presents a challenge, despite recent advances through machine learning. A major hurdle, especially for strategies utilizing machine learning, is that backmappings can only approximately recover the atomistic ensemble of interest. We demonstrate conditions for which backmapped configurations may be reweighted to exactly recover the desired atomistic ensemble. By training separate decoding models for each sidechain type, we develop an algorithm based on normalizing flows and geometric algebra attention to autoregressively propose backmapped configurations for any protein sequence. Critical for reweighting with modern protein force fields, our trained models include all hydrogen atoms in the backmapping and make probabilities associated with atomistic configurations directly accessible. We also demonstrate, however, that reweighting is extremely challenging despite state-of-the-art performance on recently developed metrics and generation of configurations with low energies in atomistic protein force fields. Through detailed analysis of configurational weights, we show that machine-learned backmappings must not only generate configurations with reasonable energies, but also correctly assign relative probabilities under the generative model. These are broadly important considerations in generative modeling of atomistic molecular configurations.

Monroe, Jacob I. [Univ. of Arkansas, Fayetteville,

Anisotropic interactions for continuum modeling of protein–membrane systems

In this work, a model for anisotropic interactions between proteins and cellular membranes is proposed for large-scale continuum simulations. The framework of the model is based on dynamic density functional theory, which provides a formalism to describe the lipid densities within the membrane as continuum fields while still maintaining the fidelity of the underlying molecular interactions. Within this framework, we extend recent results to include the anisotropic effects of protein–lipid interactions. As applications, we consider two membrane proteins of biological interest: a RAS–RAF complex tethered to the membrane and a membrane embedded G protein-coupled receptor. A strong qualitative and quantitative agreement is found between the numerical results and the corresponding molecular dynamics simulations. Combining the scope of continuum level simulations with the details from molecular level particle simulations enables research into protein–membrane behaviors at a more biologically relevant scale, which crucially can also be accessed via experiment.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH

De novo atomic protein structure modeling for cryoEM density maps using 3D transformer and HMM

Accurately building 3D atomic structures from cryo-EM density maps is a crucial step in cryo-EM-based protein structure determination. Converting density maps into 3D atomic structures for proteins lacking accurate homologous or predicted structures as templates remains a significant challenge. Here, we introduce Cryo2Struct, a fully automated de novo cryo-EM structure modeling method. Cryo2Struct utilizes a 3D transformer to identify atoms and amino acid types in cryo-EM density maps, followed by an innovative Hidden Markov Model (HMM) to connect predicted atoms and build protein backbone structures. Cryo2Struct produces substantially more accurate and complete protein structural models than the widely used ab initio method Phenix. Additionally, its performance in building atomic structural models is robust against changes in the resolution of density maps and the size of protein structures.

59 BASIC BIOLOGICAL SCIENCES

Quick-and-Easy Validation of Protein–Ligand Binding Models Using Fragment-Based Semiempirical Quantum Chemistry

Electronic structure calculations in enzymes converge very slowly with respect to the size of the model region that is described using quantum mechanics (QM), requiring hundreds of atoms to obtain converged results and exhibiting substantial sensitivity (at least in smaller models) to which amino acids are included in the QM region. As such, there is considerable interest in developing automated procedures to construct a QM model region based on well-defined criteria. However, testing such procedures is burdensome due to the cost of large-scale electronic structure calculations. Here, we show that semiempirical methods can be used as alternatives to density functional theory (DFT) to assess convergence in sequences of models generated by various automated protocols. The cost of these convergence tests is reduced even further by means of a many-body expansion. We use this approach to examine convergence (with respect to model size) of protein–ligand binding energies. Fragment-based semiempirical calculations afford well-converged interaction energies in a tiny fraction of the cost required for DFT calculations. Two-body interactions between the ligand and single-residue amino acid fragments afford a low-cost way to construct a “QM-informed” enzyme model of reduced size, furnishing an automatable active-site model-building procedure. This provides a streamlined, user-friendly approach for constructing ligand binding-site models that needs neither a priori information nor manual adjustments. Extension to model-building for thermochemical calculations should be straightforward.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH

Smart culture medium optimization for recombinant protein production: Experimental, modeling, and AI/ML-driven strategies

Recombinant protein production (RPP) is central to biotechnology, where recombinant proteins are used as either end products or catalysts in the synthesis of chemicals, fuels, and materials. Among the major cost drivers, culture medium plays a pivotal role in determining protein yield and quality. This review presents a comprehensive perspective on the critical stages of “smart” culture medium optimization: planning, screening, modeling, optimization, and validation. In the planning stage, we examine the nutritional and energetic roles of medium components, including carbon, nitrogen, amino acids, salts, and trace metals, and their impacts on culture parameters such as pH, oxidative state, and osmolality. We highlight the variability in trace metal content due to water sources, culture vessels, and raw materials, which can substantially influence RPP. The screening stage covers Design of Experiments (DoE) approaches, assessing their theoretical basis, implementation, and limitations. For modeling, we describe methods that integrate experimental data to develop predictive models for smart medium formulation. Model-based optimization strategies can then be employed to select optimal media compositions for a given application. The validation stage aims to evaluate model predictions and provide feedback for model training and refinement. Finally, we survey mechanistic and artificial intelligence/machine learning (AI/ML)-driven models as integrated, transformational tools for predictive modeling of bioprocess conditions, nutrient availability, cellular metabolism, and protein quality, with the goal of optimizing culture media to enhance protein yields while reducing costs and environmental impact. We conclude by addressing the challenges of translating laboratory-scale medium optimization to industrial-scale settings and exploring future AI/ML-driven approaches that may overcome current bottlenecks and accelerate medium design for RPP. Overall, this review provides a unified framework for advancing smart medium design in RPP.

Artificial Intelligence/Machine Learning (AI/ML)

DNA Origami Incorporated into Solid-State Nanopores Enables Enhanced Sensitivity for Precise Analysis of Protein Translocations

The rapidly advancing field of nanotechnology is driving the development of precise sensing methods at the nanoscale, with solid-state nanopores emerging as promising tools for biomolecular sensing. Here, this study investigates the increased sensitivity of solid-state nanopores achieved by integrating DNA origami structures, leading to the improved analysis of protein translocations. Using holo human serum transferrin (holo-hSTf) as a model protein, we compared hybrid nanopores incorporating DNA origami with open solid-state nanopores. Results show a significant enhancement in holo-hSTf detection sensitivity with DNA origami integration, suggesting a unique role of DNA interactions beyond confinement. This approach holds potential for ultrasensitive protein detection in biosensing applications, offering advancements in biomedical research and diagnostic tool development for diseases with low-abundance protein biomarkers. Further exploration of origami designs and nanopore configurations promises even greater sensitivity and versatility in the detection of a wider range of proteins, paving the way for advanced biosensing technologies.

77 NANOSCIENCE AND NANOTECHNOLOGY

Protein Adhesion on Semi-Fluorinated Polystyrene Surfaces in Static and Dynamic Measurements

Reducing protein adhesion is a critical strategy in fouling-resistant material innovation, with broad applications spanning biomedical and healthcare devices, biosensors, industrial and environmental systems, and other important technological domains. Here, in this study, we elucidated protein adhesion behavior on polystyrene-based thin films by neutron reflectometry (NR) and quartz crystal microbalance with dissipation (QCM-D), using both lysozyme and bovine serum albumin (BSA) as model proteins. To this end, semifluorinated polystyrene thin films with gradient wettability and surface energy were fabricated through dry processing using plasma oxidation and gas-phase deposition. Although it is believed that a fully fluorinated alkyl chain offers extremely low surface energy, thus rejecting foulants, and has been used in many fouling-resistant surface designs, enhanced protein–surface interactions were observed consistently in NR and QCM-D results, due to the combined effects of surface morphology and chemistry. On the contrary, depositing shorter fluorinated silane onto a hydrophilic PS surface contributed to a more homogeneous nanoscale fluorine coating, resulting in less initial protein adsorption and improved surface recovery. Comparative analysis of proteins with different sizes on the nanopatterned semifluorinated surface revealed the influence of molecular characteristics on surface interactions. Lysozyme, being smaller and more compact, showed faster adsorption kinetics and higher surface coverage but largely reversible binding, whereas BSA, with its larger and more flexible structure, formed broader and more stable interfacial layers. This study fills the gap in understanding protein adhesion within the range of hydrophobicity (water contact angle ∼90°), as current strategies often associate with extreme hydrophilic and superhydrophobic surfaces due to hydration or low-surface-energy rejection mechanisms, respectively. It also provides in-depth insights into current combinatorial fouling-resistant surface design.

Yuan, Yue [Oak Ridge National Laboratory (ORNL), O

Native Top-Down Mass Spectrometry Characterization of Model Integral Membrane Protein Bacteriorhodopsin

Bacteriorhodopsin (bR) from Halobacterium salinarum has been a model system for structural biology and is a structural template for the characterization of membrane G-protein couple receptors (GPCRs) in particular. Here, in this study, wild-type bacteriorhodopsin and two single-residue mutants were characterized by native top-down mass spectrometry (nTD-MS) with Orbitrap-based high-energy collision dissociation (HCD) and electron capture dissociation (ECD). After in-source dissociation ejected the membrane protein from detergent micelles, high-resolution native MS measurement allowed for identification of multiple proteoforms as well as lipid-bound forms. Further top-down MS measurements by HCD produced a large number of product ions for in-depth sequencing and unambiguous localization of post-translational modifications. For the first time, native TD-MS with ECD was used to characterize an integral membrane protein. ECD yielded fragments originating from all helices and loop regions, even accessing a sequence stretch that HCD could not. Combining HCD and ECD fragmentation patterns significantly enhanced the sequence coverage of bR. We propose bR to be a model analyte for testing nTD-MS performance for membrane proteins.

crystal cleavage

Mechanisms of Polyethylene Terephthalate Pellet Fragmentation into Nanoplastics and Assimilable Carbons by Wastewater Comamonas

Comamonadaceae bacteria are enriched on poly(ethylene terephthalate) (PET) microplastics in wastewaters and urban rivers, but the PET-degrading mechanisms remain unclear. Here, we investigated these mechanisms with Comamonas testosteroniKF-1, a wastewater isolate, by combining microscopy, spectroscopy, proteomics, protein modeling, and genetic engineering. Compared to minor dents on PET films, scanning electron microscopy revealed significant fragmentation of PET pellets, resulting in a 3.5-fold increase in the abundance of small nanoparticles (<100 nm) during 30-day cultivation. Infrared spectroscopy captured primarily hydrolytic cleavage in the fragmented pellet particles. Solution analysis further demonstrated double hydrolysis of a PET oligomer, bis(2-hydroxyethyl) terephthalate, to the bioavailable monomer terephthalate. Supplementation with acetate, a common wastewater co-substrate, promoted cell growth and PET fragmentation. Of the multiple hydrolases encoded in the genome, intracellular proteomics detected only one, which was found in both acetate-only and PET-only conditions. Homology modeling of this hydrolase structure illustrated substrate binding analogous to reported PET hydrolases, despite dissimilar sequences. Mutants lacking this hydrolase gene were incapable of PET oligomer hydrolysis and had a 21% decrease in PET fragmentation; re-insertion of the gene restored both functions. Thus, we have identified constitutive production of a key PET-degrading hydrolase in wastewater Comamonas, which could be exploited for plastic bioconversion.

54 ENVIRONMENTAL SCIENCES

Polyketide synthase–like functionality acquired by plant fatty acid elongase

Fatty acid elongation typically proceeds through a four-step cycle of condensation, reduction, dehydration, and reduction for each two-carbon extension. Here, we describe a variation of this pathway in Orychophragmus limprichtianus, whose seed oil contains previously unknown C24-C28 keto-hydroxy fatty acids that account for ~25% of total fatty acids. These compounds are produced through an endoplasmic reticulum–localized discontinuous elongation process in which a 3-keto-hydroxy intermediate bypasses full reduction and is extended through a polyketide synthase–like mechanism. Transcriptomic and functional assays identified two divergent enzymes, a variant fatty acid elongase 1 (FAE1) and a low-activity 3-ketoacyl-CoA reductase (KCR1), as central to this process. Protein modeling and mutant analysis suggest that specific amino acid substitutions underlie altered KCR1 activity, enabling accumulation of keto intermediates. Our findings reveal unexpected flexibility in plant fatty acid elongation and provide innovative tools for engineering plants and microbes to produce renewable oils with tailored industrial functions.

59 BASIC BIOLOGICAL SCIENCES

Device and Method for Parallel Measurement of Phosphoproteome and Proteome from Single Cells

We present the development of an immobilized metal affinity chromatography (IMAC) chip designed to enable nanoscale phosphopeptide enrichment within microfabricated nanowells. This novel platform leverages surface chemistry to immobilize high-density Nickel-Nitrilotriacetic Acid (Ni-NTA) molecules on nanowells, followed by applying Fe 3+ . The nanowell surface serves as a capture media to enrich phosphopeptides based on IMAC. The system's efficiency was validated using ß-casein as a model protein, demonstrating the chip’s capability to significantly enrich phosphopeptides. Future applications of this technology are anticipated to enable the detection of over 100 phosphopeptides from individual cells and more than 500 phosphopeptides from pools of 100 cells, offering exciting potential for single-cell phosphoproteomics. We will next apply an integrated proteomics workflow to perform multi-omics measurements, including single-cell isolation, protein digestion, and phosphopeptide enrichment, followed by LC-MS analysis of both the global proteome and phosphoproteome. Future research will explore the use of this technology to study phosphorylation dynamics in cancer cells, enhancing our understanding of cellular signaling and disease mechanisms.

59 BASIC BIOLOGICAL SCIENCES

Assessing Metal Ion Assignment Accuracy in Protein Data Bank Models via Elemental Spectroscopy

Accurate representation of metal ions in macromolecular structures is critical for chemical interpretation, computational modeling, and machine-learning methods that rely on Protein Data Bank (PDB) entries. However, the elemental identity of metals modeled in crystallographic structures is often inferred indirectly and rarely validated experimentally. Here, we combine Particle Induced X-ray Emission (PIXE) and X-ray Fluorescence Spectroscopy (XRFS) to determine the elemental composition of protein samples used to generate 70 deposited metalloprotein crystal structures. By analyzing the original protein material employed for crystallization, but before the addition of crystallization buffer solutions, we assess whether the modeled metal ions in deposited structures are consistent with experimentally detectable elemental content. We find that in a majority of cases, the metals modeled in the corresponding PDB entries are inconsistent with the metals present in the protein samples before crystallization, or that additional metals are present but not represented in the structural models. Spectroscopic results were integrated with automated crystallographic validation metrics, including real-space Z-difference (RSZD) analysis and systematic rerefinement, to evaluate atomic-number mismatch at metal sites. PIXE and XRFS show strong agreement for dominant elemental signals and provide complementary, scalable approaches for identifying suspect metal assignments. This work does not address physiological or functional metalation but instead highlights a widespread data integrity issue in deposited macromolecular structures, PDB-wide. These results establish an experimentally corroborated link between elemental identity and crystallographic validation metrics, enabling the large-scale detection of chemically inconsistent annotations in structural databases used for computational modeling and machine learning.

Crystallization

QM Investigation of Rare Earth Ion Interactions with First Hydration Shell Waters and Protein-Based Coordination Models

Here, conventional methods for extracting rare earth metals (REMs) from mined mineral ores are inefficient, expensive, and environmentally damaging. Recent discovery of lanmodulin (LanM), a protein that coordinates REMs with high-affinity and selectivity over competing ions, provides inspiration for new REM refinement methods. Here, we used quantum mechanical (QM) methods to investigate trivalent lanthanide cation (Ln 3+ ) interactions with coordination systems representing bulk solvent water and protein binding sites. Energy decomposition analysis (EDA) showed differences in the energetic components of Ln 3+ interaction with representatives of solvent (water, H 2 O) and protein binding sites (acetate, CH 3 COO – ), highlighting the importance of accurate description of electrostatics and polarization in computational modeling of REM interactions with biological and bioinspired molecules. Relative binding free energies were obtained for Ln 3+ with coordination complexes originating from binding sites in PDB structures of a lanthanum binding peptide (PDB entry 7CCO) and LanM, with explicit consideration of the first hydration shell waters, according to quasi-chemical theory (QCT). Beyond the first shell, the bulk solvent environment was represented with an implicit continuum model. Ln 3+ interactions with (H 2 O) 9 and both binding site models became more favorable, moving down the periodic series. This trend was more pronounced with the protein binding site models than with water, resulting in affinity increasing with periodic number, except for the last REM, Lu 3+ , which bound less favorably than the preceding element, Yb 3+ . Using the truncated 7CCO binding site model, the magnitude and trend of the experimental Ln 3+ relative binding free energies for the whole 7CCO peptide were reproduced. Conversely, the previously reported experimental data for LanM show a preference for the earlier lanthanides; this is likely due to longer-range interactions and cooperative effects, which are not represented by the reduced models. Using the truncated 7CCO binding site model, the magnitude and trend of the experimental Ln 3+ relative binding free energies for the whole 7CCO peptide were reproduced. In contrast to the previously reported experimental data for LanM, the peptide preferentially binds the earlier lanthanides. This difference likely arises due to longer-range interactions and cooperative effects not represented by the peptide. Further investigation of Ln 3+ interactions with whole proteins using polarizable molecular mechanics models with explicit solvent is warranted to understand the influence of longer-ranged interactions, cooperativity, and bulk solvent. Nevertheless, the present work provides new insights into Ln 3+ interactions with biomolecules and presents an effective computational platform for designing specific single-site REM binding peptides more efficiently.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH

Covalent Drug Binding in Live Cells Monitored by Mid-Infrared Quantum Cascade Laser Spectroscopy: Photoactive Yellow Protein as a Model System

The detection of drug-target interactions in live cells enables analysis of therapeutic compounds in a native cellular environment. Recent advances in spectroscopy and molecular biology have facilitated the development of genetically encoded vibrational probes like nitriles that can sensitively report on molecular interactions. Nitriles are powerful tools for measuring electrostatic environments within condensed media like proteins, but such measurements in live cells have been hindered by low signal-to-noise ratios. In this study, we design a spectrometer based on a double-beam quantum cascade laser (QCL)-based transmission infrared (IR) source with balanced detection that can significantly enhance sensitivity to nitrile vibrational probes embedded in proteins within cells compared to a conventional FTIR spectrometer. Here, using this approach, we detect small-molecule binding in Escherichia coli, with particular focus on the interaction between para-Coumaric acid (pCA) and nitrile-incorporated photoactive yellow protein (PYP). This system effectively serves as a model for investigating covalent drug binding in a cellular environment. Notably, we observe large spectral shifts of up to 15 cm –1 for nitriles embedded in PYP between the unbound and drug-bound states directly within bacteria, in agreement with observations for purified proteins. Such large spectral shifts are ascribed to the changes in the hydrogen-bonding environment around the local environment of nitriles, accurately modeled through high-level molecular dynamics simulations using the AMOEBA force field. Our findings underscore the QCL spectrometer’s ability to enhance sensitivity for monitoring drug–protein interactions, offering new opportunities for advanced methodologies in drug development and biochemical research.

chromophores