Search NASASearch

SEARCH · Search NASA

Results for “Protein Conformation”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

Ca X ML: Chemistry‐informed machine learning explains mutual changes between protein conformations and calcium ions in calcium‐binding proteins using structural and topological features

Proteins' flexibility is a feature in communicating changes in cell signaling instigated by binding with secondary messengers, such as calcium ions, associated with the coordination of muscle contraction, neurotransmitter release, and gene expression. When binding with the disordered parts of a protein, calcium ions must balance their charge states with the shape of calcium-binding proteins and their versatile pool of partners depending on the circumstances they transmit. Accurately determining the ionic charges of those ions is essential for understanding their role in such processes. However, it is unclear whether the limited experimental data available can be effectively used to train models to accurately predict the charges of calcium-binding protein variants. Here, we developed a chemistry-informed, machine-learning algorithm that implements a game theoretic approach to explain the output of a machine-learning model without the prerequisite of an excessively large database for high-performance prediction of atomic charges. We used the ab initio electronic structure data representing calcium ions and the structures of the disordered segments of calcium-binding peptides with surrounding water molecules to train several explainable models. Network theory was used to extract the topological features of atomic interactions in the structurally complex data dictated by the coordination chemistry of a calcium ion, a potent indicator of its charge state in protein. Our design created a computational tool of Ca X ML, which provided a framework of explainable machine learning model to annotate ionic charges of calcium ions in calcium-binding proteins in response to the chemical changes in an environment. Our framework will provide new insights into protein design for engineering functionality based on the limited size of scientific data in a genome space.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH

Tracking the protein conformational motions driving HIV-1 membrane fusion

HIV-1 Env (trimeric gp120/gp41) is the surface protein responsible for membrane fusion. The Env binds to the receptor proteins, which induces gp120 shedding leading to conformational changes of gp41 from the pre-fusion to post-fusion state, allowing its fusion peptide to embed in the host cell membrane and bringing the viral and host cell membranes together. The gp41 refolding is a target of several peptide inhibitors. Yet, the molecular mechanism of this dynamic process is still not well understood. In this study, we successfully simulate the conformational change of gp41 from pre-fusion to post-fusion state in atomistic resolution using all-atom structure-based models. We reveal that maintaining the directionality of protomer interactions in both pre-fusion and post-fusion states is crucial for gp41 refolding. Additionally, we find that HR1 inherently extends as a three-helical bundle toward the host-cell membrane without any bias. Importantly, we identify native contacts in the pre-fusion state that are critical for the proper refolding of gp41 towards the post-fusion state. Lastly, by incorporating the membrane-fusion inhibitors, T20 and SFT, we identify the most vulnerable stage in the fusion pathway that exhibits the greatest sensitivity to these drugs, which could aid in a better understanding of drug resistance mechanisms.

59 BASIC BIOLOGICAL SCIENCES

LHCSR1 Functions as a Dimmer Switch for Light Harvesting

In oxygenic photosynthesis, high light leads to a set of photoprotective processes, known as nonphotochemical quenching, that are required for fitness. In moss and algae, the pigment−protein complex, light-harvesting stress-related (LHCSR), is crucial for photoprotection. Acidification of the thylakoid lumen under high light triggers the activation of LHCSR and the conversion of the xanthophyll violaxanthin into zeaxanthin, which is found within LHCSR. These interrelated molecular components combine to turn on a safety valve that dissipates excess energy as heat. Previous studies of detergent-solubilized LHCSR1 found that pH and zeaxanthin regulate distinct quenching sites via the protein conformation. However, protein function in the native membrane environment can be significantly different. In this work, we applied single-molecule fluorescence spectroscopy to LHCSR1 in membrane nanodiscs. The membrane environment enhanced total quenching, regardless of pH or zeaxanthin-binding. pH-dependent quenching was still observed whereas zeaxanthin-dependent quenching was suppressed. Conformational changes also increased in the membrane, establishing that the local environment can change the nature of the equilibrium between light harvesting and photoprotection.

Hoffmann, Madeline P. [Massachusetts Inst. of Tech

From sequence to protein structure and conformational dynamics with artificial intelligence/machine learning

The 2024 Nobel Prize in Chemistry was awarded in part for de novo protein structure prediction using AlphaFold2, an artificial intelligence/machine learning (AI/ML) model trained on vast amounts of sequence and three-dimensional structure data. AlphaFold2 and related models, including RoseTTAFold and ESMFold, employ specialized neural network architectures driven by attention mechanisms to infer relationships between sequence and structure. At a fundamental level, these AI/ML models operate on the long-standing hypothesis that the structure of a protein is determined by its amino acid sequence. More recently, AlphaFold2 has been adapted for the prediction of multiple protein conformations by subsampling multiple sequence alignments. Herein, we provide an overview of the deterministic relationship between sequence and structure, which was hypothesized over half a century ago with profound implications for the biological sciences ever since. We postulate that protein conformational dynamics are also determined, at least in part, by amino acid sequence and that this relationship may be leveraged for construction of AI/ML models dedicated to predicting protein conformational ensembles. Accordingly, we describe a conceptual model architecture, which may be trained on sequence data in combination with conformationally sensitive structural information, coming primarily from nuclear magnetic resonance (NMR) spectroscopy. Notwithstanding certain limitations in this context, NMR offers abundant structural heterogeneity conducive to conformational ensemble prediction. As NMR and other data continue to accumulate, sequence-informed prediction of protein structural dynamics with AI/ML has the potential to emerge as a transformative capability across the biological sciences.

Artificial intelligence

Integrating Ultra-Coarse-Grained Protein Models into Accessible Workflows for Multiscale Molecular Dynamics

To capture protein conformational transitions using molecular dynamics (MD), several simulation resolutions covering different spatial and temporal scales are typically needed. All-atom (AA) simulations provide fine resolution, but are computationally infeasible for large systems over longer durations. Coarse-grained (CG) and ultra-coarse-grained (UCG) models have a lower resolution and computational cost while still being able to conserve essential protein features. Prior work on a Multiscale Machinelearned Modeling Infrastructure (MuMMI) combined both AA and CG simulations to study RAS-RAF protein interactions, leveraging CG models for longer time scales and using AA to investigate unusual conformations in greater detail. However, MuMMI is still resource-intensive, and this study aims to maximize exploration of the protein conformational space while reducing computational cost. In this paper, we build on prior work that integrates UCG models based on heterogeneous elastic network modeling (hENM) into the MuMMI workflow. We demonstrate that UCG models enable accurate sampling of protein conformations, focusing on simulating RAS-RAF protein interactions. Using higher-resolution CG Martini simulation data, we can automatically refine intramolecular interactions in UCG models. We present a scalable Python package that uses fluctuations observed in higher-resolution CG Martini simulations to estimate bond coefficients of the UCG model. We built novel machine learning-based backmapping methods to recover more detailed CG Martini structures from UCG structures, using diffusion models to learn the mapping between scales. Finally, we present UCG-mini-MuMMI, an accessible and less compute-intensive version of MuMMI as a resource for the scientific community. Incorporating UCG models into MD studies is applicable to a broad range of systems and proteins, and our study offers insights into the advantages and limitations of these methods.

Chemical structure

Discovery of Multiple Light-Harvesting States of the Photosynthetic Protein PE545

Cryptophytes are photosynthetic microalga that flourish in a remarkable diversity of natural environments by using pigment-containing proteins with absorption maxima tuned to each ecological niche. While this diversity in the absorption has been well established, the subsequent photophysics is highly sensitive to the local protein environment and so may exhibit similar variation. Thermal fluctuations of the protein conformation are expected to introduce photophysical heterogeneity of the pigments that may have evolved important functional properties in a manner similar to that of the absorption. However, such heterogeneity is averaged out in ensemble measurements and, therefore, has not yet been probed. Here, we report single-molecule measurements of phycoerythrin 545 (PE545), the prototypical cryptophyte antenna protein, in its native dimeric form. A conformational ensemble was resolved consisting of distinct photophysical states with different light-harvesting properties. Proteins that did not quench, partially quenched, or fully quenched absorbed light were observed. Light intensity increased the quenched-state population of the dimer, potentially as a mechanism to deal with the extreme light intensities found in aqueous environments. Cross-linking, which mimics local interactions, introduces this light-dependent functionality while also suppressing other conformational dynamics. The cellular organization can, therefore, actively modulate the protein conformation and dynamics, selecting for distinct levels of light harvesting. Furthermore, the complex conformational equilibrium provides an additional mechanism for cryptophytes and likely other photosynthetic organisms to optimize solar energy capture and conversion.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH

Challenges of conventional iterative all-atom and coarse-grained multiscale molecular dynamics

In this work, we evaluate the biomolecular dynamics behaviors when conventionally iterating between all-atom (AA) and coarse-grained (CG) molecular dynamics (MD) simulations over multiple cycles. We implemented the workflow to iterate between AA and CG in OpenMM, namely the iterative multiscale MD (iMMD) simulation workflow. In particular, we aim to identify practical applications for iterating between AA and CG simulations in a conventional manner without any constraints or model modifications. We evaluate the iMMD workflow on four representative systems, spanning folding of two soluble proteins and protein-protein as well as protein-lipid interactions of two membrane proteins. We observe that iteration between AA and CG representations could help the soluble proteins exit undesirable metastable states to fold, resulting from random protein structural distortions due to cycling. Consequently, the most reliable use of iterative AA and CG simulations appears to be to accelerating complex lipid mixing for membrane-bound protein systems rather than sampling protein conformational space. Our work explores the practical usages and limitations for iterative AA and CG simulations using readily available AA and CG force fields. The evaluated iMMD workflow in OpenMM is made available at https://github.com/lanl/iMMD.

59 BASIC BIOLOGICAL SCIENCES

Functional protein mining with conformal guarantees

Molecular structure prediction and homology detection offer promising paths to discovering protein function and evolutionary relationships. However, current approaches lack statistical reliability assurances, limiting their practical utility for selecting proteins for further experimental and in-silico characterization. To address this challenge, we introduce a statistically principled approach to protein search leveraging principles from conformal prediction, offering a framework that ensures statistical guarantees with user-specified risk and provides calibrated probabilities (rather than raw ML scores) for any protein search model. Our method (1) lets users select many biologically-relevant loss metrics (i.e. false discovery rate) and assigns reliable functional probabilities for annotating genes of unknown function; (2) achieves state-of-the-art performance in enzyme classification without training new models; and (3) robustly and rapidly pre-filters proteins for computationally intensive structural alignment algorithms. Our framework enhances the reliability of protein homology detection and enables the discovery of uncharacterized proteins with likely desirable functional properties.

59 BASIC BIOLOGICAL SCIENCES

Heterogeneous Multilayer Nanopores via Chemically Tuned Dielectric Breakdown for Single‐Molecule Sensing

Solid-state nanopores are powerful platforms for single-molecule sensing, yet their performance is often constrained by fabrication complexity, noise, and limited control over surface properties. Here we report a direct method to fabricate heterogeneous multilayer nanopores using chemically tuned controlled dielectric breakdown (CT-CDB). We integrate hBN, MoS 2 , or graphene atop a silicon nitride membrane to form five distinct bilayer and tri-layer architectures, with bare SiN x nanopore as a control. CT-CDB achieves pore formation reproducibly through material-stacks with high efficiency, good pore size control, and strong yield, validated by various characterizations. Transferrin protein translocation experiments, supported by simulations, reveal that multilayer configurations modulate protein conformations, ionic current blockade and dwell time distributions, reflecting combined effects of membrane type, interfacial chemistry, and local electric field gradients. A supervised machine learning framework is implemented to assist identifying multilayer structure effects embedded in signal signatures, with over 96% accuracy. This work presents a modular and scalable framework for functional nanopore engineering with complex structural integration, thereby expanding the potential of 2D materials in single-molecule sensing applications.

2D materials

Repetitive proteins that undergo large conformational changes evade structural prediction algorithms

Protein structure prediction algorithms, such as AlphaFold, have accelerated protein design and advanced the understanding of the relationship between amino acid sequence and protein structure. However, these algorithms are limited in their ability to predict the structures of conformationally dynamic, intrinsically disordered, and stimuli-responsive proteins. To evaluate sequence-to-structure predictions of such challenging proteins, we explored a class of conformationally dynamic, repeats-in-toxin (RTX) proteins. RTX proteins adopt intrinsically disordered conformations in the absence of calcium and undergo reversible folding into β-roll structures upon binding to calcium. RTX proteins are characterized by tandem repeats of the sequence GGXGXDXUX, in which X can be any amino acid and U is an aliphatic amino acid. We designed RTX sequence variants with global substitutions of nonconserved amino acids, tandem repeats of consensus sequences GGAGXDTLY, and tandem repeats of scrambled sequences GGAGXDTYL. AlphaFold2 and AlphaFold3 predicted that all of these RTX variants adopt β-roll structures, characteristic of wild-type RTX bound to calcium. However, modeling the predicted structures with molecular dynamics simulations and characterizing the protein variants with circular dichroism spectroscopy, small-angle x-ray scattering, and x-ray crystallography revealed that variants adopt diverse, sequence-dependent structures in the absence and presence of calcium. To better design proteins for applications in biotechnology and sustainability, it is critical to build predictive tools that consider intrinsically disordered protein states and validate these tools with multi-mode, multi-scale experimental data.

Chang, Marina P. [Stanford Univ., CA (United State

Conformational Dynamics and Catalytic Backups in a Hyper-thermostable Engineered Archaeal Protein Tyrosine Phosphatase

Protein tyrosine phosphatases (PTPs) are a family of enzymes that play important roles in regulating cellular signaling pathways. The activity of these enzymes is regulated by the motion of a catalytic loop that places a critical conserved aspartic acid side chain into the active site for acid–base catalysis upon loop closure. These enzymes also have a conserved phosphate-binding loop that is typically highly rigid and forms a well-defined anion-binding nest. The intimate links between loop dynamics and chemistry in these enzymes make PTPs an excellent model system for understanding the role of loop dynamics in protein function and evolution. In this context, archaeal PTPs, which have often evolved in extremophilic organisms, are highly understudied, despite their unusual biophysical properties. We present here an engineered chimeric PTP (ShufPTP) generated by shuffling the amino acid sequence of five extant hyperthermophilic archaeal PTPs. Despite ShufPTP’s high sequence similarity to its natural counterparts, it presents a suite of unique properties, including high flexibility of the phosphate binding P-loop, facile oxidation of the active-site cysteine, mechanistic promiscuity, and, most notably, hyperthermostability, with a denaturation temperature likely >130 °C (>8 °C higher than the highest recorded growth temperature of any archaeal strain). Our combined structural, biochemical, biophysical, and computational analysis provides insight both into how small steps in evolutionary space can radically modulate the biophysical properties of an enzyme and showcases the tremendous potential of archaeal enzymes for biotechnology, to generate novel enzymes capable of operating under extreme conditions.

archaea

A curated benchmark for cofolding models on kinase conformational states

Abstract Protein kinases are critical drug targets, requiring therapeutics that can modulate their active and inactive conformational states. While cofolding models can generate global folds directly from kinase sequences and ligand SMILES strings, these models have not yet been tested on their ability to recover ligand-induced-fit conformational states of the kinase proteins. Here, we introduce KinConfBench, a curated benchmark of 2225 high-quality human kinase chains to evaluate the ability of four state-of-the-art cofolding models—Boltz-2, Chai-1, Protenix, and RoseTTAFold-All-Atom—to recover both canonical and rare conformational states. We show that geometric success metrics of a ligand pose in the active site do not correlate strongly with the correct kinase conformational state, motivating a new set of dynamical benchmarks for assessing cofolding models. While all four cofolding models achieve ~60–80% prediction accuracy for kinase conformational classification, they exhibit severe mode collapse when performing multiple inferences, show negligible structural diversity in sampling induced-fit motions, and display a prevalent “apo-drift” in which most cofolding models predominantly predict the kinase to be in its ligand-free state. Our results highlight that capturing ligand-induced protein conformational diversity, not just geometric fit, is critical for next-generation structure-based drug discovery.

Sun, Kunyang

SEC ‐ SAXS / MC Ensemble Structural Studies of the Microtubule Binding Protein Cdt1 Show Monomeric, Folded‐Over Conformations

ABSTRACT Cdt1 is a mixed folded protein critical for DNA replication licensing and it also has a “moonlighting” role at the kinetochore via direct binding to microtubules and the Ndc80 complex. However, it is unknown how the structure and conformations of Cdt1 could allow it to participate in these multiple, unique sets of protein complexes. While robust methods exist to study entirely folded or unfolded proteins, structure–function studies of combined, mixed folded/disordered proteins remain challenging. In this work, we employ orthogonal biophysical and computational techniques to provide structural characterization of mitosis‐competent human Cdt1. Thermal stability analyses shows that both folded winged helix domains1 are unstable. CD and NMR show that the N‐terminal and linker regions are intrinsically disordered. DLS shows that Cdt1 is monomeric and polydisperse, while SEC‐MALS confirms that it is monomeric at high concentrations, but without any apparent inter‐molecular self‐association. SEC‐SAXS enabled computational modeling of the protein structures. Using the program SASSIE, we performed rigid body Monte Carlo simulations to generate a conformational ensemble of structures. We observe that neither fully extended nor extremely compact Cdt1 conformations are consistent with SAXS. The best‐fit models have the N‐terminal and linker disordered regions extended into the solution and the two folded domains close to each other in apparent “folded over” conformations. We hypothesize the best‐fit Cdt1 conformations could be consistent with a function as a scaffold protein that may be sterically blocked without binding partners. Our study also provides a template for combining experimental and computational techniques to study mixed‐folded proteins.

Cell Biology

Structural, biophysical, and biochemical insights into C–S bond cleavage by dimethylsulfone monooxygenase

Sulfur is an essential element for life. Bacteria can obtain sulfur from inorganic sulfate; but in the sulfur starvation–induced response,Pseudomonadsemploy two-component flavin-dependent monooxygenases (TC-FMOs) from themsuandsfnoperons to assimilate sulfur from environmental compounds including alkanesulfonates and dialkylsulfones. Here, we report binding studies of oxidized FMN to enzymes involved within theP. fluorescensenzymatic pathway responsible for converting dimethylsulfone (DMSO 2 ) to sulfite. In this catabolic pathway, SfnG serves as the initial TC-FMO for sulfur assimilation, which is investigated in detail by solving the 2.6-Å resolution crystal structure of unliganded SfnG and the 1.75-Å resolution crystal structure of the SfnG ternary complex containing FMN and DMSO 2 . We find that SfnG adopts a (β/α) 8 barrel fold with a distinct quaternary configuration from other tetrameric class C TC-FMOs. To probe the unexpected tetramer arrangement, structural heterogeneity is assessed by chromatography and light scattering to confirm ligand binding correlates with a tetramer. Binding of FMN and DMSO 2 accompanies ordering of the active site, with DMSO 2 bound on thesi-face of the flavin. A previously unobserved protein backbone conformation is found within the oxygen-binding site on there-face of the flavin. Functional assays and the positioning of ligands with respect to the oxygen-binding site are consistent with use of an N5-(hydro)peroxyflavin pathway. Biochemical endpoint assays and docking studies reveal SfnG breaks the C–S bond of a range of dialkylsulfones.

Science & Technology - Other Topics