Search NASA⌕ Search

SEARCH · Search NASA

Results for “protein folding ensemble”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

Population-based heteropolymer design to mimic protein mixtures

Biological fluids, the most complex blends, have compositions that constantly vary and cannot be molecularly defined. Despite these uncertainties, proteins fluctuate, fold, function and evolve as programmed. We propose that in addition to the known monomeric sequence requirements, protein sequences encode multi-pair interactions at the segmental level to navigate random encounters; synthetic heteropolymers capable of emulating such interactions can replicate how proteins behave in biological fluids individually and collectively. Here, we extracted the chemical characteristics and sequential arrangement along a protein chain at the segmental level from natural protein libraries and used the information to design heteropolymer ensembles as mixtures of disordered, partially folded and folded proteins. For each heteropolymer ensemble, the level of segmental similarity to that of natural proteins determines its ability to replicate many functions of biological fluids including assisting protein folding during translation, preserving the viability of fetal bovine serum without refrigeration, enhancing the thermal stability of proteins and behaving like synthetic cytosol under biologically relevant conditions. Molecular studies further translated protein sequence information at the segmental level into intermolecular interactions with a defined range, degree of diversity and temporal and spatial availability. This framework provides valuable guiding principles to synthetically realize protein properties, engineer bio/abiotic hybrid materials and, ultimately, realize matter-to-life transformations.

59 BASIC BIOLOGICAL SCIENCES↗

SEC ‐ SAXS / MC Ensemble Structural Studies of the Microtubule Binding Protein Cdt1 Show Monomeric, Folded‐Over Conformations

ABSTRACT Cdt1 is a mixed folded protein critical for DNA replication licensing and it also has a “moonlighting” role at the kinetochore via direct binding to microtubules and the Ndc80 complex. However, it is unknown how the structure and conformations of Cdt1 could allow it to participate in these multiple, unique sets of protein complexes. While robust methods exist to study entirely folded or unfolded proteins, structure–function studies of combined, mixed folded/disordered proteins remain challenging. In this work, we employ orthogonal biophysical and computational techniques to provide structural characterization of mitosis‐competent human Cdt1. Thermal stability analyses shows that both folded winged helix domains1 are unstable. CD and NMR show that the N‐terminal and linker regions are intrinsically disordered. DLS shows that Cdt1 is monomeric and polydisperse, while SEC‐MALS confirms that it is monomeric at high concentrations, but without any apparent inter‐molecular self‐association. SEC‐SAXS enabled computational modeling of the protein structures. Using the program SASSIE, we performed rigid body Monte Carlo simulations to generate a conformational ensemble of structures. We observe that neither fully extended nor extremely compact Cdt1 conformations are consistent with SAXS. The best‐fit models have the N‐terminal and linker disordered regions extended into the solution and the two folded domains close to each other in apparent “folded over” conformations. We hypothesize the best‐fit Cdt1 conformations could be consistent with a function as a scaffold protein that may be sterically blocked without binding partners. Our study also provides a template for combining experimental and computational techniques to study mixed‐folded proteins.

Cell Biology↗

CryoFold: Determining protein structures and data-guided ensembles from cryo-EM density maps

Cryoelectron microscopy requires molecular modeling for refinement of structures. Ensemble models arrive at low free-energy molecular structures, but are computationally expensive and limited to resolving only small proteins. Here, we introduce CryoFold, a pipeline of molecular dynamics simulations that determines ensembles of protein structures by integrating density data of varying sparsity at 3–5 Å resolution with sequence information and coarse-grained topological knowledge of the protein folds. We present six examples, folding proteins between 72 and 2,000 residues, including large membrane and multi-domain systems, and results from two Electron Microscopy Data Bank (EMDB) competitions. Driven by data from a single state, CryoFold discovers ensembles of common low-energy models together with rare low-probability structures that capture the equilibrium distribution of proteins constrained by the density maps. Many of these conformations are experimentally validated and functionally relevant. We arrive at a set of best practices for data-guided protein folding that are controlled using a Python graphical user interface (GUI).

59 BASIC BIOLOGICAL SCIENCES↗

Properties of protein unfolded states suggest broad selection for expanded conformational ensembles

Much attention is being paid to conformational biases in the ensembles of intrinsically disordered proteins. However, it is currently unknown whether or how conformational biases within the disordered ensembles of foldable proteins affect function in vivo. Recently, we demonstrated that water can be a good solvent for unfolded polypeptide chains, even those with a hydrophobic and charged sequence composition typical of folded proteins. These results run counter to the generally accepted model that protein folding begins with hydrophobicity-driven chain collapse. Here we investigate what other features, beyond amino acid composition, govern chain collapse. We found that local clustering of hydrophobic and/or charged residues leads to significant collapse of the unfolded ensemble of pertactin, a secreted autotransporter virulence protein from Bordetella pertussis , as measured by small angle X-ray scattering (SAXS). Sequence patterns that lead to collapse also correlate with increased intermolecular polypeptide chain association and aggregation. Crucially, sequence patterns that support an expanded conformational ensemble enhance pertactin secretion to the bacterial cell surface. Similar sequence pattern features are enriched across the large and diverse family of autotransporter virulence proteins, suggesting sequence patterns that favor an expanded conformational ensemble are under selection for efficient autotransporter protein secretion, a necessary prerequisite for virulence. More broadly, we found that sequence patterns that lead to more expanded conformational ensembles are enriched across water-soluble proteins in general, suggesting protein sequences are under selection to regulate collapse and minimize protein aggregation, in addition to their roles in stabilizing folded protein structures.

59 BASIC BIOLOGICAL SCIENCES↗

A General Framework to Learn Tertiary Structure for Protein Sequence Characterization

During the past five years, deep-learning algorithms have enabled ground-breaking progress towards the prediction of tertiary structure from a protein sequence. Very recently, we developed SAdLSA, a new computational algorithm for protein sequence comparison via deep-learning of protein structural alignments. SAdLSA shows significant improvement over established sequence alignment methods. In this contribution, we show that SAdLSA provides a general machine-learning framework for structurally characterizing protein sequences. By aligning a protein sequence against itself, SAdLSA generates a fold distogram for the input sequence, including challenging cases whose structural folds were not present in the training set. About 70% of the predicted distograms are statistically significant. Although at present the accuracy of the intra-sequence distogram predicted by SAdLSA self-alignment is not as good as deep-learning algorithms specifically trained for distogram prediction, it is remarkable that the prediction of single protein structures is encoded by an algorithm that learns ensembles of pairwise structural comparisons, without being explicitly trained to recognize individual structural folds. As such, SAdLSA can not only predict protein folds for individual sequences, but also detects subtle, yet significant, structural relationships between multiple protein sequences using the same deep-learning neural network. The former reduces to a special case in this general framework for protein sequence annotation.

59 BASIC BIOLOGICAL SCIENCES↗

RNA target highlights in CASP15 : Evaluation of predicted models by structure providers

Abstract The first RNA category of the Critical Assessment of Techniques for Structure Prediction competition was only made possible because of the scientists who provided experimental structures to challenge the predictors. In this article, these scientists offer a unique and valuable analysis of both the successes and areas for improvement in the predicted models. All 10 RNA‐only targets yielded predictions topologically similar to experimentally determined structures. For one target, experimentalists were able to phase their x‐ray diffraction data by molecular replacement, showing a potential application of structure predictions for RNA structural biologists. Recommended areas for improvement include: enhancing the accuracy in local interaction predictions and increased consideration of the experimental conditions such as multimerization, structure determination method, and time along folding pathways. The prediction of RNA–protein complexes remains the most significant challenge. Finally, given the intrinsic flexibility of many RNAs, we propose the consideration of ensemble models.

59 BASIC BIOLOGICAL SCIENCES↗

Unfolding of the Villin Headpiece Domain: Revealing Structural Heterogeneity with Time‐Resolved X‐Ray Solution Scattering and Markov State Modeling

Understanding protein folding pathways is crucial to deciphering the principles of protein structure and function. Here, the unfolding dynamics of the 35‐residue villin headpiece (HP35) and a norleucine‐substituted variant (2F4K) using a combination of experimental and computational techniques is investigated. Time‐resolved X‐ray solution scattering coupled with equilibrium molecular dynamics simulations and Markov state modeling reveals distinct unfolding mechanisms between the two variants: HP35 and 2F4K. Specifically, HP35 exhibits a two‐state unfolding process, whereas an intermediate state is identified for the 2F4K mutant. A Markov state model constructed from simulations is used to map atomic‐level transitions to experimental observations, providing insights into the role of sequence variations in modulating folding pathways. The findings underscore the importance of integrating experimental and computational approaches to unravel protein unfolding mechanisms between heterogenous structural ensembles.

Nijhawan, Adam K. [Department of Chemistry Northwe↗

Consequences of the failure of equipartition for the p – V behavior of liquid water and the hydration free energy components of a small protein

Earlier we showed that in the molecular dynamics simulation of a rigid model of water it is necessary to use an integration time-step δ t ≤ 0.5 fs to ensure equipartition between translational and rotational modes. Here we extend that study in the NVT ensemble to NpT conditions and to an aqueous protein. We study neat liquid water with the rigid, SPC/E model and the protein BBA (PDB ID: 1FME) solvated in the rigid, TIP3P model. We examine integration time-steps ranging from 0.5 fs to 4.0 fs for various thermostat plus barostat combinations. We find that a small δ t is necessary to ensure consistent prediction of the simulation volume. Hydrogen mass repartitioning alleviates the problem somewhat, but is ineffective for the typical time-step used with this approach. The compressibility, a measure of volume fluctuations, and the dielectric constant, a measure of dipole moment fluctuations, are also seen to be sensitive to δ t . Using the mean volume estimated from the NpT simulation, we examine the electrostatic and van der Waals contribution to the hydration free energy of the protein in the NVT ensemble. These contributions are also sensitive to δ t . In going from δ t = 2 fs to δ t = 0.5 fs, the change in the net electrostatic plus van der Waals contribution to the hydration of BBA is already in excess of the folding free energy reported for this protein.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Effects of pH on an IDP conformational ensemble explored by molecular dynamics simulation

The conformational ensemble of intrinsically disordered proteins, such as α-synuclein, are responsible for their function and malfunction. Misfolding of α-synuclein can lead to neurodegenerative diseases, and the ability to study their conformations and those of other intrinsically disordered proteins under varying physiological conditions can be crucial to understanding and preventing pathologies. In contrast to well-folded peptides, a consensus feature of IDPs is their low hydropathy and high charge, which makes their conformations sensitive to pH perturbation. We examine a prominent member of this subset of IDPs, α-synuclein, using a divide-and-conquer scheme that provides enhanced sampling of IDP structural ensembles. We constructed conformational ensembles of α-synuclein under neutral (pH ~ 7) and low (pH ~ 3) pH conditions and compared our results with available information obtained from smFRET, SAXS, and NMR studies. Specifically, α-synuclein has been found to in a more compact state at low pH conditions and the structural changes observed are consistent with those from experiments. We also characterize the conformational and dynamic differences between these ensembles and discussed the implication on promoting pathogenic fibril formation. We find that under low pH conditions, neutralization of negatively charged residues leads to compaction of the C-terminal portion of α-synuclein while internal reorganization allows α-synuclein to maintain its overall end-to-end distance. Here, we also observe different levels of intra-protein interaction between three regions of α-synuclein at varying pH and a shift towards more hydrophilic interactions with decreasing pH.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Dataset for manuscript "Consequences of the failure of equipartition for the p-V behavior of liquid water and the hydration free energy components of a small protein"

Previously, we showed that in the molecular dynamics simulation of a rigid model of water it is necessary to use an integration time-step dt that is less than or equal to 0.5 fs to ensure equipartition between translational and rotational modes. We extended that study in the NVT ensemble to NpT conditions and to an aqueous protein. We study neat liquid water with the rigid, SPC/E model and the protein BBA (PDB ID: 1FME) solvated in the rigid, TIP3P model. We examined integration time-steps ranging from 0.5 fs to 4.0 fs for various thermostat plus barostat combinations. We find that a small time-step, dt, is necessary to ensure consistent prediction of the simulation volume. Hydrogen mass repartitioning alleviates the problem somewhat, but is ineffective for the typical time-step used with this approach. The compressibility, a measure of volume fluctuations, is seen to be sensitive to dt. Using the mean volume estimated from the NpT simulation, we examined the electrostatic and van der Waals contribution to the hydration free energy of the protein in the NVT ensemble. These contributions are also sensitive to dt. In going from a time-step of 2 fs to a time-step of 0.5 fs, the change in the net electrostatic plus van der Waals contribution to the hydration of BBA is already in excess of the folding free energy reported for this protein. The data-set contains the simulation metadata and log files that support the claims noted above.

59 BASIC BIOLOGICAL SCIENCES↗

Early events in G-quadruplex folding captured by time-resolved small-angle X-ray scattering

Abstract Time-resolved small-angle X-ray experiments are reported here that capture and quantify a previously unknown rapid collapse of the unfolded oligonucleotide as an early step in the folding of hybrid 1 and hybrid 2 telomeric G-quadruplex structures. The rapid collapse, initiated by a pH jump, is characterized by an exponential decrease in the radius of gyration from 24.3 to 12.6 Å. The collapse is monophasic and is complete in <600 ms. Additional hand-mixing pH-jump kinetic studies show that slower kinetic steps follow the collapse. The folded and unfolded states at equilibrium were further characterized by SAXS studies and other biophysical tools, showing that G4 unfolding was complete at alkaline pH, but not in LiCl solution as is often claimed. The SAXS Ensemble Optimization Method analysis reveals models of the unfolded state as a dynamic ensemble of flexible oligonucleotide chains with a variety of transient hairpin structures. These results suggest a G4 folding pathway in which a rapid collapse, analogous to molten globule formation seen in proteins, is followed by a confined conformational search within the collapsed particle to form the native contacts ultimately found in the stable folded form.

Biochemistry & Molecular Biology↗

Real-time tracking of protein unfolding with time-resolved x-ray solution scattering

The correct folding of proteins is of paramount importance for their function, and protein misfolding is believed to be the primary cause of a wide range of diseases. Protein folding has been investigated with time-averaged methods and time-resolved spectroscopy, but observing the structural dynamics of the unfolding process in real-time is challenging. Here, we demonstrate an approach to directly reveal the structural changes in the unfolding reaction. We use nano- to millisecond time-resolved x-ray solution scattering to probe the unfolding of apomyoglobin. The unfolding reaction was triggered using a temperature jump, which was induced by a nanosecond laser pulse. We demonstrate a new strategy to interpret time-resolved x-ray solution scattering data, which evaluates ensembles of structures obtained from molecular dynamics simulations. We find that apomyoglobin passes three states when unfolding, which we characterize as native, molten globule, and unfolded. The molten globule dominates the population under the conditions investigated herein, whereas native and unfolded structures primarily contribute before the laser jump and 30 μs after it, respectively. The molten globule retains much of the native structure but shows a dynamic pattern of inter-residue contacts. Our study demonstrates a new strategy to directly observe structural changes over the cause of the unfolding reaction, providing time- and spatially resolved atomic details of the folding mechanism of globular proteins.

59 BASIC BIOLOGICAL SCIENCES↗

Unlocking the unfolded structure of ubiquitin: Combining time-resolved x-ray solution scattering and molecular dynamics to generate unfolded ensembles

The unfolding dynamics of ubiquitin were studied using a combination of x-ray solution scattering (XSS) and molecular dynamics (MD) simulations. The kinetic analysis of the XSS ubiquitin signals showed that the protein unfolds through a two-state process, independent of the presence of destabilizing salts. In order to characterize the ensemble of unfolded states in atomic detail, the experimental XSS results were used as a constraint in the MD simulations through the incorporation of x-ray scattering derived potential to drive the folded ubiquitin structure toward sampling unfolded states consistent with the XSS signals. We detail how biased MD simulations provide insight into unfolded states that are otherwise difficult to resolve and underscore how experimental XSS data can be combined with MD to efficiently sample structures away from the native state. Our results indicate that ubiquitin samples unfolded in states with a high degree of loss in secondary structure yet without a collapse to a molten globule or fully solvated extended chain. Finally, we propose how using biased-MD can significantly decrease the computational time and resources required to sample experimentally relevant nonequilibrium states.

Chemistry↗

Oxygen Deficiency in Spaceflight & its Impact on Plants’ Adaptive Changes

The goal of this study was to investigate the effects of hypoxic conditions in spaceflight. The distribution of genes involved with hypoxia in Arabidopsis thaliana and Brassica rapa were analyzed with the results from past spaceflight experiments to evaluate genes for future studies. Transcriptomes data of two different spaceflight studies of Arabidopsis thaliana from the NASA GeneLab database, GLDS-7 and GLDS-17, were compared. DNA microarrays were utilized for transcription profiling to conduct these studies. For GLDS-7, the response in spaceflight was studied with approaches that collected gene expression data. Leaves, hypocotyls, and root tissues were compared to the whole plant. For GLDS-17, seedlings and undifferentiated cultured cells were placed in the Biological Research in Canisters (BRIC), specifically BRIC-16. The genes related to hypoxia in Arabidopsis thaliana from these two studies were compared to genes in Brassica rapa with the TOAST database to evaluate similarities. When transcriptomes were analyzed for GLDS-7 and 17, genes that were considered significant had p-values ≤ 0.05 and log fold change values ≤ -1 or ≥1. Sixteen genes fulfilled the criteria. The genes related to hypoxia were alcohol dehydrogenase, elongation factor, ethylene-responsive factor, GUS, heat-shock proteins, NAP, RAP2.12, and RD20. The genes most impacted by spaceflight were heat-shock proteins. These genes were compared with Brassica rapa through Arabidopsis Ensemble Orthology from the TOAST Database. Similarities were seen in alcohol dehydrogenase, elongation factor, ethylene-responsive factor, heat-shock proteins, NAP, and RAP2.12. Overall, transcription profiling indicates that plants’ survival in spaceflight is dependent on adaptive changes with gene expression. This study also indicates that there are similarities in gene expression between Arabidopsis thaliana and Brassica rapa with comparable gene expression. Future studies could include analyzing additional species to understand which genes could be modified to ensure better yield of space crops amid hypoxic conditions.

hypoxia↗

A Deep Learning-Driven Sampling Technique to Explore the Phase Space of an RNA Stem-Loop

The folding and unfolding of RNA stem-loops are critical biological processes; however, their computational studies are often hampered by the ruggedness of their folding landscape, necessitating long simulation times at the atomistic scale. Here, we adapted DeepDriveMD (DDMD), an advanced deep learning-driven sampling technique originally developed for protein folding, to address the challenges of RNA stem-loop folding. Although tempering- and order parameter-based techniques are commonly used for similar rare-event problems, the computational costs or the need for a priori knowledge about the system often present a challenge in their effective use. DDMD overcomes these challenges by adaptively learning from an ensemble of running MD simulations using generic contact maps as the raw input. DeepDriveMD enables on-the-fly learning of a low-dimensional latent representation and guides the simulation toward the undersampled regions while optimizing the resources to explore the relevant parts of the phase space. We showed that DDMD estimates the free energy landscape of the RNA stem-loop reasonably well at room temperature. Our simulation framework runs at a constant temperature without external biasing potential, hence preserving the information on transition rates, with a computational cost much lower than that of the simulations performed with external biasing potentials. Here, we also introduced a reweighting strategy for obtaining unbiased free energy surfaces and presented a qualitative analysis of the latent space. This analysis showed that the latent space captures the relevant slow degrees of freedom for the RNA folding problem of interest. Finally, throughout the manuscript, we outlined how different parameters are selected and optimized to adapt DDMD for this system. We believe this compendium of decision-making processes will help new users adapt this technique for the rare-event sampling problems of their interest.

Gupta, Ayush↗

Cryo-electron tomography related radiation-damage parameters for individual-molecule 3D structure determination

To understand the dynamic structure–function relationship of soft- and biomolecules, the determination of the three-dimensional (3D) structure of each individual molecule (nonaveraged structure) in its native state is sought-after. Cryo-electron tomography (cryo-ET) is a unique tool for imaging an individual object from a series of tilted views. However, due to radiation damage from the incident electron beam, the tolerable electron dose limits image contrast and the signal-to-noise ratio (SNR) of the data, preventing the 3D structure determination of individual molecules, especially at high-resolution. Although recently developed technologies and techniques, such as the direct electron detector, phase plate, and computational algorithms, can partially improve image contrast/SNR at the same electron dose, the high-resolution structure, such as tertiary structure of individual molecules, has not yet been resolved. Here, we review the cryo-electron microscopy (cryo-EM) and cryo-ET experimental parameters to discuss how these parameters affect the extent of radiation damage. This discussion can guide us in optimizing the experimental strategy to increase the imaging dose or improve image SNR without increasing the radiation damage. With a higher dose, a higher image contrast/SNR can be achieved, which is crucial for individual-molecule 3D structure. With 3D structures determined from an ensemble of individual molecules in different conformations, the molecular mechanism through their biochemical reactions, such as self-folding or synthesis, can be elucidated in a straightforward manner.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗