Generative artificial intelligence performs rudimentary structural biology modeling
Not Available
SEARCH · Search NASA
Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.
Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.
Not Available
NASA is developing a systems biology approach to improve the assessment of health risks associated with space radiation. The primary toxic and mutagenic lesion following radiation exposure is the DNA double strand break (DSB), thus a model incorporating proteins and pathways important in response and repair of this lesion is critical. One key protein heterodimer for systems models of radiation effects is the Ku70/80 complex. The Ku70/80 complex is important in the initial binding of DSB ends following DNA damage, and is a component of nonhomologous end joining repair, the primary pathway for DSB repair in mammalian cells. The SAP domain of Ku70 (residues 556-609), contains an a helix-extended strand-helix motif and similar motifs have been found in other nucleic acid-binding proteins critical for DNA repair. However, the exact mechanism of damage recognition and substrate specificity for the Ku heterodimer remains unclear in part due to the absence of a high-resolution structure of the SAP/DNA complex. We performed a series of molecular dynamics (MD) simulations on a system with the SAP domain of Ku70 and a 10 base pairs DNA duplex. Large-scale conformational changes were observed and some putative binding modes were suggested based on energetic analysis. These modes are consistent with previous experimental investigations. In addition, the results indicate that cooperation of SAP with other domains of Ku70/80 is necessary to explain the high affinity of binding as observed in experiments.
NASA is developing a systems biology approach to improve the assessment of health risks associated with space radiation. The primary toxic and mutagenic lesion following radiation exposure is the DNA double strand break (DSB), thus a model incorporating proteins and pathways important in response and repair of this lesion is critical. One key protein heterodimer for systems models of radiation effects is the Ku(sub 70/80) complex. The Ku70/80 complex is important in the initial binding of DSB ends following DNA damage, and is a component of nonhomologous end joining repair, the primary pathway for DSB repair in mammalian cells. The C-terminal domain of Ku70 (Ku70c, residues 559-609), contains an helix-extended strand-helix motif and similar motifs have been found in other nucleic acid-binding proteins critical for DNA repair. However, the exact mechanism of damage recognition and substrate specificity for the Ku heterodimer remains unclear in part due to the absence of a high-resolution structure of the Ku70c/DNA complex. We performed a series of molecular dynamics (MD) simulations on a system with the subunit Ku70c and a 14 base pairs DNA duplex, whose starting structures are designed to be variable so as to mimic their different binding modes. By analyzing conformational changes and energetic properties of the complex during MD simulations, we found that interactions are preferred at DNA ends, and within the major groove, which is consistent with previous experimental investigations. In addition, the results indicate that cooperation of Ku70c with other subunits of Ku(sub 70/80) is necessary to explain the high affinity of binding as observed in experiments.
Not Available
Understanding protein folding pathways is crucial to deciphering the principles of protein structure and function. Here, the unfolding dynamics of the 35‐residue villin headpiece (HP35) and a norleucine‐substituted variant (2F4K) using a combination of experimental and computational techniques is investigated. Time‐resolved X‐ray solution scattering coupled with equilibrium molecular dynamics simulations and Markov state modeling reveals distinct unfolding mechanisms between the two variants: HP35 and 2F4K. Specifically, HP35 exhibits a two‐state unfolding process, whereas an intermediate state is identified for the 2F4K mutant. A Markov state model constructed from simulations is used to map atomic‐level transitions to experimental observations, providing insights into the role of sequence variations in modulating folding pathways. The findings underscore the importance of integrating experimental and computational approaches to unravel protein unfolding mechanisms between heterogenous structural ensembles.
With the advances in next generation sequencing technologies, the number of sequenced genomes is growing exponentially, resulting in a technology bottleneck for the translation of sequence information into usable hypotheses about the function of each gene. We have proposed leveraging our leadership high-performance computing (HPC) resources to help break this annotation bottleneck. Here we design an HPC-based framework to infer gene function from gene sequence by incorporating information about protein structure and interactions predicted by deep learning approaches. Accurate functional prediction and gene annotation using computational methods will facilitate breakthroughs in the genomic sciences essential to understanding and harnessing life processes in bacteria, fungi and plants. The development and applications of the state-of-the-art deep neural networks to protein structural modeling, interaction prediction, sequence comparison, and quality assessment of protein structural models will be made possible by leadership computational resources. These HPC-enabled bioinformatics and molecular modeling tools will provide powerful insights into molecular functions of genes.
The proteinoid model for the coordination of protein synthesis with nucleic acid coding within the evolving protocell is discussed. Evidence for the self-ordering of amino acid chains, which would enhance the catalytic activity of a lysine-rich proteinoid, is presented, along with that for the preferential formation of microparticles, particularly proteinoid microparticles, in various solutions. Demonstrations of the catalytic activity of lysine-rich proteinoids in the synthesis of peptide and internucleotide bonds are pointed out. The view of evolution as a two stage sequence in which the geological synthesis of peptides evolved to the protocellular synthesis of peptides and oligonucleotides is discussed, and contrasted with the alternative view, in accord with the central dogma, that nucleic acids arose first then governed the production of proteins and protocells.
Repeat proteins are made with tandem copies of similar amino acid stretches that fold into elongated architectures. These proteins constitute excellent model systems to investigate how evolution relates to structure, folding, and function. Here, we propose a scheme to map evolutionary information at the sequence level to a coarse-grained model for repeat-protein folding and use it to investigate the folding of thousands of repeat proteins. We model the energetics by a combination of an inverse Potts-model scheme with an explicit mechanistic model of duplications and deletions of repeats to calculate the evolutionary parameters of the system at the single-residue level. These parameters are used to inform an Ising-like model that allows for the generation of folding curves, apparent domain emergence, and occupation of intermediate states that are highly compatible with experimental data in specific case studies. We analyzed the folding of thousands of natural Ankyrin repeat proteins and found that a multiplicity of folding mechanisms are possible. Fully cooperative all-or-none transitions are obtained for arrays with enough sequence-similar elements and strong interactions between them, while noncooperative element-by-element intermittent folding arose if the elements are dissimilar and the interactions between them are energetically weak. Additionally, we characterized nucleation-propagation and multidomain folding mechanisms. We show that the global stability and cooperativity of the repeating arrays can be predicted from simple sequence scores.
A model for the origin of protein synthesis. The essential features of the model are that 5'-AMP and perhaps other monoribonucleotides can serve as catalysts for the selective synthesis of L-based peptides. A unique set of characteristics of 5'-AMP is responsible for the selective catalysts and these characteristics are described in detail. The model involves the formation of diesters as intermediates and selectivity for use of the L-isomer occurs principally at the step of forming the diester. However, in the formation of acetyl phenylalanine-AMP monoester there is a selectivity for esterification by the D-isomer. Data showing this selectivity is presented. This selectivity for D-isomer disappears after the first step. The identity was confirmed of all four of possible diesters of acetyl-D- and -L phenylaline with 5'-AMP by nuclear magnetic resonance (NMR). The data using flourescence and NMR show the Trp ring can associate with the adenine ring more strongly when the D-isomer is in the 2' position than it can when in the 3' position. These same data also suggest a molecular mechanisim for the faster esterificaton of 5'-AMP by acetyl-D-phenylaline. Some new data is also presented on the possible structure of the 2' isomer of acetyl-D-tryptophan-AMP monoester. The HPLC elution times of all four possible acetyl diphenylalanine esters of 5'-AMP were studied, these peptidyl esters will be products in the studies of peptide formation on the ribose of 5'-AMP. Other studies were on the rate of synthesis and the identity of the product when producing 3'Ac-Phe-2'tBOC-Phe-AMP diester. HPLC purification and identification of this product were accomplished.
In long-term space travel, the crew is exposed to microgravity and radiation that invoke potential hazards to the immune system. T cell activation is a critical step in the immune response. Receptor-mediated signaling is inhibited in both microgravity and modeled microgravity (MMG) as reflected by diminished DNA synthesis in peripheral blood lymphocytes and their locomotion through gelled type I collagen. Direct activation of protein kinase C (PKC) bypassing cell surface events using the phorbol ester PMA rescues MMG-inhibited lymphocyte activation and locomotion, whereas the calcium ionophore ionomycin had no rescue effect. Thus calcium-independent PKC isoforms may be affected in MMG-induced locomotion inhibition and rescue. Both calcium-dependent isoforms and calcium-independent PKC isoforms were investigated to assess their expression in lymphocytes in 1 g and MMG culture. Human lymphocytes were cultured and harvested at 24, 48, 72, and 96 h, and serial samples were assessed for locomotion by using type I collagen and expression of PKC isoforms. Expression of PKC-alpha, -delta, and -epsilon was assessed by RT-PCR, flow cytometry, and immunoblotting. Results indicated that PKC isoforms delta and epsilon were downregulated by >50% at the transcriptional and translational levels in MMG-cultured lymphocytes compared with 1-g controls. Events upstream of PKC, such as phosphorylation of phospholipase Cgamma in MMG, revealed accumulation of inactive enzyme. Depressed calcium-independent PKC isoforms may be a consequence of an upstream lesion in the signal transduction pathway. The differential response among calcium-dependent and calcium-independent isoforms may actually result from MMG intrusion events earlier than PKC, but after ligand-receptor interaction.
In long-term space travel, the crew is exposed to microgravity and radiation that invoke potential hazards to the immune system. T cell activation is a critical step in the immune response. Receptor-mediated signaling is inhibited both in microgravity and modeled microgravity (MMG) as reflected in diminished DNA synthess in peripheral blood lymphocytes and their locomotion through gelled type 1 collagen. Direct activation of Protein Kinase C (PKC) bypassing cell surface events using the phorbol ester PMA rescues MMG-inhibited lymphocyte activation and locomotion, whereas calcium ionophore ionomycin had no rescue effect. Thus calcium-independent PKC isoforms may be affected in MMG-induced locomotion inhibition and rescue. Both calcium-dependent isoforms and calcium-independent PKC isoforms were investigated to assess their expression in lymphocytes in 19 and MMG-culture. Human lymphocytes were cultured and harvested at 24, 48, 72 and 96 hours and serial samples assessed for locomotion using type I collagen and expression of PKC isoforms. Expression of PKC-alpha, -delta and -epsilon was assessed by RT-PCR, flow cytometry and immunoblotting. Results indicated that PKC isoforms delta and epsilon were down-regulated by more than 50% at the transcriptional and translational levels in MMG-cultured lymphocytes compared with 19 controls. Events upstream of PKC such as phosphorylation of Phospholipase C(gamma) (PLC-gamma) in MMG, revealed accumulation of inactive enzyme. Depressed Ca++ -independent PKC isoforms may be a consequence of an upstream lesion in the signal transduction pathway. The differential response among calcium-dependent and calcium-independent isoforms may actually result from MMG intrusion events earlier than, but after ligand-receptor interaction. Keywords: Signal transduction, locomotion, immunity
Protein language models (PLMs) have recently demonstrated potential to supplant classical protein database search methods based on sequence alignment, but are slower than common alignment-based tools and appear to be prone to a high rate of false labeling. Here, we present NEAR, a method based on neural representation learning that is designed to improve both speed and accuracy of search for likely homologs in a large protein sequence database. NEAR’s ResNet embedding model is trained using contrastive learning guided by trusted sequence alignments. It computes per-residue embeddings for target and query protein sequences, and identifies alignment candidates with a pipeline consisting of residue-level k-NN search and a simple neighbor aggregation scheme. Tests on a benchmark consisting of trusted remote homologs and randomly shuffled decoy sequences reveal that NEAR substantially improves accuracy relative to state-of-the-art PLMs, with lower memory requirements and faster embedding and search speed. While these results suggest that the NEAR model may be useful for standalone homology detection with increased sensitivity over standard alignment-based methods, in this manuscript we focus on a more straightforward analysis of the model’s value as a high-speed pre-filter for sensitive annotation. In that context, NEAR is at least 5x faster than the pre-filter currently used in the widely-used profile hidden Markov model (pHMM) search tool HMMER3, and also outperforms the pre-filter used in our fast pHMM tool, nail.
Abstract The Protein Data Bank (PDB), established in 1971, is the primary global, open‐access archive for experimentally determined 3D macromolecular structures (proteins, RNA, DNA). The research‐focused RCSB.org web‐portal provides access to these data alongside more than one million machine‐learning‐predicted structure models, greatly expanding the available structural landscape. Rapid growth of both experimental and computational structures has increased the need for powerful yet accessible search tools that serve a broad and diverse scientific community. Herein, we describe a redesigned RCSB Protein Data Bank RCSB.org Advanced Search capability that supports intuitive discovery of 3D structures through a unified interface. This interface integrates annotation‐, sequence‐, and 3D structure‐based searches, embeds an interactive 3D viewer, and incorporates curated biological knowledge, such as catalytic site definitions from Mechanism and Catalytic Site Atlas and ligand‐guided structural motifs, for constructing geometry‐driven queries. A new Chemical Search tool allows definition of chemical queries via an integrated drawing tool or standard identifiers, seamlessly combining them with annotation filters. By allowing query definition directly within spatial and chemical contexts, these search interfaces reduce the need for detailed knowledge of residue numbering, chain identifiers, or external cheminformatics software. This capability enables efficient exploration of structures, chemical diversity, and structure–function relationships across all life domains. The redesigned interfaces can be accessed directly at rcsb.org/search/advanced for Advanced Search and rcsb.org/search/chemical for Chemical Search.
The spatial organization of chromatin is governed by epigenetic factors, including epigenetic marks and the reader proteins that bind them. By dictating the accessibility of genomic loci, epigenetic factors contribute to the physical regulation of gene expression, enabling diverse cellular phenotypes to be encoded by a shared genome in an individual. Epigenetic dysregulation can lead to aberrations in chromatin architecture, contributing to diseases such as neurological disorders and cancers. Despite the known importance of chromatin organization for human health, the physical mechanisms governing chromatin folding remain underspecified. In this work, we develop a physical model of chromatin organization based on contributions from multiple epigenetic factors. Using our model, we evaluate how conditions in the nuclear environment and crosstalk between epigenetic marks affect the compartmentalization of chromatin into dense heterochromatin and loose euchromatin. Our results emphasize the role of reader protein binding in chromatin compartmentalization. We show that reader proteins interact through an indirect mechanism facilitated by the shared chromatin “scaffold” to which they bind. Under a scenario where reader proteins compete for binding sites, we find that indirect interactions affect the program adopted by the chromatin fiber. By isolating indirect modes of epigenetic crosstalk, we demonstrate how the interplay between epigenetic patterning and environmental factors influences chromatin architecture.
NonDarwinian evolution of protein and DNA, comparing expectations of evolution models for protein and amino acid changes
Calmodulin (CaM) was used as an affinity tail to facilitate the purification of the green fluorescent protein (GFP), which was used as a model target protein. The protein GFP was fused to the C-terminus of CaM, and a factor Xa cleavage site was introduced between the two proteins. A CaM-GFP fusion protein was expressed in E. coli and purified on a phenothiazine-derivatized silica column. CaM binds to the phenothiazine on the column in a Ca(2+)-dependent fashion and it was, therefore, used as an affinity tail for the purification of GFP. The fusion protein bound to the affinity column was then subjected to a proteolytic digestion with factor Xa. Pure GFP was eluted with a Ca(2+)-containing buffer, while CaM was eluted later with a buffer containing the Ca(2+)-chelating agent EGTA. The purity of the isolated GFP was verified by SDS-PAGE, and the fluorescence properties of the purified GFP were characterized.
The N-heptad repeat (NHR) of the HIV-1 gp41 prehairpin intermediate (PHI) is an attractive potential vaccine target with high sequence conservation across diverse strains. However, despite the potency of NHR-targeting peptides and clinical efficacy of the NHR-targeting entry inhibitor enfuvirtide, no potently neutralizing NHR-directed monoclonal antibodies (mAbs) nor antisera have been identified or elicited to date. The lack of potent NHR-binding mAbs both dampens enthusiasm for vaccine development efforts at this target and presents a barrier to performing passive immunization experiments with NHR-targeting antibodies. To address this challenge, we previously developed an improved variant of the NHR-directed mAb D5, called D5_AR, which is capable of neutralizing diverse tier-2 viruses. Building on that work, here we present the 2.7Å-crystal structure of D5_AR bound to NHR mimetic peptide IQN17. We then utilize protein language models and supervised machine learning to generate small (n < 100) libraries of D5_AR variants that are subsequently screened for improved neutralization potency. We identify a variant with 5-fold improved neutralization potency, D5_FI, which is the most potent NHR-directed monoclonal antibody characterized to date and exhibits broad neutralization of tier-2 and −3 pseudoviruses as well as replicating R5 and X4 challenge strains. Additionally, our work highlights the ability of protein language models to efficiently identify improved mAb variants from relatively small libraries.
Free transition metal ions oxidize lipids and lipoproteins in vitro; however, recent evidence suggests that free metal ion-independent mechanisms are more likely in vivo. We have shown previously that human ceruloplasmin (Cp), a serum protein containing seven Cu atoms, induces low density lipoprotein oxidation in vitro and that the activity depends on the presence of a single, chelatable Cu atom. We here use biochemical and molecular approaches to determine the site responsible for Cp prooxidant activity. Experiments with the His-specific reagent diethylpyrocarbonate (DEPC) showed that one or more His residues was specifically required. Quantitative [14C]DEPC binding studies indicated the importance of a single His residue because only one was exposed upon removal of the prooxidant Cu. Plasmin digestion of [14C]DEPC-treated Cp (and N-terminal sequence analysis of the fragments) showed that the critical His was in a 17-kDa region containing four His residues in the second major sequence homology domain of Cp. A full length human Cp cDNA was modified by site-directed mutagenesis to give His-to-Ala substitutions at each of the four positions and was transfected into COS-7 cells, and low density lipoprotein oxidation was measured. The prooxidant site was localized to a region containing His426 because CpH426A almost completely lacked prooxidant activity whereas the other mutants expressed normal activity. These observations support the hypothesis that Cu bound at specific sites on protein surfaces can cause oxidative damage to macromolecules in their environment. Cp may serve as a model protein for understanding mechanisms of oxidant damage by copper-containing (or -binding) proteins such as Cu, Zn superoxide dismutase, and amyloid precursor protein.