Search NASASearch

SEARCH · Search NASA

Results for “Protein structure”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 37 records · Page 2

CryoTEN: efficiently enhancing cryo-EM density maps using transformers

Abstract Motivation Cryogenic electron microscopy (cryo-EM) is a core experimental technique used to determine the structure of macromolecules such as proteins. However, the effectiveness of cryo-EM is often hindered by the noise and missing density values in cryo-EM density maps caused by experimental conditions such as low contrast and conformational heterogeneity. Although various global and local map-sharpening techniques are widely employed to improve cryo-EM density maps, it is still challenging to efficiently improve their quality for building better protein structures from them. Results In this study, we introduce CryoTEN—a 3D UNETR++ style transformer to improve cryo-EM maps effectively. CryoTEN is trained using a diverse set of 1295 cryo-EM maps as inputs and their corresponding simulated maps generated from known protein structures as targets. An independent test set containing 150 maps is used to evaluate CryoTEN, and the results demonstrate that it can robustly enhance the quality of cryo-EM density maps. In addition, automatic de novo protein structure modeling shows that protein structures built from the density maps processed by CryoTEN have substantially better quality than those built from the original maps. Compared to the existing state-of-the-art deep learning methods for enhancing cryo-EM density maps, CryoTEN ranks second in improving the quality of density maps, while running >10 times faster and requiring much less GPU memory than them. Availability and implementation The source code and data are freely available at https://github.com/jianlin-cheng/cryoten.

Biochemistry & Molecular Biology

Ca X ML: Chemistry‐informed machine learning explains mutual changes between protein conformations and calcium ions in calcium‐binding proteins using structural and topological features

Proteins' flexibility is a feature in communicating changes in cell signaling instigated by binding with secondary messengers, such as calcium ions, associated with the coordination of muscle contraction, neurotransmitter release, and gene expression. When binding with the disordered parts of a protein, calcium ions must balance their charge states with the shape of calcium-binding proteins and their versatile pool of partners depending on the circumstances they transmit. Accurately determining the ionic charges of those ions is essential for understanding their role in such processes. However, it is unclear whether the limited experimental data available can be effectively used to train models to accurately predict the charges of calcium-binding protein variants. Here, we developed a chemistry-informed, machine-learning algorithm that implements a game theoretic approach to explain the output of a machine-learning model without the prerequisite of an excessively large database for high-performance prediction of atomic charges. We used the ab initio electronic structure data representing calcium ions and the structures of the disordered segments of calcium-binding peptides with surrounding water molecules to train several explainable models. Network theory was used to extract the topological features of atomic interactions in the structurally complex data dictated by the coordination chemistry of a calcium ion, a potent indicator of its charge state in protein. Our design created a computational tool of Ca X ML, which provided a framework of explainable machine learning model to annotate ionic charges of calcium ions in calcium-binding proteins in response to the chemical changes in an environment. Our framework will provide new insights into protein design for engineering functionality based on the limited size of scientific data in a genome space.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH

De novo design of buttressed loops for sculpting protein functions

In natural proteins, structured loops have central roles in molecular recognition, signal transduction and enzyme catalysis. However, because of the intrinsic flexibility and irregularity of loop regions, organizing multiple structured loops at protein functional sites has been very difficult to achieve by de novo protein design. Here we describe a solution to this problem that designs tandem repeat proteins with structured loops (9–14 residues) buttressed by extensive hydrogen bonding interactions. Experimental characterization shows that the designs are monodisperse, highly soluble, folded and thermally stable. Crystal structures are in close agreement with the design models, with the loops structured and buttressed as designed. We demonstrate the functionality afforded by loop buttressing by designing and characterizing binders for extended peptides in which the loops form one side of an extended binding pocket. The ability to design multiple structured loops should contribute generally to efforts to design new protein functions.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH

Comparative genomic analysis of thermophilic fungi reveals convergent evolutionary adaptations and gene losses

Thermophily is a trait scattered across the fungal tree of life, with its highest prevalence within three fungal families (Chaetomiaceae, Thermoascaceae, and Trichocomaceae), as well as some members of the phylum Mucoromycota. We examined 37 thermophilic and thermotolerant species and 42 mesophilic species for this study and identified thermophily as the ancestral state of all three prominent families of thermophilic fungi. Thermophilic fungal genomes were found to encode various thermostable enzymes, including carbohydrate-active enzymes such as endoxylanases, which are useful for many industrial applications. At the same time, the overall gene counts, especially in gene families responsible for microbial defense such as secondary metabolism, are reduced in thermophiles compared to mesophiles. We also found a reduction in the core genome size of thermophiles in both the Chaetomiaceae family and the Eurotiomycetes class. The Gene Ontology terms lost in thermophilic fungi include primary metabolism, transporters, UV response, and O-methyltransferases. Comparative genomics analysis also revealed higher GC content in the third base of codons (GC3) and a lower effective number of codons in fungal thermophiles than in both thermotolerant and mesophilic fungi. Furthermore, using the Support Vector Machine classifier, we identified several Pfam domains capable of discriminating between genomes of thermophiles and mesophiles with 94% accuracy. Using AlphaFold2 to predict protein structures of endoxylanases (GH10), we built a similarity network based on the structures. We found that the number of disulfide bonds appears important for protein structure, and the network clusters based on protein structures correlate with the optimal activity temperature. Thus, comparative genomics offers new insights into the biology, adaptation, and evolutionary history of thermophilic fungi while providing a parts list for bioengineering applications.

59 BASIC BIOLOGICAL SCIENCES

Predicting metal-binding proteins and structures through integration of evolutionary-scale and physics-based modeling

Metals are essential elements in all living organisms, binding to approximately 50% of proteins. They serve to stabilize proteins, catalyze reactions, regulate activities, and fulfill various physiological and pathological functions. While there have been many advancements in determining the structures of protein-metal complexes, numerous metal-binding proteins still need to be identified through computational methods and validated through experiments. Here, to address this need, we have developed the ESMBind workflow, which combines evolutionary scale modeling (ESM) for metal-binding prediction and physics-based protein-metal modeling. Our approach utilizes the ESM-2 and ESM-IF models to predict metal-binding probability at the residue level. In addition, we have designed a metal-placement method and energy minimization technique to generate detailed 3D structures of protein-metal complexes. Our workflow outperforms other models in terms of residue and 3D-level predictions. To demonstrate its effectiveness, we applied the workflow to 142 uncharacterized fungal pathogen proteins and predicted metal-binding proteins involved in fungal infection and virulence.

59 BASIC BIOLOGICAL SCIENCES

The protein structurome of Orthornavirae and its dark matter

Metatranscriptomics is uncovering more and more diverse families of viruses with RNA genomes comprising the viral kingdom Orthornavirae in the realm Riboviria. Thorough protein annotation and comparison are essential to get insights into the functions of viral proteins and virus evolution. In addition to sequence- and hmm profile-based methods, protein structure comparison adds a powerful tool to uncover protein functions and relationships. We constructed an Orthornavirae “structurome” consisting of already annotated as well as unannotated (“dark matter”) proteins and domains encoded in viral genomes. We used protein structure modeling and similarity searches to illuminate the remaining dark matter in hundreds of thousands of orthornavirus genomes. The vast majority of the dark matter domains showed either “generic” folds, such as single α-helices, or no high confidence structure predictions. Nevertheless, a variety of lineage-specific globular domains that were new either to orthornaviruses in general or to particular virus families were identified within the proteomic dark matter of orthornaviruses, including several predicted nucleic acid-binding domains and nucleases. In addition, we identified a case of exaptation of a cellular nucleoside monophosphate kinase as an RNA-binding protein in several virus families. Notwithstanding the continuing discovery of numerous orthornaviruses, it appears that all the protein domains conserved in large groups of viruses have already been identified. The rest of the viral proteome seems to be dominated by poorly structured domains including intrinsically disordered ones that likely mediate specific virus-host interactions.

59 BASIC BIOLOGICAL SCIENCES

High throughput, accurate gene annotation through AI and HPC-enabled structural analysis

With the advances in next generation sequencing technologies, the number of sequenced genomes is growing exponentially, resulting in a technology bottleneck for the translation of sequence information into usable hypotheses about the function of each gene. We have proposed leveraging our leadership high-performance computing (HPC) resources to help break this annotation bottleneck. Here we design an HPC-based framework to infer gene function from gene sequence by incorporating information about protein structure and interactions predicted by deep learning approaches. Accurate functional prediction and gene annotation using computational methods will facilitate breakthroughs in the genomic sciences essential to understanding and harnessing life processes in bacteria, fungi and plants. The development and applications of the state-of-the-art deep neural networks to protein structural modeling, interaction prediction, sequence comparison, and quality assessment of protein structural models will be made possible by leadership computational resources. These HPC-enabled bioinformatics and molecular modeling tools will provide powerful insights into molecular functions of genes.

59 BASIC BIOLOGICAL SCIENCES

Characterization of Gas-Phase Native(-like) Proteins Using Structures for Lossless Ion Manipulations

High-resolution mobility-based ion separations in Structures for Lossless Ion Manipulations (SLIM) have been useful for ion mobility separations for a variety of molecular classes in the gas phase. Here, in this study, we present multipass SLIM separations for gas-phase proteins in their near-native state exhibiting charge-state-dependent arrival time distributions using carbonic anhydrase (29 kDa), alcohol dehydrogenase (148 kDa), and apo-transferrin (79 kDa). The experimental CCS values were obtained from calibration curves for the arrival times of Agilent Tune Mix ions. For multipass separations, the ATDs were converted to CCS values by deconvoluting the multipass arrival times into accurate single-pass values amenable to the single-pass calibration curves. Mass spectra of carbonic anhydrase (CA) showed three different charge states (z = 9+ to 11+). Their corresponding mobility peaks were baseline-separated by using 8-m single-pass separations. When compared to the corresponding drift tube ion mobility (DTIMS) measurements, the CCS values obtained from DTIMS and SLIM were in agreement within experimental error. Single-pass analysis of alcohol dehydrogenase (ADH) exhibits three predominant charge states (z = 23+ to 25+) with mobility overlap between adjacent charge states. The mobility peak resolution for ADH improved with multipass separations (up to 24-m path length). In addition, CCS distributions obtained for charge states z = 16+ to 18+ of apo-transferrin reveal a transition from a compact unimodal form (z = 18+ and 19+) to broader multimodal CCS distributions for z = 16+. For apo-transferrin, 40-m multipass separations were performed allowing for complete isolation of the selected mobility range corresponding to z = 17+, leading to selective isolation of a narrow arrival time window. The extended mobility separations provided minimal alterations to the structure of the proteins, and the experimentally derived CCS values showed minimal change as a function of the separation time or number of passes. Mobility-based ion separations for native-like proteins, using SLIM, open opportunities for native-IMS applications as well as other manipulations enabled by SLIM-like mobility-selective isolation and collection.

charge state distribution

PNNL-Predictive-Phenomics/ProCaliper

ProCaliper is a Python library that curates, organizes, and computes protein structure features in a way that easily interfaces with user-provided experimental data. It extracts or computes protein binding site, active site, charge, pLDDT (order/disorder), acid dissociation, protonation, solvent accessible surface area, disulfide bond distance, and protein secondary structure data using precomputed protein structures and publicly available databases. It provides a unified API for integrating additional residue-level data and for visualizing residue features in 3D.

Rozum, Jordan [Pacific Northwest National Lab]

Geometry, spin coupling, and dielectric control of redox potentials in [4Fe–4S] Clusters

Iron–sulfur (Fe–S) clusters are common biological cofactors that facilitate vital redox reactions. Despite extensive research, the molecular basis of redox potential tuning in ferredoxin-like proteins remains an active area of debate. In this study, we combine statistical analysis of over one thousand [4Fe–4S]-containing protein structures from the Protein Data Bank (PDB) with broken-symmetry and extended broken-symmetry density functional theory to examine how cysteine ligand orientations and environmental screening affect redox properties of the clusters. We identified five main ligand configurations, three of which are predominant in natural structures. Among these, the adiabatic electron affinity differs by less than 0.1 V, indicating that, while geometry plays a secondary role, it allows localized fine-tuning of redox properties. In contrast, electrostatic and solvation effects primarily determine the overall potential range.

Computational Chemistry

SLAB: simultaneous labeling and binding affinity prediction for protein–ligand structures

Machine learning models are often used as scoring functions to predict the binding affinity of a protein–ligand complex. These models are trained with limited amounts of data with experimentally measured binding affinity values. A large number of compounds are labeled inactive through single-concentration screens without measuring binding affinities. These inactive compounds, along with the active ones, can be used to train binary classification models, while regression models are trained using compounds with binding affinities only. However, the classification and regression tasks are often handled separately, without sharing the learned feature representations. In this paper, we propose a novel model architecture that jointly performs regression and classification objectives, aiming to maximize data utilization and improve predictive performance by leveraging two complementary tasks. In our setup, the regression yields the binding affinity, whereas the classification task yields the label as active or inactive. We demonstrate our method using PDBbind, the standard 3D structure database, as well as a dataset of flavivirus protease compounds with binding affinity data. Our experiments show that the new joint training strategy improves the accuracy of the model, increasing applicability in various practical drug screening scenarios.

Biological and medical sciences

Isolation and characterization of 24 phages infecting the plant growth-promoting rhizobacterium Klebsiella sp. M5al

Bacteriophages largely impact bacterial communities via lysis, gene transfer, and metabolic reprogramming and thus are increasingly thought to alter nutrient and energy cycling across many of Earth’s ecosystems. However, there are few model systems to mechanistically and quantitatively study phage-bacteria interactions, especially in soil systems. Here, we isolated, sequenced, and genomically characterized 24 novel phages infectingKlebsiellasp. M5al, a plant growth-promoting, nonencapsulated rhizosphere-associated bacterium, and compared many of their features against all 565 sequenced, dsDNAKlebsiellaphage genomes. Taxonomic analyses revealed that theseKlebsiellaphages belong to three known phage families (Autographiviridae,Drexlerviridae, andStraboviridae) and two newly proposed phage families (CandidatusMavericviridaeand Ca.Rivulusviridae). At the phage family level, we found that core genes were often phage-centric proteins, such as structural proteins for the phage head and tail and DNA packaging proteins. In contrast, genes involved in transcription, translation, or hypothetical proteins were commonly not shared or flexible genes. Ecologically, we assessed the phages’ ubiquity in recent large-scale metagenomic datasets, which revealed they were not widespread, as well as a possible direct role in reprogramming specific metabolisms during infection by screening their genomes for phage-encoded auxiliary metabolic genes (AMGs). Even though AMGs are common in the environmental literature, only one of our phage families,Straboviridae, contained AMGs, and the types of AMGs were correlated at the genus level. Host range phenotyping revealed the phages had a wide range of infectivity, infecting between 1–14 of our 22 bacterial strain panel that included pathogenicKlebsiellaandRaoultellastrains. This indicates that not all capsule-independent Klebsiella phages have broad host ranges. Together, these isolates, with corresponding genome, AMG, and host range analyses, help build theKlebsiellamodel system for studying phage-host interactions of rhizosphere-associated bacteria.

Science & Technology - Other Topics

Designed 2D protein crystals as dynamic molecular gatekeepers for a solid-state device

The sensitivity and responsiveness of living cells to environmental changes are enabled by dynamic protein structures, inspiring efforts to construct artificial supramolecular protein assemblies. However, despite their sophisticated structures, designed protein assemblies have yet to be incorporated into macroscale devices for real-life applications. We report a 2D crystalline protein assembly of C98/E57/E66 L-rhamnulose-1-phosphate aldolase ( CEE RhuA) that selectively blocks or passes molecular species when exposed to a chemical trigger. CEE RhuA crystals are engineered via cobalt(II) coordination bonds to undergo a coherent conformational change from a closed state (pore dimensions <1 nm) to an ajar state (pore dimensions ~4 nm) when exposed to an HCN(g) trigger. When layered onto a mesoporous silicon (pSi) photonic crystal optical sensor configured to detect HCN (g) , the 2D CEE RhuA crystal layer effectively blocks interferents that would otherwise result in a false positive signal. The 2D CEE RhuA crystal layer opens in selective response to low-ppm levels of HCN (g) , allowing analyte penetration into the pSi sensor layer for detection. These findings illustrate that designed protein assemblies can function as dynamic components of solid-state devices in non-aqueous environments.

36 MATERIALS SCIENCE

Birth of protein folds and functions in the virome

The rapid evolution of viruses generates proteins that are essential for infectivity and replication but with unknown functions, due to extreme sequence divergence. Here, using a database of 67,715 newly predicted protein structures from 4,463 eukaryotic viral species, we found that 62% of viral proteins are structurally distinct and lack homologues in the AlphaFold database. Among the remaining 38% of viral proteins, many have non-viral structural analogues that revealed surprising similarities between human pathogens and their eukaryotic hosts. Structural comparisons suggested putative functions for up to 25% of unannotated viral proteins, including those with roles in the evasion of innate immunity. In particular, RNA ligase T-like phosphodiesterases were found to resemble phage-encoded proteins that hydrolyse the host immune-activating cyclic dinucleotides 3',3'- and 2',3'-cyclic GMP-AMP (cGAMP). Experimental analysis showed that RNA ligase T homologues encoded by avian poxviruses similarly hydrolyse cGAMP, showing that RNA ligase T-mediated targeting of cGAMP is an evolutionarily conserved mechanism of immune evasion that is present in both bacteriophage and eukaryotic viruses. Together, the viral protein structural database and analyses presented here afford new opportunities to identify mechanisms of virus–host interactions that are common across the virome.

59 BASIC BIOLOGICAL SCIENCES

Assembly of respiratory syncytial virus matrix protein lattice and its coordination with fusion glycoprotein trimers

Respiratory syncytial virus (RSV) is an enveloped, filamentous, negative-strand RNA virus that causes significant respiratory illness worldwide. RSV vaccines are available, however there is still significant need for research to support the development of vaccines and therapeutics against RSV and related Mononegavirales viruses. Individual virions vary in size, with an average diameter of ~130 nm and ranging from ~500 nm to over 10 µm in length. Though the general arrangement of structural proteins in virions is known, we use cryo-electron tomography and sub-tomogram averaging to determine the molecular organization of RSV structural proteins. We show that the peripheral membrane-associated RSV matrix (M) protein is arranged in a packed helical-like lattice of M-dimers. We report that RSV F glycoprotein is frequently observed as pairs of trimers oriented in an anti-parallel conformation to support potential interactions between trimers. Our sub-tomogram averages indicate the positioning of F-trimer pairs is correlated with the underlying M lattice. These results provide insight into RSV virion organization and may aid in the development of RSV vaccines and anti-viral targets.

59 BASIC BIOLOGICAL SCIENCES

Repetitive proteins that undergo large conformational changes evade structural prediction algorithms

Protein structure prediction algorithms, such as AlphaFold, have accelerated protein design and advanced the understanding of the relationship between amino acid sequence and protein structure. However, these algorithms are limited in their ability to predict the structures of conformationally dynamic, intrinsically disordered, and stimuli-responsive proteins. To evaluate sequence-to-structure predictions of such challenging proteins, we explored a class of conformationally dynamic, repeats-in-toxin (RTX) proteins. RTX proteins adopt intrinsically disordered conformations in the absence of calcium and undergo reversible folding into β-roll structures upon binding to calcium. RTX proteins are characterized by tandem repeats of the sequence GGXGXDXUX, in which X can be any amino acid and U is an aliphatic amino acid. We designed RTX sequence variants with global substitutions of nonconserved amino acids, tandem repeats of consensus sequences GGAGXDTLY, and tandem repeats of scrambled sequences GGAGXDTYL. AlphaFold2 and AlphaFold3 predicted that all of these RTX variants adopt β-roll structures, characteristic of wild-type RTX bound to calcium. However, modeling the predicted structures with molecular dynamics simulations and characterizing the protein variants with circular dichroism spectroscopy, small-angle x-ray scattering, and x-ray crystallography revealed that variants adopt diverse, sequence-dependent structures in the absence and presence of calcium. To better design proteins for applications in biotechnology and sustainability, it is critical to build predictive tools that consider intrinsically disordered protein states and validate these tools with multi-mode, multi-scale experimental data.

Chang, Marina P. [Stanford Univ., CA (United State

Packaging “vegetable oils”: Insights into plant lipid droplet proteins

Abstract Plant neutral lipids, also known as “vegetable oils”, are synthesized within the endoplasmic reticulum (ER) membrane and packaged into subcellular compartments called lipid droplets (LDs) for stable storage in the cytoplasm. The biogenesis, modulation, and degradation of cytoplasmic LDs in plant cells are orchestrated by a variety of proteins localized to the ER, LDs, and peroxisomes. Recent studies of these LD-related proteins have greatly advanced our understanding of LDs not only as steady oil depots in seeds but also as dynamic cell organelles involved in numerous physiological processes in different tissues and developmental stages of plants. In the past 2 decades, technology advances in proteomics, transcriptomics, genome sequencing, cellular imaging and protein structural modeling have markedly expanded the inventory of LD-related proteins, provided unprecedented structural and functional insights into the protein machinery modulating LDs in plant cells, and shed new light on the functions of LDs in nonseed plant tissues as well as in unicellular algae. Here, we review critical advances in revealing new LD proteins in various plant tissues, point out structural and mechanistic insights into key proteins in LD biogenesis and dynamic modulation, and discuss future perspectives on bridging our knowledge gaps in plant LD biology.

Cai, Yingqi (ORCID:0000000203575809)

Structural analysis of extracellular ATP-independent chaperones of streptococcal species and protein substrate interactions

ABSTRACT During infection, bacterial pathogens rely on secreted virulence factors to manipulate the host cell. However, in gram-positive bacteria, the molecular mechanisms underlying the folding and activity of these virulence factors after membrane translocation are not clear. Here, we solved the protein structures of two secreted parvulin and two secreted cyclophilin-like peptidyl-prolyl isomerase (PPIase) ATP-independent chaperones found in gram-positive streptococcal species. The extracellular parvulin-type PPIase, PrsA inStreptococcus pneumoniaeandStreptococcus mutansmaintain dimeric crystal structures reminiscent of folding catalysts that consist of two domains, a PPIase and foldase domain. Structural comparison of the two cyclophilin-like extracellular chaperones fromS. pneumoniaeandStreptococcus pyogeneswith other cyclophilins demonstrates that this group of cyclophilin-like chaperones has novel structural appendages formed by 9- and 24-residue insertions. Furthermore, we demonstrate that deletion ofprsAandslrAgenes impairs the secretion of the cholesterol-dependent pore-forming toxin, pneumolysin inS. pneumoniae. Using protein pull-down and biophysical assays, we demonstrate a direct interaction between PrsA and SlrA with Ply. Then, we developed chaperone-assisted folding assays that show that theS. pneumoniaePrsA and SlrA extracellular chaperones accelerate pneumolysin folding. In addition, we demonstrate that SlrA and, for the first time,S. pyogenes PpiA exhibit PPIase activity and can bind the immunosuppressive drug, cyclosporine A. Altogether, these findings suggest a mechanistic role for streptococcal PPIase chaperones in the activity and folding of secreted virulence factors such as pneumolysin. IMPORTANCE Streptococcal species are a leading cause of lower respiratory infections that annually affect millions of people worldwide. During infection, streptococcal species secrete a medley of virulence factors that allow the bacteria to colonize and translocate to deeper tissues. In many gram-positive bacteria, virulence factors are secreted from the cytosol across the bacterial membrane in an unfolded state. The bacterial membrane-cell wall interface is exposed to the potentially harsh extracellular environment, making it difficult for native virulence factors to fold before being released into the host. ATP-independent PPIase-type chaperones, PrsA and SlrA, are thought to facilitate folding and stabilization of several unfolded proteins to promote the colonization and spread of streptococci. Here, we present crystal structures of the molecular chaperones of PrsA and SlrA homologs from streptococcal species. We provide evidence that theStreptococcus pyogenesSlrA homolog, PpiA, has PPIase activity and binds to cyclosporine A. In addition, we show thatStreptococcus pneumoniaePrsA and SlrA directly interact and fold the cholesterol-dependent pore-forming toxin and critical virulence determinant, pneumolysin.

Microbiology