Search NASASearch

SEARCH · Search NASA

Results for “Protein structure predictions”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

Repetitive proteins that undergo large conformational changes evade structural prediction algorithms

Protein structure prediction algorithms, such as AlphaFold, have accelerated protein design and advanced the understanding of the relationship between amino acid sequence and protein structure. However, these algorithms are limited in their ability to predict the structures of conformationally dynamic, intrinsically disordered, and stimuli-responsive proteins. To evaluate sequence-to-structure predictions of such challenging proteins, we explored a class of conformationally dynamic, repeats-in-toxin (RTX) proteins. RTX proteins adopt intrinsically disordered conformations in the absence of calcium and undergo reversible folding into β-roll structures upon binding to calcium. RTX proteins are characterized by tandem repeats of the sequence GGXGXDXUX, in which X can be any amino acid and U is an aliphatic amino acid. We designed RTX sequence variants with global substitutions of nonconserved amino acids, tandem repeats of consensus sequences GGAGXDTLY, and tandem repeats of scrambled sequences GGAGXDTYL. AlphaFold2 and AlphaFold3 predicted that all of these RTX variants adopt β-roll structures, characteristic of wild-type RTX bound to calcium. However, modeling the predicted structures with molecular dynamics simulations and characterizing the protein variants with circular dichroism spectroscopy, small-angle x-ray scattering, and x-ray crystallography revealed that variants adopt diverse, sequence-dependent structures in the absence and presence of calcium. To better design proteins for applications in biotechnology and sustainability, it is critical to build predictive tools that consider intrinsically disordered protein states and validate these tools with multi-mode, multi-scale experimental data.

Chang, Marina P. [Stanford Univ., CA (United State

Structure-aware annotation of leucine-rich repeat domains

Protein domain annotation is typically done by predictive models such as HMMs trained on sequence motifs. However, sequence-based annotation methods are prone to error, particularly in calling domain boundaries and motifs within them. These methods are limited by a lack of structural information accessible to the model. With the advent of deep learning-based protein structure prediction, existing sequenced-based domain annotation methods can be improved by taking into account the geometry of protein structures. We develop dimensionality reduction methods to annotate repeat units of the Leucine Rich Repeat solenoid domain. The methods are able to correct mistakes made by existing machine learning-based annotation tools and enable the automated detection of hairpin loops and structural anomalies in the solenoid. The methods are applied to 127 predicted structures of LRR-containing intracellular innate immune proteins in the model plant Arabidopsis thaliana and validated against a benchmark dataset of 172 manually-annotated LRR domains.

Xu, Boyan

Artificial intelligence methods for protein structure and interaction prediction: Recent advances and challenges

Recent advances in artificial intelligence have introduced novel methods for high-accuracy prediction of protein tertiary structures, protein complex structures, and interactions between proteins and other biomolecules, such as small molecules and nucleic acids. Such advancements are accelerating biomedical research and the development of new protein design and bioengineering methods among many other important biotechnology applications. Here, in this review, we outline the recent advances in protein-centric biomolecular structure and interaction prediction, highlight some major challenges in the field, and discuss potential directions to address them.

Morehead, Alex [Lawrence Berkeley National Laborat

From sequence to protein structure and conformational dynamics with artificial intelligence/machine learning

The 2024 Nobel Prize in Chemistry was awarded in part for de novo protein structure prediction using AlphaFold2, an artificial intelligence/machine learning (AI/ML) model trained on vast amounts of sequence and three-dimensional structure data. AlphaFold2 and related models, including RoseTTAFold and ESMFold, employ specialized neural network architectures driven by attention mechanisms to infer relationships between sequence and structure. At a fundamental level, these AI/ML models operate on the long-standing hypothesis that the structure of a protein is determined by its amino acid sequence. More recently, AlphaFold2 has been adapted for the prediction of multiple protein conformations by subsampling multiple sequence alignments. Herein, we provide an overview of the deterministic relationship between sequence and structure, which was hypothesized over half a century ago with profound implications for the biological sciences ever since. We postulate that protein conformational dynamics are also determined, at least in part, by amino acid sequence and that this relationship may be leveraged for construction of AI/ML models dedicated to predicting protein conformational ensembles. Accordingly, we describe a conceptual model architecture, which may be trained on sequence data in combination with conformationally sensitive structural information, coming primarily from nuclear magnetic resonance (NMR) spectroscopy. Notwithstanding certain limitations in this context, NMR offers abundant structural heterogeneity conducive to conformational ensemble prediction. As NMR and other data continue to accumulate, sequence-informed prediction of protein structural dynamics with AI/ML has the potential to emerge as a transformative capability across the biological sciences.

Artificial intelligence

The crystal structure of Grindelia robusta 7,13-copalyl diphosphate synthase reveals active site features controlling catalytic specificity

Diterpenoid natural products serve critical functions in plant development and ecological adaptation and many diterpenoids have economic value as bioproducts. The family of class II diterpene synthases catalyzes the committed reactions in diterpenoid biosynthesis, converting a common geranylgeranyl diphosphate precursor into different bicyclic prenyl diphosphate scaffolds. Enzymatic rearrangement and modification of these precursors generate the diversity of bioactive diterpenoids. We report the crystal structure of Grindelia robusta 7,13-copalyl diphosphate synthase, GrTPS2, at 2.1 Å of resolution. GrTPS2 catalyzes the committed reaction in the biosynthesis of grindelic acid, which represents the signature metabolite in species of gumweed (Grindelia spp., Asteraceae). Grindelic acid has been explored as a potential source for drug leads and biofuel production. The GrTPS2 crystal structure adopts the conserved three-domain fold of class II diterpene synthases featuring a functional active site in the γβ-domain and a vestigial α-domain. Substrate docking into the active site of the GrTPS2 apo protein structure predicted catalytic amino acids. Biochemical characterization of protein variants identified residues with impact on enzyme activity and catalytic specificity. Specifically, mutagenesis of Y457 provided mechanistic insight into the position-specific deprotonation of the intermediary carbocation to form the characteristic 7,13 double bond of 7,13-copalyl diphosphate.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH

De novo atomic protein structure modeling for cryoEM density maps using 3D transformer and HMM

Accurately building 3D atomic structures from cryo-EM density maps is a crucial step in cryo-EM-based protein structure determination. Converting density maps into 3D atomic structures for proteins lacking accurate homologous or predicted structures as templates remains a significant challenge. Here, we introduce Cryo2Struct, a fully automated de novo cryo-EM structure modeling method. Cryo2Struct utilizes a 3D transformer to identify atoms and amino acid types in cryo-EM density maps, followed by an innovative Hidden Markov Model (HMM) to connect predicted atoms and build protein backbone structures. Cryo2Struct produces substantially more accurate and complete protein structural models than the widely used ab initio method Phenix. Additionally, its performance in building atomic structural models is robust against changes in the resolution of density maps and the size of protein structures.

59 BASIC BIOLOGICAL SCIENCES

Structure and mechanism of human vesicular polyamine transporter

Polyamines play essential roles in gene expression and modulate neuronal transmission in mammals. Vesicular polyamine transporters (VPAT) from the SLC18 family exploit the transmembrane H + gradient to translocate polyamines into secretory vesicles, enabling the quantal release of polyamine neuromodulators and underpinning learning and memory formation. Here, we report the cryo-electron microscopy structures of human VPAT in complex with spermine, spermidine, H + , or tetrabenazine, elucidating discrete lumen-facing states of the antiporter and pivotal interactions between VPAT and its substrate or inhibitor. Leveraging structure-inspired mutagenesis studies and protein structure prediction, we deduce an unforeseen mechanism whereby the polyamine and H + compete for multiple acidic protein residues both directly and indirectly, and rationalize how the antidopaminergic therapeutic tetrabenazine impedes vesicular transport of polyamines. This study unravels the mechanism of an H + -coupled polyamine antiporter, reveals mechanistic diversity between VPAT and other SLC18 antiporters, and raises new prospects for combating human disorders of polyamine homeostasis.

59 BASIC BIOLOGICAL SCIENCES

SLAB: simultaneous labeling and binding affinity prediction for protein–ligand structures

Machine learning models are often used as scoring functions to predict the binding affinity of a protein–ligand complex. These models are trained with limited amounts of data with experimentally measured binding affinity values. A large number of compounds are labeled inactive through single-concentration screens without measuring binding affinities. These inactive compounds, along with the active ones, can be used to train binary classification models, while regression models are trained using compounds with binding affinities only. However, the classification and regression tasks are often handled separately, without sharing the learned feature representations. In this paper, we propose a novel model architecture that jointly performs regression and classification objectives, aiming to maximize data utilization and improve predictive performance by leveraging two complementary tasks. In our setup, the regression yields the binding affinity, whereas the classification task yields the label as active or inactive. We demonstrate our method using PDBbind, the standard 3D structure database, as well as a dataset of flavivirus protease compounds with binding affinity data. Our experiments show that the new joint training strategy improves the accuracy of the model, increasing applicability in various practical drug screening scenarios.

Biological and medical sciences

African Swine Fever Virus Protein–Protein Interaction Prediction

The African swine fever virus (ASFV) is an often deadly disease in swine and poses a threat to swine livestock and swine producers. With its complex genome containing more than 150 coding regions, developing effective vaccines for this virus remains a challenge due to a lack of basic knowledge about viral protein function and protein–protein interactions between viral proteins and between viral and host proteins. In this work, we identified ASFV-ASFV protein–protein interactions (PPIs) using artificial intelligence-powered protein structure prediction tools. We benchmarked our PPI identification workflow on the Vaccinia virus, a widely studied nucleocytoplasmic large DNA virus, and found that it could identify gold-standard PPIs that have been validated in vitro in a genome-wide computational screening. We applied this workflow to more than 18,000 pairwise combinations of ASFV proteins and were able to identify seventeen novel PPIs, many of which have corroborating experimental or bioinformatic evidence for their protein–protein interactions, further validating their relevance. Two protein–protein interactions, I267L and I8L, I267L__I8L, and B175L and DP79L, B175L__DP79L, are novel PPIs involving viral proteins known to modulate host immune response.

59 BASIC BIOLOGICAL SCIENCES

ZMPY3D: accelerating protein structure volume analysis through vectorized 3D Zernike moments and Python-based GPU integration

Abstract Motivation Volumetric 3D object analyses are being applied in research fields such as structural bioinformatics, biophysics, and structural biology, with potential integration of artificial intelligence/machine learning (AI/ML) techniques. One such method, 3D Zernike moments, has proven valuable in analyzing protein structures (e.g., protein fold classification, protein–protein interaction analysis, and molecular dynamics simulations). Their compactness and efficiency make them amenable to large-scale analyses. Established methods for deriving 3D Zernike moments, however, can be inefficient, particularly when higher order terms are required, hindering broader applications. As the volume of experimental and computationally-predicted protein structure information continues to increase, structural biology has become a “big data” science requiring more efficient analysis tools. Results This application note presents a Python-based software package, ZMPY3D, to accelerate computation of 3D Zernike moments by vectorizing the mathematical formulae and using graphical processing units (GPUs). The package offers popular GPU-supported libraries such as CuPy and TensorFlow together with NumPy implementations, aiming to improve computational efficiency, adaptability, and flexibility in future algorithm development. The ZMPY3D package can be installed via PyPI, and the source code is available from GitHub. Volumetric-based protein 3D structural similarity scores and transform matrix of superposition functionalities have both been implemented, creating a powerful computational tool that will allow the research community to amalgamate 3D Zernike moments with existing AI/ML tools, to advance research and education in protein structure bioinformatics. Availability and implementation ZMPY3D, implemented in Python, is available on GitHub (https://github.com/tawssie/ZMPY3D) and PyPI, released under the GPL License.

Lai, Jhih-Siang (ORCID:0000000156775890)

Protein Data Bank (PDB): Fifty-three years young and having a transformative impact on science and society

This review article describes the co-evolution of structural biology as a discipline and the Protein Data Bank (PDB), established in 1971 as the first open-access data resource in biology by like-minded structural scientists. As the PDB archive grew in size and scope to encompass macromolecular crystallography, NMR spectroscopy, and cryo-electron microscopy, new technologies were developed to ingest, validate, curate, store, and distribute the information. Community engagement ensured that the needs of structural biologists (data depositors) and data consumers were met. Today, the archive houses more than 230,000 experimentally determined structures of proteins, nucleic acids, and macromolecular machines and their complexes with one another and small-molecule ligands. Aggregate costs of PDB data preservation are ~1% of the cost of structure determination. The enormous impact of PDB data on basic and applied research and education across the natural and medical sciences is presented and highlighted with illustrative examples. Enablement of de novo protein structure prediction (AlphaFold2, RoseTTAfold, OpenFold, etc.) is the most widely appreciated benefit of having a corpus of rigorously validated, expertly curated 3D biostructure data.

bioinformatics

Unraveling the Molecular Origin of Prey-Wrapping Spider Silk's Unique Mechanical Properties and Assembly Process Using NMR

Prey wrapping spider silk's unique mechanical properties are investigated confirming the silk's high degree of extensibility and superior toughness compared to other types of spider silk. For the first time, the pre-spinning dope phase is studied in isotope-enriched intact aciniform (AC) silk glands using solution NMR that reveals a combination of α-helical domains linked by disordered random coil chains consistent with previously proposed “beads-on-a-string” models. The model is further refined through the AlphaFold2 protein structure prediction tool. Finally, extensive magic angle spinning (MAS) solid-state (SS) NMR data for isotopically-enriched fibers is used to refine the structural model for AC silk from two species, A. aurantia and A. argentata. The SSNMR data shows that the AC silk fibers are highly α-helical, coiled-coil in structure but, also exhibit significant β-sheet components that can be traced back to the Gly-rich disordered linker regions in the pre-spinning dope phase that are converted to β-sheet structures during fiber formation. This combination of mechanical and structural characterization enhances the understanding of AC silk's liquid-to-solid transition and structure-mechanics relationship. In conclusion, these prey wrap silk results and models will provide the basis for the design of biomimetic materials inspired by the AC spider silk system.

36 MATERIALS SCIENCE

Birth of protein folds and functions in the virome

The rapid evolution of viruses generates proteins that are essential for infectivity and replication but with unknown functions, due to extreme sequence divergence. Here, using a database of 67,715 newly predicted protein structures from 4,463 eukaryotic viral species, we found that 62% of viral proteins are structurally distinct and lack homologues in the AlphaFold database. Among the remaining 38% of viral proteins, many have non-viral structural analogues that revealed surprising similarities between human pathogens and their eukaryotic hosts. Structural comparisons suggested putative functions for up to 25% of unannotated viral proteins, including those with roles in the evasion of innate immunity. In particular, RNA ligase T-like phosphodiesterases were found to resemble phage-encoded proteins that hydrolyse the host immune-activating cyclic dinucleotides 3',3'- and 2',3'-cyclic GMP-AMP (cGAMP). Experimental analysis showed that RNA ligase T homologues encoded by avian poxviruses similarly hydrolyse cGAMP, showing that RNA ligase T-mediated targeting of cGAMP is an evolutionarily conserved mechanism of immune evasion that is present in both bacteriophage and eukaryotic viruses. Together, the viral protein structural database and analyses presented here afford new opportunities to identify mechanisms of virus–host interactions that are common across the virome.

59 BASIC BIOLOGICAL SCIENCES

Comparative genomic analysis of thermophilic fungi reveals convergent evolutionary adaptations and gene losses

Thermophily is a trait scattered across the fungal tree of life, with its highest prevalence within three fungal families (Chaetomiaceae, Thermoascaceae, and Trichocomaceae), as well as some members of the phylum Mucoromycota. We examined 37 thermophilic and thermotolerant species and 42 mesophilic species for this study and identified thermophily as the ancestral state of all three prominent families of thermophilic fungi. Thermophilic fungal genomes were found to encode various thermostable enzymes, including carbohydrate-active enzymes such as endoxylanases, which are useful for many industrial applications. At the same time, the overall gene counts, especially in gene families responsible for microbial defense such as secondary metabolism, are reduced in thermophiles compared to mesophiles. We also found a reduction in the core genome size of thermophiles in both the Chaetomiaceae family and the Eurotiomycetes class. The Gene Ontology terms lost in thermophilic fungi include primary metabolism, transporters, UV response, and O-methyltransferases. Comparative genomics analysis also revealed higher GC content in the third base of codons (GC3) and a lower effective number of codons in fungal thermophiles than in both thermotolerant and mesophilic fungi. Furthermore, using the Support Vector Machine classifier, we identified several Pfam domains capable of discriminating between genomes of thermophiles and mesophiles with 94% accuracy. Using AlphaFold2 to predict protein structures of endoxylanases (GH10), we built a similarity network based on the structures. We found that the number of disulfide bonds appears important for protein structure, and the network clusters based on protein structures correlate with the optimal activity temperature. Thus, comparative genomics offers new insights into the biology, adaptation, and evolutionary history of thermophilic fungi while providing a parts list for bioengineering applications.

59 BASIC BIOLOGICAL SCIENCES

Enhanced polymorph metastability drives glycine nucleation in aqueous salt solutions

Crystal nucleation from aqueous solutions influences countless geological, biochemical, astrophysical, environmental, and materials science–related phenomena, including ice formation, the manufacturing of active pharmaceutical ingredients, development of diseases such as Alzheimer’s and the origin of life itself. Understanding and controlling nucleation is essential for designing materials with specific properties, developing strategies to inhibit or promote crystallization in various contexts and preventing pathological aggregation in neurodegenerative diseases. Similar to the protein structure prediction problem—where a single amino acid sequence can in theory adopt one most stable conformation but in practice may sample multiple competing conformations—crystal nucleation faces a parallel challenge: the same chemical species can form diverse polymorphs under different environmental conditions (e.g., temperature, pressure, solvent). Each polymorph presents its own set of physical and chemical properties, highlighting the importance of understanding and controlling polymorph selection in fields ranging from pharmaceuticals to materials design. Despite advances in experimental and computational methods for studying phase transitions and polymorph stability, nucleation remains challenging due to its nanoscale nature. Furthermore, in practical settings, salts and impurities can further influence crystal nucleation in diverse contexts, from scaling in pipelines and desalination plants to the durability of concrete and the efficiency of battery materials. This can lead to the formation of polymorphs that may differ from the most stable phase in pure solutions. Or, even though the final structure might appear same irrespective of whether the environment contained impurities or not, the mechanism through which it was formed might be completely different and not intuitive.

Wang, Ruiyu [University of Maryland, College Park,

MAL33 drives natural variation in maltose metabolism in Saccharomyces eubayanus

Maltose is one of the most abundant sugars in brewer’s wort, and its efficient utilization is critical for successful fermentation. However, maltose consumption varies naturally among Saccharomyces eubayanus strains isolated from different host trees, such as Quercus and Nothofagus. To identify the genetic determinants underlying these phenotypic differences, we performed bulk segregant analysis (BSA) and quantitative trait loci (QTL) mapping using an F 2 offspring derived from QC18 (Quercus-associated) and CL467.1 (Nothofagus-associated) strains. QTL mapping identified two significant genomic regions on subtelomeric loci of chromosomes V-R and XVI-L, each containing complete MAL loci composed of MAL32 (encoding maltase), MAL31 (transporter), and MAL33 (transcriptional activator) genes. Comparative polymorphism analyses identified mutations in MAL32 and MAL33 of QC18, including frameshift mutations resulting in premature stop codons. Functional validation demonstrated that the heterologous expression of MAL33 ChrV from CL467.1 fully restored maltose utilization in QC18, indicating the functional presence of MAL33 cis-regulatory sequences and MAL32 and MAL31 genes in QC18. While structural protein predictions identified truncation and impaired functionality in the maltose-responsive activation domain of Mal33p from QC18, overexpression of QC18’s own MAL33 ChrV allele also improved maltose metabolism, suggesting dosage-dependent transcriptional limitations rather than complete functional loss. These results indicate that allelic variations in the maltose-responsive activation domain of Mal33p result in differences in maltose consumption between strains. Here, we hypothesized that reduced maltose metabolism in QC18 is an adaptive response to the distinct sugar composition in Quercus robur bark, contrasting with the starch-rich environment of Nothofagus pumilio. These findings highlight subtelomeric MAL gene diversity as a reservoir of genetic variation, representing a key evolutionary mechanism that influences maltose adaptation among natural Saccharomyces isolates.

evolutionary plasticity

SFold v0.1

This is a scientific software package to integrate Small Angle X-ray Scattering (SAXS) experimental data into OpenFold deep learning models to improve protein structure prediction.

Prince, Stephanie [Lawrence Berkeley National Labo