Search NASA⌕ Search

SEARCH · Search NASA

Results for “Protein structure predictions”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 73 records · Page 4

An evolutionarily conserved tryptophan cage promotes folding of the extended RNA recognition motif in the hnRNPR ‐like protein family

Abstract The heterogeneous nuclear ribonucleoprotein (hnRNP) R‐like family is a class of RNA binding proteins in the hnRNP superfamily with diverse functions in RNA processing. Here, we present the 1.90 Å X‐ray crystal structure and solution NMR studies of the first RNA recognition motif (RRM) of human hnRNPR. We find that this domain adopts an extended RRM (eRRM1) featuring a canonical RRM with a structured N‐terminal extension (N ext ) motif that docks against the RRM and extends the β‐sheet surface. The adjoining loop is structured and forms a tryptophan cage motif to position the N ext motif for docking to the RRM. Combining mutagenesis, solution NMR spectroscopy, and thermal denaturation studies, we evaluate the importance of residues in the N ext –RRM interface and adjoining loop on eRRM folding and conformational dynamics. We find that these sites are essential for protein solubility, conformational ordering, and thermal stability. Consistent with their importance, mutations in the N ext –RRM interface and loop are associated with several cancers in a survey of somatic mutations in cancer studies. Sequence and structure comparison of the human hnRNPR eRRM1 to experimentally verified and predicted hnRNPR‐like proteins reveals conserved features in the eRRM.

Biochemistry & Molecular Biology↗

PNNL-Predictive-Phenomics/ProCaliper

ProCaliper is a Python library that curates, organizes, and computes protein structure features in a way that easily interfaces with user-provided experimental data. It extracts or computes protein binding site, active site, charge, pLDDT (order/disorder), acid dissociation, protonation, solvent accessible surface area, disulfide bond distance, and protein secondary structure data using precomputed protein structures and publicly available databases. It provides a unified API for integrating additional residue-level data and for visualizing residue features in 3D.

Rozum, Jordan [Pacific Northwest National Lab]↗

Light-modulated abundance of an mRNA encoding a calmodulin-regulated, chromatin-associated NTPase in pea

A CDNA encoding a 47 kDa nucleoside triphosphatase (NTPase) that is associated with the chromatin of pea nuclei has been cloned and sequenced. The translated sequence of the cDNA includes several domains predicted by known biochemical properties of the enzyme, including five motifs characteristic of the ATP-binding domain of many proteins, several potential casein kinase II phosphorylation sites, a helix-turn-helix region characteristic of DNA-binding proteins, and a potential calmodulin-binding domain. The deduced primary structure also includes an N-terminal sequence that is a predicted signal peptide and an internal sequence that could serve as a bipartite-type nuclear localization signal. Both in situ immunocytochemistry of pea plumules and immunoblots of purified cell fractions indicate that most of the immunodetectable NTPase is within the nucleus, a compartment proteins typically reach through nuclear pores rather than through the endoplasmic reticulum pathway. The translated sequence has some similarity to that of human lamin C, but not high enough to account for the earlier observation that IgG against human lamin C binds to the NTPase in immunoblots. Northern blot analysis shows that the NTPase MRNA is strongly expressed in etiolated plumules, but only poorly or not at all in the leaf and stem tissues of light-grown plants. Accumulation of NTPase mRNA in etiolated seedlings is stimulated by brief treatments with both red and far-red light, as is characteristic of very low-fluence phytochrome responses. Southern blotting with pea genomic DNA indicates the NTPase is likely to be encoded by a single gene.

NASA Discipline Number 40-50↗

Improved deep learning prediction of antigen–antibody interactions

Identifying antibodies that neutralize specific antigens is crucial for developing effective immunotherapies, but this task remains challenging for many target antigens. The rise of deep learning–based computational approaches presents a promising avenue to address this challenge. Here, we assess the performance of a deep learning approach through two benchmark tests aimed at predicting antibodies for the receptor-binding domain of the severe acute respiratory syndrome coronavirus 2 (SARS-CoV-2) spike protein. Three different strategies for constructing input sequence alignments are employed for predicting structural models of antigen–antibody complexes. In our initial testing set, which comprises known experimental structures, these strategies collectively yield a significant top-ranked prediction for 61% of cases and a success rate of 47%. Notably, one strategy that utilizes the sequences of known antigen binders outperforms the other two, achieving a precision of 90% in a subsequent test set of ~1,000 antibodies, balanced between true and control antibodies for the antigen, albeit with a lower recall of 25%. Our results underscore the potential of integrating deep learning methods with single B cell sequencing techniques to enhance the prediction accuracy of antigen–antibody interactions.

Science & Technology - Other Topics↗

MarK, a Novosphingobium aromaticivorans kinase required for catabolism of multiple aromatic monomers

The aromatic compounds used in a variety of industrial products are currently obtained from nonrenewable petroleum sources. Alternatively, the plant polymer lignin is an abundant renewable source of aromatics, and its depolymerization generates a variety of products that can include acetovanillone, a vanillin derivative containing an acetyl side chain. The Alphaproteobacterium Novosphingobium aromaticivorans DSM12444 can metabolize several chemically modified aromatics in deconstructed lignin, but not acetovanillone. In this work, adaptive laboratory evolution identified a single amino acid change in the previously uncharacterized gene product Saro_1862 that is necessary and sufficient for N. aromaticivorans growth with acetovanillone as a sole growth substrate, as well as other aromatic monomers not metabolized by wild-type cells. We show that a glutamate (E) to lysine (K) substitution at amino acid residue 16 of Saro_1862 results in a ~1600-fold increase in the rate of ATP-dependent acetovanillone phosphorylation. We also find that recombinant Saro_1862 E16K phosphorylates several other aromatic compounds in vitro , defining the first reported catalytic activity for the widespread UPF0261 protein domain contained in Saro_1862. Thus, we propose naming Saro_1862 MarK, for multiple aromatic kinase. A 1.57 Å crystal structure of MarK E16K predicts that the E16K substitution lies in a potential ATP binding site, suggesting how this amino acid change increased catalytic activity. A search for homologs of MarK and other proteins required for acetovanillone degradation predicts that this pathway for aromatic metabolism exists throughout the bacterial phylogeny.

Novosphingobium↗

Correlating Protein Dynamics and Catalytic Activity of a Model Hydrogenase Using Paramagnetic and Biological Nuclear Magnetic Resonance Spectroscopy

Rational catalyst design remains a significant challenge, with electronic structure, steric, and electrostatic effects known to contribute to activity. Recently, dynamics has been recognized as another factor that impacts catalysis, though identifying and predicting these effects has remained out of reach. Nickel-substituted rubredoxin (NiRd), a protein-based mimic of a hydrogenase enzyme, serves as a model catalytic system in which dynamics can be systematically investigated with respect to activity. While over 30 secondary-sphere mutants of NiRd have been shown to be catalytically active, no significant correlation was observed between the rates and catalytic overpotential or electronic structure, prompting questions about the protein-derived factors that modulate activity. Here, in this work, NMR spectroscopy was used to investigate the roles of substrate accessibility, protein dynamics, and protein stability in controlling catalysis. Significant paramagnetic effects from the nickel center (S = 1) isolate the methylene proton resonances of the metal-coordinating cysteine residues. The sensitivity of resonance positions and linewidths to local environment offers an opportunity to study dynamical molecular changes around the metal center with high resolution. Machine learning algorithms were employed to identify correlations between the catalytic activity and the paramagnetic NMR spectra. These analyses revealed spectroscopic features of specific cysteine protons that report on catalytic overpotential and increased turnover rates, which are further supported by the results obtained using high-field NMR techniques. Collectively, these studies indicate the potential for multifrequency NMR techniques to resolve key contributors to catalytic activity and highlight the importance of local and outer-sphere dynamics.

Protein Engineering↗

metagRoot: a comprehensive database of protein families associated with plant root microbiomes

The plant root microbiome is vital in plant health, nutrient uptake, and environmental resilience. To explore and harness this diversity, we present metagRoot, a specialized and enriched database focused on the protein families of the plant root microbiome. MetagRoot integrates metagenomic, metatranscriptomic, and reference genome-derived protein data to characterize 71 091 enriched protein families, each containing at least 100 sequences. These families are annotated with multiple sequence alignments, CRISPR elements, hidden Markov models, taxonomic and functional classifications, ecosystem and geolocation metadata, and predicted 3D structures using AlphaFold2. MetagRoot is a powerful tool for decoding the molecular landscape of root-associated microbial communities and advancing microbiome-informed agricultural practices by enriching protein family information with ecological and structural context. The database is available at https://pavlopoulos-lab.org/metagroot/ or https://www.metagroot.org.

Chasapi, Maria N↗

A curated benchmark for cofolding models on kinase conformational states

Abstract Protein kinases are critical drug targets, requiring therapeutics that can modulate their active and inactive conformational states. While cofolding models can generate global folds directly from kinase sequences and ligand SMILES strings, these models have not yet been tested on their ability to recover ligand-induced-fit conformational states of the kinase proteins. Here, we introduce KinConfBench, a curated benchmark of 2225 high-quality human kinase chains to evaluate the ability of four state-of-the-art cofolding models—Boltz-2, Chai-1, Protenix, and RoseTTAFold-All-Atom—to recover both canonical and rare conformational states. We show that geometric success metrics of a ligand pose in the active site do not correlate strongly with the correct kinase conformational state, motivating a new set of dynamical benchmarks for assessing cofolding models. While all four cofolding models achieve ~60–80% prediction accuracy for kinase conformational classification, they exhibit severe mode collapse when performing multiple inferences, show negligible structural diversity in sampling induced-fit motions, and display a prevalent “apo-drift” in which most cofolding models predominantly predict the kinase to be in its ligand-free state. Our results highlight that capturing ligand-induced protein conformational diversity, not just geometric fit, is critical for next-generation structure-based drug discovery.

Sun, Kunyang↗

Fast myosin binding protein C knockout in skeletal muscle alters length-dependent activation and myofilament structure

In striated muscle, the sarcomeric protein myosin-binding protein-C (MyBP-C) is bound to the myosin thick filament and is predicted to stabilize myosin heads in a docked position against the thick filament, which limits crossbridge formation. Here, we use the homozygous Mybpc2 knockout (C2 -/- ) mouse line to remove the fast-isoform MyBP-C from fast skeletal muscle and then conduct mechanical functional studies in parallel with small-angle X-ray diffraction to evaluate the myofilament structure. We report that C2 -/- fibers present deficits in force production and calcium sensitivity. Structurally, passive C2 -/- fibers present altered sarcomere length-independent and -dependent regulation of myosin head conformations, with a shift of myosin heads towards actin. At shorter sarcomere lengths, the thin filament is axially extended in C2 -/- , which we hypothesize is due to increased numbers of low-level crossbridges. These findings provide testable mechanisms to explain the etiology of debilitating diseases associated with MyBP-C.

59 BASIC BIOLOGICAL SCIENCES↗

The Factors Governing Metal Dependence of an Emergent Superfamily of Bimetallic Oxygenases

Metalloenzyme superfamilies are typically defined by their protein scaffolds and active sites. Owing to the high tunability of protein structures, members of a single superfamily can catalyze diverse reactions with the same metallocofactor. Some superfamilies, such as amidohydrolase-related dinuclear oxygenases (AROs), display further versatility by utilizing multiple metallocofactors. We have shown that certain AROs catalyze monooxygenation reactions with diiron, dimanganese, and/or mixed manganese−iron cofactors, but the molecular factors governing the selection of a particular cofactor remain unknown, and the extent of this superfamily in biology is unclear. Here, we report bioinformatic analyses that expand the ARO superfamily to approximately 17,000 unique UniProt sequences, far exceeding the number of previously characterized enzymes. Through the integration of structural, spectroscopic, and thermodynamic analyses of representative proteins with a bioinformatic pipeline that identifies key secondary- and tertiary-sphere residues, we can predict in silico the metal preference for the majority of reported ARO sequences. These annotations were validated via the characterization of multiple new AROs, including ones implicated in key oxidative steps of natural product biosyntheses. This study establishes the key structure−function relationships governing metal preferences in AROs and highlights their vastly underappreciated role in myriad biological processes.

Liu, Chang [University of California, Berkeley, CA↗

A characterization of recombinant Arabidopsis FRIABLE1 (FRB1) reveals robust rhamnogalacturonan-I rhamnosyltransferase activity and critical catalytic residues

Plant cell walls are glycan-rich extracellular matrices that fundamentally impact essential cellular processes, such as growth, adhesion, and cell shape acquisition. Understanding plant cell wall glycans requires the identification and characterization of the biosynthetic enzymes that produce these polymers. Most successful in vitro protein expression studies of plant cell wall glycosyltransferases have relied on insect, fungal/yeast, or human cell expression systems, whereas prokaryotic expression systems have been generally unsuccessful. Here, we show that Arabidopsis FRIABLE1 (FRB1)/rhamnogalacturonan-I rhamnosyltransferase 8 (RRT8) can be produced in Escherichia coli RosettaGami2 cells as N-terminal maltose-binding protein fusion proteins containing C-terminal 6X-His-tags. We also report the catalytic constants of FRB1/RRT8 with apparent K M and K cat values of 226 μM and 33 min -1 for UDP-Rhamnose and 117 μM and 28.7 min -1 for rhamnogalacturonan-I (RG-I), respectively. We examine the catalytic activities of mutated FRB1/RRT8 proteins based on an AlphaFold 3-generated FRB1/RRT8 protein structural model with a virtually docked UDP-Rha donor. Enzymatic characterization of the mutated and wildtype FRB1/RRT8 protein confirmed that mutation of predicted catalytic site amino acid residues resulted in a 20-fold reduction in RRT activity. FRB1 also robustly polymerizes RG-I in combination with RG-I galacturonosyltransferase 1. These results show how a robust E. coli expression system combined with artificial intelligence tools can be used to increase understanding of plant cell wall glycosyltransferase structure and function.

glycosyltransferase↗

Contact-dependent growth inhibition (CDI) systems deploy a large family of polymorphic ionophoric toxins for inter-bacterial competition

Contact-dependent growth inhibition (CDI) is a widespread form of inter-bacterial competition mediated by CdiA effector proteins. CdiA is presented on the inhibitor cell surface and delivers its toxic C-terminal region (CdiA-CT) into neighboring bacteria upon contact. Inhibitor cells also produce CdiI immunity proteins, which neutralize CdiA-CT toxins to prevent auto-inhibition. Here, we describe a diverse group of CDI ionophore toxins that dissipate the transmembrane potential in target bacteria. These CdiA-CT toxins are composed of two distinct domains based on AlphaFold2 modeling. The C-terminal ionophore domains are all predicted to form five-helix bundles capable of spanning the cell membrane. The N-terminal "entry" domains are variable in structure and appear to hijack different integral membrane proteins to promote toxin assembly into the lipid bilayer. The CDI ionophores deployed by E. coli isolates partition into six major groups based on their entry domain structures. Comparative sequence analyses led to the identification of receptor proteins for ionophore toxins from groups 1 & 3 (AcrB), group 2 (SecY) and groups 4 (YciB). Using forward genetic approaches, we identify novel receptors for the group 5 and 6 ionophores. Group 5 exploits homologous putrescine import proteins encoded by puuP and plaP, and group 6 toxins recognize di/tripeptide transporters encoded by paralogous dtpA and dtpB genes. Finally, we find that the ionophore domains exhibit significant intra-group sequence variation, particularly at positions that are predicted to interact with CdiI. Accordingly, the corresponding immunity proteins are also highly polymorphic, typically sharing only ~30% sequence identity with members of the same group. Competition experiments confirm that the immunity proteins are specific for their cognate ionophores and provide no protection against other toxins from the same group. The specificity of this protein interaction network provides a mechanism for self/nonself discrimination between E. coli isolates.

59 BASIC BIOLOGICAL SCIENCES↗

Evolutionary, structural and biochemical evidence for a new interaction site of the leptin obesity protein

The Leptin protein is central to the regulation of energy metabolism in mammals. By integrating evolutionary, structural, and biochemical information, a surface segment, outside of its known receptor contacts, is predicted as a second interaction site that may help to further define its roles in energy balance and its functional differences between humans and other mammals.

Evolution, Molecular↗

Force Field X: A computational microscope to study genetic variation and organic crystals using theory and experiment

Force Field X (FFX) is an open-source software package for atomic resolution modeling of genetic variants and organic crystals that leverages advanced potential energy functions and experimental data. FFX currently consists of nine modular packages with novel algorithms that include global optimization via a many-body expansion, acid–base chemistry using polarizable constant-pH molecular dynamics, estimation of free energy differences, generalized Kirkwood implicit solvent models, and many more. Applications of FFX focus on the use and development of a crystal structure prediction pipeline, biomolecular structure refinement against experimental datasets, and estimation of the thermodynamic effects of genetic variants on both proteins and nucleic acids. The use of Parallel Java and OpenMM combines to offer shared memory, message passing, and graphics processing unit parallelization for high performance simulations. Overall, the FFX platform serves as a computational microscope to study systems ranging from organic crystals to solvated biomolecular systems.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Machine learning guided selection of broad-spectrum epitope-specific functional antibodies for "Disease X"

Our project established and demonstrated a transfer learning framework that enables prediction of antibody–antigen interactions across related viruses. The approach focused on three major activities: 1. Conserved region and epitope identification – We compared viral protein structures and sequences to identify shared receptor-binding domains and neutralizing epitope regions across variants and related viruses. These conserved features formed the foundation for discovering broadly functional antibodies. 2. Machine learning model development – We built neural network–based models that integrate epitope features with antibody sequence information. Instead of relying solely on structural or physical properties, the models learned transferable patterns that describe antibody binding potential across different viral families. 3. Transfer learning and validation – Using SARS-CoV-2 and Ebola as source systems, we successfully transferred learned epitope features to predict antibody interactions for SARS CoV-1 and Marburg virus. Iterative cycles of dataset generation, retraining, and evaluation improved generalization and predictive power, ensuring the framework can adapt to new threats.

59 BASIC BIOLOGICAL SCIENCES↗

Study of Fluid Flow Control in Protein Crystallization using Strong Magnetic Fields

An important component in biotechnology, particularly in the area of protein engineering and rational drug design is the knowledge of the precise three-dimensional molecular structure of proteins. The quality of structural information obtained from X-ray diffraction methods is directly dependent on the degree of perfection of the protein crystals. As a consequence, the growth of high quality macromolecular crystals for diffraction analyses has been the central focus for biochemists, biologists, and bioengineers. Macromolecular crystals are obtained from solutions that contain the crystallizing species in equilibrium with higher aggregates, ions, precipitants, other possible phases of the protein, foreign particles, the walls of the container, and a likely host of other impurities. By changing transport modes in general, i.e., reduction of convection and sedimentation, as is achieved in "microgravity", researchers have been able to dramatically affect the movement and distribution of macromolecules in the fluid, and thus their transport, formation of crystal nuclei, and adsorption to the crystal surface. While a limited number of high quality crystals from space flights have been obtained, as the recent National Research Council (NRC) review of the NASA microgravity crystallization program pointed out, the scientific approach and research in crystallization of proteins has been mainly empirical yielding inconclusive results. We postulate that we can reduce convection in ground-based experiments and we can understand the different aspects of convection control through the use of strong magnetic fields and field gradients. Whether this limited convection in a magnetic field will provide the environment for the growth of high quality crystals is still a matter of conjecture that our research will address. The approach exploits the variation of fluid magnetic susceptibility with concentration for this purpose and the convective damping is realized by appropriately positioning the crystal growth cell so that the magnetic susceptibility force counteracts terrestrial gravity. The general objective is to test the hypothesis of convective control using a strong magnetic field and magnetic field gradient and to understand the nature of the various forces that come into play. Specifically we aim to delineate causative factors and to quantify them through experiments, analysis and numerical modeling. Once the basic understanding is obtained, the study will focus on testing the hypothesis on proteins of pyruvate dehydrogenase complex (PDC), proteins E1 and E3. Obtaining high crystal quality of these proteins is of great importance to structural biologists since their structures need to be determined. Specific goals for the investigation are: 1. To develop an understanding of convection control in diamagnetic fluids with concentration gradients through experimentation and numerical modeling. Specifically solutal buoyancy driven convection due to crystal growth will be considered. 2. To develop predictive measures for successful crystallization in a magnetic field using analyses and numerical modeling for use in future protein crystal growth experiments. This will establish criteria that can be used to estimate the efficacy of magnetic field flow damping on crystallization of candidate proteins. 3. To demonstrate the understanding of convection damping by high magnetic fields to a class of proteins that is of interest and whose structure is as yet not determined. 4. To compare quantitatively, the quality of the grown crystals with and without a magnetic field. X-ray diffraction techniques will be used for the comparative studies. In a preliminary set of experiments, we studied crystal dissolution effects in a 5 Tesla magnet available at NASA Marshall Space Flight Center (MSFC). Using a Schlieren setup, a 1mm crystal of Alum (Aluminum-Potassium Sulfate) was introduced in a 75% saturated solution and the resulting dissolution plume was observed. The experiment was conducted both in the presence and absence of a magnetic field gradient. The magnet produces a gradient field of approx. 1 Tesla2/cm. Image analysis of the recorded images indicated an enhanced plume velocity that was of the order of the measurement limit. For this experiment, both the gradient and gravity fields are in the same direction resulting in an enhanced effective gravity that tends to accelerate the observed plume velocity. While the results are not conclusive, pending further tests, it clearly points out the inadequacy of the MSFC magnet for conducting protein crystallization experiments and the need for a stronger magnet. In spacebased experiments, however, where the gravitational effects are small, only a weak magnetic field will be required to control or mitigate the effects of convective contamination.

Ramachandran, Narayanan↗

Free Energy Landscapes for Elucidating the Structural Consequences of Exon-20 mutations on the ErbB Family of Protein Kinases

The ErbB family of protein kinases plays an important role in major cellular functions and consequently mutations in the functional regions of these proteins are implicated in several types of cancer growths. To envision rational design of small molecule drugs that target the diseased proteins it is important to quantify the structural effects of the mutations, as some of these mutants render the protein resistant to tyrosine kinase inhibitors (TKIs). Herein we use advanced sampling techniques and long-timescale molecular dynamics simulations to predict the effect of major exon 20 mutations on the ErbB family, specifically EGFR and HER2 proteins. Exon 20 mutations have been clinically known to induce TKI resistance, though the mechanisms of such an effect is poorly understood. By mapping out the free energy landscape of the mutants and comparing them against the wild-type, we elucidate the structural differences in the binding pocket region that alter the nature of drug-protein interactions. We believe that these insights will play a pivotal role in developing small molecule drugs that overcome the TKI resistance.

Ashwin Ravichandran↗

Structure-Aware Unsupervised, Transformational Machine Learning for Drug Discovery (DTRA Basic Research Final Report)

The major goal of this project is to develop machine learning (ML) methods to enable improved predictive power on real drug discovery for novel targets. More specifically, we planned to demonstrate the capability and effectiveness of ML tools utilizing unlabeled large-volume protein-ligand datasets. We investigated multiple pre-training approaches for 3D protein-ligand structure-based foundation models, without relying on experimental binding data. We also addressed scenarios in which crystal structures are unavailable or binding data are limited. We also planned to develop a complete pipeline to screen novel compounds as well as to demonstrate the capability and effectiveness of the developed methods by testing on a realistic drug discovery task such as SARS-CoV-2. While the major goals and milestones remain consistent with the original proposal, certain technical details have been adjusted, based on the experimental results and related outcomes.

97 MATHEMATICS AND COMPUTING↗