Search NASASearch

SEARCH · Search NASA

Results for “protein”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

Modeling Protein–Protein and Protein–Ligand Interactions by the ClusPro Team in CASP16

ABSTRACT In the CASP16 experiment, our team employed hybrid computational strategies to predict both protein–protein and protein–ligand complex structures. For protein–protein docking, we combined physics‐based sampling—using ClusPro FFT docking and molecular dynamics—with AlphaFold (AF)‐based sampling, followed by AF‐based refinement. Our method produced numerous high‐accuracy complex models, including cases where AF alone failed, underscoring the critical role of physics‐based sampling alongside deep learning‐based refinement. For protein–ligand docking, we integrated the ClusPro LigTBM template‐based approach with a machine learning‐based confidence model for rescoring. The method preserves conserved interaction fragments derived from homologous complexes, followed by local resampling using physics‐based sampling and a diffusion model. Our template‐based strategy achieved a mean lDDT‐PLI of 0.69 across 233 targets, which was highly competitive. These results demonstrate that combining physics‐based modeling with AI‐driven refinement can significantly enhance the accuracy of both protein–protein and protein–ligand structure predictions.

Ashizawa, Ryota [Department of Applied Mathematics

Sequence and structural implications of a bovine corneal keratan sulfate proteoglycan core protein. Protein 37B represents bovine lumican and proteins 37A and 25 are unique

Amino acid sequence from tryptic peptides of three different bovine corneal keratan sulfate proteoglycan (KSPG) core proteins (designated 37A, 37B, and 25) showed similarities to the sequence of a chicken KSPG core protein lumican. Bovine lumican cDNA was isolated from a bovine corneal expression library by screening with chicken lumican cDNA. The bovine cDNA codes for a 342-amino acid protein, M(r) 38,712, containing amino acid sequences identified in the 37B KSPG core protein. The bovine lumican is 68% identical to chicken lumican, with an 83% identity excluding the N-terminal 40 amino acids. Location of 6 cysteine and 4 consensus N-glycosylation sites in the bovine sequence were identical to those in chicken lumican. Bovine lumican had about 50% identity to bovine fibromodulin and 20% identity to bovine decorin and biglycan. About two-thirds of the lumican protein consists of a series of 10 amino acid leucine-rich repeats that occur in regions of calculated high beta-hydrophobic moment, suggesting that the leucine-rich repeats contribute to beta-sheet formation in these proteins. Sequences obtained from 37A and 25 core proteins were absent in bovine lumican, thus predicting a unique primary structure and separate mRNA for each of the three bovine KSPG core proteins.

NASA Discipline Cell Biology

African Swine Fever Virus Protein–Protein Interaction Prediction

The African swine fever virus (ASFV) is an often deadly disease in swine and poses a threat to swine livestock and swine producers. With its complex genome containing more than 150 coding regions, developing effective vaccines for this virus remains a challenge due to a lack of basic knowledge about viral protein function and protein–protein interactions between viral proteins and between viral and host proteins. In this work, we identified ASFV-ASFV protein–protein interactions (PPIs) using artificial intelligence-powered protein structure prediction tools. We benchmarked our PPI identification workflow on the Vaccinia virus, a widely studied nucleocytoplasmic large DNA virus, and found that it could identify gold-standard PPIs that have been validated in vitro in a genome-wide computational screening. We applied this workflow to more than 18,000 pairwise combinations of ASFV proteins and were able to identify seventeen novel PPIs, many of which have corroborating experimental or bioinformatic evidence for their protein–protein interactions, further validating their relevance. Two protein–protein interactions, I267L and I8L, I267L__I8L, and B175L and DP79L, B175L__DP79L, are novel PPIs involving viral proteins known to modulate host immune response.

59 BASIC BIOLOGICAL SCIENCES

Engineering a new tripartite split-ccGFP system from Corynactis californica for detecting protein–protein interactions

Protein-protein interactions (PPIs) are critical to a range of biological processes and, consequently, aberrant interactions are implicated in many disorders. The study of the complex networks of PPIs promises to elucidate undiscovered roles in cellular processes and the mechanisms of disease. To accomplish this, tools to effectively sense PPIs are necessary. Effective PPI sensors must rapidly detect interactions in real-time with high sensitivity without perturbing the proteins of interest (POIs) under study. Split fluorescent proteins have previously been used to successfully monitor PPIs, in part due to the small size of the tags. Here, we developed an optimized tripartite split GFP system based on Corynactis californica GFP (ccGFP) to detect PPIs in vitro. In this sensor system, ccGFP fragments ccGFP10 and ccGFP11 are tagged to two POIs. PPIs can then be detected via fluorescence by complementation to the third fragment, ccGFP1-9, which reconstitutes functional ccGFP. The optimized ccGFP system shows improved detection kinetics and pH and temperature stability compared to a previous system. We then validated the sensor by monitoring PPIs in two model systems: attractive/repulsive coiled-coils and rapamycin-inducible FRB/FKBP heterodimerization. Finally, we developed an anti-tripartite ccGFP single-chain variable fragment (scFv), which could enable versatile detection of identified protein-protein complexes.

59 BASIC BIOLOGICAL SCIENCES

The impact of curation errors in the PDBBind Database on machine learning predictions of protein–protein binding affinity

The PDBBind database has been widely utilized for the computational prediction of protein–protein binding affinities. While the accuracy of the PDBBind-curated equilibrium dissociation constants (K D ) has been reported for the protein–ligand subset of the PDBBind database, the curation accuracy has not been reported for the protein–protein subset. Here, we present a detailed manual analysis for the subset of PDBBind records with PubMed Central Open Access primary publications and find that ~19% of these records had K D values that were not supported by their primary publications. The impact of these putative curation errors on the machine learning-based prediction of K D from experimental protein–protein 3D structures was evaluated and correcting the curation errors improved the Pearson correlation coefficient between measured and random forest-predicted log 10 (K D ) values by ~8 percentage points. This finding underscores the importance of dataset accuracy for computational modelling and highlights the need for more stringent curation processes when extracting information from the scientific literature.

59 BASIC BIOLOGICAL SCIENCES

An Arabidopsis Ran-binding protein, AtRanBP1c, is a co-activator of Ran GTPase-activating protein and requires the C-terminus for its cytoplasmic localization

Ran-binding proteins (RanBPs) are a group of proteins that bind to Ran (Ras-related nuclear small GTP-binding protein), and thus either control the GTP/GDP-bound states of Ran or help couple the Ran GTPase cycle to a cellular process. AtRanBP1c is a Ran-binding protein from Arabidopsis thaliana (L.) Heynh. that was recently shown to be critically involved in the regulation of auxin-induced mitotic progression [S.-H. Kim et al. (2001) Plant Cell 13:2619-2630]. Here we report that AtRanBP1c inhibits the EDTA-induced release of GTP from Ran and serves as a co-activator of Ran-GTPase-activating protein (RanGAP) in vitro. Transient expression of AtRanBP1c fused to a beta-glucuronidase (GUS) reporter reveals that the protein localizes primarily to the cytosol. Neither the N- nor C-terminus of AtRanBP1c, which flank the Ran-binding domain (RanBD), is necessary for the binding of PsRan1-GTP to the protein, but both are needed for the cytosolic localization of GUS-fused AtRanBP1c. These findings, together with a previous report that AtRanBP1c is critically involved in root growth and development, imply that the promotion of GTP hydrolysis by the Ran/RanGAP/AtRanBP1c complex in the cytoplasm, and the resulting concentration gradient of Ran-GDP to Ran-GTP across the nuclear membrane could be important in the regulation of auxin-induced mitotic progression in root tips of A. thaliana.

NASA Discipline Plant Biology

Suppression of muscle protein turnover and amino acid degradation by dietary protein deficiency

To define the adaptations that conserve amino acids and muscle protein when dietary protein intake is inadequate, rats (60-70 g final wt) were fed a normal or protein-deficient (PD) diet (18 or 1% lactalbumin), and their muscles were studied in vitro. After 7 days on the PD diet, both protein degradation and synthesis fell 30-40% in skeletal muscles and atria. This fall in proteolysis did not result from reduced amino acid supply to the muscle and preceded any clear decrease in plasma amino acids. Oxidation of branched-chain amino acids, glutamine and alanine synthesis, and uptake of alpha-aminoisobutyrate also fell by 30-50% in muscles and adipose tissue of PD rats. After 1 day on the PD diet, muscle protein synthesis and amino acid uptake decreased by 25-40%, and after 3 days proteolysis and leucine oxidation fell 30-45%. Upon refeeding with the normal diet, protein synthesis also rose more rapidly (+30% by 1 day) than proteolysis, which increased significantly after 3 days (+60%). These different time courses suggest distinct endocrine signals for these responses. The high rate of protein synthesis and low rate of proteolysis during the first 3 days of refeeding a normal diet to PD rats contributes to the rapid weight gain ("catch-up growth") of such animals.

Non-NASA Center

Ca X ML: Chemistry‐informed machine learning explains mutual changes between protein conformations and calcium ions in calcium‐binding proteins using structural and topological features

Proteins' flexibility is a feature in communicating changes in cell signaling instigated by binding with secondary messengers, such as calcium ions, associated with the coordination of muscle contraction, neurotransmitter release, and gene expression. When binding with the disordered parts of a protein, calcium ions must balance their charge states with the shape of calcium-binding proteins and their versatile pool of partners depending on the circumstances they transmit. Accurately determining the ionic charges of those ions is essential for understanding their role in such processes. However, it is unclear whether the limited experimental data available can be effectively used to train models to accurately predict the charges of calcium-binding protein variants. Here, we developed a chemistry-informed, machine-learning algorithm that implements a game theoretic approach to explain the output of a machine-learning model without the prerequisite of an excessively large database for high-performance prediction of atomic charges. We used the ab initio electronic structure data representing calcium ions and the structures of the disordered segments of calcium-binding peptides with surrounding water molecules to train several explainable models. Network theory was used to extract the topological features of atomic interactions in the structurally complex data dictated by the coordination chemistry of a calcium ion, a potent indicator of its charge state in protein. Our design created a computational tool of Ca X ML, which provided a framework of explainable machine learning model to annotate ionic charges of calcium ions in calcium-binding proteins in response to the chemical changes in an environment. Our framework will provide new insights into protein design for engineering functionality based on the limited size of scientific data in a genome space.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH

Predicting compatibility between ferredoxins and the Fe protein of nitrogenase using in silico protein modeling

Biological nitrogen fixation is the process by which certain bacteria and archaea use the enzyme nitrogenase to reduce atmospheric nitrogen into bioavailable ammonium. Engineering non‐nitrogen‐fixing organisms, like plants, to use nitrogenase could reduce dependency on synthetic fertilizer and mitigate the environmental impacts of industrial fertilizer production. However, nitrogenase activity requires delivery of reducing power by small electron carrying proteins known as ferredoxins and flavodoxins, and successfully engineering nitrogenase into new systems will require a mechanistic understanding of electron delivery by these proteins. Most organisms often have multiple ferredoxins, raising the question of which ferredoxin can support nitrogenase activity. The purpose of this study is to gain insight into how we can predict which ferredoxin is compatible with the Fe protein, the component of nitrogenase that interacts with ferredoxin or flavodoxin. Our in silico protein–protein docking simulations reveal that most ferredoxins and flavodoxins involved in nitrogen fixation have the shortest distance (≤10 Å) between their redox cofactor and the [4Fe‐4S] cluster of the Fe protein. We found shorter cofactor distance contributes to faster intermolecular electron tunneling rates. Bacterial ferredoxins that play a role in nitrogen fixation also exhibit more complementary interactions with the Fe protein than bacterial and plant ferredoxins not involved in this process. Heterologous expression of a set of ferredoxins from both nitrogen‐fixing and non‐nitrogen‐fixing bacteria in the diazotroph Rhodopseudomonas palustris supports our model‐derived prediction that shorter distances between the electron‐carrying cofactors favor nitrogenase compatibility. These findings offer a framework to predict and potentially enhance ferredoxin–nitrogenase compatibility, which will help to improve our ability to engineer nitrogen fixation into non‐nitrogen‐fixing organisms like plants.

59 BASIC BIOLOGICAL SCIENCES

Data for A Generalized Platform for Artificial Intelligence-powered Autonomous Protein Engineering

Proteins are the molecular machines of life with numerous applications in energy, health, and sustainability. However, engineering proteins with desired functions for practical applications remains slow, expensive, and specialist-dependent. Here we report a generally applicable platform for autonomous enzyme engineering that integrates machine learning and large language models with biofoundry automation to eliminate the need for human intervention, judgement, and domain expertise. Requiring only an input protein sequence and a quantifiable way to measure fitness, this automated platform can be applied to engineer a wide array of proteins. As a proof of concept, we engineer Arabidopsis thaliana halide methyltransferase (AtHMT) for a 90-foldimprovement in substrate preference and 16-fold improvement in ethyl-transferase activity, along with developing a Yersinia mollaretii phytase (YmPhytase) variant with 26-fold improvement in activity at neutral pH. This is accomplished in four rounds over 4 weeks, while requiring construction and characterization of fewer than 500 variants for each enzyme. This platform for autonomous experimentation paves the way for rapid advancements across diverse industries, from medicine and biotechnology to renewable energy and sustainable chemistry.

AI/ML

Photoactivable analogs for labeling 25-hydroxyvitamin D3 serum binding protein and for 1,25-dihydroxyvitamin D3 intestinal receptor protein

3-Azidobenzoates and 3-azidonitrobenzoates of 25-hydroxyvitamin D3 as well as 3-deoxy-3-azido-25-hydroxyvitamin D3 and 3-deoxy-3-azido-1,25-dihydroxyvitamin D3 were prepared as photoaffinity labels for vitamin D serum binding protein and 1,25-dihydroxyvitamin D3 intestinal receptor protein. The compounds prepared were easily activated by short- or long-wavelength uv light, as monitored by uv and ir spectrometry. The efficacy of the compounds to compete with 25-hydroxyvitamin D3 or 1,25-dihydroxyvitamin D3 for the binding site of serum binding protein and receptor, respectively, was studied to evaluate the vitamin D label with the highest affinity for the protein. The presence of an azidobenzoate or azidonitrobenzoate substituent at the C-3 position of 25-OH-D3 significantly decreased (10(4)- to 10(6)-fold) the binding activity. However, the labels containing the azido substituent attached directly to the vitamin D skeleton at the C-3 position showed a high affinity, only 20- to 150-fold lower than that of the parent compounds with their respective proteins. Therefore, 3-deoxy-3-azidovitamins present potential ligands for photolabeling of vitamin D proteins and for studying the structures of the protein active sites.

NASA Discipline Musculoskeletal

The 'tubulin-like' S1 protein of Spirochaeta is a member of the hsp65 stress protein family

A 65-kDa protein (called S1) from Spirochaeta bajacaliforniensis was identified as 'tubulin-like' because it cross-reacted with at least four different antisera raised against tubulin and was isolated, with a co-polymerizing 45-kDa protein, by warm-cold cycling procedures used to purify tubulin from mammalian brain. Furthermore, at least three genera of non-cultivable symbiotic spirochetes (Pillotina, Diplocalyx, and Hollandina) that contain conspicuous 24-nm cytoplasmic tubules displayed a strong fluorescence in situ when treated with polyclonal antisera raised against tubulin. Here we summarize results that lead to the conclusion that this 65-kDa protein has no homology to tubulin. S1 is an hsp65 stress protein homologue. Hsp65 is a highly immunogenic family of hsp60 proteins which includes the 65-kDa antigens of Mycobacterium tuberculosis (an active component of Freund's complete adjuvant), Borrelia, Treponema, Chlamydia, Legionella, and Salmonella. The hsp60s, also known as chaperonins, include E. coli GroEL, mitochondrial and chloroplast chaperonins, the pea aphid 'symbionin' and many other proteins involved in protein folding and the stress response.

NASA Discipline Exobiology

An Integral Activity-Based Protein Profiling Method for Higher Throughput Determination of Protein Target Sensitivity to Small Molecules

Activity-based protein profiling (ABPP) is a chemoproteomic technique that uses small molecule probes to label active enzymes selectively and covalently in complex proteomes. Competitive ABPP, which involves treatment of the active proteome with an analyte of interest, is especially powerful for profiling how small molecules impact specific protein activities. Advances in higher throughput workflows have made it possible to generate extensive competitive ABPP data across diverse biological samples, making this approach highly appealing for characterizing shared and unique proteins affected by perturbations such as drug or chemical exposures. To use the competitive ABPP approach effectively to understand potential adverse effects of chemicals of concern (CoC), a wide range of concentrations may be needed, particularly for chemicals that lack potency or toxicity data. In this work, we present an integral competitive ABPP method that enables target sensitivity determination for different organophosphate (OP) pesticides as model toxicants. Using previously developed OP-ABPs, we optimized conditions for tandem mass tag (TMT) multiplexing of ABPP samples and compared conventional competitive ABPP involving samples at discrete paraoxon concentrations to pooled samples across that same concentration range. We then expanded our approach to compare protein target sensitivities toward two additional OP pesticides, chlorpyrifos oxon and malaoxon. The results showed that differences in integral intensities for the pooled competition sample can be used to evaluate the relative sensitivity of specific proteins without increasing the overall number of samples. For 8 CoC concentrations of interest, this strategy reduced the number of TMT plexes and the corresponding number of LC–MS/MS analyses 3-fold. In conclusion, we envision the integral ABPP (IABPP) method will provide a means to screen diverse chemicals more rapidly to identify both high and low sensitivity protein targets.

activity-based probes

Casein kinase II protein kinase is bound to lamina-matrix and phosphorylates lamin-like protein in isolated pea nuclei

A casein kinase II (CK II)-like protein kinase was identified and partially isolated from a purified envelope-matrix fraction of pea (Pisum sativum L.) nuclei. When [gamma-32P]ATP was directly added to the envelope-matrix preparation, the three most heavily labeled protein bands had molecular masses near 71, 48, and 46 kDa. Protein kinases were removed from the preparation by sequential extraction with Triton X-100, EGTA, 0.3 M NaCl, and a pH 10.5 buffer, but an active kinase still remained bound to the remaining lamina-matrix fraction after these treatments. This kinase had properties resembling CK II kinases previously characterized from animal and plant sources: it preferred casein as an artificial substrate, could use GTP as efficiently as ATP as the phosphoryl donor, was stimulated by spermine, was calcium independent, and had a catalytic subunit of 36 kDa. Some animal and plant CK II kinases have regulatory subunits near 29 kDa, and a lamina-matrix-bound protein of this molecular mass was recognized on immunoblot by anti-Drosophila CK II polyclonal antibodies. Also found associated with the envelope-matrix fraction of pea nuclei were p34cdc2-like and Ca(2+)-dependent protein kinases, but their properties could not account for the protein kinase activity bound to the lamina. The 71-kDa substrate of the CK II-like kinase was lamin A-like, both in its molecular mass and in its cross-reactivity with anti-intermediate filament antibodies. Lamin phosphorylation is considered a crucial early step in the entry of cells into mitosis, so lamina-bound CK II kinases may be important control points for cellular proliferation.

NASA Discipline Plant Biology

Purification method for recombinant proteins based on a fusion between the target protein and the C-terminus of calmodulin

Calmodulin (CaM) was used as an affinity tail to facilitate the purification of the green fluorescent protein (GFP), which was used as a model target protein. The protein GFP was fused to the C-terminus of CaM, and a factor Xa cleavage site was introduced between the two proteins. A CaM-GFP fusion protein was expressed in E. coli and purified on a phenothiazine-derivatized silica column. CaM binds to the phenothiazine on the column in a Ca(2+)-dependent fashion and it was, therefore, used as an affinity tail for the purification of GFP. The fusion protein bound to the affinity column was then subjected to a proteolytic digestion with factor Xa. Pure GFP was eluted with a Ca(2+)-containing buffer, while CaM was eluted later with a buffer containing the Ca(2+)-chelating agent EGTA. The purity of the isolated GFP was verified by SDS-PAGE, and the fluorescence properties of the purified GFP were characterized.

Non-NASA Center

Challenges in predicting protein-protein interactions of understudied viruses: Arenavirus-human interactions

Understanding protein-protein interactions (PPIs) between viruses and host organisms is crucial for uncovering infection mechanisms and identifying potential therapeutic targets. The ability to generalize PPI predictive models across understudied viruses presents a significant challenge. In this work, we use arenavirus-human PPIs to illustrate the difficulties associated with model generalization, which are compounded by a lack of both positive and negative data. We employ a Transfer Learning approach to investigate arenavirus-human PPIs by utilizing models trained on better-studied virus-human and human-human PPIs. Additionally, we curate and assess four types of negative sampling datasets to evaluate their impact on model performance. Despite the overall high accuracies (93–99 %) and AUPRC scores (0.8–0.9) appearing promising, further analysis indicates that these performance metrics can be misleading due to data leakage, data bias, and overfitting, especially concerning under-represented viral proteins. We reveal these gaps and assess the impact of data imbalance using standard k-fold cross-validation and Independent Blind Testing with a Balanced Dataset, resulting in a drop in accuracy below 50 %. We propose a viral protein-specific evaluation framework that categorizes viral proteins into majority and minority classes based on their representation in the dataset, enabling comparison of model performance across these groups using balanced accuracies. This framework offers a more robust evaluation of model generalizability, addressing biases inherent in standard evaluation techniques and paving the way for more reliable PPI prediction models for understudied viruses.

59 BASIC BIOLOGICAL SCIENCES

Predicting protein functions from redundancies in large-scale protein interaction networks

Interpreting data from large-scale protein interaction experiments has been a challenging task because of the widespread presence of random false positives. Here, we present a network-based statistical algorithm that overcomes this difficulty and allows us to derive functions of unannotated proteins from large-scale interaction data. Our algorithm uses the insight that if two proteins share significantly larger number of common interaction partners than random, they have close functional associations. Analysis of publicly available data from Saccharomyces cerevisiae reveals >2,800 reliable functional associations, 29% of which involve at least one unannotated protein. By further analyzing these associations, we derive tentative functions for 81 unannotated proteins with high certainty. Our method is not overly sensitive to the false positives present in the data. Even after adding 50% randomly generated interactions to the measured data set, we are able to recover almost all (approximately 89%) of the original associations.

Proteins/chemistry/metabolism