Search NASASearch

SEARCH · Search NASA

Results for “Protein modeling”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 91 records · Page 5

Size effects in models for mechanically-stressed protein crystals and aggregates

As protein aggregates increase in size, they become easier to disrupt mechanically. Using the scaling properties of models proposed to govern protein aggregation, the effect of thermal vibrations and gravity are investigated as deforming forces. For typical protein assemblies made of 30 A proteins, the assembled diameter must remain less than 100-10,000 times the molecular radius to survive in finite thermal and gravity fields. The analysis predicts the following experimental outcomes: (1) reductions in gravitational strain should favor larger protein aggregates; (2) in comparing the aggregate stability of different proteins, the addition of peptide chains should stabilize against thermal strain, but should not affect gravitational strain; (3) critical aggregate sizes should show significant (exponential) sensitivity to cluster geometry, solution preparation and growth conditions. The analysis is extended to consider qualitative size effects in crystal damage during X-ray exposure.

Noever, David A.

Functional protein mining with conformal guarantees

Molecular structure prediction and homology detection offer promising paths to discovering protein function and evolutionary relationships. However, current approaches lack statistical reliability assurances, limiting their practical utility for selecting proteins for further experimental and in-silico characterization. To address this challenge, we introduce a statistically principled approach to protein search leveraging principles from conformal prediction, offering a framework that ensures statistical guarantees with user-specified risk and provides calibrated probabilities (rather than raw ML scores) for any protein search model. Our method (1) lets users select many biologically-relevant loss metrics (i.e. false discovery rate) and assigns reliable functional probabilities for annotating genes of unknown function; (2) achieves state-of-the-art performance in enzyme classification without training new models; and (3) robustly and rapidly pre-filters proteins for computationally intensive structural alignment algorithms. Our framework enhances the reliability of protein homology detection and enables the discovery of uncharacterized proteins with likely desirable functional properties.

59 BASIC BIOLOGICAL SCIENCES

ProtNHF: Neural Hamiltonian Flows for Controllable Protein Sequence Generation

This dataset accompanies the publication "ProtNHF: Neural Hamiltonian Flows for Controllable Protein Sequence Generation". This paper introduces a new AI model for protein sequence generation. This dataset contains data related to experiments discussed in the publication. This includes generated sequences and evaluation metrics supporting all unconditional and bias-controlled experiments in the ProtNHF paper.

60 APPLIED LIFE SCIENCES

Can protein expression be ‘solved’?

Recombinant protein expression is central to biotechnology’s application in academic exploration as well as human health, climate applications and the bioeconomy in general. However, not all proteins can be expressed in all organisms, and the field lacks a predictive model of soluble protein overexpression that could replace laborious experimental trial-and-error. Here, we discuss the state of the field and identify the lack of large, high-fidelity datasets as the primary bottleneck to progress. We review possible assays that could be used for data collection to identify a path toward an extensible experimental platform for collecting soluble recombinant protein overexpression data across organisms. We suggest that the resulting dataset should be used to train increasingly generalizable predictive models of protein expression to answer the question: “How can predictive protein expression be solved?”.

59 BASIC BIOLOGICAL SCIENCES

Packaging “vegetable oils”: Insights into plant lipid droplet proteins

Abstract Plant neutral lipids, also known as “vegetable oils”, are synthesized within the endoplasmic reticulum (ER) membrane and packaged into subcellular compartments called lipid droplets (LDs) for stable storage in the cytoplasm. The biogenesis, modulation, and degradation of cytoplasmic LDs in plant cells are orchestrated by a variety of proteins localized to the ER, LDs, and peroxisomes. Recent studies of these LD-related proteins have greatly advanced our understanding of LDs not only as steady oil depots in seeds but also as dynamic cell organelles involved in numerous physiological processes in different tissues and developmental stages of plants. In the past 2 decades, technology advances in proteomics, transcriptomics, genome sequencing, cellular imaging and protein structural modeling have markedly expanded the inventory of LD-related proteins, provided unprecedented structural and functional insights into the protein machinery modulating LDs in plant cells, and shed new light on the functions of LDs in nonseed plant tissues as well as in unicellular algae. Here, we review critical advances in revealing new LD proteins in various plant tissues, point out structural and mechanistic insights into key proteins in LD biogenesis and dynamic modulation, and discuss future perspectives on bridging our knowledge gaps in plant LD biology.

Cai, Yingqi (ORCID:0000000203575809)

Machine learning-driven descriptions of protein dynamics at solid-liquid interfaces

This chapter has described how ML has enabled quantitative analysis of HS-AFM data to discover the physical phenomena governing protein dynamics and ordering at solid-liquid interfaces. The research detailed in this chapter modeled the rotation models of protein nanorods, the discovery of which would otherwise not be possible. By tracking the trajectories of individual protein rods from frame to frame, it was possible to model Brownian type motion and behaviors and Levy-flight dynamics that had not previously been shown. We also described the application of the Python package AtomAI, which has been developed specifically to analyze and extract physical phenomena, providing exemplar code for training an ensemble of deep neural networks to produce the semantic segmentation of AFM data and functions for encoding and decoding local environments. We last described a combinatorial approach to analyze very noisy data with a densely covered substrate where the emergence of order for the protein liquid crystals could be elucidated. By combining the methods from Case 1 and 2, it was possible to obtain the center of mass and angle for each rod in the images and track the assembly of the rods over time into a 2D liquid crystal array on the surface of mica.

protein dynamics, solid-liquid interfaces, atomic

Artificial intelligence tools for enzyme engineering and metabolic engineering

Enzyme engineering and metabolic engineering drive innovation in energy biotechnology. In recent years, artificial intelligence (AI) has supported successful applications in designing effective enzymes and productive microbial cell factories. This review summarizes recent advances in enzyme redesign using protein language models, de novo enzyme design with generative models, and AI tools for engineering metabolism and related cellular phenotypes. Across these areas, AI models are shifting from single modality inputs to integrated representations of protein function, metabolic pathways, and cell states. We emphasize that unifying the diverse data representations across scales will be necessary for advancements in energy biotechnology.

Volk, Michael [Univ. of Illinois at Urbana-Champai

Activation Domain Hunter (ADhunter) v2.0

ADhunter is a software program that enables accurate identification and quantification of transcriptional activation domains. Unlike previous software, ADhunter uses protein representations from a pre-trained protein language model, model ensembling, and a training dataset from a diverse sampling of protein sequence space for state-of-the-art performance. These advantages enable improved perception of transcriptional activation domains across sequence space that can be used for mapping natural genetic circuits and engineering synthetic genetic circuits. In particular, ADhunter enables fine-tuned control of gene expression through synthetic transcription factors that can be used for complex control of cellular programs.

Waldburger, Lucas [Lawrence Berkeley National Labo

A characterization of recombinant Arabidopsis FRIABLE1 (FRB1) reveals robust rhamnogalacturonan-I rhamnosyltransferase activity and critical catalytic residues

Plant cell walls are glycan-rich extracellular matrices that fundamentally impact essential cellular processes, such as growth, adhesion, and cell shape acquisition. Understanding plant cell wall glycans requires the identification and characterization of the biosynthetic enzymes that produce these polymers. Most successful in vitro protein expression studies of plant cell wall glycosyltransferases have relied on insect, fungal/yeast, or human cell expression systems, whereas prokaryotic expression systems have been generally unsuccessful. Here, we show that Arabidopsis FRIABLE1 (FRB1)/rhamnogalacturonan-I rhamnosyltransferase 8 (RRT8) can be produced in Escherichia coli RosettaGami2 cells as N-terminal maltose-binding protein fusion proteins containing C-terminal 6X-His-tags. We also report the catalytic constants of FRB1/RRT8 with apparent K M and K cat values of 226 μM and 33 min -1 for UDP-Rhamnose and 117 μM and 28.7 min -1 for rhamnogalacturonan-I (RG-I), respectively. We examine the catalytic activities of mutated FRB1/RRT8 proteins based on an AlphaFold 3-generated FRB1/RRT8 protein structural model with a virtually docked UDP-Rha donor. Enzymatic characterization of the mutated and wildtype FRB1/RRT8 protein confirmed that mutation of predicted catalytic site amino acid residues resulted in a 20-fold reduction in RRT activity. FRB1 also robustly polymerizes RG-I in combination with RG-I galacturonosyltransferase 1. These results show how a robust E. coli expression system combined with artificial intelligence tools can be used to increase understanding of plant cell wall glycosyltransferase structure and function.

glycosyltransferase

Nucleation and Growth According to Lysozyme

How does one take a molecule, strongly asymmetric in both shape and charge distribution, and assemble it into a crystal? We propose a model for the nucleation and crystal growth process for tetragonal lysozyme that may be very germane to other monomeric proteins. The first species formed is postulated to be a dimer. Through repeating associations involving the same intermolecular interactions this becomes the 4(sub 3) helix, that in turn serves as the basic unit for nucleation and crystal growth. High salt attenuates surface charges while promoting hydrophobic interactions. Symmetry facilitates helix self-association. Assembly stability is enhanced when a four helix structure is obtained, with each bound to two neighbors. Only two unique interactions are required. The first are those for helix formation, where the dominant interaction is the intermolecular bridging anion. The second is the anti-parallel side-by-side helix-helix interaction, guided by alternating pairs of symmetry related salt bridges along each side. At this stage all eight unique positions of the P4(sub 3)2(sub 1)2(sub 1) unit cell are filled. From the above, the process is one of a) attenuating the most strongly interacting groups, such that b) the molecules begin to self-associate in defined patterns, so that c) symmetry is obtained, which d) propagates as a growing crystal. Simple and conceptually obvious in hindsight, this tells much about what we are empirically doing when we crystallize macromolecules. By adjusting the solution parameters we are empirically balancing the intermolecular interactions, preferentially attenuating the dominant strong (for lysozyme the charged groups) while strengthening the lesser strong (hydrophobic) interactions. Lysozyme is atypical in the breadth of its crystallization conditions; many proteins only crystallize under narrowly defined conditions, pointing to the criticality of the empirical balancing process. Lack of a singularly defined association pathway leads to formation of multiple species, i.e., amorphous precipitation. Weak interactions, such as hydrogen bonds, are promiscuous, serving to strengthen rather than define specific interactions. Participation in an interaction sequesters that surface from subsequent interactions, and we expect the strongest bonds to form first. When two molecules self associate the resulting species will have an axis of symmetry. Subsequent interactions between two associated species having equivalent interactions will also have symmetry. Only a few unique sets of interactions are required to give any of the commonly found space groups for monomeric proteins. This model and what it suggests will be discussed.

Pusey, Marc L.

Convective diffusion in protein crystal growth

A protein crystal modeled as a flat plate suspended in the parent solution, with the normal to the largest face perpendicular to gravity and the protein concentration in the solution adjacent to the plate taken to be the equilibrium solubility, is studied. The Navier-Stokes equation and the equation for convective diffusion in the boundary layer next to the plate are solved to calculate the flow velocity and the protein mass flux. The local rate of growth of the plate is shown to vary significantly with depth due to the convection. For an aqueous solution of lysozyme at a concentration of 40 mg/ml, the boundary layer at the top of a 1-mm-high crystal has a thickness of 80 microns at 1 g, and 2570 microns at 10 to the -6th g.

Baird, J. K.

Mechanistic insights into a heterobifunctional degrader-induced PTPN2/N1 complex

PTPN2 (protein tyrosine phosphatase non-receptor type 2, or TC-PTP) and PTPN1 are attractive immuno-oncology targets, with the deletion of Ptpn1 and Ptpn2 improving response to immunotherapy in disease models. Targeted protein degradation has emerged as a promising approach to drug challenging targets including phosphatases. We developed potent PTPN2/N1 dual heterobifunctional degraders (Cmpd-1 and Cmpd-2) which facilitate efficient complex assembly with E3 ubiquitin ligase CRL4 CRBN , and mediate potent PTPN2/N1 degradation in cells and mice. To provide mechanistic insights into the cooperative complex formation introduced by degraders, we employed a combination of structural approaches. Our crystal structure reveals how PTPN2 is recognized by the tri-substituted thiophene moiety of the degrader. We further determined a high-resolution structure of DDB1-CRBN/Cmpd-1/PTPN2 using single-particle cryo-electron microscopy (cryo-EM). This structure reveals that the degrader induces proximity between CRBN and PTPN2, albeit the large conformational heterogeneity of this ternary complex. The molecular dynamic (MD)-simulations constructed based on the cryo-EM structure exhibited a large rigid body movement of PTPN2 and illustrated the dynamic interactions between PTPN2 and CRBN. Together, our study demonstrates the development of PTPN2/N1 heterobifunctional degraders with potential applications in cancer immunotherapy. Furthermore, the developed structural workflow could help to understand the dynamic nature of degrader-induced cooperative ternary complexes.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH

Concerning the production of free radicals in proteins by ultraviolet light.

The response to UV light of several solid proteins and model compounds has been studied in vacuum and at low temperature, using electron paramagnetic resonance techniques. The results indicate that the details of amino acid composition and sequence, and the tertiary structure of a protein are important in determining both the rate of, and the mechanism for, the production of free radicals, and in determining the conditions under which sulfur-type radicals can be produced. The results presented are related to enzyme inactivation and to the UV stability of proteins generally.

Androes, G. M.

A foundation model for atomistic materials chemistry

Atomistic simulations of matter, especially those that leverage first-principles (ab initio) electronic structure theory, provide a microscopic view of the world, underpinning much of our understanding of chemistry and materials science. Over the last decade or so, machine-learned force fields have transformed atomistic modeling by enabling simulations of ab initio quality over unprecedented time and length scales. However, early machine-learning (ML) force fields have largely been limited by (i) the substantial computational and human effort required to develop and validate potentials for each particular system of interest and (ii) a general lack of transferability from one chemical system to the next. Here, we show that it is possible to create a general-purpose atomistic ML model, trained on a public dataset of moderate size, that is capable of running stable molecular dynamics for a wide range of molecules and materials. We demonstrate the power of the MACE-MP-0 model-and its qualitative and at times quantitative accuracy-on a diverse set of problems in the physical sciences, including properties of solids, liquids, gases, chemical reactions, interfaces, and even the dynamics of a small protein. The model can be applied out of the box as a starting or "foundation" model for any atomistic system of interest and, when desired, can be fine-tuned on just a handful of application-specific data points to reach ab initio accuracy. Establishing that a stable force-field model can cover almost all materials changes atomistic modeling in a fundamental way: experienced users obtain reliable results much faster, and beginners face a lower barrier to entry. Foundation models thus represent a step toward democratizing the revolution in atomic-scale modeling that has been brought about by ML force fields.

Batatia, Ilyes

A large iris-like expansion of a mechanosensitive channel protein induced by membrane tension

MscL, a bacterial mechanosensitive channel of large conductance, is the first structurally characterized mechanosensor protein. Molecular models of its gating mechanisms are tested here. Disulfide crosslinking shows that M1 transmembrane alpha-helices in MscL of resting Escherichia coli are arranged similarly to those in the crystal structure of MscL from Mycobacterium tuberculosis. An expanded conformation was trapped in osmotically shocked cells by the specific bridging between Cys 20 and Cys 36 of adjacent M1 helices. These bridges stabilized the open channel. Disulfide bonds engineered between the M1 and M2 helices of adjacent subunits (Cys 32-Cys 81) do not prevent channel gating. These findings support gating models in which interactions between M1 and M2 of adjacent subunits remain unaltered while their tilts simultaneously increase. The MscL barrel, therefore, undergoes a large concerted iris-like expansion and flattening when perturbed by membrane tension.

NASA Discipline Cell Biology

Detection of non‐native species formed during fibrillization of the myocilin olfactomedin domain

Abstract Glaucoma is a group of neurodegenerative diseases that together are the leading cause of irreversible blindness worldwide. Myocilin‐associated glaucoma is an inherited form of this disease, caused by intracellular aggregation of misfolded mutant myocilin. In vitro, the myocilin C‐terminal olfactomedin domain (OLF), the relevant domain for glaucoma pathogenesis, can be driven to form amyloid‐like fibrils under mild conditions. Here we characterize a species present during in vitro fibrillization. Purified OLF was subjected to fibrillization at concentrations required for downstream electron microscopy imaging and NMR spectroscopy. Additional biophysical techniques, including analytical ultracentrifugation and X‐ray crystallography, were employed to further characterize the multicomponent mixture. Negative stain transmission electron microscopy (TEM) shows a non‐native species reminiscent of known prefibrillar oligomers from other amyloid systems, NMR indicates a minor population of partially misfolded species is present in solution, and cryo‐EM imaging shows two‐dimensional protein arrays. The predominant soluble species remaining in solution after the fibril reaction is natively folded, as evidenced by X‐ray crystallography. In summary, after incubating OLF under fibrillization‐promoting conditions, there is a heterogeneous mixture consisting of soluble folded protein, mature amyloid‐like fibrils, and partially misfolded intermediate species that at present belie additional molecular detail. The characterization of OLF fibrillar species illustrates the challenges associated with developing a comprehensive understanding of the fibrillization process for large, non‐model amyloidogenic proteins.

Scelsi, Hailee F. [School of Chemistry and Biochem

Peripheral positions encode transport specificity in the small multidrug resistance exporters

In secondary active transporters, a relatively limited set of protein folds have evolved diverse solute transport functions. Because of the conformational changes inherent to transport, altering substrate specificity typically involves remodeling the entire structural landscape, limiting our understanding of how novel substrate specificities evolve. In the current work, we examine a structurally minimalist family of model transport proteins, the small multidrug resistance (SMR) transporters, to understand the molecular basis for the emergence of a novel substrate specificity. We engineer a selective SMR protein to promiscuously export quaternary ammonium antiseptics, similar to the activity of a clade of multidrug exporters in this family. Using combinatorial mutagenesis and deep sequencing, we identify the necessary and sufficient molecular determinants of this engineered activity. Using X-ray crystallography, solid-supported membrane electrophysiology, binding assays, and a proteoliposome-based quaternary ammonium antiseptic transport assay that we developed, we dissect the mechanistic contributions of these residues to substrate polyspecificity. We find that substrate preference changes not through modification of the residues that directly interact with the substrate but through mutations peripheral to the binding pocket. Our work provides molecular insight into substrate promiscuity among the SMRs and can be applied to understand multidrug export and the evolution of novel transport functions more generally.

Science & Technology - Other Topics