Search NASA⌕ Search

SEARCH · Search NASA

Results for “Protein modeling”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 55 records · Page 3

Packaging “vegetable oils”: Insights into plant lipid droplet proteins

Abstract Plant neutral lipids, also known as “vegetable oils”, are synthesized within the endoplasmic reticulum (ER) membrane and packaged into subcellular compartments called lipid droplets (LDs) for stable storage in the cytoplasm. The biogenesis, modulation, and degradation of cytoplasmic LDs in plant cells are orchestrated by a variety of proteins localized to the ER, LDs, and peroxisomes. Recent studies of these LD-related proteins have greatly advanced our understanding of LDs not only as steady oil depots in seeds but also as dynamic cell organelles involved in numerous physiological processes in different tissues and developmental stages of plants. In the past 2 decades, technology advances in proteomics, transcriptomics, genome sequencing, cellular imaging and protein structural modeling have markedly expanded the inventory of LD-related proteins, provided unprecedented structural and functional insights into the protein machinery modulating LDs in plant cells, and shed new light on the functions of LDs in nonseed plant tissues as well as in unicellular algae. Here, we review critical advances in revealing new LD proteins in various plant tissues, point out structural and mechanistic insights into key proteins in LD biogenesis and dynamic modulation, and discuss future perspectives on bridging our knowledge gaps in plant LD biology.

Cai, Yingqi (ORCID:0000000203575809)↗

Machine learning-driven descriptions of protein dynamics at solid-liquid interfaces

This chapter has described how ML has enabled quantitative analysis of HS-AFM data to discover the physical phenomena governing protein dynamics and ordering at solid-liquid interfaces. The research detailed in this chapter modeled the rotation models of protein nanorods, the discovery of which would otherwise not be possible. By tracking the trajectories of individual protein rods from frame to frame, it was possible to model Brownian type motion and behaviors and Levy-flight dynamics that had not previously been shown. We also described the application of the Python package AtomAI, which has been developed specifically to analyze and extract physical phenomena, providing exemplar code for training an ensemble of deep neural networks to produce the semantic segmentation of AFM data and functions for encoding and decoding local environments. We last described a combinatorial approach to analyze very noisy data with a densely covered substrate where the emergence of order for the protein liquid crystals could be elucidated. By combining the methods from Case 1 and 2, it was possible to obtain the center of mass and angle for each rod in the images and track the assembly of the rods over time into a 2D liquid crystal array on the surface of mica.

protein dynamics, solid-liquid interfaces, atomic ↗

Artificial intelligence tools for enzyme engineering and metabolic engineering

Enzyme engineering and metabolic engineering drive innovation in energy biotechnology. In recent years, artificial intelligence (AI) has supported successful applications in designing effective enzymes and productive microbial cell factories. This review summarizes recent advances in enzyme redesign using protein language models, de novo enzyme design with generative models, and AI tools for engineering metabolism and related cellular phenotypes. Across these areas, AI models are shifting from single modality inputs to integrated representations of protein function, metabolic pathways, and cell states. We emphasize that unifying the diverse data representations across scales will be necessary for advancements in energy biotechnology.

Volk, Michael [Univ. of Illinois at Urbana-Champai↗

Activation Domain Hunter (ADhunter) v2.0

ADhunter is a software program that enables accurate identification and quantification of transcriptional activation domains. Unlike previous software, ADhunter uses protein representations from a pre-trained protein language model, model ensembling, and a training dataset from a diverse sampling of protein sequence space for state-of-the-art performance. These advantages enable improved perception of transcriptional activation domains across sequence space that can be used for mapping natural genetic circuits and engineering synthetic genetic circuits. In particular, ADhunter enables fine-tuned control of gene expression through synthetic transcription factors that can be used for complex control of cellular programs.

Waldburger, Lucas [Lawrence Berkeley National Labo↗

A characterization of recombinant Arabidopsis FRIABLE1 (FRB1) reveals robust rhamnogalacturonan-I rhamnosyltransferase activity and critical catalytic residues

Plant cell walls are glycan-rich extracellular matrices that fundamentally impact essential cellular processes, such as growth, adhesion, and cell shape acquisition. Understanding plant cell wall glycans requires the identification and characterization of the biosynthetic enzymes that produce these polymers. Most successful in vitro protein expression studies of plant cell wall glycosyltransferases have relied on insect, fungal/yeast, or human cell expression systems, whereas prokaryotic expression systems have been generally unsuccessful. Here, we show that Arabidopsis FRIABLE1 (FRB1)/rhamnogalacturonan-I rhamnosyltransferase 8 (RRT8) can be produced in Escherichia coli RosettaGami2 cells as N-terminal maltose-binding protein fusion proteins containing C-terminal 6X-His-tags. We also report the catalytic constants of FRB1/RRT8 with apparent K M and K cat values of 226 μM and 33 min -1 for UDP-Rhamnose and 117 μM and 28.7 min -1 for rhamnogalacturonan-I (RG-I), respectively. We examine the catalytic activities of mutated FRB1/RRT8 proteins based on an AlphaFold 3-generated FRB1/RRT8 protein structural model with a virtually docked UDP-Rha donor. Enzymatic characterization of the mutated and wildtype FRB1/RRT8 protein confirmed that mutation of predicted catalytic site amino acid residues resulted in a 20-fold reduction in RRT activity. FRB1 also robustly polymerizes RG-I in combination with RG-I galacturonosyltransferase 1. These results show how a robust E. coli expression system combined with artificial intelligence tools can be used to increase understanding of plant cell wall glycosyltransferase structure and function.

glycosyltransferase↗

Mechanistic insights into a heterobifunctional degrader-induced PTPN2/N1 complex

PTPN2 (protein tyrosine phosphatase non-receptor type 2, or TC-PTP) and PTPN1 are attractive immuno-oncology targets, with the deletion of Ptpn1 and Ptpn2 improving response to immunotherapy in disease models. Targeted protein degradation has emerged as a promising approach to drug challenging targets including phosphatases. We developed potent PTPN2/N1 dual heterobifunctional degraders (Cmpd-1 and Cmpd-2) which facilitate efficient complex assembly with E3 ubiquitin ligase CRL4 CRBN , and mediate potent PTPN2/N1 degradation in cells and mice. To provide mechanistic insights into the cooperative complex formation introduced by degraders, we employed a combination of structural approaches. Our crystal structure reveals how PTPN2 is recognized by the tri-substituted thiophene moiety of the degrader. We further determined a high-resolution structure of DDB1-CRBN/Cmpd-1/PTPN2 using single-particle cryo-electron microscopy (cryo-EM). This structure reveals that the degrader induces proximity between CRBN and PTPN2, albeit the large conformational heterogeneity of this ternary complex. The molecular dynamic (MD)-simulations constructed based on the cryo-EM structure exhibited a large rigid body movement of PTPN2 and illustrated the dynamic interactions between PTPN2 and CRBN. Together, our study demonstrates the development of PTPN2/N1 heterobifunctional degraders with potential applications in cancer immunotherapy. Furthermore, the developed structural workflow could help to understand the dynamic nature of degrader-induced cooperative ternary complexes.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

A foundation model for atomistic materials chemistry

Atomistic simulations of matter, especially those that leverage first-principles (ab initio) electronic structure theory, provide a microscopic view of the world, underpinning much of our understanding of chemistry and materials science. Over the last decade or so, machine-learned force fields have transformed atomistic modeling by enabling simulations of ab initio quality over unprecedented time and length scales. However, early machine-learning (ML) force fields have largely been limited by (i) the substantial computational and human effort required to develop and validate potentials for each particular system of interest and (ii) a general lack of transferability from one chemical system to the next. Here, we show that it is possible to create a general-purpose atomistic ML model, trained on a public dataset of moderate size, that is capable of running stable molecular dynamics for a wide range of molecules and materials. We demonstrate the power of the MACE-MP-0 model-and its qualitative and at times quantitative accuracy-on a diverse set of problems in the physical sciences, including properties of solids, liquids, gases, chemical reactions, interfaces, and even the dynamics of a small protein. The model can be applied out of the box as a starting or "foundation" model for any atomistic system of interest and, when desired, can be fine-tuned on just a handful of application-specific data points to reach ab initio accuracy. Establishing that a stable force-field model can cover almost all materials changes atomistic modeling in a fundamental way: experienced users obtain reliable results much faster, and beginners face a lower barrier to entry. Foundation models thus represent a step toward democratizing the revolution in atomic-scale modeling that has been brought about by ML force fields.

Batatia, Ilyes↗

Detection of non‐native species formed during fibrillization of the myocilin olfactomedin domain

Abstract Glaucoma is a group of neurodegenerative diseases that together are the leading cause of irreversible blindness worldwide. Myocilin‐associated glaucoma is an inherited form of this disease, caused by intracellular aggregation of misfolded mutant myocilin. In vitro, the myocilin C‐terminal olfactomedin domain (OLF), the relevant domain for glaucoma pathogenesis, can be driven to form amyloid‐like fibrils under mild conditions. Here we characterize a species present during in vitro fibrillization. Purified OLF was subjected to fibrillization at concentrations required for downstream electron microscopy imaging and NMR spectroscopy. Additional biophysical techniques, including analytical ultracentrifugation and X‐ray crystallography, were employed to further characterize the multicomponent mixture. Negative stain transmission electron microscopy (TEM) shows a non‐native species reminiscent of known prefibrillar oligomers from other amyloid systems, NMR indicates a minor population of partially misfolded species is present in solution, and cryo‐EM imaging shows two‐dimensional protein arrays. The predominant soluble species remaining in solution after the fibril reaction is natively folded, as evidenced by X‐ray crystallography. In summary, after incubating OLF under fibrillization‐promoting conditions, there is a heterogeneous mixture consisting of soluble folded protein, mature amyloid‐like fibrils, and partially misfolded intermediate species that at present belie additional molecular detail. The characterization of OLF fibrillar species illustrates the challenges associated with developing a comprehensive understanding of the fibrillization process for large, non‐model amyloidogenic proteins.

Scelsi, Hailee F. [School of Chemistry and Biochem↗

Peripheral positions encode transport specificity in the small multidrug resistance exporters

In secondary active transporters, a relatively limited set of protein folds have evolved diverse solute transport functions. Because of the conformational changes inherent to transport, altering substrate specificity typically involves remodeling the entire structural landscape, limiting our understanding of how novel substrate specificities evolve. In the current work, we examine a structurally minimalist family of model transport proteins, the small multidrug resistance (SMR) transporters, to understand the molecular basis for the emergence of a novel substrate specificity. We engineer a selective SMR protein to promiscuously export quaternary ammonium antiseptics, similar to the activity of a clade of multidrug exporters in this family. Using combinatorial mutagenesis and deep sequencing, we identify the necessary and sufficient molecular determinants of this engineered activity. Using X-ray crystallography, solid-supported membrane electrophysiology, binding assays, and a proteoliposome-based quaternary ammonium antiseptic transport assay that we developed, we dissect the mechanistic contributions of these residues to substrate polyspecificity. We find that substrate preference changes not through modification of the residues that directly interact with the substrate but through mutations peripheral to the binding pocket. Our work provides molecular insight into substrate promiscuity among the SMRs and can be applied to understand multidrug export and the evolution of novel transport functions more generally.

Science & Technology - Other Topics↗

Correction to “COCOMO2: A Coarse-Grained Model for Interacting Folded and Disordered Proteins”

Biomolecular interactions are essential in many biological processes, including complex formation and phase separation processes. Coarse-grained computational models are especially valuable for studying such processes via simulation. Here, we present COCOMO2, an updated residue-based coarse-grained model that extends its applicability from intrinsically disordered peptides to folded proteins. This is accomplished with the introduction of a surface exposure scaling factor, which adjusts interaction strengths based on solvent accessibility, to enable the more realistic modeling of interactions involving folded domains without additional computational costs. COCOMO2 was parametrized directly with solubility and phase separation data to improve its performance on predicting concentration-dependent phase separation for a broader range of biomolecular systems compared to the original version. COCOMO2 enables new applications including the study of condensates that involve IDPs together with folded domains and the study of complex assembly processes. COCOMO2 also provides an expanded foundation for the development of multiscale approaches for modeling biomolecular interactions that span from residue-level to atomistic resolution.

Molecular interactions↗

COCOMO2: A Coarse-Grained Model for Interacting Folded and Disordered Proteins

Biomolecular interactions are essential in many biological processes, including complex formation and phase separation processes. Coarse-grained computational models are especially valuable for studying such processes via simulation. Here, we present COCOMO2, an updated residue-based coarse-grained model that extends its applicability from intrinsically disordered peptides to folded proteins. This is accomplished with the introduction of a surface exposure scaling factor, which adjusts interaction strengths based on solvent accessibility, to enable the more realistic modeling of interactions involving folded domains without additional computational costs. COCOMO2 was parametrized directly with solubility and phase separation data to improve its performance on predicting concentration-dependent phase separation for a broader range of biomolecular systems compared to the original version. COCOMO2 enables new applications including the study of condensates that involve IDPs together with folded domains and the study of complex assembly processes. COCOMO2 also provides an expanded foundation for the development of multiscale approaches for modeling biomolecular interactions that span from residue-level to atomistic resolution.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

The protein structurome of Orthornavirae and its dark matter

Metatranscriptomics is uncovering more and more diverse families of viruses with RNA genomes comprising the viral kingdom Orthornavirae in the realm Riboviria. Thorough protein annotation and comparison are essential to get insights into the functions of viral proteins and virus evolution. In addition to sequence- and hmm profile-based methods, protein structure comparison adds a powerful tool to uncover protein functions and relationships. We constructed an Orthornavirae “structurome” consisting of already annotated as well as unannotated (“dark matter”) proteins and domains encoded in viral genomes. We used protein structure modeling and similarity searches to illuminate the remaining dark matter in hundreds of thousands of orthornavirus genomes. The vast majority of the dark matter domains showed either “generic” folds, such as single α-helices, or no high confidence structure predictions. Nevertheless, a variety of lineage-specific globular domains that were new either to orthornaviruses in general or to particular virus families were identified within the proteomic dark matter of orthornaviruses, including several predicted nucleic acid-binding domains and nucleases. In addition, we identified a case of exaptation of a cellular nucleoside monophosphate kinase as an RNA-binding protein in several virus families. Notwithstanding the continuing discovery of numerous orthornaviruses, it appears that all the protein domains conserved in large groups of viruses have already been identified. The rest of the viral proteome seems to be dominated by poorly structured domains including intrinsically disordered ones that likely mediate specific virus-host interactions.

59 BASIC BIOLOGICAL SCIENCES↗

Sequence, structure prediction, and epitope analysis of the polymorphic membrane protein family in Chlamydia trachomatis

The polymorphic membrane proteins (Pmps) are a family of autotransporters that play an important role in infection, adhesion and immunity in Chlamydia trachomatis. Here we show that the characteristic GGA(I,L,V) and FxxN tetrapeptide repeats fit into a larger repeat sequence, which correspond to the coils of a large beta-helical domain in high quality structure predictions. Analysis of the protein using structure prediction algorithms provided novel insight to the chlamydial Pmp family of proteins. While the tetrapeptide motifs themselves are predicted to play a structural role in folding and close stacking of the beta-helical backbone of the passenger domain, we found many of the interesting features of Pmps are localized to the side loops jutting out from the beta helix including protease cleavage, host cell adhesion, and B-cell epitopes; while T-cell epitopes are predominantly found in the beta-helix itself. This analysis more accurately defines the Pmp family of Chlamydia and may better inform rational vaccine design and functional studies.

59 BASIC BIOLOGICAL SCIENCES↗

Towards time-resolved MicroED grid preparation using mix-and-inject gas dynamic virtual nozzles

Recent progress in gas dynamic virtual nozzle (GDVN) technologies in combination with high-brilliance synchrotron and X-ray free-electron lasers (XFELs) has allowed the visualization of protein dynamics in crystallo by mixing macromolecular protein crystals with a substrate using tunable mixing times on the order of milliseconds to seconds prior to serial X-ray diffraction data collection. This has become the method of choice for high-resolution structure determination of intermediate states. However, such experiments require large counts of crystals of proper sizes for high-resolution data collection, and premium beam times for screening efforts. Cryogenic microcrystal electron diffraction (MicroED) represents a complementary technique that may be a more accessible avenue for time-resolved nanocrystallography compared with serial X-ray diffraction experiments. MicroED can produce full diffraction datasets from just a few submicrometre-thick crystals, and the approach is more readily accessible, requiring standard cryogenic transmission electron microscopy (TEM) equipment available at many universities and institutes. Cryogenic MicroED, like other forms of cryo-EM, begins with rapidly freezing biological material on electron microscopy grids. In the case of MicroED, micro- to nano-crystals (<500 nm thick) are deposited onto electron microscopy grids and plunge-frozen for subsequent electron diffraction data collection. Here, we have incorporated GDVN technology developed originally for XFEL experiments into the freezing process as a first step towards time-resolved studies. We describe the limited deposition efficiency of the model MicroED protein proteinase K on TEM grids using GDVNs, preceding sample vitrification and successful MicroED data collection. We discuss both the initial results from such experiments and the methodological challenges in developing this approach into a reliable workflow for millisecond-to-second time-resolved structural studies of macromolecules. Our results promise a strategy to deposit crystals on grids using GDVNs and determine high-resolution structures by MicroED, constituting a first step towards development of time-resolved MicroED experiments.

MicroED↗

nf-core/proteinfamilies: a scalable pipeline for the generation of protein families

The growth of metagenomics-derived amino acid sequence data has transformed our understanding of protein function, microbial diversity, and evolutionary relationships. However, the vast majority of these proteins remain functionally uncharacterized. Grouping the millions of such uncharacterized sequences with the few experimentally characterized ones allows the transfer of annotations, while the inspection of conserved residues with multiple sequence alignments can provide clues to function, even in the absence of existing functional information. To address the challenges associated with this data surge and the need to group sequences, we present a scalable, open-source, parametrizable Nextflow pipeline (nf-core/proteinfamilies) that generates nascent protein families or assigns new proteins to existing families. The computational benchmarks demonstrated that resource usage scales approximately linearly with input size, and the biological benchmarks showed that the generated protein families closely resemble manually curated families in widely used databases.

Nextflow↗

A conserved chaperone protein is required for the formation of a noncanonical type VI secretion system spike tip complex

Type VI secretion systems (T6SSs) are dynamic protein nanomachines found in Gram-negative bacteria that deliver toxic effector proteins into target cells in a contact-dependent manner. Prior to secretion, many T6SS effector proteins require chaperones and/or accessory proteins for proper loading onto the structural components of the T6SS apparatus. However, despite their established importance, the precise molecular function of several T6SS accessory protein families remains unclear. In this study, we set out to characterize the DUF2169 family of T6SS accessory proteins. Using gene co-occurrence analyses, we find that DUF2169-encoding genes strictly co-occur with genes encoding T6SS spike complexes formed by valine-glycine repeat protein G (VgrG) and DUF4150 domains. Although structurally similar to Pro-Ala-Ala-Arg (PAAR) domains, “PAAR-like” DUF4150 domains lack PAAR motifs and instead contain a conserved PIPY motif, leading us to designate them PIPY domains. Next, we present both genetic and biochemical evidence that PIPY domains require a cognate DUF2169 protein to form a functional T6SS spike complex with VgrG. This contrasts with canonical PAAR proteins, which bind VgrG on their own to form functional spike complexes. By solving the first crystal structure of a DUF2169 protein, we show that this T6SS accessory protein adopts a novel protein fold. Furthermore, biophysical and structural modeling data suggest that DUF2169 contains a dynamic loop that physically interacts with a hydrophobic patch on the surface of its cognate PIPY domain. Based on these findings, we propose a model whereby DUF2169 proteins function as molecular chaperones that maintain VgrG–PIPY spike complexes in a secretion-competent state prior to their export by the T6SS apparatus.

DUF2169↗

Sequence-based generative AI design of versatile tryptophan synthases

Enzymes are powerful and sustainable catalysts, but their widespread application is limited by the difficulty of identifying functional starting points for optimization, creating a major bottleneck in early- stage biocatalyst discovery. Designing libraries of such starting enzymes remains particularly challenging. Here, we use the GenSLM protein language model to generate novel β-subunit of tryptophan synthase (TrpB) enzymes that express in Escherichia coli and are both stable and catalytically active. Many generated TrpBs also display significant substrate promiscuity, outperforming their natural counterparts on non-native substrates. Some even surpass laboratory-evolved TrpBs. Comparison of the most-active and most-promiscuous generated TrpB to its closest natural homolog confirms that the enhanced versatility is absent from the natural enzyme, highlighting the creative potential of generative models. These results demonstrate that the generated TrpBs not only preserve natural structure and function but also acquire non-natural properties, establishing generative models as powerful tools for biocatalyst discovery and engineering.

biocatalysis↗