Search NASA⌕ Search

DOE OSTI · 1827362

Understanding protein-complex assembly through grand canonical maximum entropy modeling

Abstract

Inside a cell, heterotypic proteins assemble in inhomogeneous, crowded systems where the abundance of these proteins vary with cell types. While some protein complexes form putative structures that can be visualized with imaging, there are far more protein complexes that are yet to be solved because of their dynamic associations with one another. Nevertheless, it is possible to infer these protein complexes through a physical model. However, it is often not clear to physicists what kind of data from biology is necessary for such a modeling endeavor. Here, we aim to model these clusters of coarse-grained protein assemblies from multiple subunits through the constraints of interactions among the subunits and the chemical potential of each subunit. We obtained the constraints on the interactions among subunits from the known protein structures. We inferred the chemical potential that dictates the particle number distribution of each protein subunit from the knowledge of protein abundance from experimental data. Guided by the maximum entropy principle, we formulated an inverse statistical mechanical method to infer the distribution of particle numbers from the data of protein abundance as chemical potentials for a grand canonical multicomponent mixture. Using grand canonical Monte Carlo simulations, we captured a distribution of high-order clusters in a protein complex of succinate dehydrogenase with four known subunits. The complexity of hierarchical clusters varies with the relative protein abundance of each subunit in distinctive cell types such as lung, heart, and brain. When the crowding content increases, we observed that crowding stabilizes emergent clusters that do not exist in dilute conditions. We, therefore, proposed a testable hypothesis that the hierarchical complexity of protein clusters on a molecular scale is a plausible biomarker of predicting the phenotypes of a cell.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Gasic, Andrei G., Sarkar, Atrayee, Cheung, Margaret S.. 2021-09-07. Understanding protein-complex assembly through grand canonical maximum entropy modeling. https://doi.org/10.1103/physrevresearch.3.033220

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related reports

Soil metagenomics umbrella narrative

Implementing accessible, authentic research experiences in introductory courses is challenging, particularly at institutions serving diverse student populations. To address this gap, we developed and deployed a Course-based Undergraduate Research Experience (CURE) focused on plant-microbe interactions in General Biology II at Northeastern Illinois University (NEIU), a minority-serving institution with a diverse student body. Students grew sugar beets (Beta vulgaris), extracted DNA from the rhizoplane, and used the Department of Energy Systems Biology Knowledgebase (KBase) for bioinformatic analysis to compare microbial relative abundance in fertilized versus unfertilized soil. Over five semesters, the CURE engaged 103 students and leveraged the intuitive KBase platform to make complex sequencing data accessible. Pre/post-course survey data revealed significant increases in student self-assessed research skills, including the ability to explain results and determine the types of data to collect. Furthermore, students reported significant gains in confidence related to experimental design and hypothesis development, alongside a strong increase in familiarity with KBase. Informal faculty feedback indicated high student engagement and appreciation for the real-world connections (e.g. food systems, agriculture, and health). This scalable, low-cost model effectively integrates data science tools into the foundational curriculum, demonstrating a potent strategy for boosting research skills and broadening participation in authentic scientific inquiry among diverse undergraduate students.

59 BASIC BIOLOGICAL SCIENCES↗

Genome-resolved insights into microbial diversity and elemental cycling in Winogradsky columns

We retained 18 MAGs with ≥50% completion and <10% contamination (i.e., at least medium quality). Of these, 10 had >90% completion and <5% contamination; however, only one (Paceibacteria Bin.003_MG) can be described as high-quality, as the others lacked a full suite of 5S, 16S, and 23S rRNA genes. To maximize the diversity of our recovered MAGs, we also retained one MAG (Chromatiaceae Bin.008_AM) with >40% (but less than 50%) completion and <5% contamination, as well as one (Rhodopseudomonas Bin.015_MK) with >90% completion and <20% (but>10%) contamination. Interestingly, significant chimerism was not detected in this MAG (40) , suggesting that the elevated contamination (20%) may instead reflect two closely related strains collapsing into a single bin. Consistent with this, contig coverage was bimodal, with roughly 17% of the assembly at ~115x and the remaining 83% at ~282x, while GC content remained uniform across both groups (~64%), arguing against contamination from a taxonomically distinct source.

59 BASIC BIOLOGICAL SCIENCES↗