Search NASASearch

SEARCH · Search NASA

Results for “Protein modeling”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 109 records · Page 6

Correction to “COCOMO2: A Coarse-Grained Model for Interacting Folded and Disordered Proteins”

Biomolecular interactions are essential in many biological processes, including complex formation and phase separation processes. Coarse-grained computational models are especially valuable for studying such processes via simulation. Here, we present COCOMO2, an updated residue-based coarse-grained model that extends its applicability from intrinsically disordered peptides to folded proteins. This is accomplished with the introduction of a surface exposure scaling factor, which adjusts interaction strengths based on solvent accessibility, to enable the more realistic modeling of interactions involving folded domains without additional computational costs. COCOMO2 was parametrized directly with solubility and phase separation data to improve its performance on predicting concentration-dependent phase separation for a broader range of biomolecular systems compared to the original version. COCOMO2 enables new applications including the study of condensates that involve IDPs together with folded domains and the study of complex assembly processes. COCOMO2 also provides an expanded foundation for the development of multiscale approaches for modeling biomolecular interactions that span from residue-level to atomistic resolution.

Molecular interactions

COCOMO2: A Coarse-Grained Model for Interacting Folded and Disordered Proteins

Biomolecular interactions are essential in many biological processes, including complex formation and phase separation processes. Coarse-grained computational models are especially valuable for studying such processes via simulation. Here, we present COCOMO2, an updated residue-based coarse-grained model that extends its applicability from intrinsically disordered peptides to folded proteins. This is accomplished with the introduction of a surface exposure scaling factor, which adjusts interaction strengths based on solvent accessibility, to enable the more realistic modeling of interactions involving folded domains without additional computational costs. COCOMO2 was parametrized directly with solubility and phase separation data to improve its performance on predicting concentration-dependent phase separation for a broader range of biomolecular systems compared to the original version. COCOMO2 enables new applications including the study of condensates that involve IDPs together with folded domains and the study of complex assembly processes. COCOMO2 also provides an expanded foundation for the development of multiscale approaches for modeling biomolecular interactions that span from residue-level to atomistic resolution.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH

Protein folding, protein structure and the origin of life: Theoretical methods and solutions of dynamical problems

Theoretical methods and solutions of the dynamics of protein folding, protein aggregation, protein structure, and the origin of life are discussed. The elements of a dynamic model representing the initial stages of protein folding are presented. The calculation and experimental determination of the model parameters are discussed. The use of computer simulation for modeling protein folding is considered.

Weaver, D. L.

The protein structurome of Orthornavirae and its dark matter

Metatranscriptomics is uncovering more and more diverse families of viruses with RNA genomes comprising the viral kingdom Orthornavirae in the realm Riboviria. Thorough protein annotation and comparison are essential to get insights into the functions of viral proteins and virus evolution. In addition to sequence- and hmm profile-based methods, protein structure comparison adds a powerful tool to uncover protein functions and relationships. We constructed an Orthornavirae “structurome” consisting of already annotated as well as unannotated (“dark matter”) proteins and domains encoded in viral genomes. We used protein structure modeling and similarity searches to illuminate the remaining dark matter in hundreds of thousands of orthornavirus genomes. The vast majority of the dark matter domains showed either “generic” folds, such as single α-helices, or no high confidence structure predictions. Nevertheless, a variety of lineage-specific globular domains that were new either to orthornaviruses in general or to particular virus families were identified within the proteomic dark matter of orthornaviruses, including several predicted nucleic acid-binding domains and nucleases. In addition, we identified a case of exaptation of a cellular nucleoside monophosphate kinase as an RNA-binding protein in several virus families. Notwithstanding the continuing discovery of numerous orthornaviruses, it appears that all the protein domains conserved in large groups of viruses have already been identified. The rest of the viral proteome seems to be dominated by poorly structured domains including intrinsically disordered ones that likely mediate specific virus-host interactions.

59 BASIC BIOLOGICAL SCIENCES

Sequence, structure prediction, and epitope analysis of the polymorphic membrane protein family in Chlamydia trachomatis

The polymorphic membrane proteins (Pmps) are a family of autotransporters that play an important role in infection, adhesion and immunity in Chlamydia trachomatis. Here we show that the characteristic GGA(I,L,V) and FxxN tetrapeptide repeats fit into a larger repeat sequence, which correspond to the coils of a large beta-helical domain in high quality structure predictions. Analysis of the protein using structure prediction algorithms provided novel insight to the chlamydial Pmp family of proteins. While the tetrapeptide motifs themselves are predicted to play a structural role in folding and close stacking of the beta-helical backbone of the passenger domain, we found many of the interesting features of Pmps are localized to the side loops jutting out from the beta helix including protease cleavage, host cell adhesion, and B-cell epitopes; while T-cell epitopes are predominantly found in the beta-helix itself. This analysis more accurately defines the Pmp family of Chlamydia and may better inform rational vaccine design and functional studies.

59 BASIC BIOLOGICAL SCIENCES

Towards time-resolved MicroED grid preparation using mix-and-inject gas dynamic virtual nozzles

Recent progress in gas dynamic virtual nozzle (GDVN) technologies in combination with high-brilliance synchrotron and X-ray free-electron lasers (XFELs) has allowed the visualization of protein dynamics in crystallo by mixing macromolecular protein crystals with a substrate using tunable mixing times on the order of milliseconds to seconds prior to serial X-ray diffraction data collection. This has become the method of choice for high-resolution structure determination of intermediate states. However, such experiments require large counts of crystals of proper sizes for high-resolution data collection, and premium beam times for screening efforts. Cryogenic microcrystal electron diffraction (MicroED) represents a complementary technique that may be a more accessible avenue for time-resolved nanocrystallography compared with serial X-ray diffraction experiments. MicroED can produce full diffraction datasets from just a few submicrometre-thick crystals, and the approach is more readily accessible, requiring standard cryogenic transmission electron microscopy (TEM) equipment available at many universities and institutes. Cryogenic MicroED, like other forms of cryo-EM, begins with rapidly freezing biological material on electron microscopy grids. In the case of MicroED, micro- to nano-crystals (<500 nm thick) are deposited onto electron microscopy grids and plunge-frozen for subsequent electron diffraction data collection. Here, we have incorporated GDVN technology developed originally for XFEL experiments into the freezing process as a first step towards time-resolved studies. We describe the limited deposition efficiency of the model MicroED protein proteinase K on TEM grids using GDVNs, preceding sample vitrification and successful MicroED data collection. We discuss both the initial results from such experiments and the methodological challenges in developing this approach into a reliable workflow for millisecond-to-second time-resolved structural studies of macromolecules. Our results promise a strategy to deposit crystals on grids using GDVNs and determine high-resolution structures by MicroED, constituting a first step towards development of time-resolved MicroED experiments.

MicroED

nf-core/proteinfamilies: a scalable pipeline for the generation of protein families

The growth of metagenomics-derived amino acid sequence data has transformed our understanding of protein function, microbial diversity, and evolutionary relationships. However, the vast majority of these proteins remain functionally uncharacterized. Grouping the millions of such uncharacterized sequences with the few experimentally characterized ones allows the transfer of annotations, while the inspection of conserved residues with multiple sequence alignments can provide clues to function, even in the absence of existing functional information. To address the challenges associated with this data surge and the need to group sequences, we present a scalable, open-source, parametrizable Nextflow pipeline (nf-core/proteinfamilies) that generates nascent protein families or assigns new proteins to existing families. The computational benchmarks demonstrated that resource usage scales approximately linearly with input size, and the biological benchmarks showed that the generated protein families closely resemble manually curated families in widely used databases.

Nextflow

A conserved chaperone protein is required for the formation of a noncanonical type VI secretion system spike tip complex

Type VI secretion systems (T6SSs) are dynamic protein nanomachines found in Gram-negative bacteria that deliver toxic effector proteins into target cells in a contact-dependent manner. Prior to secretion, many T6SS effector proteins require chaperones and/or accessory proteins for proper loading onto the structural components of the T6SS apparatus. However, despite their established importance, the precise molecular function of several T6SS accessory protein families remains unclear. In this study, we set out to characterize the DUF2169 family of T6SS accessory proteins. Using gene co-occurrence analyses, we find that DUF2169-encoding genes strictly co-occur with genes encoding T6SS spike complexes formed by valine-glycine repeat protein G (VgrG) and DUF4150 domains. Although structurally similar to Pro-Ala-Ala-Arg (PAAR) domains, “PAAR-like” DUF4150 domains lack PAAR motifs and instead contain a conserved PIPY motif, leading us to designate them PIPY domains. Next, we present both genetic and biochemical evidence that PIPY domains require a cognate DUF2169 protein to form a functional T6SS spike complex with VgrG. This contrasts with canonical PAAR proteins, which bind VgrG on their own to form functional spike complexes. By solving the first crystal structure of a DUF2169 protein, we show that this T6SS accessory protein adopts a novel protein fold. Furthermore, biophysical and structural modeling data suggest that DUF2169 contains a dynamic loop that physically interacts with a hydrophobic patch on the surface of its cognate PIPY domain. Based on these findings, we propose a model whereby DUF2169 proteins function as molecular chaperones that maintain VgrG–PIPY spike complexes in a secretion-competent state prior to their export by the T6SS apparatus.

DUF2169

Sequence-based generative AI design of versatile tryptophan synthases

Enzymes are powerful and sustainable catalysts, but their widespread application is limited by the difficulty of identifying functional starting points for optimization, creating a major bottleneck in early- stage biocatalyst discovery. Designing libraries of such starting enzymes remains particularly challenging. Here, we use the GenSLM protein language model to generate novel β-subunit of tryptophan synthase (TrpB) enzymes that express in Escherichia coli and are both stable and catalytically active. Many generated TrpBs also display significant substrate promiscuity, outperforming their natural counterparts on non-native substrates. Some even surpass laboratory-evolved TrpBs. Comparison of the most-active and most-promiscuous generated TrpB to its closest natural homolog confirms that the enhanced versatility is absent from the natural enzyme, highlighting the creative potential of generative models. These results demonstrate that the generated TrpBs not only preserve natural structure and function but also acquire non-natural properties, establishing generative models as powerful tools for biocatalyst discovery and engineering.

biocatalysis

Understanding the selectivity of nonsteroidal anti-inflammatory drugs for cyclooxygenases using quantum crystallography and electrostatic interaction energy

Quantum crystallography methods have been employed to investigate complex formation between nonsteroidal anti-inflammatory drugs (NSAIDs) and cyclooxygenase (COX) enzymes, with particular focus on the COX-1 and COX-2 isoforms. This study analyzed the electrostatic interaction energies of selected NSAIDs (flurbiprofen, ibuprofen, meloxicam and celecoxib) with the active sites of COX-1 and COX-2, revealing significant differences in binding profiles. Flurbiprofen exhibited the strongest interactions with both COX-1 and COX-2, indicating its potent binding affinity. Celecoxib and meloxicam showed a preference for COX-2, consistent with their known selectivity for this isoform, while ibuprofen showed comparable interaction energies with both isoforms, reflecting its nonselective inhibition pattern. Key amino-acid residues, including Arg120, Arg/His513 and Tyr355, were identified as critical determinants of NSAID selectivity and binding affinity. The findings highlight the complex interplay between interaction energy and selectivity, suggesting that while electrostatic interactions play a fundamental role, additional factors such as enzyme dynamics and the hydrophobic effect also contribute to the therapeutic efficacy and safety profiles of NSAIDs. These insights provide valuable guidance for the rational design of NSAIDs with enhanced therapeutic benefits and minimized adverse effects.

60 APPLIED LIFE SCIENCES

smol_model_benchmark

Machine learning models for protein–small molecule binding prediction using graph (GAT), image-based (ResNet), and transformer-based SMILES approaches (ChemBERTa and MIST). The code includes training pipelines and experiments on the dataset from the BELKA Kaggle competition.

Gibson, Kaetlyn [Los Alamos National Lab]

Nitrogen limitation causes a seismic shift in redox state and phosphorylation of proteins implicated in carbon flux and lipidome remodeling in Rhodotorula toruloides

Background: Oleaginous yeast are prodigious producers of oleochemicals, offering alternative and secure sources for applications in foodstuff, skincare, biofuels, and bioplastics. Nitrogen starvation is the primary strategy used to induce oil accumulation in oleaginous yeast as part of a global stress response. While research has demonstrated that post-translational modifications (PTMs), including phosphorylation and protein cysteine thiol oxidation (redox PTMs), are involved in signaling pathways that regulate stress responses in metazoa and algae, their role in oleaginous yeast remain understudied and unexplored. Results: Towards linking the yeast oleaginous phenotype to protein function, we integrated lipidomics, redox proteomics, and phosphoproteomics to investigate Rhodotorula toruloides under nitrogen-rich and starved conditions over time. Our lipidomics results unearthed interactions involving sphingolipids and cardiolipins with ER stress and mitophagy. Our redox and phosphoproteomics data highlighted the roles of the AMPK, TOR, and calcium signaling pathways in regulation of lipogenesis, autophagy, and oxidative stress response. As a first, we also demonstrated that lipogenic enzymes including fatty acid synthase are modified as a consequence of shifts in cellular redox states due to nutrient availability. Conclusions: We conclude that lipid accumulation is largely a consequence of carbon rerouting and autophagy governed by changes to PTMs, and not increases in the abundance of enzymes involved in central carbon metabolism and fatty acid biosynthesis. Our systems-level approach sets the stage for acquiring multidimensional data sets for protein structural modeling and predicting the functional relevance of PTMs using Artificial Intelligence/Machine Learning (AI/ML). Coupled to those bioinformatics approaches, the putative PTM switches that we delineate will enable advanced metabolic engineering strategies to decouple lipid accumulation from nitrogen limitation.

Lipid Signalling

Genomic-based biosurveillance for avian influenza: whole genome sequencing from wild mallards sampled during autumn migration in 2022–2023 reveals a high co-infection rate on migration stopover site in Georgia

The Caucasus region, including Georgia, is an important intersection for migratory waterbirds, offering potential for avian influenza virus (AIV) transmission between populations from different geographic areas. In 2022 and 2023, wild ducks were sampled during autumn migration events in Georgia to study the genetic relationships and molecular characteristics of influenza strains. Sequencing and phylogenetic analysis were used to compare the sampled strains to reference sequences from Africa, Asia, and Europe, allowing assessment of genetic relationships and virus transmission between migratory birds. Protein language modeling identified potential co-infections. Of 225 duck samples, 128 tested positive for the influenza M gene. 55 influenza-positive samples underwent whole-genome sequencing, revealing significant diversity. Analysis of the hemagglutinin (HA) segment showed notable differences among subtypes. Most samples were H6N1 and H6N6, but co-infections with combinations like H6H3, N8N1, N6H9, N2N6, and H9H6/N1N2 were also identified. These findings demonstrate the high variability of influenza viruses in migratory waterbirds in Georgia, including a notable rate of co-infections. Some samples exhibited uncommon genetic characteristics compared to other strains from the same year, suggesting Georgia’s role as a mixing vessel for influenza viruses. This facilitates reassortment during co-infections and contributes to the genetic diversity observed across flyways.

59 BASIC BIOLOGICAL SCIENCES

Response of thyroid follicular cells to gamma irradiation compared to proton irradiation: II. The role of connexin 32

The objective of this study was to determine whether connexin 32-type gap junctions contribute to the "contact effect" in follicular thyrocytes and whether the response is influenced by radiation quality. Our previous studies demonstrated that early-passage follicular cultures of Fischer rat thyroid cells express functional connexin 32 gap junctions, with later-passage cultures expressing a truncated nonfunctional form of the protein. This model allowed us to assess the role of connexin 32 in radiation responsiveness without relying solely on chemical manipulation of gap junctions. The survival curves generated after gamma irradiation revealed that early-passage follicular cultures had significantly lower values of alpha (0.04 Gy(-1)) than later-passage cultures (0.11 Gy(-1)) (P < 0.0001, n = 12). As an additional way to determine whether connexin 32 was contributing to the difference in survival, cultures were treated with heptanol, resulting in higher alpha values, with early-passage cultures (0.10 Gy(-1)) nearly equivalent to untreated late-passage cultures (0.11 Gy(-1)) (P > 0.1, n = 9). This strongly suggests that the presence of functional connexin 32-type gap junctions was contributing to radiation resistance in gamma-irradiated thyroid follicles. Survival curves from proton-irradiated cultures had alpha values that were not significantly different whether cells expressed functional connexin 32 (0.10 Gy(-1)), did not express connexin 32 (0.09 Gy(-1)), or were down-regulated (early-passage plus heptanol, 0.09 Gy(-1); late-passage plus heptanol, 0.12 Gy(-1)) (P > 0.1, n = 19). Thus, for proton irradiation, the presence of connexin 32-type gap junctional channels did not influence their radiosensitivity. Collectively, the data support the following conclusions. (1) The lower alpha values from the gamma-ray survival curves of the early-passage cultures suggest greater repair efficiency and/or enhanced resistance to radiation-induced damage, coincident with the expression of connexin 32-type gap junctions. (2) The increased sensitivity of FRTL-5 cells to proton irradiation was independent of their ability to communicate through connexin 32 gap junctions. (3) The fact that the beta components of the survival curves from both gamma rays and proton beams were similar (average 0.022 +/- 0.008 Gy(-2), P > 0.1, n = 39) suggests that at higher doses the loss of viability occurs at a relatively constant rate and is independent of radiation quality and the presence of functional gap junctions.

Non-NASA Center

CryoSegNet: accurate cryo-EM protein particle picking by integrating the foundational AI image segmentation model and attention-gated U-Net

Picking protein particles in cryo-electron microscopy (cryo-EM) micrographs is a crucial step in the cryo-EM-based structure determination. However, existing methods trained on a limited amount of cryo-EM data still cannot accurately pick protein particles from noisy cryo-EM images. The general foundational artificial intelligence–based image segmentation model such as Meta’s Segment Anything Model (SAM) cannot segment protein particles well because their training data do not include cryo-EM images. Here, we present a novel approach (CryoSegNet) of integrating an attention-gated U-shape network (U-Net) specially designed and trained for cryo-EM particle picking and the SAM. The U-Net is first trained on a large cryo-EM image dataset and then used to generate input from original cryo-EM images for SAM to make particle pickings. CryoSegNet shows both high precision and recall in segmenting protein particles from cryo-EM micrographs, irrespective of protein type, shape and size. On several independent datasets of various protein types, CryoSegNet outperforms two top machine learning particle pickers crYOLO and Topaz as well as SAM itself. The average resolution of density maps reconstructed from the particles picked by CryoSegNet is 3.33 Å, 7% better than 3.58 Å of Topaz and 14% better than 3.87 Å of crYOLO. It is publicly available at https://github.com/jianlin-cheng/CryoSegNet

59 BASIC BIOLOGICAL SCIENCES

PRIME: Protein Representation Inference for Mutation Evaluation

Protein language machine learning models built upon existing ESM-2 model developed by Evolutionary Scale (evolutionaryscale.ai) and an in-house protein language model based on the BERT model developed by Google. The code also includes model training scripts and saved checkpoints from our own training using publicly available SARS-CoV-2 protein sequences.

Gibson, Kaetlyn [Los Alamos National Lab]

Protein crystal growth in low gravity

The solubility and growth of the protein canavalin, and the application of the schlieren technique to study fluid flow in protein crystal growth systems were investigated. These studies have resulted in the proposal of a model to describe protein crystal growth and the preliminary plans for a long-term space flight experiment. Canavalin, which may be crystallized from a basic solution by the addition of hydrogen (H+) ions, was shown to have normal solubility characteristics over the range of temperatures (5 to 25 C) and pH (5 to 7.5) studies. The solubility data combined with growth rate data gathered from the seeded growth of canavalin crystals indicated that the growth rate limiting step is a screw dislocation mechanism. A schlieren apparatus was constructed and flow patterns were observed in Rochelle salt (sodium potassium tartrate), lysozyme, and canavalin. The critical parameters were identified as the change in density with concentration (dp/dc) and the change in index of refraction with concentration (dn/dc). Some of these values were measured for the materials listed. The data for lyrozyme showed non-linearities in plots of optical properties and density vs. concentration. In conjunction with with W. A. Tiller, a model based on colloid stability theory was proposed to describe protein crystallization. The model was used to explain observations made by ourselves and others. The results of this research has lead to the development for a preliminary design for a long-term, low-g experiment. The proposed apparatus is univeral and capable of operation under microprocessor control.

Feigelson, Robert S.

A predictive theoretical model for electron tunneling pathways in proteins

A practical method is presented for calculating the dependence of electron transfer rates on details of the protein medium intervening between donor and acceptor. The method takes proper account of the relative energetics and mutual interactions of the donor, acceptor, and peptide groups. It also provides a quantitative search scheme for determining the important tunneling pathways (specific sequences of localized bonding and antibonding orbitals of the protein which dominate the donor-acceptor electronic coupling) in native and tailored proteins, a tool for designing new proteins with prescribed electron transfer rates, and a consistent description of observed electron transfer rates in existing redox labeled metalloproteins and small molecule model compounds.

Onuchic, Jose Nelson