Search NASASearch

SEARCH · Search NASA

Results for “Protein structure”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 55 records · Page 3

The protein structurome of Orthornavirae and its dark matter

Metatranscriptomics is uncovering more and more diverse families of viruses with RNA genomes comprising the viral kingdom Orthornavirae in the realm Riboviria. Thorough protein annotation and comparison are essential to get insights into the functions of viral proteins and virus evolution. In addition to sequence- and hmm profile-based methods, protein structure comparison adds a powerful tool to uncover protein functions and relationships. We constructed an Orthornavirae “structurome” consisting of already annotated as well as unannotated (“dark matter”) proteins and domains encoded in viral genomes. We used protein structure modeling and similarity searches to illuminate the remaining dark matter in hundreds of thousands of orthornavirus genomes. The vast majority of the dark matter domains showed either “generic” folds, such as single α-helices, or no high confidence structure predictions. Nevertheless, a variety of lineage-specific globular domains that were new either to orthornaviruses in general or to particular virus families were identified within the proteomic dark matter of orthornaviruses, including several predicted nucleic acid-binding domains and nucleases. In addition, we identified a case of exaptation of a cellular nucleoside monophosphate kinase as an RNA-binding protein in several virus families. Notwithstanding the continuing discovery of numerous orthornaviruses, it appears that all the protein domains conserved in large groups of viruses have already been identified. The rest of the viral proteome seems to be dominated by poorly structured domains including intrinsically disordered ones that likely mediate specific virus-host interactions.

59 BASIC BIOLOGICAL SCIENCES

Characterization of Gas-Phase Native(-like) Proteins Using Structures for Lossless Ion Manipulations

High-resolution mobility-based ion separations in Structures for Lossless Ion Manipulations (SLIM) have been useful for ion mobility separations for a variety of molecular classes in the gas phase. Here, in this study, we present multipass SLIM separations for gas-phase proteins in their near-native state exhibiting charge-state-dependent arrival time distributions using carbonic anhydrase (29 kDa), alcohol dehydrogenase (148 kDa), and apo-transferrin (79 kDa). The experimental CCS values were obtained from calibration curves for the arrival times of Agilent Tune Mix ions. For multipass separations, the ATDs were converted to CCS values by deconvoluting the multipass arrival times into accurate single-pass values amenable to the single-pass calibration curves. Mass spectra of carbonic anhydrase (CA) showed three different charge states (z = 9+ to 11+). Their corresponding mobility peaks were baseline-separated by using 8-m single-pass separations. When compared to the corresponding drift tube ion mobility (DTIMS) measurements, the CCS values obtained from DTIMS and SLIM were in agreement within experimental error. Single-pass analysis of alcohol dehydrogenase (ADH) exhibits three predominant charge states (z = 23+ to 25+) with mobility overlap between adjacent charge states. The mobility peak resolution for ADH improved with multipass separations (up to 24-m path length). In addition, CCS distributions obtained for charge states z = 16+ to 18+ of apo-transferrin reveal a transition from a compact unimodal form (z = 18+ and 19+) to broader multimodal CCS distributions for z = 16+. For apo-transferrin, 40-m multipass separations were performed allowing for complete isolation of the selected mobility range corresponding to z = 17+, leading to selective isolation of a narrow arrival time window. The extended mobility separations provided minimal alterations to the structure of the proteins, and the experimentally derived CCS values showed minimal change as a function of the separation time or number of passes. Mobility-based ion separations for native-like proteins, using SLIM, open opportunities for native-IMS applications as well as other manipulations enabled by SLIM-like mobility-selective isolation and collection.

charge state distribution

PNNL-Predictive-Phenomics/ProCaliper

ProCaliper is a Python library that curates, organizes, and computes protein structure features in a way that easily interfaces with user-provided experimental data. It extracts or computes protein binding site, active site, charge, pLDDT (order/disorder), acid dissociation, protonation, solvent accessible surface area, disulfide bond distance, and protein secondary structure data using precomputed protein structures and publicly available databases. It provides a unified API for integrating additional residue-level data and for visualizing residue features in 3D.

Rozum, Jordan [Pacific Northwest National Lab]

Geometry, spin coupling, and dielectric control of redox potentials in [4Fe–4S] Clusters

Iron–sulfur (Fe–S) clusters are common biological cofactors that facilitate vital redox reactions. Despite extensive research, the molecular basis of redox potential tuning in ferredoxin-like proteins remains an active area of debate. In this study, we combine statistical analysis of over one thousand [4Fe–4S]-containing protein structures from the Protein Data Bank (PDB) with broken-symmetry and extended broken-symmetry density functional theory to examine how cysteine ligand orientations and environmental screening affect redox properties of the clusters. We identified five main ligand configurations, three of which are predominant in natural structures. Among these, the adiabatic electron affinity differs by less than 0.1 V, indicating that, while geometry plays a secondary role, it allows localized fine-tuning of redox properties. In contrast, electrostatic and solvation effects primarily determine the overall potential range.

Computational Chemistry

SLAB: simultaneous labeling and binding affinity prediction for protein–ligand structures

Machine learning models are often used as scoring functions to predict the binding affinity of a protein–ligand complex. These models are trained with limited amounts of data with experimentally measured binding affinity values. A large number of compounds are labeled inactive through single-concentration screens without measuring binding affinities. These inactive compounds, along with the active ones, can be used to train binary classification models, while regression models are trained using compounds with binding affinities only. However, the classification and regression tasks are often handled separately, without sharing the learned feature representations. In this paper, we propose a novel model architecture that jointly performs regression and classification objectives, aiming to maximize data utilization and improve predictive performance by leveraging two complementary tasks. In our setup, the regression yields the binding affinity, whereas the classification task yields the label as active or inactive. We demonstrate our method using PDBbind, the standard 3D structure database, as well as a dataset of flavivirus protease compounds with binding affinity data. Our experiments show that the new joint training strategy improves the accuracy of the model, increasing applicability in various practical drug screening scenarios.

Biological and medical sciences

Isolation and characterization of 24 phages infecting the plant growth-promoting rhizobacterium Klebsiella sp. M5al

Bacteriophages largely impact bacterial communities via lysis, gene transfer, and metabolic reprogramming and thus are increasingly thought to alter nutrient and energy cycling across many of Earth’s ecosystems. However, there are few model systems to mechanistically and quantitatively study phage-bacteria interactions, especially in soil systems. Here, we isolated, sequenced, and genomically characterized 24 novel phages infectingKlebsiellasp. M5al, a plant growth-promoting, nonencapsulated rhizosphere-associated bacterium, and compared many of their features against all 565 sequenced, dsDNAKlebsiellaphage genomes. Taxonomic analyses revealed that theseKlebsiellaphages belong to three known phage families (Autographiviridae,Drexlerviridae, andStraboviridae) and two newly proposed phage families (CandidatusMavericviridaeand Ca.Rivulusviridae). At the phage family level, we found that core genes were often phage-centric proteins, such as structural proteins for the phage head and tail and DNA packaging proteins. In contrast, genes involved in transcription, translation, or hypothetical proteins were commonly not shared or flexible genes. Ecologically, we assessed the phages’ ubiquity in recent large-scale metagenomic datasets, which revealed they were not widespread, as well as a possible direct role in reprogramming specific metabolisms during infection by screening their genomes for phage-encoded auxiliary metabolic genes (AMGs). Even though AMGs are common in the environmental literature, only one of our phage families,Straboviridae, contained AMGs, and the types of AMGs were correlated at the genus level. Host range phenotyping revealed the phages had a wide range of infectivity, infecting between 1–14 of our 22 bacterial strain panel that included pathogenicKlebsiellaandRaoultellastrains. This indicates that not all capsule-independent Klebsiella phages have broad host ranges. Together, these isolates, with corresponding genome, AMG, and host range analyses, help build theKlebsiellamodel system for studying phage-host interactions of rhizosphere-associated bacteria.

Science & Technology - Other Topics

Fluorescent Applications to Crystallization

By covalently modifying a subpopulation, less than or equal to 1%, of a macromolecule with a fluorescent probe, the labeled material will add to a growing crystal as a microheterogeneous growth unit. Labeling procedures can be readily incorporated into the final stages of purification, and tests with model proteins have shown that labeling u to 5 percent of the protein molecules does not affect the X-ray data quality obtained . The presence of the trace fluorescent label gives a number of advantages. Since the label is covalently attached to the protein molecules, it "tracks" the protein s response to the crystallization conditions. The covalently attached probe will concentrate in the crystal relative to the solution, and under fluorescent illumination crystals show up as bright objects against a darker background. Non-protein structures, such as salt crystals, do not show up under fluorescent illumination. Crystals have the highest protein concentration and are readily observed against less bright precipitated phases, which under white light illumination may obscure the crystals. Automated image analysis to find crystals should be greatly facilitated, without having to first define crystallization drop boundaries as the protein or protein structures is all that shows up. Fluorescence intensity is a faster search parameter, whether visually or by automated methods, than looking for crystalline features. Preliminary tests, using model proteins, indicates that we can use high fluorescence intensity regions, in the absence of clear crystalline features or "hits", as a means for determining potential lead conditions. A working hypothesis is that more rapid amorphous precipitation kinetics may overwhelm and trap more slowly formed ordered assemblies, which subsequently show up as regions of brighter fluorescence intensity. Experiments are now being carried out to test this approach using a wider range, of proteins. The trace fluorescently labeled crystals will also emit with sufficient intensity to aid in the automation of crystal alignment using relatively low cost optics, further increasing throughput at synchrotrons.

Pusey, Marc L.

Designed 2D protein crystals as dynamic molecular gatekeepers for a solid-state device

The sensitivity and responsiveness of living cells to environmental changes are enabled by dynamic protein structures, inspiring efforts to construct artificial supramolecular protein assemblies. However, despite their sophisticated structures, designed protein assemblies have yet to be incorporated into macroscale devices for real-life applications. We report a 2D crystalline protein assembly of C98/E57/E66 L-rhamnulose-1-phosphate aldolase ( CEE RhuA) that selectively blocks or passes molecular species when exposed to a chemical trigger. CEE RhuA crystals are engineered via cobalt(II) coordination bonds to undergo a coherent conformational change from a closed state (pore dimensions <1 nm) to an ajar state (pore dimensions ~4 nm) when exposed to an HCN(g) trigger. When layered onto a mesoporous silicon (pSi) photonic crystal optical sensor configured to detect HCN (g) , the 2D CEE RhuA crystal layer effectively blocks interferents that would otherwise result in a false positive signal. The 2D CEE RhuA crystal layer opens in selective response to low-ppm levels of HCN (g) , allowing analyte penetration into the pSi sensor layer for detection. These findings illustrate that designed protein assemblies can function as dynamic components of solid-state devices in non-aqueous environments.

36 MATERIALS SCIENCE

Birth of protein folds and functions in the virome

The rapid evolution of viruses generates proteins that are essential for infectivity and replication but with unknown functions, due to extreme sequence divergence. Here, using a database of 67,715 newly predicted protein structures from 4,463 eukaryotic viral species, we found that 62% of viral proteins are structurally distinct and lack homologues in the AlphaFold database. Among the remaining 38% of viral proteins, many have non-viral structural analogues that revealed surprising similarities between human pathogens and their eukaryotic hosts. Structural comparisons suggested putative functions for up to 25% of unannotated viral proteins, including those with roles in the evasion of innate immunity. In particular, RNA ligase T-like phosphodiesterases were found to resemble phage-encoded proteins that hydrolyse the host immune-activating cyclic dinucleotides 3',3'- and 2',3'-cyclic GMP-AMP (cGAMP). Experimental analysis showed that RNA ligase T homologues encoded by avian poxviruses similarly hydrolyse cGAMP, showing that RNA ligase T-mediated targeting of cGAMP is an evolutionarily conserved mechanism of immune evasion that is present in both bacteriophage and eukaryotic viruses. Together, the viral protein structural database and analyses presented here afford new opportunities to identify mechanisms of virus–host interactions that are common across the virome.

59 BASIC BIOLOGICAL SCIENCES

Assembly of respiratory syncytial virus matrix protein lattice and its coordination with fusion glycoprotein trimers

Respiratory syncytial virus (RSV) is an enveloped, filamentous, negative-strand RNA virus that causes significant respiratory illness worldwide. RSV vaccines are available, however there is still significant need for research to support the development of vaccines and therapeutics against RSV and related Mononegavirales viruses. Individual virions vary in size, with an average diameter of ~130 nm and ranging from ~500 nm to over 10 µm in length. Though the general arrangement of structural proteins in virions is known, we use cryo-electron tomography and sub-tomogram averaging to determine the molecular organization of RSV structural proteins. We show that the peripheral membrane-associated RSV matrix (M) protein is arranged in a packed helical-like lattice of M-dimers. We report that RSV F glycoprotein is frequently observed as pairs of trimers oriented in an anti-parallel conformation to support potential interactions between trimers. Our sub-tomogram averages indicate the positioning of F-trimer pairs is correlated with the underlying M lattice. These results provide insight into RSV virion organization and may aid in the development of RSV vaccines and anti-viral targets.

59 BASIC BIOLOGICAL SCIENCES

Fluorescent Approaches to High Throughput Crystallography

We have shown that by covalently modifying a subpopulation, less than or equal to 1%, of a macromolecule with a fluorescent probe, the labeled material will add to a growing crystal as a microheterogeneous growth unit. Labeling procedures can be readily incorporated into the final stages of purification, and the presence of the probe at low concentrations does not affect the X-ray data quality or the crystallization behavior. The presence of the trace fluorescent label gives a number of advantages when used with high throughput crystallizations. The covalently attached probe will concentrate in the crystal relative to the solution, and under fluorescent illumination crystals show up as bright objects against a dark background. Non-protein structures, such as salt crystals, will not incorporate the probe and will not show up under fluorescent illumination. Brightly fluorescent crystals are readily found against less bright precipitated phases, which under white light illumination may obscure the crystals. Automated image analysis to find crystals should be greatly facilitated, without having to first define crystallization drop boundaries as the protein or protein structures is all that shows up. Fluorescence intensity is a faster search parameter, whether visually or by automated methods, than looking for crystalline features. We are now testing the use of high fluorescence intensity regions, in the absence of clear crystalline features or "hits", as a means for determining potential lead conditions. A working hypothesis is that kinetics leading to non-structured phases may overwhelm and trap more slowly formed ordered assemblies, which subsequently show up as regions of brighter fluorescence intensity. Preliminary experiments with test proteins have resulted in the extraction of a number of crystallization conditions from screening outcomes based solely on the presence of bright fluorescent regions. Subsequent experiments will test this approach using a wider range of proteins. The trace fluorescently labeled crystals will also emit with sufficient intensity to aid in the automation of crystal alignment using relatively low cost optics, further increasing throughput at synchrotrons.

Pusey, Marc L.

Repetitive proteins that undergo large conformational changes evade structural prediction algorithms

Protein structure prediction algorithms, such as AlphaFold, have accelerated protein design and advanced the understanding of the relationship between amino acid sequence and protein structure. However, these algorithms are limited in their ability to predict the structures of conformationally dynamic, intrinsically disordered, and stimuli-responsive proteins. To evaluate sequence-to-structure predictions of such challenging proteins, we explored a class of conformationally dynamic, repeats-in-toxin (RTX) proteins. RTX proteins adopt intrinsically disordered conformations in the absence of calcium and undergo reversible folding into β-roll structures upon binding to calcium. RTX proteins are characterized by tandem repeats of the sequence GGXGXDXUX, in which X can be any amino acid and U is an aliphatic amino acid. We designed RTX sequence variants with global substitutions of nonconserved amino acids, tandem repeats of consensus sequences GGAGXDTLY, and tandem repeats of scrambled sequences GGAGXDTYL. AlphaFold2 and AlphaFold3 predicted that all of these RTX variants adopt β-roll structures, characteristic of wild-type RTX bound to calcium. However, modeling the predicted structures with molecular dynamics simulations and characterizing the protein variants with circular dichroism spectroscopy, small-angle x-ray scattering, and x-ray crystallography revealed that variants adopt diverse, sequence-dependent structures in the absence and presence of calcium. To better design proteins for applications in biotechnology and sustainability, it is critical to build predictive tools that consider intrinsically disordered protein states and validate these tools with multi-mode, multi-scale experimental data.

Chang, Marina P. [Stanford Univ., CA (United State

Theoretical foundations for quantitative paleogenetics. III - The molecular divergence of nucleic acids and proteins for the case of genetic events of unequal probability

Theoretical equations are derived for molecular divergence with respect to gene and protein structure in the presence of genetic events with unequal probabilities: amino acid and base compositions, the frequencies of nucleotide replacements, the usage of degenerate codons, the distribution of fixed base replacements within codons and the distribution of fixed base replacements among codons. Results are presented in the form of tables relating the probabilities of given numbers of codon base changes with respect to the original codon for the alpha hemoglobin, beta hemoglobin, myoglobin, cytochrome c and parvalbumin group gene families. Application of the calculations to the rabbit alpha and beta hemoglobin mRNAs and proteins indicates that the genes are separated by about 425 fixed based replacements distributed over 114 codon sites, which is a factor of two greater than previous estimates. The theoretical results also suggest that many more base replacements are required to effect a given gene or protein structural change than previously believed.

Holmquist, R.

Packaging “vegetable oils”: Insights into plant lipid droplet proteins

Abstract Plant neutral lipids, also known as “vegetable oils”, are synthesized within the endoplasmic reticulum (ER) membrane and packaged into subcellular compartments called lipid droplets (LDs) for stable storage in the cytoplasm. The biogenesis, modulation, and degradation of cytoplasmic LDs in plant cells are orchestrated by a variety of proteins localized to the ER, LDs, and peroxisomes. Recent studies of these LD-related proteins have greatly advanced our understanding of LDs not only as steady oil depots in seeds but also as dynamic cell organelles involved in numerous physiological processes in different tissues and developmental stages of plants. In the past 2 decades, technology advances in proteomics, transcriptomics, genome sequencing, cellular imaging and protein structural modeling have markedly expanded the inventory of LD-related proteins, provided unprecedented structural and functional insights into the protein machinery modulating LDs in plant cells, and shed new light on the functions of LDs in nonseed plant tissues as well as in unicellular algae. Here, we review critical advances in revealing new LD proteins in various plant tissues, point out structural and mechanistic insights into key proteins in LD biogenesis and dynamic modulation, and discuss future perspectives on bridging our knowledge gaps in plant LD biology.

Cai, Yingqi (ORCID:0000000203575809)