Search NASA⌕ Search

SEARCH · Search NASA

Results for “Protein sequence analysis”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

Bacterial hemophilin homologs and their specific type eleven secretor proteins have conserved roles in heme capture and are diversifying as a family

Cellular life relies on enzymes that require metals, which must be acquired from extracellular sources. Bacteria utilize surface and secreted proteins to acquire such valuable nutrients from their environment. These include the cargo proteins of the type eleven secretion system (T11SS), which have been connected to host specificity, metal homeostasis, and nutritional immunity evasion. This Sec-dependent, Gram-negative secretion system is encoded by organisms throughout the phylum Proteobacteria, including human pathogens Neisseria meningitidis, Proteus mirabilis, Acinetobacter baumannii, and Haemophilus influenzae. Experimentally verified T11SS-dependent cargo include transferrin-binding protein B (TbpB), the hemophilin homologs heme receptor protein C (HrpC), hemophilin A (HphA), the immune evasion protein factor-H binding protein (fHbp), and the host symbiosis factor nematode intestinal localization protein C (NilC). Here, we examined the specificity of T11SS systems for their cognate cargo proteins using taxonomically distributed homolog pairs of T11SS and hemophilin cargo and explored the ligand binding ability of those hemophilin cargo homologs. In vivo expression in Escherichia coli of hemophilin homologs revealed that each is secreted in a specific manner by its cognate T11SS protein. Sequence analysis and structural modeling suggest that all hemophilin homologs share an N-terminal ligand-binding domain with the same topology as the ligand-binding domains of the Haemophilus haemolyticus heme binding protein (Hpl) and HphA. We term this signature feature of this group of proteins the hemophilin ligand-binding domain. Network analysis of hemophilin homologs revealed five subclusters and representatives from four of these showed variable heme-binding activities, which, combined with sequence-structure variation, suggests that hemophilins are diversifying in function.

59 BASIC BIOLOGICAL SCIENCES↗

Plant sulfate transporter protein sequences for phylogenetic analysis

Sulfur is an essential macronutrient that supports plant growth, development, and responses to environmental stress. Sulfate is the predominant inorganic form of sulfur in soils, and its uptake by roots and translocation to shoots are facilitated by the sulfate transporter (SULTR) family of proteins. Although the first plant SULTR gene was identified nearly three decades ago, several subfamily members, particularly those in the expansive and angiosperm-specific SULTR3 group, remain poorly characterized. To support comprehensive phylogenetic and sequence-based analyses, we compiled a curated dataset of 262 SULTR protein sequences from 22 plant species spanning the evolutionary breadth of land plants. This collection includes representatives from two basal lineages, two early-divergent angiosperms, six monocots, and ten dicots. All sequences were extracted from genome assemblies available in Phytozome v13 (Joint Genome Institute) and manually curated, with cross-referencing to additional databases such as NCBI when needed. This dataset provides a valuable resource for reconstructing the evolutionary history of the SULTR family, with particular emphasis on the diversification of SULTR3 transporters in flowering plants. This resource may also support functional annotation, comparative genomics, and structural modeling of sulfate transport proteins.

CBI↗

BMC Caller: a webtool to identify and analyze bacterial microcompartment types in sequence data

Bacterial microcompartments (BMCs) are protein-based organelles found across the bacterial tree of life. They consist of a shell, made of proteins that oligomerize into hexagonally and pentagonally shaped building blocks, that surrounds enzymes constituting a segment of a metabolic pathway. The proteins of the shell are unique to BMCs. They also provide selective permeability; this selectivity is dictated by the requirements of their cargo enzymes. We have recently surveyed the wealth of different BMC types and their occurrence in all available genome sequence data by analyzing and categorizing their components found in chromosomal loci using HMM (Hidden Markov Model) protein profiles. To make this a “do-it yourself” analysis for the public we have devised a webserver, BMC Caller (https://bmc-caller.prl.msu.edu), that compares user input sequences to our HMM profiles, creates a BMC locus visualization, and defines the functional type of BMC, if known. Shell proteins in the input sequence data are also classified according to our function-agnostic naming system and there are links to similar proteins in our database as well as an external link to a structure prediction website to easily generate structural models of the shell proteins, which facilitates understanding permeability properties of the shell. Additionally, the BMC Caller website contains a wealth of information on previously analyzed BMC loci with links to detailed data for each BMC protein and phylogenetic information on the BMC shell proteins. Our tools greatly facilitate BMC type identification to provide the user information about the associated organism’s metabolism and enable discovery of new BMC types by providing a reference database of all currently known examples.

59 BASIC BIOLOGICAL SCIENCES↗

Quantifying Structural Relationships of Metal-Binding Sites Suggests Origins of Biological Electron Transfer

Biological redox reactions drive planetary biogeochemical cycles. Using a novel, structure-guided sequence analysis of proteins, we explored the patterns of evolution of enzymes responsible for these reactions. Our analysis reveals that the folds that bind transition metal–containing ligands have similar structural geometry and amino acid sequences across the full diversity of proteins. Similarity across folds reflects the availability of key transition metals over geological time and strongly suggests that transition metal–ligand binding had a small number of common peptide origins. We observe that structures central to our similarity network come primarily from oxidoreductases, suggesting that ancestral peptides may have also facilitated electron transfer reactions. Last, our results reveal that the earliest biologically functional peptides were likely available before the assembly of fully functional protein domains over 3.8 billion years ago. Thus, life is a special, very complex form of motion of matter, but this form did not always exist, and it is not separated from inorganic nature by an impassable abyss; rather, it arose from inorganic nature as a new property in the process of evolution of the world. We must study the history of this evolution if we want to solve the problem of the origin of life.

Yana Bromberg↗

Chromosome-level genome assembly of Quercus variabilis provides insights into the molecular mechanism of cork thickness

Quercus variabilis is a deciduous woody species with high ecological and economic value and is a major source of cork in East Asia. Cork from thick softwood sheets have higher commercial value than those from thin sheets. It is extremely difficult to genetically improve Q. variabilis to produce high quality softwood due to the lack of genomic information. Here, we present a high-quality chromosomal genome assembly for Q. variabilis with length of 791,89 Mb and 54,606 predicted genes. Comparative analysis of protein sequences of Q. variabilis with 11 other species revealed that specific and expanded gene families were significantly enriched in the "fatty acid biosynthesis" pathway in Q. variabilis, which may contribute to the formation of its unique cork. Additionally, based on weighted correlation network analysis of time-course (i.e., five important developmental ages) gene expression data in thick-cork versus thin-cork genotypes of Q. variabilis, we identified one co-expression gene module associated with the thick-cork trait. Within this co-expression gene module, 10 hub genes were associated with suberin biosynthesis. Furthermore, we identified a total of 198 suberin biosynthesis-related new candidate genes that were up-regulated in trees with a thick cork layer relative to those with a thin cork layer. Also, we found that some genes related to cell expansion and cell division were highly expressed in trees with a thick cork layer. Collectively, our results revealed that two metabolic pathways (i.e., suberin biosynthesis, fatty acid biosynthesis), along with other genes involved in cell expansion, cell division, and transcriptional regulation, were associated with the thick-cork trait in Q. variabilis, providing insights into the molecular basis of cork development and knowledge for informing genetic improvement of cork thickness in Q. variabilis and closely related species.

59 BASIC BIOLOGICAL SCIENCES↗

Energy metric prediction for double insertion mutants via the RoseNet deep learning framework

Studying the structural and functional implications of protein mutations is an important task in computational biology and bioinformatics. We leverage our previously proposed RoseNet neural network architecture to predict energy metrics of proteins with double amino acid insertions or deletions (InDels). We train models on previously generated benchmark datasets containing the exhaustive double InDel mutations for three proteins, as well as an additional three proteins for which ∼145k random mutants, each with two InDels, have been generated. We expand on our previous work by evaluating three additional proteins and analyzing domain features that impact the prediction capabilities of RoseNet. These features include InDels into secondary structures and the solvent accessible surface area (SASA) scores of the residues. We uncover further evidence to support that RoseNet has a higher proficiency of generalizing to unseen residue combinations than unseen insertion positions. We also observe that RoseNet produces higher-quality predictions when inserting into a β-sheet over an α-helix. Additionally, when the insertions fall in an area of high SASA, RoseNet often displays better performance than inserting into areas of low SASA.

59 BASIC BIOLOGICAL SCIENCES↗

Library Screening, In Vivo Confirmation, and Structural and Bioinformatic Analysis of Pentapeptide Sequences as Substrates for Protein Farnesyltransferase

Protein farnesylation is a post-translational modification where a 15-carbon farnesyl isoprenoid is appended to the C-terminal end of a protein by farnesyltransferase (FTase). This process often causes proteins to associate with the membrane and participate in signal transduction pathways. The most common substrates of FTase are proteins that have C-terminal tetrapeptide CaaX box sequences where the cysteine is the site of modification. However, recent work has shown that five amino acid sequences can also be recognized, including the pentapeptides CMIIM and CSLMQ. In this work, peptide libraries were initially used to systematically vary the residues in those two parental sequences using an assay based on Matrix Assisted Laser Desorption Ionization–Mass Spectrometry (MALDI-MS). In addition, 192 pentapeptide sequences from the human proteome were screened using that assay to discover additional extended CaaaX-box motifs. Selected hits from that screening effort were rescreened using an in vivo yeast reporter protein assay. The X-ray crystal structure of CMIIM bound to FTase was also solved, showing that the C-terminal tripeptide of that sequence interacted with the enzyme in a similar manner as the C-terminal tripeptide of CVVM, suggesting that the tripeptide comprises a common structural element for substrate recognition in both tetrapeptide and pentapeptide sequences. Molecular dynamics simulation of CMIIM bound to FTase further shed light on the molecular interactions involved, showing that a putative catalytically competent Zn(II)-thiolate species was able to form. Bioinformatic predictions of tetrapeptide (CaaX-box) reactivity correlated well with the reactivity of pentapeptides obtained from in vivo analysis, reinforcing the importance of the C-terminal tripeptide motif. This analysis provides a structural framework for understanding the reactivity of extended CaaaX-box motifs and a method that may be useful for predicting the reactivity of additional FTase substrates bearing CaaaX-box sequences.

59 BASIC BIOLOGICAL SCIENCES↗

Prevalence and diversity of TAL effector-like proteins in fungal endosymbiotic Mycetohabitans spp.

EndofungalMycetohabitans(formerlyBurkholderia) spp. rely on a type III secretion system to deliver mostly unidentified effector proteins when colonizing their host fungus,Rhizopus microsporus. The one known secreted effector family fromMycetohabitansconsists of homologues of transcription activator-like (TAL) effectors, which are used by plant pathogenicXanthomonasandRalstoniaspp. to activate host genes that promote disease. These ‘BurkholderiaTAL-like (Btl)’ proteins bind corresponding specific DNA sequences in a predictable manner, but their genomic target(s) and impact on transcription in the fungus are unknown. Recent phenotyping of Btl mutants of twoMycetohabitansstrains revealed that the single Btl in oneMycetohabitans endofungorumstrain enhances fungal membrane stress tolerance, while others in aMycetohabitans rhizoxinicastrain promote bacterial colonization of the fungus. The phenotypic diversity underscores the need to assess the sequence diversity and, given that sequence diversity translates to DNA targeting specificity, the functional diversity of Btl proteins. Using a dual approach to maximize capture of Btl protein sequences for our analysis, we sequenced and assembled nineMycetohabitansspp. genomes using long-read PacBio technology and also mined available short-read Illumina fungal–bacterial metagenomes. We show thatbtlgenes are present across diverseMycetohabitansstrains from Mucoromycota fungal hosts yet vary in sequences and predicted DNA binding specificity. Phylogenetic analysis revealed distinct clades of Btl proteins and suggested thatMycetohabitansmight contain more species than previously recognized. Within our data set, Btl proteins were more conserved acrossM. rhizoxinicastrains than acrossM. endofungorum, but there was also evidence of greater overall strain diversity within the latter clade. Overall, the results suggest that Btl proteins contribute to bacterial–fungal symbioses in myriad ways.

Genetics & Heredity↗

Sequence, structure prediction, and epitope analysis of the polymorphic membrane protein family in Chlamydia trachomatis

The polymorphic membrane proteins (Pmps) are a family of autotransporters that play an important role in infection, adhesion and immunity in Chlamydia trachomatis. Here we show that the characteristic GGA(I,L,V) and FxxN tetrapeptide repeats fit into a larger repeat sequence, which correspond to the coils of a large beta-helical domain in high quality structure predictions. Analysis of the protein using structure prediction algorithms provided novel insight to the chlamydial Pmp family of proteins. While the tetrapeptide motifs themselves are predicted to play a structural role in folding and close stacking of the beta-helical backbone of the passenger domain, we found many of the interesting features of Pmps are localized to the side loops jutting out from the beta helix including protease cleavage, host cell adhesion, and B-cell epitopes; while T-cell epitopes are predominantly found in the beta-helix itself. This analysis more accurately defines the Pmp family of Chlamydia and may better inform rational vaccine design and functional studies.

59 BASIC BIOLOGICAL SCIENCES↗

cWINNOWER algorithm for finding fuzzy dna motifs

The cWINNOWER algorithm detects fuzzy motifs in DNA sequences rich in protein-binding signals. A signal is defined as any short nucleotide pattern having up to d mutations differing from a motif of length l. The algorithm finds such motifs if a clique consisting of a sufficiently large number of mutated copies of the motif (i.e., the signals) is present in the DNA sequence. The cWINNOWER algorithm substantially improves the sensitivity of the winnower method of Pevzner and Sze by imposing a consensus constraint, enabling it to detect much weaker signals. We studied the minimum detectable clique size qc as a function of sequence length N for random sequences. We found that qc increases linearly with N for a fast version of the algorithm based on counting three-member sub-cliques. Imposing consensus constraints reduces qc by a factor of three in this case, which makes the algorithm dramatically more sensitive. Our most sensitive algorithm, which counts four-member sub-cliques, needs a minimum of only 13 signals to detect motifs in a sequence of length N = 12,000 for (l, d) = (15, 4). Copyright Imperial College Press.

Evaluation Studies↗

A Re-Evaluation of African Swine Fever Genotypes Based on p72 Sequences Reveals the Existence of Only Six Distinct p72 Groups

The African swine fever virus (ASFV) is currently causing a world-wide pandemic of a highly lethal disease in domestic swine and wild boar. Currently, recombinant ASF live-attenuated vaccines based on a genotype II virus strain are commercially available in Vietnam. With 25 reported ASFV genotypes in the literature, it is important to understand the molecular basis and usefulness of ASFV genotyping, as well as the true significance of genotypes in the epidemiology, transmission, evolution, control, and prevention of ASFV. Historically, genotyping of ASFV was used for the epidemiological tracking of the disease and was based on the analysis of small fragments that represent less than 1% of the viral genome. The predominant method for genotyping ASFV relies on the sequencing of a fragment within the gene encoding the structural p72 protein. Genotype assignment has been accomplished through automated phylogenetic trees or by comparing the target sequence to the most closely related genotyped p72 gene. To evaluate its appropriateness for the classification of genotypes by p72, we reanalyzed all available genomic data for ASFV. We conclude that the majority of p72-based genotypes, when initially created, were neither identified under any specific methodological criteria nor correctly compared with the already existing ASFV genotypes. Based on our analysis of the p72 protein sequences, we propose that the current twenty-five genotypes, created exclusively based on the p72 sequence, should be reduced to only six genotypes. To help differentiate between the new and old genotype classification systems, we propose that Arabic numerals (1, 2, 8, 9, 15, and 23) be used instead of the previously used Roman numerals. Furthermore, we discuss the usefulness of genotyping ASFV isolates based only on the p72 gene sequence.

59 BASIC BIOLOGICAL SCIENCES↗

CRITICA: coding region identification tool invoking comparative analysis

Gene recognition is essential to understanding existing and future DNA sequence data. CRITICA (Coding Region Identification Tool Invoking Comparative Analysis) is a suite of programs for identifying likely protein-coding sequences in DNA by combining comparative analysis of DNA sequences with more common noncomparative methods. In the comparative component of the analysis, regions of DNA are aligned with related sequences from the DNA databases; if the translation of the aligned sequences has greater amino acid identity than expected for the observed percentage nucleotide identity, this is interpreted as evidence for coding. CRITICA also incorporates noncomparative information derived from the relative frequencies of hexanucleotides in coding frames versus other contexts (i.e., dicodon bias). The dicodon usage information is derived by iterative analysis of the data, such that CRITICA is not dependent on the existence or accuracy of coding sequence annotations in the databases. This independence makes the method particularly well suited for the analysis of novel genomes. CRITICA was tested by analyzing the available Salmonella typhimurium DNA sequences. Its predictions were compared with the DNA sequence annotations and with the predictions of GenMark. CRITICA proved to be more accurate than GenMark, and moreover, many of its predictions that would seem to be errors instead reflect problems in the sequence databases. The source code of CRITICA is freely available by anonymous FTP (rdp.life.uiuc.edu in/pub/critica) and on the World Wide Web (http:/(/)rdpwww.life.uiuc.edu).

Non-NASA Center↗

Snekmer: a scalable pipeline for protein sequence fingerprinting based on amino acid recoding

Abstract Motivation The vast expansion of sequence data generated from single organisms and microbiomes has precipitated the need for faster and more sensitive methods to assess evolutionary and functional relationships between proteins. Representing proteins as sets of short peptide sequences (kmers) has been used for rapid, accurate classification of proteins into functional categories; however, this approach employs an exact-match methodology and thus may be limited in terms of sensitivity and coverage. We have previously used similarity groupings, based on the chemical properties of amino acids, to form reduced character sets and recode proteins. This amino acid recoding (AAR) approach simplifies the construction of protein representations in the form of kmer vectors, which can link sequences with distant sequence similarity and provide accurate classification of problematic protein families. Results Here, we describe Snekmer, a software tool for recoding proteins into AAR kmer vectors and performing either (i) construction of supervised classification models trained on input protein families or (ii) clustering for de novo determination of protein families. We provide examples of the operation of the tool against a set of nitrogen cycling families originally collected using both standard hidden Markov models and a larger set of proteins from Uniprot and demonstrate that our method accurately differentiates these sequences in both operation modes. Availability and implementation Snekmer is written in Python using Snakemake. Code and data used in this article, along with tutorial notebooks, are available at http://github.com/PNNL-CompBio/Snekmer under an open-source BSD-3 license. Supplementary information Supplementary data are available at Bioinformatics Advances online.

59 BASIC BIOLOGICAL SCIENCES↗

NGPINT V3: a containerized orchestration Python software for discovery of next-generation protein–protein interactions

Abstract Summary Batch yeast two-hybrid (Y2H) assays, leveraged with next-generation sequencing, have afforded successful innovations for the analysis of protein–protein interactions. NGPINT is a Conda-based software designed to process the millions of raw sequencing reads resulting from Y2H–next-generation interaction screens. Over time, increasing compatibility and dependency issues have prevented clean NGPINT installation and operation. A system-wide update was essential to continue effective use with its companion software, Y2H-SCORES. We present NGPINT V3, a containerized implementation built with both Singularity and Docker, allowing accessibility across virtually any operating system and computing environment. Availability and implementation This update includes streamlined dependencies and container images hosted on Sylabs (https://cloud.sylabs.io/library/schuyler/ngpint/ngpint) and Dockerhub (https://hub.docker.com/r/schuylerds/ngpint), facilitating easier adoption and integration into high-throughput and cloud-computing workflows. Full instructions and software can be also found in the GitHub repository https://github.com/Wiselab2/NGPINT_V3 and Zenodo https://doi.org/10.5281/zenodo.15256036.

Biochemistry & Molecular Biology↗

Sequence Design of Random Heteropolymers as Protein Mimics

Random heteropolymers (RHPs) have been computationally designed and experimentally shown to recapitulate protein-like phase behavior and function. However, unlike proteins, RHP sequences are only statistically defined and cannot be sequenced. Recent developments in reversible-deactivation radical polymerization allowed simulated polymer sequences based on the well-established Mayo–Lewis equation to more accurately reflect ground-truth sequences that are experimentally synthesized. This led to opportunities to perform bioinformatics-inspired analysis on simulated sequences to guide the design, synthesis, and interpretation of RHPs. We compared batches on the order of 10000 simulated RHP sequences that vary by synthetically controllable and measurable RHP characteristics such as chemical heterogeneity and average degree of polymerization. Our analysis spans across 3 levels: segments along a single chain, sequences within a batch, and batch-averaged statistics. We discuss simulator fidelity and highlight the importance of robust segment definition. Examples are presented that demonstrate the use of simulated sequence analysis for in-silico iterative design to mimic protein hydrophobic/hydrophilic segment distributions in RHPs and compare RHP and protein sequence segments to explain experimental results of RHPs that mimic protein function. To facilitate the community use of this workflow, the simulator and analysis modules have been made available through an open source toolkit, the RHPapp.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Purification and sequence analysis of two rat tissue inhibitors of metalloproteinases

Two protein inhibitors of metalloproteinases (TIMP) were isolated from medium conditioned by the clonal rat osteosarcoma line UMR 106-01. Initial purification of both a 30-kDa inhibitor and a 20-kDa inhibitor was accomplished using heparin-Sepharose chromatography with dextran sulfate elution followed by DEAE-Sepharose and CM-Sepharose chromatography. Purification of the 20-kDa inhibitor to homogeneity was completed with reverse-phase high-performance liquid chromatography. The 20-kDa inhibitor was identified as rat TIMP-2. The 30-kDa inhibitor, although not purified to homogeneity, was identified as rat TIMP-1. Amino terminal amino acid sequence analysis of the 30-kDa inhibitor demonstrated 86% identity to human TIMP-1 for the first 22 amino acids while the sequence of the 20-kDa inhibitor was identical to that of human TIMP-2 for the first 22 residues. Treatment with peptide:N-glycosidase F indicated that the 30-kDa rat inhibitor is glycosylated while the 20-kDa inhibitor is apparently unglycosylated. Inhibition of both rat and human interstitial collagenase by rat TIMP-2 was stoichiometric, with a 1:1 molar ratio required for complete inhibition. Exposure of UMR 106-01 cells to 10(-7) M parathyroid hormone resulted in approximately a 40% increase in total inhibitor production over basal levels.

NASA Discipline Musculoskeletal↗