Search NASA⌕ Search

SEARCH · Search NASA

Results for “Protein sequence analysis”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

Quantifying Structural Relationships of Metal-Binding Sites Suggests Origins of Biological Electron Transfer

Biological redox reactions drive planetary biogeochemical cycles. Using a novel, structure-guided sequence analysis of proteins, we explored the patterns of evolution of enzymes responsible for these reactions. Our analysis reveals that the folds that bind transition metal–containing ligands have similar structural geometry and amino acid sequences across the full diversity of proteins. Similarity across folds reflects the availability of key transition metals over geological time and strongly suggests that transition metal–ligand binding had a small number of common peptide origins. We observe that structures central to our similarity network come primarily from oxidoreductases, suggesting that ancestral peptides may have also facilitated electron transfer reactions. Last, our results reveal that the earliest biologically functional peptides were likely available before the assembly of fully functional protein domains over 3.8 billion years ago. Thus, life is a special, very complex form of motion of matter, but this form did not always exist, and it is not separated from inorganic nature by an impassable abyss; rather, it arose from inorganic nature as a new property in the process of evolution of the world. We must study the history of this evolution if we want to solve the problem of the origin of life.

Yana Bromberg↗

cWINNOWER algorithm for finding fuzzy dna motifs

The cWINNOWER algorithm detects fuzzy motifs in DNA sequences rich in protein-binding signals. A signal is defined as any short nucleotide pattern having up to d mutations differing from a motif of length l. The algorithm finds such motifs if a clique consisting of a sufficiently large number of mutated copies of the motif (i.e., the signals) is present in the DNA sequence. The cWINNOWER algorithm substantially improves the sensitivity of the winnower method of Pevzner and Sze by imposing a consensus constraint, enabling it to detect much weaker signals. We studied the minimum detectable clique size qc as a function of sequence length N for random sequences. We found that qc increases linearly with N for a fast version of the algorithm based on counting three-member sub-cliques. Imposing consensus constraints reduces qc by a factor of three in this case, which makes the algorithm dramatically more sensitive. Our most sensitive algorithm, which counts four-member sub-cliques, needs a minimum of only 13 signals to detect motifs in a sequence of length N = 12,000 for (l, d) = (15, 4). Copyright Imperial College Press.

Evaluation Studies↗

CRITICA: coding region identification tool invoking comparative analysis

Gene recognition is essential to understanding existing and future DNA sequence data. CRITICA (Coding Region Identification Tool Invoking Comparative Analysis) is a suite of programs for identifying likely protein-coding sequences in DNA by combining comparative analysis of DNA sequences with more common noncomparative methods. In the comparative component of the analysis, regions of DNA are aligned with related sequences from the DNA databases; if the translation of the aligned sequences has greater amino acid identity than expected for the observed percentage nucleotide identity, this is interpreted as evidence for coding. CRITICA also incorporates noncomparative information derived from the relative frequencies of hexanucleotides in coding frames versus other contexts (i.e., dicodon bias). The dicodon usage information is derived by iterative analysis of the data, such that CRITICA is not dependent on the existence or accuracy of coding sequence annotations in the databases. This independence makes the method particularly well suited for the analysis of novel genomes. CRITICA was tested by analyzing the available Salmonella typhimurium DNA sequences. Its predictions were compared with the DNA sequence annotations and with the predictions of GenMark. CRITICA proved to be more accurate than GenMark, and moreover, many of its predictions that would seem to be errors instead reflect problems in the sequence databases. The source code of CRITICA is freely available by anonymous FTP (rdp.life.uiuc.edu in/pub/critica) and on the World Wide Web (http:/(/)rdpwww.life.uiuc.edu).

Non-NASA Center↗

Purification and sequence analysis of two rat tissue inhibitors of metalloproteinases

Two protein inhibitors of metalloproteinases (TIMP) were isolated from medium conditioned by the clonal rat osteosarcoma line UMR 106-01. Initial purification of both a 30-kDa inhibitor and a 20-kDa inhibitor was accomplished using heparin-Sepharose chromatography with dextran sulfate elution followed by DEAE-Sepharose and CM-Sepharose chromatography. Purification of the 20-kDa inhibitor to homogeneity was completed with reverse-phase high-performance liquid chromatography. The 20-kDa inhibitor was identified as rat TIMP-2. The 30-kDa inhibitor, although not purified to homogeneity, was identified as rat TIMP-1. Amino terminal amino acid sequence analysis of the 30-kDa inhibitor demonstrated 86% identity to human TIMP-1 for the first 22 amino acids while the sequence of the 20-kDa inhibitor was identical to that of human TIMP-2 for the first 22 residues. Treatment with peptide:N-glycosidase F indicated that the 30-kDa rat inhibitor is glycosylated while the 20-kDa inhibitor is apparently unglycosylated. Inhibition of both rat and human interstitial collagenase by rat TIMP-2 was stoichiometric, with a 1:1 molar ratio required for complete inhibition. Exposure of UMR 106-01 cells to 10(-7) M parathyroid hormone resulted in approximately a 40% increase in total inhibitor production over basal levels.

NASA Discipline Musculoskeletal↗

Genes encoding calmodulin-binding proteins in the Arabidopsis genome

Analysis of the recently completed Arabidopsis genome sequence indicates that approximately 31% of the predicted genes could not be assigned to functional categories, as they do not show any sequence similarity with proteins of known function from other organisms. Calmodulin (CaM), a ubiquitous and multifunctional Ca(2+) sensor, interacts with a wide variety of cellular proteins and modulates their activity/function in regulating diverse cellular processes. However, the primary amino acid sequence of the CaM-binding domain in different CaM-binding proteins (CBPs) is not conserved. One way to identify most of the CBPs in the Arabidopsis genome is by protein-protein interaction-based screening of expression libraries with CaM. Here, using a mixture of radiolabeled CaM isoforms from Arabidopsis, we screened several expression libraries prepared from flower meristem, seedlings, or tissues treated with hormones, an elicitor, or a pathogen. Sequence analysis of 77 positive clones that interact with CaM in a Ca(2+)-dependent manner revealed 20 CBPs, including 14 previously unknown CBPs. In addition, by searching the Arabidopsis genome sequence with the newly identified and known plant or animal CBPs, we identified a total of 27 CBPs. Among these, 16 CBPs are represented by families with 2-20 members in each family. Gene expression analysis revealed that CBPs and CBP paralogs are expressed differentially. Our data suggest that Arabidopsis has a large number of CBPs including several plant-specific ones. Although CaM is highly conserved between plants and animals, only a few CBPs are common to both plants and animals. Analysis of Arabidopsis CBPs revealed the presence of a variety of interesting domains. Our analyses identified several hypothetical proteins in the Arabidopsis genome as CaM targets, suggesting their involvement in Ca(2+)-mediated signaling networks.

NASA Discipline Plant Biology↗

A novel beta-glucosidase from the cell wall of maize (Zea mays L.): rapid purification and partial characterization

Plants have a variety of glycosidic conjugates of hormones, defense compounds, and other molecules that are hydrolyzed by beta-glucosidases (beta-D-glucoside glucohydrolases, E.C. 3.2.1.21). Workers have reported several beta-glucosidases from maize (Zea mays L.; Poaceae), but have localized them mostly by indirect means. We have purified and partly characterized a 58-Ku beta-glucosidase from maize, which we conclude from a partial sequence analysis, from kinetic data, and from its localization is not identical to any of those already reported. A monoclonal antibody, mWP 19, binds this enzyme, and localizes it in the cell walls of maize coleoptiles. An earlier report showed that mWP19 inhibits peroxidase activity in crude cell wall extracts and can immunoprecipitate peroxidase activity from these extracts, yet purified preparations of the 58 Ku protein had little or no peroxidase activity. The level of sequence similarity between beta-glucosidases and peroxidases makes it unlikely that these enzymes share epitopes in common. Contrary to a previous conclusion, these results suggest that the enzyme recognized by mWP19 is not a peroxidase, but there is a wall peroxidase closely associated with the 58 Ku beta-glucosidase in crude preparations. Other workers also have co-purified distinct proteins with beta-glucosidases. We found no significant charge in the level of immunodetectable beta-glucosidase in mesocotyls or coleoptiles that precedes the red light-induced changes in the growth rate of these tissues.

Non-NASA Center↗

Mutational analysis of photosystem I polypeptides in the cyanobacterium Synechocystis sp. PCC 6803. Targeted inactivation of psaI reveals the function of psaI in the structural organization of psaL

We cloned, characterized, and inactivated the psaI gene encoding a 4-kDa hydrophobic subunit of photosystem I from the cyanobacterium Synechocystis sp. PCC 6803. The psaI gene is located 90 base pairs downstream from psaL, and is transcribed on 0.94- and 0.32-kilobase transcripts. To identify the function of PsaI, we generated a cyanobacterial strain in which psaI has been interrupted by a gene for chloramphenicol resistance. The wild-type and the mutant cells showed comparable rates of photoautotrophic growth at 25 degrees C. However, the mutant cells grew slower and contained less chlorophyll than the wild-type cells, when grown at 40 degrees C. The PsaI-less membranes from cells grown at either temperature showed a small decrease in NADP+ photoreduction rate when compared to the wild-type membranes. Inactivation of psaI led to an 80% decrease in the PsaL level in the photosynthetic membranes and to a complete loss of PsaL in the purified photosystem I preparations, but had little effect on the accumulation of other photosystem I subunits. Upon solubilization with nonionic detergents, photosystem I trimers could be obtained from the wild-type, but not from the PsaI-less membranes. The PsaI-less photosystem I monomers did not contain detectable levels of PsaL. Therefore, a structural interaction between PsaL and PsaI may stabilize the association of PsaL with the photosystem I core. PsaL in the wild-type and PsaI-less membranes showed equal resistance to removal by chaotropic agents. However, PsaL in the PsaI-less strain exhibited an increased susceptibility to proteolysis. From these data, we conclude that PsaI has a crucial role in aiding normal structural organization of PsaL within the photosystem I complex and the absence of PsaI alters PsaL organization, leading to a small, but physiologically significant, defect in photosystem I function.

NASA Discipline Cell Biology↗

Characterization of the proteins comprising the integral matrix of Strongylocentrotus purpuratus embryonic spicules

In the present study, we enumerate and characterize the proteins that comprise the integral spicule matrix of the Strongylocentrotus purpuratus embryo. Two-dimensional gel electrophoresis of [35S]methionine radiolabeled spicule matrix proteins reveals that there are 12 strongly radiolabeled spicule matrix proteins and approximately three dozen less strongly radiolabeled spicule matrix proteins. The majority of the proteins have acidic isoelectric points; however, there are several spicule matrix proteins that have more alkaline isoelectric points. Western blotting analysis indicates that SM50 is the spicule matrix protein with the most alkaline isoelectric point. In addition, two distinct SM30 proteins are identified in embryonic spicules, and they have apparent molecular masses of approximately 43 and 46 kDa. Comparisons between embryonic spicule matrix proteins and adult spine integral matrix proteins suggest that the embryonic 43-kDa SM30 protein is an embryonic isoform of SM30. An adult 49-kDa spine matrix protein is also identified as a possible adult isoform of SM30. Analysis of the SM30 amino acid sequences indicates that a portion of SM30 proteins is very similar to the carbohydrate recognition domain of C-type lectin proteins.

Non-NASA Center↗

Functional and evolutionary relationships between bacteriorhodopsin and halorhodopsin in the archaebacterium, halobacterium halobium

The archaebacteria occupy a unique place in phylogenetic trees constructed from analyses of sequences from key informational macromolecules, and their study continues to yield interesting ideas on the early evolution and divergence of biological forms. It is now known that the halobacteria among these species contain various retinal-proteins, resembling eukaryotic rhodopsins, but with different functions. Two of these pigments, located in the cytoplasmic membranes of the bacteria, are bacteriorhodopsin (a light-driven proton pump) and halorhodopsin (a light-driven chloride pump). Comparison of these systems is expected to reveal structure/function relationships in these simple (primitive?) energy transducing membrane components and evolutionary relationships which had produced the structural features which allow the divergent functions. Findings indicate that very different primary structures are needed for these proteins to accomplish their different functions. Indeed, analysis of partial amino acid sequences from halo-opsin shows already that few if any long segments exist which are homologous to bacterio-opsin. Either these proteins diverged a very long time ago to allow for the observed differences, or the evolutionary clock in the halobacteria runs faster than usual.

Lanyi, J. K.↗

Streptococcus pneumoniae PstS production is phosphate responsive and enhanced during growth in the murine peritoneal cavity

Differential display-PCR (DDPCR) was used to identify a Streptococcus pneumoniae gene with enhanced transcription during growth in the murine peritoneal cavity. Northern dot blot analysis and comparative densitometry confirmed a 1.8-fold increase in expression of the encoded sequence following murine peritoneal culture (MPC) versus laboratory culture or control culture (CC). Sequencing and basic local alignment search tool analysis identified the DDPCR fragment as pstS, the phosphate-binding protein of a high-affinity phosphate uptake system. PCR amplification of the complete pstS gene followed by restriction analysis and sequencing suggests a high level of conservation between strains and serotypes. Quantitative immunodot blotting using antiserum to recombinant PstS (rPstS) demonstrated an approximately twofold increase in PstS production during MPC from that during CCs, a finding consistent with the low levels of phosphate observed in the peritoneum. Moreover, immunodot blot and Northern analysis demonstrated phosphate-dependent production of PstS in six of seven strains examined. These results identify pstS expression as responsive to the MPC environment and extracellular phosphate concentrations. Presently, it remains unclear if phosphate concentrations in vivo contribute to the regulation of pstS. Finally, polyclonal antiserum to rPstS did not inhibit growth of the pneumococcus in vitro, suggesting that antibodies do not block phosphate uptake; moreover, vaccination of mice with rPstS did not protect against intraperitoneal challenge as assessed by the 50% lethal dose.

Non-NASA Center↗

Sequence data - Magnitude and implications of some ambiguities.

A stochastic model is applied to the divergence of the horse-pig lineage from a common ansestor in terms of the alpha and beta chains of hemoglobin and fibrinopeptides. The results are compared with those based on the minimum mutation distance model of Fitch (1972). Buckwheat and cauliflower cytochrome c sequences are analyzed to demonstrate their ambiguities. A comparative analysis of evolutionary rates for various proteins of horses and pigs shows that errors of considerable magnitude are introduced by Glx and Asx ambiguities into evolutionary conclusions drawn from sequences of incompletely analyzed proteins.

Holmquist, R.↗

Structural characterization and regulatory element analysis of the heart isoform of cytochrome c oxidase VIa

In order to investigate the mechanism(s) governing the striated muscle-specific expression of cytochrome c oxidase VIaH we have characterized the murine gene and analyzed its transcriptional regulatory elements in skeletal myogenic cell lines. The gene is single copy, spans 689 base pairs (bp), and is comprised of three exons. The 5'-ends of transcripts from the gene are heterogeneous, but the most abundant transcript includes a 5'-untranslated region of 30 nucleotides. When fused to the luciferase reporter gene, the 3.5-kilobase 5'-flanking region of the gene directed the expression of the heterologous protein selectively in differentiated Sol8 cells and transgenic mice, recapitulating the pattern of expression of the endogenous gene. Deletion analysis identified a 300-bp fragment sufficient to direct the myotube-specific expression of luciferase in Sol8 cells. The region lacks an apparent TATA element, and sequence motifs predicted to bind NRF-1, NRF-2, ox-box, or PPAR factors known to regulate other nuclear genes encoding mitochondrial proteins are not evident. Mutational analysis, however, identified two cis-elements necessary for the high level expression of the reporter protein: a MEF2 consensus element at -90 to -81 bp and an E-box element at -147 to -142 bp. Additional E-box motifs at closely located positions were mutated without loss of transcriptional activity. The dependence of transcriptional activation of cytochrome c oxidase VIaH on cis-elements similar to those found in contractile protein genes suggests that the striated muscle-specific expression is coregulated by mechanisms that control the lineage-specific expression of several contractile and cytosolic proteins.

Non-NASA Center↗

Gene Fusion: A Genome Wide Survey

As a well known fact, organisms form larger and complex multimodular (composite or chimeric) and mostly multi-functional proteins through gene fusion of two or more individual genes which have independent evolution histories and functions. We call each of these components a module. The existence of multimodular proteins may improves the efficiency in gene regulation and in cellular functions, and thus may give the host organism advantages in adaptation to environments. Analysis of all gene fusions in present-day organisms should allow us to examine the patterns of gene fusion in context with cellular functions, to trace back the evolution processes from the ancient smaller and uni-functional proteins to the present-day larger and complex multi-functional proteins, and to estimate the minimal number of ancestor proteins that existed in the last common ancestor for all life on earth. Although many multimodular proteins have been experimentally known, identification of gene fusion events systematically at genome scale had not been possible until recently when large number of completed genome sequences have been becoming available. In addition, technical difficulties for such analysis also exist due to the complexity of this biological and evolutionary process. We report from this study a new strategy to computationally identify multimodular proteins using completed genome sequences and the results surveyed from 22 organisms with the data from over 40 organisms to be presented during the meeting. Additional information is contained in the original extended abstract.

Liang, Ping↗

Alterations of p53 in tumorigenic human bronchial epithelial cells correlate with metastatic potential

The cellular and molecular mechanisms of radiation-induced lung cancer are not known. In the present study, alterations of p53 in tumorigenic human papillomavirus-immortalized human bronchial epithelial (BEP2D) cells induced by a single low dose of either alpha-particles or 1 GeV/nucleon (56)Fe were analyzed by PCR-single-stranded conformation polymorphism (SSCP) coupled with sequencing analysis and immunoprecipitation assay. A total of nine primary and four secondary tumor cell lines, three of which were metastatic, together with the parental BEP2D and primary human bronchial epithelial (NHBE) cells were studied. The immunoprecipitation assay showed overexpression of mutant p53 proteins in all the tumor lines but not in NHBE and BEP2D cells. PCR-SSCP and sequencing analysis found band shifts and gene mutations in all four of the secondary tumors. A G-->T transversion in codon 139 in exon 5 that replaced Lys with Asn was detected in two tumor lines. One mutation each, involving a G-->T transversion in codon 215 in exon 6 (Ser-->lle) and a G-->A transition in codon 373 in exon 8 (Arg-->His), was identified in the remaining two secondary tumors. These results suggest that p53 alterations correlate with tumorigenesis in the BEP2D cell model and that mutations in the p53 gene may be indicative of metastatic potential.

Non-NASA Center↗

Sequence, overproduction and purification of Vibrio proteolyticus ribosomal protein L18 for in vitro and in vivo studies

A strategy suggested by comparative genomic studies was used to amplify the entire Vibrio proteolyticus (Vp) gene for ribosomal protein L18. Vp L18 and its flanking regions were sequenced and compared with the deduced amino acid (aa) sequences of other known L18 proteins. A 26-aa residue segment at the carboxy terminus contains many strongly conserved residues and may be critical for the L18 interaction with 5S rRNA. This approach should allow rapid characterization of L18 from large numbers of bacteria. Both Vp L18 and Escherichia coli (Ec) L18 were overproduced and purified using a T7 expression vector which fuses an N-terminal peptide segment (His-tag) containing 6 histidine residues to the recombinant protein. The purified fusion proteins, Vp His::L18 and Ec His::L18, were both found to bind to either the Vp 5S or Ec 5S rRNAs in vitro. Vp His::L18 protein was also shown to incorporate into Ec ribosomes in vivo. This His-tag strategy likely will have general applicability for the study of ribosomal proteins in vitro and in vivo.

Non-NASA Center↗