Search NASASearch

SEARCH · Search NASA

Results for “Sequence Alignment”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

Implied alignment: a synapomorphy-based multiple-sequence alignment method and its use in cladogram search

A method to align sequence data based on parsimonious synapomorphy schemes generated by direct optimization (DO; earlier termed optimization alignment) is proposed. DO directly diagnoses sequence data on cladograms without an intervening multiple-alignment step, thereby creating topology-specific, dynamic homology statements. Hence, no multiple-alignment is required to generate cladograms. Unlike general and globally optimal multiple-alignment procedures, the method described here, implied alignment (IA), takes these dynamic homologies and traces them back through a single cladogram, linking the unaligned sequence positions in the terminal taxa via DO transformation series. These "lines of correspondence" link ancestor-descendent states and, when displayed as linearly arrayed columns without hypothetical ancestors, are largely indistinguishable from standard multiple alignment. Since this method is based on synapomorphy, the treatment of certain classes of insertion-deletion (indel) events may be different from that of other alignment procedures. As with all alignment methods, results are dependent on parameter assumptions such as indel cost and transversion:transition ratios. Such an IA could be used as a basis for phylogenetic search, but this would be questionable since the homologies derived from the implied alignment depend on its natal cladogram and any variance, between DO and IA + Search, due to heuristic approach. The utility of this procedure in heuristic cladogram searches using DO and the improvement of heuristic cladogram cost calculations are discussed. c2003 The Willi Hennig Society. Published by Elsevier Science (USA). All rights reserved.

Non-NASA Center

Size and Structure of the Sequence Space of Repeat Proteins

The coding space of protein sequences is shaped by evolutionary constraints set by requirements of function and stability. We show that the coding space of a given protein family— the total number of sequences in that family—can be estimated using models of maximum entropy trained on multiple sequence alignments of naturally occurring amino acid sequences. We analyzed and calculated the size of three abundant repeat proteins families, whose members are large proteins made of many repetitions of conserved portions of *30 amino acids. While amino acid conservation at each position of the alignment explains most of the reduction of diversity relative to completely random sequences, we found that correlations between amino acid usage at different positions significantly impact that diversity. We quantified the impact of different types of correlations, functional and evolutionary, on sequence diversity. Analysis of the detailed structure of the coding space of the families revealed a rugged landscape, with many local energy minima of varying sizes with a hierarchical structure, reminiscent of frustrated energy landscapes of spin glass in physics. This clustered structure indicates a multiplicity of subtypes within each family and suggests new strategies for protein design.

Jacopo Marchi

Cloning and characterization of ftsZ and pyrF from the archaeon Thermoplasma acidophilum

To characterize cytoskeletal components of archaea, the ftsZ gene from Thermoplasma acidophilum was cloned and sequenced. In T. acidophilum ftsZ, which is involved in cell division, was found to be in an operon with the pyrF gene, which encodes orotidine-5'-monophosphate decarboxylase (ODC), an essential enzyme in pyrimidine biosynthesis. Both ftsZ and pyrF from T. acidophilum were expressed in Escherichia coli and formed functional proteins. FtsZ expression in wild-type E. coli resulted in the filamentous phenotype characteristic of ftsZ mutants. T. acidophilum pyrF expression in an E. coli mutant lacking pyrF complemented the mutation and rescued the strain. Sequence alignments of ODCs from archaea, bacteria, and eukarya reveal five conserved regions, two of which have homology to 3-hexulose-6-phosphate synthase (HPS), suggesting a common substrate recognition and binding motif. Copyright 2000 Academic Press.

Bacterial Proteins/chemistry/genetics/metabolism

Origins of prokaryotes, eukaryotes, mitochondria, and chloroplasts

A computer branching model is used to analyze cellular evolution. Attention is given to certain key amino acids and nucleotide residues (ferredoxin, 5s ribosomal RNA, and c-type cytochromes) because of their commonality over a wide variety of cell types. Each amino acid or nucleotide residue is a sequence in an inherited biological trait; and the branching method is employed to align sequences so that changes reflect substitution of one residue for another. Based on the computer analysis, the symbiotic theory of cellular evolution is considered the most probable. This theory holds that organelles, e.g., mitochondria and chloroplasts invaded larger bodies, e.g., bacteria, and combined functions to form eucaryotic cells.

Schwartz, R. M.

Search-based optimization

The problem of determining the minimum cost hypothetical ancestral sequences for a given cladogram is known to be NP-complete (Wang and Jiang, 1994). Traditionally, point estimations of hypothetical ancestral sequences have been used to gain heuristic, upper bounds on cladogram cost. These include procedures with such diverse approaches as non-additive optimization of multiple sequence alignment, direct optimization (Wheeler, 1996), and fixed-state character optimization (Wheeler, 1999). A method is proposed here which, by extending fixed-state character optimization, replaces the estimation process with a search. This form of optimization examines a diversity of potential state solutions for cost-efficient hypothetical ancestral sequences and can result in greatly more parsimonious cladograms. Additionally, such an approach can be applied to other NP-complete phylogenetic optimization problems such as genomic break-point analysis. c2003 The Willi Hennig Society. Published by Elsevier Science (USA). All rights reserved.

Non-NASA Center

Two Strategies for Microbial Production of an Industrial Enzyme-Alpha-Amylase

Extremophiles are microorganisms that thrive in, from an anthropocentric view, extreme environments including hot springs, soda lakes and arctic water. This ability of survival at extreme conditions has rendered extremophiles to be of interest in astrobiology, evolutionary biology as well as in industrial applications. Of particular interest to the biotechnology industry are the biological catalysts of the extremophiles, the extremozymes, whose unique stabilities at extreme conditions make them potential sources of novel enzymes in industrial applications. There are two major approaches to microbial enzyme production. This entails enzyme isolation directly from the natural host or creating a recombinant expression system whereby the targeted enzyme can be overexpressed in a mesophilic host. We are employing both methods in the effort to produce alpha-amylases from a hyperthermophilic archaeon (Thermococcus) isolated from a hydrothermal vent in the Atlantic Ocean, as well as from alkaliphilic bacteria (Bacillus) isolated from a soda lake in Tanzania. Alpha-amylases catalyze the hydrolysis of internal alpha-1,4-glycosidic linkages in starch to produce smaller sugars. Thermostable alpha-amylases are used in the liquefaction of starch for production of fructose and glucose syrups, whereas alpha-amylases stable at high pH have potential as detergent additives. The alpha-amylase encoding gene from Thermococcus was PCR amplified using carefully designed primers and analyzed using bioinformatics tools such as BLAST and Multiple Sequence Alignment for cloning and expression in E.coli. Four strains of Bacillus were grown in alkaline starch-enriched medium of which the culture supernatant was used as enzyme source. Amylolytic activity was detected using the starch-iodine method.

Bernhardsdotter, Eva C. M. J.

The Potato virus X TGBp3 protein associates with the ER network for virus cell-to-cell movement

Potato virus X (PVX) TGBp3 is required for virus cell-to-cell movement. Cell-to-cell movement of TGBp3 was studied using biolistic bombardment of plasmids expressing GFP:TGBp3. TGBp3 moves between cells in Nicotiana benthamiana, but requires TGBp1 to move in N. tabacum leaves. In tobacco leaves GFP:TGBp3 accumulated in a pattern resembling the endoplasmic reticulum (ER). To determine if the ER network is important for GFP:TGBp3 and for PVX cell-to-cell movement, a single mutation inhibiting membrane binding of TGBp3 was introduced into GFP:TGBp3 and into PVX. This mutation disrupted movement of GFP:TGBp3 and PVX. Brefeldin A, which disrupts the ER network, also inhibited GFP:TGBp3 movement in both Nicotiana species. Two deletion mutations, that do not affect membrane binding, hindered GFP:TGBp3 and PVX cell-to-cell movement. Plasmids expressing GFP:TGBp2 and GFP:TGBp3 were bombarded to several other PVX hosts and neither protein moved between adjacent cells. In most hosts, TGBp2 or TGBp3 cannot move cell-to-cell.

NASA Discipline Plant Biology

An amphioxus nodal gene (AmphiNodal) with early symmetrical expression in the organizer and mesoderm and later asymmetrical expression associated with left-right axis formation

The full-length sequence and zygotic expression of an amphioxus nodal gene are described. Expression is first detected in the early gastrula just within the dorsal lip of the blastopore in a region of hypoblast that is probably comparable with the vertebrate Spemann's organizer. In the late gastrula and early neurula, expression remains bilaterally symmetrical, limited to paraxial mesoderm and immediately overlying regions of the neural plate. Later in the neurula stage, all neural expression disappears, and mesodermal expression disappears from the right side. All along the left side of the neurula, mesodermal expression spreads into the left side of the gut endoderm. Soon thereafter, all expression is down-regulated except near the anterior and posterior ends of the animal, where transcripts are still found in the mesoderm and endoderm on the left side. At this time, expression also begins in the ectoderm on the left side of the head, in the region where the mouth later forms. These results suggest that amphioxus and vertebrate nodal genes play evolutionarily conserved roles in establishing Spemann's organizer, patterning the mesoderm rostrocaudally and setting up the asymmetrical left-right axis of the body.

Non-NASA Center

The western red cedar (Thuja plicata) 8-8' DIRIGENT family displays diverse expression patterns and conserved monolignol coupling specificity

The isolation and characterization of a multigene family of the first class of dirigent proteins (namely that mainly involved in 8-8' coupling leading to (+)-pinoresinol in this case) is reported, this comprising of nine western red cedar (Thuja plicata) DIRIGENT genes (DIR1-9) of 72-99.5% identity to each other. Their corresponding cDNA clones had coding regions for 180-183 amino acids with each having a predicted molecular mass of ca. 20 kDa including the signal peptide. Real time-PCR established that the DIRIGENT isovariants were differentially expressed during growth and development of T. plicata (P < 0.05). The phylogenetic relationships and the rates and patterns of nucleotide substitution suggest that the DIRIGENT gene may have evolved via paralogous expansion at an early stage of vascular plant diversification. Thereafter, western red cedar paralogues have maintained an high homogeneity presumably via a concerted evolutionary mode. This, in turn, is assumed to be the driving force for the differential formation of 8-8'-linked pinoresinol derived (poly)lignans in the needles, stems, bark and branches, as well as for massive accumulation of 8-8'-linked plicatic acid-derived (poly)lignans in heartwood.

Non-NASA Center

Role of hypoxia-inducible factor-1 in transcriptional activation of ceruloplasmin by iron deficiency

A role of the copper protein ceruloplasmin (Cp) in iron metabolism is suggested by its ferroxidase activity and by the tissue iron overload in hereditary Cp deficiency patients. In addition, plasma Cp increases markedly in several conditions of anemia, e.g. iron deficiency, hemorrhage, renal failure, sickle cell disease, pregnancy, and inflammation. However, little is known about the cellular and molecular mechanism(s) involved. We have reported that iron chelators increase Cp mRNA expression and protein synthesis in human hepatocarcinoma HepG2 cells. Furthermore, we have shown that the increase in Cp mRNA is due to increased rate of transcription. We here report the results of new studies designed to elucidate the molecular mechanism underlying transcriptional activation of Cp by iron deficiency. The 5'-flanking region of the Cp gene was cloned from a human genomic library. A 4774-base pair segment of the Cp promoter/enhancer driving a luciferase reporter was transfected into HepG2 or Hep3B cells. Iron deficiency or hypoxia increased luciferase activity by 5-10-fold compared with untreated cells. Examination of the sequence showed three pairs of consensus hypoxia-responsive elements (HREs). Deletion and mutation analysis showed that a single HRE was necessary and sufficient for gene activation. The involvement of hypoxia-inducible factor-1 (HIF-1) was shown by gel-shift and supershift experiments that showed HIF-1alpha and HIF-1beta binding to a radiolabeled oligonucleotide containing the Cp promoter HRE. Furthermore, iron deficiency (and hypoxia) did not activate Cp gene expression in Hepa c4 hepatoma cells deficient in HIF-1beta, as shown functionally by the inactivity of a transfected Cp promoter-luciferase construct and by the failure of HIF-1 to bind the Cp HRE in nuclear extracts from these cells. These results are consistent with in vivo findings that iron deficiency increases plasma Cp and provides a molecular mechanism that may help to understand these observations.

NASA Discipline Cardiopulmonary

Chimeric calcium/calmodulin-dependent protein kinase in tobacco: differential regulation by calmodulin isoforms

cDNA clones of chimeric Ca2+/calmodulin-dependent protein kinase (CCaMK) from tobacco (TCCaMK-1 and TCCaMK-2) were isolated and characterized. The polypeptides encoded by TCCaMK-1 and TCCaMK-2 have 15 different amino acid substitutions, yet they both contain a total of 517 amino acids. Northern analysis revealed that CCaMK is expressed in a stage-specific manner during anther development. Messenger RNA was detected when tobacco bud sizes were between 0.5 cm and 1.0 cm. The appearance of mRNA coincided with meiosis and became undetectable at later stages of anther development. The reverse polymerase chain reaction (RT-PCR) amplification assay using isoform-specific primers showed that both of the CCaMK mRNAs were expressed in anther with similar expression patterns. The CCaMK protein expressed in Escherichia coli showed Ca2+-dependent autophosphorylation and Ca2+/calmodulin-dependent substrate phosphorylation. Calmodulin isoforms (PCM1 and PCM6) had differential effects on the regulation of autophosphorylation and substrate phosphorylation of tobacco CCaMK, but not lily CCaMK. The evolutionary tree of plant serine/threonine protein kinases revealed that calmodulin-dependent kinases form one subgroup that is distinctly different from Ca2+-dependent protein kinases (CDPKs) and other serine/threonine kinases in plants.

NASA Discipline Plant Biology

Archaeal translation initiation revisited: the initiation factor 2 and eukaryotic initiation factor 2B alpha-beta-delta subunit families

As the amount of available sequence data increases, it becomes apparent that our understanding of translation initiation is far from comprehensive and that prior conclusions concerning the origin of the process are wrong. Contrary to earlier conclusions, key elements of translation initiation originated at the Universal Ancestor stage, for homologous counterparts exist in all three primary taxa. Herein, we explore the evolutionary relationships among the components of bacterial initiation factor 2 (IF-2) and eukaryotic IF-2 (eIF-2)/eIF-2B, i.e., the initiation factors involved in introducing the initiator tRNA into the translation mechanism and performing the first step in the peptide chain elongation cycle. All Archaea appear to posses a fully functional eIF-2 molecule, but they lack the associated GTP recycling function, eIF-2B (a five-subunit molecule). Yet, the Archaea do posses members of the gene family defined by the (related) eIF-2B subunits alpha, beta, and delta, although these are not specifically related to any of the three eukaryotic subunits. Additional members of this family also occur in some (but by no means all) Bacteria and even in some eukaryotes. The functional significance of the other members of this family is unclear and requires experimental resolution. Similarly, the occurrence of bacterial IF-2-like molecules in all Archaea and in some eukaryotes further complicates the picture of translation initiation. Overall, these data lend further support to the suggestion that the rudiments of translation initiation were present at the Universal Ancestor stage.

NASA Discipline Exobiology

The beetle Tribolium castaneum has a fushi tarazu homolog expressed in stripes during segmentation

The genetic control of embryonic organization is far better understood for the fruit fly Drosophila melanogaster than for any other metazoan. A gene hierarchy acts during oogenesis and embryogenesis to regulate the establishment of segmentation along the anterior-posterior axis, and homeotic selector genes define developmental commitments within each parasegmental unit delineated. One of the most intensively studied Drosophila segmentation genes is fushi tarazu (ftz), a pair-rule gene expressed in stripes that is important for the establishment of the parasegmental boundaries. Although ftz is flanked by homeotic selector genes conserved throughout the metazoa, there is no evidence that it was part of the ancestral homeotic complex, and it has been unclear when the gene arose and acquired a role in segmentation. We show here that the beetle Tribolium castaneum has a ftz homolog located in its Homeotic complex and expressed in a pair-rule fashion, albeit in a register differing from that of the fly gene. These and other observations demonstrate that a ftz gene preexisted the radiation of holometabolous insects and suggest that it has a role in beetle embryogenesis which differs somewhat from that described in flies.

NASA Discipline Developmental Biology

Toward a New Paradigm for the Unification of Radio Loud AGN and its Connection to Accretion

We recently argued [21J that the collective properties. of radio loud active galactic nuclei point to the existence of two families of sources, one of powerful sources with single velocity jets and one of weaker jets with significant velocity gradients in the radiating plasma. These families also correspond to different accretion modes and therefore different thermal and emission line intrinsic properties: powerful sources have radiatively efficient accretion disks, while in weak sources accretion must be radiatively inefficient. Here, after we briefly review of our recent work, we present the following findings that support our unification scheme: (i) along the broken sequence of aligned objects, the jet kinetic power increases. (ii) in the powerful branch of the sequence of aligned objects the fraction of BLLs decreases with increasing jet power. (iii) for powerful sources, the fraction of BLLs increases for more un-aligned objects, as measured by the core to extended radio emission. Our results are also compatible with the possibility that a given accretion power produces jets of comparable kinetic power.

Georganpoulos, Markos

Sequence information signal processor

An electronic circuit is used to compare two sequences, such as genetic sequences, to determine which alignment of the sequences produces the greatest similarity. The circuit includes a linear array of series-connected processors, each of which stores a single element from one of the sequences and compares that element with each successive element in the other sequence. For each comparison, the processor generates a scoring parameter that indicates which segment ending at those two elements produces the greatest degree of similarity between the sequences. The processor uses the scoring parameter to generate a similar scoring parameter for a comparison between the stored element and the next successive element from the other sequence. The processor also delivers the scoring parameter to the next processor in the array for use in generating a similar scoring parameter for another pair of elements. The electronic circuit determines which processor and alignment of the sequences produce the scoring parameter with the highest value.

Peterson, John C.

CRITICA: coding region identification tool invoking comparative analysis

Gene recognition is essential to understanding existing and future DNA sequence data. CRITICA (Coding Region Identification Tool Invoking Comparative Analysis) is a suite of programs for identifying likely protein-coding sequences in DNA by combining comparative analysis of DNA sequences with more common noncomparative methods. In the comparative component of the analysis, regions of DNA are aligned with related sequences from the DNA databases; if the translation of the aligned sequences has greater amino acid identity than expected for the observed percentage nucleotide identity, this is interpreted as evidence for coding. CRITICA also incorporates noncomparative information derived from the relative frequencies of hexanucleotides in coding frames versus other contexts (i.e., dicodon bias). The dicodon usage information is derived by iterative analysis of the data, such that CRITICA is not dependent on the existence or accuracy of coding sequence annotations in the databases. This independence makes the method particularly well suited for the analysis of novel genomes. CRITICA was tested by analyzing the available Salmonella typhimurium DNA sequences. Its predictions were compared with the DNA sequence annotations and with the predictions of GenMark. CRITICA proved to be more accurate than GenMark, and moreover, many of its predictions that would seem to be errors instead reflect problems in the sequence databases. The source code of CRITICA is freely available by anonymous FTP (rdp.life.uiuc.edu in/pub/critica) and on the World Wide Web (http:/(/)rdpwww.life.uiuc.edu).

Non-NASA Center