Search NASA⌕ Search

SEARCH · Search NASA

Results for “Sequence databases”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

A 5.8S nuclear ribosomal RNA gene sequence database: applications to ecology and evolution

We complied a 5.8S nuclear ribosomal gene sequence database for animals, plants, and fungi using both newly generated and GenBank sequences. We demonstrate the utility of this database as an internal check to determine whether the target organism and not a contaminant has been sequenced, as a diagnostic tool for ecologists and evolutionary biologists to determine the placement of asexual fungi within larger taxonomic groups, and as a tool to help identify fungi that form ectomycorrhizae.

Databases, Factual↗

VIZARD: analysis of Affymetrix Arabidopsis GeneChip data

SUMMARY: The Affymetrix GeneChip Arabidopsis genome array has proved to be a very powerful tool for the analysis of gene expression in Arabidopsis thaliana, the most commonly studied plant model organism. VIZARD is a Java program created at the University of California, Berkeley, to facilitate analysis of Arabidopsis GeneChip data. It includes several integrated tools for filtering, sorting, clustering and visualization of gene expression data as well as tools for the discovery of regulatory motifs in upstream sequences. VIZARD also includes annotation and upstream sequence databases for the majority of genes represented on the Affymetrix Arabidopsis GeneChip array. AVAILABILITY: VIZARD is available free of charge for educational, research, and not-for-profit purposes, and can be downloaded at http://www.anm.f2s.com/research/vizard/ CONTACT: moseyko@uclink4.berkeley.edu.

Non-NASA Center↗

Characterization of Two Microbial Isolates from Andean Lakes in Bolivia

We are currently investigating the biological population present in the highest and least explored perennial lakes on earth in the Bolivian and Chilean Andes, including several volcanic crater lakes of more than 6000 m elevation, in combination of microbiological and molecular biological methods. Our samples were collected in saline lakes of the Laguna Blanca Laguna Verde area in the Bolivian Altiplano and in the Licancabur volcano crater (27 deg. 47 min S/67 deg. 47 min. W) in the ongoing project studying high altitude lakes. The main goal of the project is to look for analogies with Martian paleolakes. These Bolivian lakes can be described as Andean lakes following the classification of Chong. We have attempted to isolate pure cultures and phylogenetically characterize prokaryotes that grew under laboratory conditions. Sediment samples taken from the Licancabur crater lake (LC), Laguna Verde (LV), and Laguna Blanca (LB) were analyzed and cultured using enriched liquid media under both aerobic and anaerobic conditions. All cultures were incubated at room temperature (15 to 20 C) and under light exposure. For the reported isolates, 36 hours incubation were necessary for reaching optimal optical densities to consider them viable cultures. Ten serial dilutions starting from 1% inoculum were required to obtain a suitable enriched cell culture to transfer into solid media. Cultures on solid medium were necessary to verify the formation of colonies in order to isolate pure cultures. Different solid media were prepared using several combinations of both trace minerals and carbohydrates sources in order to fit their nutrient requirements. The microorganisms formed individual colonies on solid media enriched with tryptone, yeast extract and sodium chloride. Cells morphology was studied by optical and electronic microscopy. Rodshape morphologies were observed in most cases. Total bacterial genomic DNA was isolated from 50 ml late-exponential phase culture by using the CTAB miniprep protocol. The 16S rRNA genes were amplified by PCR using both Bacteria- and Archaeauniversal primer sets: 27f and 1492r, 21f and 1492r respectively. Sequences of 16S rRNA gene were determined and initially compared with reference sequences contained in the EMBL nucleotide sequence database by using the BLAST program and were subsequently aligned with 16S rRNA reference sequences in the ARB package (http://www.mikro.biologie.tu-muenchen.de). Aligned sequences were inserted within a stable phylogenetic tree by using the ARB parsimony tool. In this work we report the morphology and phylogenetic characterization of two isolates belonged to Laguna Blanca sediments.

Demergasso, C.↗

Molecular motors and their functions in plants

Molecular motors that hydrolyze ATP and use the derived energy to generate force are involved in a variety of diverse cellular functions. Genetic, biochemical, and cellular localization data have implicated motors in a variety of functions such as vesicle and organelle transport, cytoskeleton dynamics, morphogenesis, polarized growth, cell movements, spindle formation, chromosome movement, nuclear fusion, and signal transduction. In non-plant systems three families of molecular motors (kinesins, dyneins, and myosins) have been well characterized. These motors use microtubules (in the case of kinesines and dyneins) or actin filaments (in the case of myosins) as tracks to transport cargo materials intracellularly. During the last decade tremendous progress has been made in understanding the structure and function of various motors in animals. These studies are yielding interesting insights into the functions of molecular motors and the origin of different families of motors. Furthermore, the paradigm that motors bind cargo and move along cytoskeletal tracks does not explain the functions of some of the motors. Relatively little is known about the molecular motors and their roles in plants. In recent years, by using biochemical, cell biological, molecular, and genetic approaches a few molecular motors have been isolated and characterized from plants. These studies indicate that some of the motors in plants have novel features and regulatory mechanisms. The role of molecular motors in plant cell division, cell expansion, cytoplasmic streaming, cell-to-cell communication, membrane trafficking, and morphogenesis is beginning to be understood. Analyses of the Arabidopsis genome sequence database (51% of genome) with conserved motor domains of kinesin and myosin families indicates the presence of a large number (about 40) of molecular motors and the functions of many of these motors remain to be discovered. It is likely that many more motors with novel regulatory mechanisms that perform plant-specific functions are yet to be discovered. Although the identification of motors in plants, especially in Arabidopsis, is progressing at a rapid pace because of the ongoing plant genome sequencing projects, only a few plant motors have been characterized in any detail. Elucidation of function and regulation of this multitude of motors in a given species is going to be a challenging and exciting area of research in plant cell biology. Structural features of some plant motors suggest calcium, through calmodulin, is likely to play a key role in regulating the function of both microtubule- and actin-based motors in plants.

Non-NASA Center↗

The growing world of expansins

Expansins are cell wall proteins that induce pH-dependent wall extension and stress relaxation in a characteristic and unique manner. Two families of expansins are known, named alpha- and beta-expansins, and they comprise large multigene families whose members show diverse organ-, tissue- and cell-specific expression patterns. Other genes that bear distant sequence similarity to expansins are also represented in the sequence databases, but their biological and biochemical functions have not yet been uncovered. Expansin appears to weaken glucan-glucan binding, but its detailed mechanism of action is not well established. The biological roles of expansins are diverse, but can be related to the action of expansins to loosen cell walls, for example during cell enlargement, fruit softening, pollen tube and root hair growth, and abscission. Expansin-like proteins have also been identified in bacteria and fungi, where they may aid microbial invasion of the plant body.

NASA Discipline Plant Biology↗

Methods for determining the genetic affinity of microorganisms and viruses

Selecting which sub-sequences in a database of nucleic acid such as 16S rRNA are highly characteristic of particular groupings of bacteria, microorganisms, fungi, etc. on a substantially phylogenetic tree. Also applicable to viruses comprising viral genomic RNA or DNA. A catalogue of highly characteristic sequences identified by this method is assembled to establish the genetic identity of an unknown organism. The characteristic sequences are used to design nucleic acid hybridization probes that include the characteristic sequence or its complement, or are derived from one or more characteristic sequences. A plurality of these characteristic sequences is used in hybridization to determine the phylogenetic tree position of the organism(s) in a sample. Those target organisms represented in the original sequence database and sufficient characteristic sequences can identify to the species or subspecies level. Oligonucleotide arrays of many probes are especially preferred. A hybridization signal can comprise fluorescence, chemiluminescence, or isotopic labeling, etc.; or sequences in a sample can be detected by direct means, e.g. mass spectrometry. The method's characteristic sequences can also be used to design specific PCR primers. The method uniquely identifies the phylogenetic affinity of an unknown organism without requiring prior knowledge of what is present in the sample. Even if the organism has not been previously encountered, the method still provides useful information about which phylogenetic tree bifurcation nodes encompass the organism.

Fox, George E.↗

Identification of characteristic oligonucleotides in the bacterial 16S ribosomal RNA sequence dataset

MOTIVATION: The phylogenetic structure of the bacterial world has been intensively studied by comparing sequences of 16S ribosomal RNA (16S rRNA). This database of sequences is now widely used to design probes for the detection of specific bacteria or groups of bacteria one at a time. The success of such methods reflects the fact that there are local sequence segments that are highly characteristic of particular organisms or groups of organisms. It is not clear, however, the extent to which such signature sequences exist in the 16S rRNA dataset. A better understanding of the numbers and distribution of highly informative oligonucleotide sequences may facilitate the design of hybridization arrays that can characterize the phylogenetic position of an unknown organism or serve as the basis for the development of novel approaches for use in bacterial identification. RESULTS: A computer-based algorithm that characterizes the extent to which any individual oligonucleotide sequence in 16S rRNA is characteristic of any particular bacterial grouping was developed. A measure of signature quality, Q(s), was formulated and subsequently calculated for every individual oligonucleotide sequence in the size range of 5-11 nucleotides and for 15mers with reference to each cluster and subcluster in a 929 organism representative phylogenetic tree. Subsequently, the perfect signature sequences were compared to the full set of 7322 sequences to see how common false positives were. The work completed here establishes beyond any doubt that highly characteristic oligonucleotides exist in the bacterial 16S rRNA sequence dataset in large numbers. Over 16,000 15mers were identified that might be useful as signatures. Signature oligonucleotides are available for over 80% of the nodes in the representative tree.

NASA Discipline Life Sciences Technologies↗

CRITICA: coding region identification tool invoking comparative analysis

Gene recognition is essential to understanding existing and future DNA sequence data. CRITICA (Coding Region Identification Tool Invoking Comparative Analysis) is a suite of programs for identifying likely protein-coding sequences in DNA by combining comparative analysis of DNA sequences with more common noncomparative methods. In the comparative component of the analysis, regions of DNA are aligned with related sequences from the DNA databases; if the translation of the aligned sequences has greater amino acid identity than expected for the observed percentage nucleotide identity, this is interpreted as evidence for coding. CRITICA also incorporates noncomparative information derived from the relative frequencies of hexanucleotides in coding frames versus other contexts (i.e., dicodon bias). The dicodon usage information is derived by iterative analysis of the data, such that CRITICA is not dependent on the existence or accuracy of coding sequence annotations in the databases. This independence makes the method particularly well suited for the analysis of novel genomes. CRITICA was tested by analyzing the available Salmonella typhimurium DNA sequences. Its predictions were compared with the DNA sequence annotations and with the predictions of GenMark. CRITICA proved to be more accurate than GenMark, and moreover, many of its predictions that would seem to be errors instead reflect problems in the sequence databases. The source code of CRITICA is freely available by anonymous FTP (rdp.life.uiuc.edu in/pub/critica) and on the World Wide Web (http:/(/)rdpwww.life.uiuc.edu).

Non-NASA Center↗

Analysis of xylem formation in pine by cDNA sequencing

Secondary xylem (wood) formation is likely to involve some genes expressed rarely or not at all in herbaceous plants. Moreover, environmental and developmental stimuli influence secondary xylem differentiation, producing morphological and chemical changes in wood. To increase our understanding of xylem formation, and to provide material for comparative analysis of gymnosperm and angiosperm sequences, ESTs were obtained from immature xylem of loblolly pine (Pinus taeda L.). A total of 1,097 single-pass sequences were obtained from 5' ends of cDNAs made from gravistimulated tissue from bent trees. Cluster analysis detected 107 groups of similar sequences, ranging in size from 2 to 20 sequences. A total of 361 sequences fell into these groups, whereas 736 sequences were unique. About 55% of the pine EST sequences show similarity to previously described sequences in public databases. About 10% of the recognized genes encode factors involved in cell wall formation. Sequences similar to cell wall proteins, most known lignin biosynthetic enzymes, and several enzymes of carbohydrate metabolism were found. A number of putative regulatory proteins also are represented. Expression patterns of several of these genes were studied in various tissues and organs of pine. Sequencing novel genes expressed during xylem formation will provide a powerful means of identifying mechanisms controlling this important differentiation pathway.

Non-NASA Center↗

System, method and apparatus for generating phrases from a database

A phrase generation is a method of generating sequences of terms, such as phrases, that may occur within a database of subsets containing sequences of terms, such as text. A database is provided and a relational model of the database is created. A query is then input. The query includes a term or a sequence of terms or multiple individual terms or multiple sequences of terms or combinations thereof. Next, several sequences of terms that are contextually related to the query are assembled from contextual relations in the model of the database. The sequences of terms are then sorted and output. Phrase generation can also be an iterative process used to produce sequences of terms from a relational model of a database.

McGreevy, Michael W.↗

Numerical Investigation and Optimization of a Flushwall Injector for Scramjet Applications at Hypervelocity Flow Conditions

An investigation utilizing Reynolds-averaged simulations (RAS) was performed in order to demonstrate the use of design and analysis of computer experiments (DACE) methods in Sandia’s DAKOTA software package for surrogate modeling and optimization. These methods were applied to a flow- path fueled with an interdigitated flushwall injector suitable for scramjet applications at hyper- velocity conditions and ascending along a constant dynamic pressure flight trajectory. The flight Mach number, duct height, spanwise width, and injection angle were the design variables selected to maximize two objective functions: the thrust potential and combustion efficiency. Because the RAS of this case are computationally expensive, surrogate models are used for optimization. To build a surrogate model a RAS database is created. The sequence of the design variables comprising the database were generated using a Latin hypercube sampling (LHS) method. A methodology was also developed to automatically build geometries and generate structured grids for each design point. The ensuing RAS analysis generated the simulation database from which the two objective functions were computed using a one-dimensionalization (1D) of the three-dimensional simulation data. The data were fitted using four surrogate models: an artificial neural network (ANN), a cubic polynomial, a quadratic polynomial, and a Kriging model. Variance-based decomposition showed that both objective functions were primarily driven by changes in the duct height. Multiobjective design optimization was performed for all four surrogate models via a genetic algorithm method. Optimal solutions were obtained at the upper and lower bounds of the flight Mach number range. The Kriging model predicted an optimal solution set that exhibited high values for both objective functions. Additionally, three challenge points were selected to assess the designs on the Pareto fronts. Further sampling among the designs of the Pareto fronts may be required to lower the surrogate model errors and perform more accurate surrogate-model-based optimization.

Shenoy, Rajiv R.↗

System, Method and Apparatus for Discovering Phrases in a Database

A phrase discovery is a method of identifying sequences of terms in a database. First, a selection of one or more relevant sequences of terms. such as relevant text, is provided. Next, several shorter sequences of terms, such as phrases, are extracted from the provided relevant sequences of terms. The extracted sequences of terms are then reduced through a culling process. A gathering process then emphasizes the more relevant of the extracted and culled sequences of terms and de-emphasizes the more generic of the extracted and culled sequences of terms. The gathering process can also include iteratively retrieving additional selections of relevant sequences (e.g.. text). extracting and culling additional sequences of terms (e.g.. phrases). emphasizing and de-emphasizing extracted and culled sequences of terms and accumulating all gathered sequences of terms. The resulting gathered sequences of terms are then output.

Michael W McGreevy↗

Hemichordates and the Origin of Chordates

At the start of the period of the NASA grant three years ago, we had no information on the organization and development of the body axis of the hemichordate, Saccoglossus kowalevskii. Now we have substantial findings about the anteroposterior axis and dorsoventral axis, and based on this information, we have new insights about the origin of chordates from ancestral deuterostomes. We found ways to obtain and preserve large numbers of embryos and hatched juveniles. We can now collect about 40,000 embryos in the month of September, the time of S. kowalevskii spawning at Woods Hole. Excellent cDNA libraries were prepared from three developmental stages. From these libraries, we directly isolated about 30 gene ortholog sequences by screening and pcr techniques, all of these sequences of interest in the inquiry about the animal's organization and development. We also performed a mid-sized EST project (60,000 randomly picked clones, many of these arrayed). About half of these have been analyzed so far by blastx and are suitable for direct use of clones. We have obtained about 50 interesting sequences from this set. The rest still await analysis. Thus, at this time we have isolated orthologs of 80 genes that are known to be expressed in chordates in conserved domains and known to have interesting roles in chordate organization and development. The orthology of the S. kowalevskii sequences has been verified by neighbor joining and parsimony methods, with bootstrap estimates of validity. The S. kowalevskii sequences cluster with other deuterostome sequences, namely, other hemichordates, echinoderms, ascidians, amphioxus, or vertebrates, depending on what sequences are available in the database for comparison. We have used these sequences to do high quality in situ hybridization on S. kowalevskii embryos, and the results can be divided into three sections-those concerning the anteroposterior axis of S. kowalevskii in comparison to the same axis of chordates, those concerning the dorsoventral axis of S. kowalevskii in comparison to the same axis of chordates, and those concerning the signals and transcription factors found in the endoderm, of S. kowalevskii compared to the signals and transcription factors in the endo-mesodermal cells of Spemann's organizer of chordates.

Gerhart, John↗

A new version of the RDP (Ribosomal Database Project)

The Ribosomal Database Project (RDP-II), previously described by Maidak et al. [ Nucleic Acids Res. (1997), 25, 109-111], is now hosted by the Center for Microbial Ecology at Michigan State University. RDP-II is a curated database that offers ribosomal RNA (rRNA) nucleotide sequence data in aligned and unaligned forms, analysis services, and associated computer programs. During the past two years, data alignments have been updated and now include >9700 small subunit rRNA sequences. The recent development of an ObjectStore database will provide more rapid updating of data, better data accuracy and increased user access. RDP-II includes phylogenetically ordered alignments of rRNA sequences, derived phylogenetic trees, rRNA secondary structure diagrams, and various software programs for handling, analyzing and displaying alignments and trees. The data are available via anonymous ftp (ftp.cme.msu. edu) and WWW (http://www.cme.msu.edu/RDP). The WWW server provides ribosomal probe checking, approximate phylogenetic placement of user-submitted sequences, screening for possible chimeric rRNA sequences, automated alignment, and a suggested placement of an unknown sequence on an existing phylogenetic tree. Additional utilities also exist at RDP-II, including distance matrix, T-RFLP, and a Java-based viewer of the phylogenetic trees that can be used to create subtrees.

Non-NASA Center↗

Evaluation of nearest-neighbor methods for detection of chimeric small-subunit rRNA sequences

Detection of chimeric artifacts formed when PCR is used to retrieve naturally occurring small-subunit (SSU) rRNA sequences may rely on demonstrating that different sequence domains have different phylogenetic affiliations. We evaluated the CHECK_CHIMERA method of the Ribosomal Database Project and another method which we developed, both based on determining nearest neighbors of different sequence domains, for their ability to discern artificially generated SSU rRNA chimeras from authentic Ribosomal Database Project sequences. The reliability of both methods decreases when the parental sequences which contribute to chimera formation are more than 82 to 84% similar. Detection is also complicated by the occurrence of authentic SSU rRNA sequences that behave like chimeras. We developed a naive statistical test based on CHECK_CHIMERA output and used it to evaluate previously reported SSU rRNA chimeras. Application of this test also suggests that chimeras might be formed by retrieving SSU rRNAs as cDNA. The amount of uncertainty associated with nearest-neighbor analyses indicates that such tests alone are insufficient and that better methods are needed.

NASA Discipline Exobiology↗

RECOVIR Software for Identifying Viruses

Most single-stranded RNA (ssRNA) viruses mutate rapidly to generate a large number of strains with highly divergent capsid sequences. Determining the capsid residues or nucleotides that uniquely characterize these strains is critical in understanding the strain diversity of these viruses. RECOVIR (an acronym for "recognize viruses") software predicts the strains of some ssRNA viruses from their limited sequence data. Novel phylogenetic-tree-based databases of protein or nucleic acid residues that uniquely characterize these virus strains are created. Strains of input virus sequences (partial or complete) are predicted through residue-wise comparisons with the databases. RECOVIR uses unique characterizing residues to identify automatically strains of partial or complete capsid sequences of picorna and caliciviruses, two of the most highly diverse ssRNA virus families. Partition-wise comparisons of the database residues with the corresponding residues of more than 300 complete and partial sequences of these viruses resulted in correct strain identification for all of these sequences. This study shows the feasibility of creating databases of hitherto unknown residues uniquely characterizing the capsid sequences of two of the most highly divergent ssRNA virus families. These databases enable automated strain identification from partial or complete capsid sequences of these human and animal pathogens.

Chakravarty, Sugoto↗

Automated Identification of Nucleotide Sequences

STITCH is a computer program that processes raw nucleotide-sequence data to automatically remove unwanted vector information, perform reverse-complement comparison, stitch shorter sequences together to make longer ones to which the shorter ones presumably belong, and search against the user s choice of private and Internet-accessible public 16S rRNA databases. ["16S rRNA" denotes a ribosomal ribonucleic acid (rRNA) sequence that is common to all organisms.] In STITCH, a template 16S rRNA sequence is used to position forward and reverse reads. STITCH then automatically searches known 16S rRNA sequences in the user s chosen database(s) to find the sequence most similar to (the sequence that lies at the smallest edit distance from) each spliced sequence. The result of processing by STITCH is the identification of the most similar well-described bacterium. Whereas previously commercially available software for analyzing genetic sequences operates on one sequence at a time, STITCH can manipulate multiple sequences simultaneously to perform the aforementioned operations. A typical analysis of several dozen sequences (length of the order of 103 base pairs) by use of STITCH is completed in a few minutes, whereas such an analysis performed by use of prior software takes hours or days.

Osman, Shariff↗