Search NASA⌕ Search

SEARCH · Search NASA

Results for “Base Sequence”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 433 records · Page 24

Modeling the Activity of Single Genes

The central dogma of molecular biology states that information is stored in DNA, transcribed to messenger RNA (mRNA) and then translated into proteins. This picture is significantly augmentated when we consider the action of certain proteins in regulating transcription. These transcription factors provide a feedback pathway by which genes can regulate one another's expression as mRNA and then as protein. To review: DNA, RNA and proteins have different functions. DNA is the molecular storehouse of genetic information. When cells divide, the DNA is replicated, so that each daughter cell maintains the same genetic information as the mother cell. RNA acts as a go-between from DNA to proteins. Only a single copy of DNA is present, but multiple copies of the same piece of RNA may be present, allowing cells to make huge amounts of protein. In eukaryotes (organisms with a nucleus), DNA is found in the nucleus only. RNA is copied in the nucleus then translocates(moves) outside the nucleus, where it is transcribed into proteins. Along the way, the RNA may be spliced, i.e., may have pieces cut out. RNA then attaches to ribosomes and is translated to proteins. Proteins are the machinery of the cell other than DNA and RNA, all the complex molecules of the cell are proteins. Proteins are specialized machines, each of which fulfills its own task, which may be transporting oxygen, catalyzing reactions, or responding to extracellular signals, just to name a few. One of the more interesting functions a protein may have is binding directly or indirectly to DNA to perform transcriptional regulation, thus forming a closed feedback loop of gene regulation. The structure of DNA and the central dogma were understood in the 50s; in the early 80s it became possible to make arbitrary modifications to DNA and use cellular machinery to transcribe and translate the resulting genes; more recently, genomes (i.e., the complete DNA sequence) of many organisms have been sequenced. This large-scale sequencing began with simple organisms, viruses and bacteria, progressed to eukaryotes such as yeast, and more recently (1998) progressed to a multi-cellular animal, the nematode Caenorhabditis elegans. Sequencers have now moved on to the fruit fly Drosophila melanogaster, whose sequence is slated for completion by the end of 1999. The human genome project is expected to determine the complete sequence of all 3 billion bases of human DNA within the next five years. In the wake of genome-scale sequencing, further instrumentation is being developed to assay gene expression and function on a comparably large scale. Much of the work in computational biology focuses on computational tools used in sequencing, finding genes that are related to a particular gene, finding which parts of the DNA code for proteins and which do not, understanding what proteins will be formed from a given length of DNA, predicting how the proteins will fold from a one-dimensional structure into a three dimensional structure, and so on. Much less computational work has been done regarding the function of proteins. One reason for this is that different proteins function very differently, and so work on protein function is very specific to certain classes of proteins. There are, for example, proteins such enzymes that catalyze various intracellular reactions, receptors that respond to extracellular signals and ion channels that regulate the flow of charged particles into and out of the cell. In this chapter, we will consider a particular class of proteins called transcription factors(TFs), which are responsible for regulating when a certain gene is expressed in a certain cell, which cells it is express in, and how much is expressed. Understanding these processes will involve developing a deeper understanding of transcription, translation, and the cellular processes that control those processes. All of these elements fall under the aegis of gene regulation or more narrowly transcriptional regulation. Some of the key questions in gene regulation are: What genes are expressed in a certain cell at a certain time? How does gene expression differ from cell to cell in a multicellular organism? Which proteins act as transcription factors, i.e., are important in regulating gene expression? From questions like these, we hope to understand which genes are important for various macroscopic processes. Nearly all of the cells of a multicellular organism contain the same DNA. Yet this same genetic information yields a large number of different cell types. The fundamental difference between a neuron and a liver cell, for example, is which genes are expressed. Thus understanding gene regulation is an important step in understanding development. Furthermore, understanding the usual genes that are expressed in cells may give important clues about various diseases. Some diseases, such as sickle cell anemia and cystic fibrosis, are caused by defects in single, non-regulatory genes; others, such as certain cancers, are caused when the cellular control circuitry malfunctions - an understanding of these diseases will involve pathways of multiple interacting gene products. There are numerous challenges in the area of understanding and modeling gene regulation. First and foremost, biologists would like to develop a deeper understanding of the processes involved, including which genes and families of genes are important, how they interact, etc. From a computation point of view, there has been embarrassingly little work done. In this chapter there are many areas in which we can phrase meaningful, non-trivial computational questions, but questions that have not been addressed. Some of these are purely computational (what is a good algorithm for dealing with a model of type X) and others are more mathematical (given a system with certain characteristics, what sort of model can one use? How does one find biochemical parameters from system-level behavior using as few experiments as possible?). In addition to biological and algorithmic problems, there is also the ever-present issue of theoretical biology - what general principles can be derived from these systems, what can one do with models other than just simulate time-courses, what can be deduced about a class of systems without knowing all the details? The fundamental challenge to computationalists and theorists is to add value to the biology - to use models, modeling techniques and algorithms to understand the biology in new ways.

Mjolsness, Eric↗

Automated rover sequence report generation

A web-based rover mission operations report and its various elements are described. The system was used for documentation of the Field Integrated Development and Operations (FIDO) rover May 2000 field test and results from the field test are provided. Implementation of automated sequence report generation for the operations report is explained in detail.

rovers mission operations↗

Eye movements - On-line measurement, analysis, and control

The paper outlines the plans and progress related to the development of a Programmed Eye-track Recording System and Eye-coupled Ubiquitous Scene-generator known by the acronym PERSEUS. Particular attention is given to the design and implementation of a computer-based real-time eye-tracking system with associated digital scenic display capability. The accurate eye-tracker developed by Cornsweet and Crane (1973) is selected for this purpose. The discussion covers automatic detection of fixations and saccades, automatic scanpath analysis, fixation-conditional stimulation, and digital scene generation. The all-digital approach to scenic simulation not only eliminates the camera optics and electromechanical servomechanisms of TV-model systems of simulation but also opens the way to the virtually unlimited sequencing of data-base contents and perspectives thereof.

Anliker, J.↗

Structural characterization and regulatory element analysis of the heart isoform of cytochrome c oxidase VIa

In order to investigate the mechanism(s) governing the striated muscle-specific expression of cytochrome c oxidase VIaH we have characterized the murine gene and analyzed its transcriptional regulatory elements in skeletal myogenic cell lines. The gene is single copy, spans 689 base pairs (bp), and is comprised of three exons. The 5'-ends of transcripts from the gene are heterogeneous, but the most abundant transcript includes a 5'-untranslated region of 30 nucleotides. When fused to the luciferase reporter gene, the 3.5-kilobase 5'-flanking region of the gene directed the expression of the heterologous protein selectively in differentiated Sol8 cells and transgenic mice, recapitulating the pattern of expression of the endogenous gene. Deletion analysis identified a 300-bp fragment sufficient to direct the myotube-specific expression of luciferase in Sol8 cells. The region lacks an apparent TATA element, and sequence motifs predicted to bind NRF-1, NRF-2, ox-box, or PPAR factors known to regulate other nuclear genes encoding mitochondrial proteins are not evident. Mutational analysis, however, identified two cis-elements necessary for the high level expression of the reporter protein: a MEF2 consensus element at -90 to -81 bp and an E-box element at -147 to -142 bp. Additional E-box motifs at closely located positions were mutated without loss of transcriptional activity. The dependence of transcriptional activation of cytochrome c oxidase VIaH on cis-elements similar to those found in contractile protein genes suggests that the striated muscle-specific expression is coregulated by mechanisms that control the lineage-specific expression of several contractile and cytosolic proteins.

Non-NASA Center↗

Statistical and linguistic features of DNA sequences

We present evidence supporting the idea that the DNA sequence in genes containing noncoding regions is correlated, and that the correlation is remarkably long range--indeed, base pairs thousands of base pairs distant are correlated. We do not find such a long-range correlation in the coding regions of the gene. We resolve the problem of the "non-stationary" feature of the sequence of base pairs by applying a new algorithm called Detrended Fluctuation Analysis (DFA). We address the claim of Voss that there is no difference in the statistical properties of coding and noncoding regions of DNA by systematically applying the DFA algorithm, as well as standard FFT analysis, to all eukaryotic DNA sequences (33 301 coding and 29 453 noncoding) in the entire GenBank database. We describe a simple model to account for the presence of long-range power-law correlations which is based upon a generalization of the classic Levy walk. Finally, we describe briefly some recent work showing that the noncoding sequences have certain statistical features in common with natural languages. Specifically, we adapt to DNA the Zipf approach to analyzing linguistic texts, and the Shannon approach to quantifying the "redundancy" of a linguistic text in terms of a measurable entropy function. We suggest that noncoding regions in plants and invertebrates may display a smaller entropy and larger redundancy than coding regions, further supporting the possibility that noncoding regions of DNA may carry biological information.

Non-NASA Center↗

WFPC2 Observations of the URSA Minor Dwarf Spheroidal Galaxy

We present our analysis of archival Hubble Space Telescope Wide Field Planetary Camera 2 (WFPC2) observations in F555W (approximately V) and F814W (approximately I) of the central region of the Ursa Minor dwarf spheroidal galaxy. The V versus V - I color-magnitude diagram features a sparsely populated blue horizontal branch, a steep thin red giant branch, and a narrow subgiant branch. The main sequence reaches approximately 2 magnitudes below the main-sequence turnoff (V(sup UMi, sub TO) approximately equals 23.27 +/- 0.11 mag) of the median stellar population. We compare the fiducial sequence of the Galactic globular cluster M92 (NGC 6341). The excellent match between Ursa Minor and M92 confirms that the median stellar population of the UMi dSph galaxy is metal poor ([Fe/H](sub UMi) approximately equals [Fe/H](sub M92) approximately equals -2.2 dex) and ancient (age(sub UMi)approximately equalsage(sub M92) approximately equals 14 Gyr). The B - V reddening and the absorption in V are estimated to be E(B - V) = 0.03 +/- 0.01 mag and A(sup UMi, sub V) = 0.09 +/- 0.03 mag. A new estimate of the distance modulus of Ursa Minor, (m - M)(sup UMi, sub 0) = 19.18 +/- 0.12 mag, has been derived based on fiducial-sequence fitting M92 [DELTA.V(sub UMi - M92) = 4.60 +/- 0.03 mag and DELTA(V - I)(sub UMi - M92) = 0.010 +/- 0.005 mag] and the adoption of the apparent V distance modulus for M92 of (m - M)(sup M92, sub V) = 14.67 +/- 0.08 mag (Pont et al. 1998, A&A, 329, 87). The Ursa Minor dwarf spheroidal galaxy is then at a distance of 69 +/- 4 kpc from the Sun. These HST observations indicate that Ursa Minor has had a very simple star formation history consisting mainly of a single major burst of star formation about 14 Gyr ago which lasted approximately < 2 Gyr. While we may have missed minor younger stellar populations due to the small field-of-view of the WFPC2 instrument, these observations clearly show that most of the stars in the central region Ursa Minor dwarf spheroidal galaxy are ancient. If the ancient Galactic globular clusters, like M92, formed concurrently with the early formation of the Milky Way galaxy itself, then the Ursa Minor dwarf spheroidal is probably as old as the Milky Way.

Mighell, Kenneth J.↗

“PowerCell”: The Interface Between Mars Resources and Human Exploration

The barriers to forming human settlements on Mars are high but surmountable within our lifetime. While the Apollo astronauts carried their life support with them, our success in exploring and forming settlements on Mars depends on our ability to use local Martian resources to generate the materials and conditions humans need to survive, so-called in situ resource utilization (ISRU). On Earth, biology provides us with food, shelter, oxygen, and other materials. Off-planet, synthetic biology will enable numerous parallel productions: optimized food production, water treatment, air treatment, environmental monitoring, regolith biomining, waste management, cell based biomaterial production, biocementation, and in situ synthesis based on received DNA sequences. How will the organisms responsible for these synthetic production systems obtain organic carbon and fixed nitrogen in the hostile Martian environment? We envision a synthetic-biology enabled Martian colony and introduce here the critical intermediate component a biological power source needed to transform the in situ resources found on Mars into biological feedstocks to enable growth of production organisms. Here, we present our first PowerCell, a photosynthetic and nitrogen-fixing filamentous cyanobacterium engineered to provide a carbon-rich fuel source for a biological life support system on Mars. We provide a vision of how the PowerCell system will operate in a Martian colony based on ground experiments and preparations for testing in space as a NASA secondary payload aboard the upcoming DLR Eu:CROPIS satellite mission experiments.

Rothschild, Lynn J.↗

Discovering methylated DNA motifs in bacterial nanopore sequencing data with MIJAMP

Abstract Bacterial DNA methylation is involved in diverse cellular functions, including modulation of gene expression, DNA repair, and restriction–modification systems for defense against viruses and other foreign DNA. Restriction systems hinder efforts to engineer organisms to produce fuels and chemicals from waste and renewable feedstocks by degrading DNA during transformation. Methylome analysis allows identification of motifs within a bacterial chromosome that may be targeted by native restriction enzymes. Further expression of the corresponding methyltransferases in Escherichia coli allows plasmid DNA to be protected from restriction in the target organism, thereby drastically enhancing transformation efficiency. Nanopore sequencing can detect methylated bases, but software is needed to transform modified base coordinates into methylated motifs. Here, we develop MIJAMP (MIJAMP Is Just A MethylBED Parser), a software package that was developed to discover methylated motifs from the output of ONT’s Modkit or other data in the methylBED format. MIJAMP employs a human-driven refinement strategy that empirically validates all motifs against genome-wide methylation data, thus eliminating incorrect motifs. MIJAMP also reports methylation data on specific, user-defined motifs. Using MIJAMP, we determined the methylated motifs both in a control strain (wild-type E. coli) and in Synecococcus sp. strain PCC7002, laying the foundation for improved transformation in this organism. MIJAMP is available at https://code.ornl.gov/alexander-public/mijamp/. One Sentence Summary: Here we describe software written to discover DNA methylation motifs from nanopore sequencing data.

59 BASIC BIOLOGICAL SCIENCES↗

Visualizing and analyzing 3D biomolecular structures using Mol* at RCSB.org: Influenza A H5N1 virus proteome case study

The easiest and often most useful way to work with experimentally determined or computationally predicted structures of biomolecules is by viewing their three-dimensional (3D) shapes using a molecular visualization tool. Mol* was collaboratively developed by RCSB Protein Data Bank (RCSB PDB, RCSB.org) and Protein Data Bank in Europe (PDBe, PDBe.org) as an open-source, web-based, 3D visualization software suite for examination and analyses of biostructures. It is capable of displaying atomic coordinates and related experimental data of biomolecular structures together with a variety of annotations, facilitating basic and applied research, training, education, and information dissemination. Across RCSB.org, the RCSB PDB research-focused web portal, Mol* has been implemented to support single-mouse-click atomic-level visualization of biomolecules (e.g., proteins, nucleic acids, carbohydrates) with bound cofactors, small-molecule ligands, ions, water molecules, or other macromolecules. RCSB.org Mol* can seamlessly display 3D structures from various sources, allowing structure interrogation, superimposition, and comparison. Using influenza A H5N1 virus as a topical case study of an important pathogen, we exemplify how Mol* has been embedded within various RCSB.org tools—allowing users to view polymer sequence and structure-based annotations integrated from trusted bioinformatics data resources, assess patterns and trends in groups of structures, and view structures of any size and compositional complexity. In addition to being linked to every experimentally determined biostructure and Computed Structure Model made available at RCSB.org, Standalone Mol* is freely available for visualizing any atomic-level or multi-scale biostructure at rcsb.org/3d-view.

3D biostructure↗

Phylogenetic placement of the Spirosomaceae

Comparative analysis of 16S rRNA sequences shows that the family Spirosomaceae belongs within the eubacterial phylum defined by the flavobacteria and bacteriodes. Its constituent genera, Spirosoma, Flectobacillus, and Runella form a monophyletic grouping therein. The phylogenetic assignment is based not only upon evolutionary distance analysis, but also upon sequence signatures and higher order structural synapomorphies in 16S rRNA. Another genus peripherally associated with the Spirosomaceae, Ancylobacter ("Microcyclus"), does not cluster with the flavobacteria and their relatives, but rather belongs to the alpha subdivision of the purple bacteria.

NASA Discipline Number 52-30↗

Functional characteristics of the calcium modulated proteins seen from an evolutionary perspective

We have constructed dendrograms relating 173 EF-hand proteins of known amino acid sequence. We aligned all of these proteins by their EF-hand domains, omitting interdomain regions. Initial dendrograms were computed by minimum mutation distance methods. Using these as starting points, we determined the best dendrogram by the method of maximum parsimony, scored by minimum mutation distance. We identified 14 distinct subfamilies as well as 6 unique proteins that are perhaps the sole representatives of other subfamilies. This information is given in tabular form. Within subfamilies one can easily align interdomain regions. The resulting dendrograms are very similar to those computed using domains only. Dendrograms constructed using pairs of domains show general congruence. However, there are enough exceptions to caution against an overly simple scheme in which one pair of gene duplications leads from one domain precurser to a four domain prototype from which all other forms evolved. The ability to bind calcium was lost and acquired several times during evolution. The distribution of introns does not conform to the dendrogram based on amino acid sequences. The rates of evolution appear to be much slower within subfamilies, especially within calmodulin, than those prior to the definition of subfamily.

Kretsinger, R. H.↗

The final WaZP galaxy cluster catalog of the Dark Energy Survey and comparison with SZE data

In this work, we present and characterize the galaxy cluster catalog detected by the WaZP cluster finder, which is not based on red-sequence identification, on the full six years of observations of the Dark Energy Survey (DES-Y6). The full catalog contains over 400k detected clusters with richnesses, Ngals, above 5 and that reach redshifts up to 1.3. We also provide a version of the catalog where the observation depth and richness computation are homogenized to be used for cosmology, containing 33k rich (Ngals >25) clusters. We compare our results with the previous WaZP catalog obtained from the DES first-year data release (DES-Y1). We find that essentially all clusters within the common footprint and depth limit are recovered. The deeper observations on DES-Y6 and the more complete available spectroscopic redshift sample lead to improvements in the redshifts of the clusters, resulting in an average scatter of 1.4% and offset of 0.2%. The optical clusters are also cross-matched with Sunyaev Zel'dovich Effect (SZE) cluster samples detected by the South Pole Telescope (SPT) and the Atacama Cosmology Telescope (ACT). We find that essentially all SZE clusters with reasonable overlapping footprint have a corresponding WaZP cluster. Conversely, 90% of the optical detections with richness greater than 150 have a counterpart in the deeper regions of the SZE surveys. Based on cross-match with the SZE catalogs, we also find that 15-20% of the SZE matched systems have more than one possible WaZP counterpart at the same redshift and within the SZE R500c, indicating possible interacting or unrelaxed systems. Finally, given the optical and SZE beams, WaZP and SZE centerings are found to be consistent. A more detailed study of the SZE-WaZP mass-richness relation will be presented in a separate paper.

Benoist, C. [OCA, Nice, Lab. Lagrange; LIneA, Rio ↗

Sequency Hierarchy Truncation (SeqHT) for Adiabatic State Preparation and Time Evolution in Quantum Simulations

We introduce the Sequency Hierarchy Truncation (SeqHT) scheme for reducing the resources required for state preparation and time evolution in quantum simulations, based upon a truncation in sequency. For the λϕ 4 interaction in scalar field theory, or any interaction with a polynomial expansion, upper bounds on the contributions of operators of a given sequency are derived. For the systems we have examined, observables computed in sequency-truncated wavefunctions, including quantum correlations as measured by magic, are found to step-wise converge to their exact values with increasing cutoff sequency. The utility of SeqHT is demonstrated in the adiabatic state preparation of the λϕ 4 anharmonic oscillator ground state using IBM's quantum computer ibm_sherbrooke. Using SeqHT, the depth of the required quantum circuits is reduced by ∼ 30 % , leading to significantly improved determinations of observables in the quantum simulations. More generally, SeqHT is expected to lead to a reduction in required resources for quantum simulations of systems with a hierarchy of length scales.

Li, Zhiyao [Univ. of Washington, Seattle, WA (Unit↗

Ground data systems resource allocation process

The Ground Data Systems Resource Allocation Process at the Jet Propulsion Laboratory provides medium- and long-range planning for the use of Deep Space Network and Mission Control and Computing Center resources in support of NASA's deep space missions and Earth-based science. Resources consist of radio antenna complexes and associated data processing and control computer networks. A semi-automated system was developed that allows operations personnel to interactively generate, edit, and revise allocation plans spanning periods of up to ten years (as opposed to only two or three weeks under the manual system) based on the relative merit of mission events. It also enhances scientific data return. A software system known as the Resource Allocation and Planning Helper (RALPH) merges the conventional methods of operations research, rule-based knowledge engineering, and advanced data base structures. RALPH employs a generic, highly modular architecture capable of solving a wide variety of scheduling and resource sequencing problems. The rule-based RALPH system has saved significant labor in resource allocation. Its successful use affirms the importance of establishing and applying event priorities based on scientific merit, and the benefit of continuity in planning provided by knowledge-based engineering. The RALPH system exhibits a strong potential for minimizing development cycles of resource and payload planning systems throughout NASA and the private sector.

Berner, Carol A.↗

RCSB protein data Bank: Next‐generation advanced search for exploration of experimental structures and computed structure models

Abstract The Protein Data Bank (PDB), established in 1971, is the primary global, open‐access archive for experimentally determined 3D macromolecular structures (proteins, RNA, DNA). The research‐focused RCSB.org web‐portal provides access to these data alongside more than one million machine‐learning‐predicted structure models, greatly expanding the available structural landscape. Rapid growth of both experimental and computational structures has increased the need for powerful yet accessible search tools that serve a broad and diverse scientific community. Herein, we describe a redesigned RCSB Protein Data Bank RCSB.org Advanced Search capability that supports intuitive discovery of 3D structures through a unified interface. This interface integrates annotation‐, sequence‐, and 3D structure‐based searches, embeds an interactive 3D viewer, and incorporates curated biological knowledge, such as catalytic site definitions from Mechanism and Catalytic Site Atlas and ligand‐guided structural motifs, for constructing geometry‐driven queries. A new Chemical Search tool allows definition of chemical queries via an integrated drawing tool or standard identifiers, seamlessly combining them with annotation filters. By allowing query definition directly within spatial and chemical contexts, these search interfaces reduce the need for detailed knowledge of residue numbering, chain identifiers, or external cheminformatics software. This capability enables efficient exploration of structures, chemical diversity, and structure–function relationships across all life domains. The redesigned interfaces can be accessed directly at rcsb.org/search/advanced for Advanced Search and rcsb.org/search/chemical for Chemical Search.

Rose, Yana [Research Collaboratory for Structural ↗

Seeding Advanced Treated Wastewater for Purposes of Direct Potable Reuse

Direct potable reuse (DPR) is a promising solution to address water scarcity. However, a better understanding of how introducing advanced treated water (ATW) affects microbial communities present in distribution systems is needed. Here, in this study, we measured changes to the microbial water quality in simulated distribution systems that were conditioned using treated, unimpaired surface water (SW) and then transitioned to ATW. In addition, we investigated whether adding a biological filtration step would seed the microbial community of the ATW and whether the influence would persist in the simulated distribution systems. We found that the bulk water in the ATW-fed distribution systems had lower cell counts and ATP concentrations and a distinct microbial community (based on 16S amplicon sequencing) compared to the SW-fed or the seeded ATW-fed systems. However, biofilm community composition and biomass remained consistent regardless of the feedwater. Increased microbial biomass and diversity were present in the seeded ATW, with several amplicon sequence variants identified as being introduced by the biological filter. Our results suggest that directly introducing ATW to distribution systems could disturb the existing microbial community. Preparing ATW for distribution via biological filtration may deliver more predictable and stable microbial water quality than introducing unseeded ATW.

16S↗