Search NASASearch

SEARCH · Search NASA

Results for “Exons”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

44 records · Page 3

Multi-Omics Study of the Effect of Redox-Active Metalloporphyrin on Murine Retina During Spaceflight

Astronauts returning from spaceflight have experienced eye problems, which may decrease retinal performance and lead to long-term effects on visual acuity. This study leverages the collected data from spaceflown murine retinas that were treated with redox-active metalloporphyrin (BuOE) to mitigate spaceflight-induced changes and respective ground controls. 10-week-old adult C57BL/6 male mice (n=5 in each of BuOE treated and saline control groups for spaceflown and ground control samples) were flown on Space-X 24 to the ISS national lab, kept in low earth orbit for 35 days and returned to Earth alive. Our multi-omics analysis of RNA-sequencing and reduced representation bisulfite sequencing (RRBS) data generated from subsequent murine retina tissues uncovered genes, pathways, and epigenetic modifications consistent with therapeutic potential of BuOE. From RNA-Seq analysis of spaceflown murine samples, the treatment group show differentially expressed genes relative to saline controls that reached significance (adjusted p-value < 0.05) and included genes Gpx3 and Crhbp, which are related to protection against cell oxidative damage and cellular response to organonitrogen compounds. Ranked fold-changes from the same contrast were used for gene set enrichment analysis, which showed biological processes reaching significance (adjusted p-value < 0.05) including glutathione metabolic processes and cellular response to xenobiotic stimulus. RRBS data of the spaceflown murine samples found 139 hyper or hypo differentially methylated sites spread across chromosomes 1-19 (20% promoters, 21% exons, 43% introns | 20 CpG islands, 7 CpG shores) with a 10% methylation difference (q-value < 0.05).The findings from this investigation have the potential to provide valuable insights into the molecular mechanisms underlying conditions like spaceflight associated neuro-ocular syndrome and assess the effectiveness of BuOE as a countermeasure for astronauts experiencing neuro-ophthalmic abnormalities, which can lead to long-term effects on visual acuity.

Biostatistics

Statistical and linguistic features of DNA sequences

We present evidence supporting the idea that the DNA sequence in genes containing noncoding regions is correlated, and that the correlation is remarkably long range--indeed, base pairs thousands of base pairs distant are correlated. We do not find such a long-range correlation in the coding regions of the gene. We resolve the problem of the "non-stationary" feature of the sequence of base pairs by applying a new algorithm called Detrended Fluctuation Analysis (DFA). We address the claim of Voss that there is no difference in the statistical properties of coding and noncoding regions of DNA by systematically applying the DFA algorithm, as well as standard FFT analysis, to all eukaryotic DNA sequences (33 301 coding and 29 453 noncoding) in the entire GenBank database. We describe a simple model to account for the presence of long-range power-law correlations which is based upon a generalization of the classic Levy walk. Finally, we describe briefly some recent work showing that the noncoding sequences have certain statistical features in common with natural languages. Specifically, we adapt to DNA the Zipf approach to analyzing linguistic texts, and the Shannon approach to quantifying the "redundancy" of a linguistic text in terms of a measurable entropy function. We suggest that noncoding regions in plants and invertebrates may display a smaller entropy and larger redundancy than coding regions, further supporting the possibility that noncoding regions of DNA may carry biological information.

Non-NASA Center

Analysis of the myosins encoded in the recently completed Arabidopsis thaliana genome sequence

BACKGROUND: Three types of molecular motors play an important role in the organization, dynamics and transport processes associated with the cytoskeleton. The myosin family of molecular motors move cargo on actin filaments, whereas kinesin and dynein motors move cargo along microtubules. These motors have been highly characterized in non-plant systems and information is becoming available about plant motors. The actin cytoskeleton in plants has been shown to be involved in processes such as transportation, signaling, cell division, cytoplasmic streaming and morphogenesis. The role of myosin in these processes has been established in a few cases but many questions remain to be answered about the number, types and roles of myosins in plants. RESULTS: Using the motor domain of an Arabidopsis myosin we identified 17 myosin sequences in the Arabidopsis genome. Phylogenetic analysis of the Arabidopsis myosins with non-plant and plant myosins revealed that all the Arabidopsis myosins and other plant myosins fall into two groups - class VIII and class XI. These groups contain exclusively plant or algal myosins with no animal or fungal myosins. Exon/intron data suggest that the myosins are highly conserved and that some may be a result of gene duplication. CONCLUSIONS: Plant myosins are unlike myosins from any other organisms except algae. As a percentage of the total gene number, the number of myosins is small overall in Arabidopsis compared with the other sequenced eukaryotic genomes. There are, however, a large number of class XI myosins. The function of each myosin has yet to be determined.

NASA Discipline Plant Biology

Systematic analysis of coding and noncoding DNA sequences using methods of statistical linguistics

We compare the statistical properties of coding and noncoding regions in eukaryotic and viral DNA sequences by adapting two tests developed for the analysis of natural languages and symbolic sequences. The data set comprises all 30 sequences of length above 50 000 base pairs in GenBank Release No. 81.0, as well as the recently published sequences of C. elegans chromosome III (2.2 Mbp) and yeast chromosome XI (661 Kbp). We find that for the three chromosomes we studied the statistical properties of noncoding regions appear to be closer to those observed in natural languages than those of coding regions. In particular, (i) a n-tuple Zipf analysis of noncoding regions reveals a regime close to power-law behavior while the coding regions show logarithmic behavior over a wide interval, while (ii) an n-gram entropy measurement shows that the noncoding regions have a lower n-gram entropy (and hence a larger "n-gram redundancy") than the coding regions. In contrast to the three chromosomes, we find that for vertebrates such as primates and rodents and for viral DNA, the difference between the statistical properties of coding and noncoding regions is not pronounced and therefore the results of the analyses of the investigated sequences are less conclusive. After noting the intrinsic limitations of the n-gram redundancy analysis, we also briefly discuss the failure of the zeroth- and first-order Markovian models or simple nucleotide repeats to account fully for these "linguistic" features of DNA. Finally, we emphasize that our results by no means prove the existence of a "language" in noncoding DNA.

NASA Discipline Number 14-10

Long-range correlation properties of coding and noncoding DNA sequences: GenBank analysis

An open question in computational molecular biology is whether long-range correlations are present in both coding and noncoding DNA or only in the latter. To answer this question, we consider all 33301 coding and all 29453 noncoding eukaryotic sequences--each of length larger than 512 base pairs (bp)--in the present release of the GenBank to dtermine whether there is any statistically significant distinction in their long-range correlation properties. Standard fast Fourier transform (FFT) analysis indicates that coding sequences have practically no correlations in the range from 10 bp to 100 bp (spectral exponent beta=0.00 +/- 0.04, where the uncertainty is two standard deviations). In contrast, for noncoding sequences, the average value of the spectral exponent beta is positive (0.16 +/- 0.05) which unambiguously shows the presence of long-range correlations. We also separately analyze the 874 coding and the 1157 noncoding sequences that have more than 4096 bp and find a larger region of power-law behavior. We calculate the probability that these two data sets (coding and noncoding) were drawn from the same distribution and we find that it is less than 10(-10). We obtain independent confirmation of these findings using the method of detrended fluctuation analysis (DFA), which is designed to treat sequences with statistical heterogeneity, such as DNA's known mosaic structure ("patchiness") arising from the nonstationarity of nucleotide concentration. The near-perfect agreement between the two independent analysis methods, FFT and DFA, increases the confidence in the reliability of our conclusion.

Non-NASA Center

Mosaic organization of DNA nucleotides

Long-range power-law correlations have been reported recently for DNA sequences containing noncoding regions. We address the question of whether such correlations may be a trivial consequence of the known mosaic structure ("patchiness") of DNA. We analyze two classes of controls consisting of patchy nucleotide sequences generated by different algorithms--one without and one with long-range power-law correlations. Although both types of sequences are highly heterogenous, they are quantitatively distinguishable by an alternative fluctuation analysis method that differentiates local patchiness from long-range correlations. Application of this analysis to selected DNA sequences demonstrates that patchiness is not sufficient to account for long-range correlation properties.

NASA Discipline Number 14-10

Modeling study on the cleavage step of the self-splicing reaction in group I introns

A three-dimensional model of the Tetrahymena thermophila group I intron is used to further explore the catalytic mechanism of the transphosphorylation reaction of the cleavage step. Based on the coordinates of the catalytic core model proposed by Michel and Westhof (Michel, F., Westhof, E. J. Mol. Biol. 216, 585-610 (1990)), we first converted their ligation step model into a model of the cleavage step by the substitution of several bases and the removal of helix P9. Next, an attempt to place a trigonal bipyramidal transition state model in the active site revealed that this modified model for the cleavage step could not accommodate the transition state due to insufficient space. A lowering of P1 helix relative to surrounding helices provided the additional space required. Simultaneously, it provided a better starting geometry to model the molecular contacts proposed by Pyle et al. (Pyle, A. M., Murphy, F. L., Cech, T. R. Nature 358, 123-128. (1992)), based on mutational studies involving the J8/7 segment. Two hydrated Mg2+ complexes were placed in the active site of the ribozyme model, using the crystal structure of the functionally similar Klenow fragment (Beese, L.S., Steitz, T.A. EMBO J. 10, 25-33 (1991)) as a guide. The presence of two metal ions in the active site of the intron differs from previous models, which incorporate one metal ion in the catalytic site to fulfill the postulated roles of Mg2+ in catalysis. The reaction profile is simulated based on a trigonal bipyramidal transition state, and the role of the hydrated Mg2+ complexes in catalysis is further explored using molecular orbital calculations.

NASA Discipline Exobiology

Numerical classification of coding sequences

DNA sequences coding for protein may be represented by counts of nucleotides or codons. A complete reading frame may be abbreviated by its base count, e.g. A76C158G121T74, or with the corresponding codon table, e.g. (AAA)0(AAC)1(AAG)9 ... (TTT)0. We propose that these numerical designations be used to augment current methods of sequence annotation. Because base counts and codon tables do not require revision as knowledge of function evolves, they are well-suited to act as cross-references, for example to identify redundant GenBank entries. These descriptors may be compared, in place of DNA sequences, to extract homologous genes from large databases. This approach permits rapid searching with good selectivity.

Non-NASA Center