Search NASASearch

SEARCH · Search NASA

Results for “Nucleotide Mapping”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

Combining genome-wide association studies and expression quantitative trait nucleotide mapping with molecular and genetic validations to identify transcriptional networks regulating drought tolerance in Populus

Objectives: (i). To deploy a large-scale experimental drought trial for up to 1000 unique genotypes of Populus equipping the sites with controlled irrigation and drought treatments that are fully automated and monitored. FULLY COMPLETED (ii) To test the hypothesis that a suite of traits identified for drought tolerance in P. nigra can be measured in drought and control treatments in the wide germplasm collection of P. trichocarpa. FULLY COMPLETED (iii) To use established and novel GWAS model approaches to identify gene loci linked to drought tolerance traits on interest in P. trichocarpa. FULLY COMPLETED (iv) To undertake comparative analysis of GWAS results for drought tolerance traits in P. nigra and P. trichocarpa. PARTIALLY COMPLETED – remains active (v) Using RNAseq in P. trichocarpa, in droughted and control treatments to identify cis- and trans-regulated eQTN. FULLY COMPLETED (vi) Validate up to 50 cis-QTNs, from network hubs using transient protoplast assays. FULLY COMPLETED (vii) To establish Agrobacterium-based gene editing protocols in Populus. FULLY COMPLETED (viii) To utilize early leads from previous research to investigate at least 6 candidate genes for drought tolerance in Populus. FULLY COMPLETED (ix) To validate up to 20 candidate genes for drought tolerance in P. trichocarpa refined from the long-list tested in the transient assays for cis-acting hub gene targets. PARTIALLY COMPLETED- remains active.

60 APPLIED LIFE SCIENCES

Fractal landscape analysis of DNA walks

By mapping nucleotide sequences onto a "DNA walk", we uncovered remarkably long-range power law correlations [Nature 356 (1992) 168] that imply a new scale invariant property of DNA. We found such long-range correlations in intron-containing genes and in non-transcribed regulatory DNA sequences, but not in cDNA sequences or intron-less genes. In this paper, we present more explicit evidences to support our findings.

NASA Program Space Physiology and Countermeasures

Long-range correlations in nucleotide sequences

DNA sequences have been analysed using models, such as an n-step Markov chain, that incorporate the possibility of short-range nucleotide correlations. We propose here a method for studying the stochastic properties of nucleotide sequences by constructing a 1:1 map of the nucleotide sequence onto a walk, which we term a 'DNA walk'. We then use the mapping to provide a quantitative measure of the correlation between nucleotides over long distances along the DNA chain. Thus we uncover in the nucleotide sequence a remarkably long-range power law correlation that implies a new scale-invariant property of DNA. We find such long-range correlations in intron-containing genes and in nontranscribed regulatory DNA sequences, but not in complementary DNA sequences or intron-less genes.

NASA Discipline Cardiopulmonary

Statistical properties of DNA sequences

We review evidence supporting the idea that the DNA sequence in genes containing non-coding regions is correlated, and that the correlation is remarkably long range--indeed, nucleotides thousands of base pairs distant are correlated. We do not find such a long-range correlation in the coding regions of the gene. We resolve the problem of the "non-stationarity" feature of the sequence of base pairs by applying a new algorithm called detrended fluctuation analysis (DFA). We address the claim of Voss that there is no difference in the statistical properties of coding and non-coding regions of DNA by systematically applying the DFA algorithm, as well as standard FFT analysis, to every DNA sequence (33301 coding and 29453 non-coding) in the entire GenBank database. Finally, we describe briefly some recent work showing that the non-coding sequences have certain statistical features in common with natural and artificial languages. Specifically, we adapt to DNA the Zipf approach to analyzing linguistic texts. These statistical properties of non-coding sequences support the possibility that non-coding regions of DNA may carry biological information.

Non-NASA Center

Mosaic organization of DNA nucleotides

Long-range power-law correlations have been reported recently for DNA sequences containing noncoding regions. We address the question of whether such correlations may be a trivial consequence of the known mosaic structure ("patchiness") of DNA. We analyze two classes of controls consisting of patchy nucleotide sequences generated by different algorithms--one without and one with long-range power-law correlations. Although both types of sequences are highly heterogenous, they are quantitatively distinguishable by an alternative fluctuation analysis method that differentiates local patchiness from long-range correlations. Application of this analysis to selected DNA sequences demonstrates that patchiness is not sufficient to account for long-range correlation properties.

NASA Discipline Number 14-10

Human RNome Project draft human RNome sequence of GM12878, B-cell line, obtained by mass-spectrometry sequencing, long-read sequencing and short-read sequencing.

Here we report the first draft of the human RNome sequence, a reference map of RNA chemical modifications in a human B-cell line. RNA carries a diverse repertoire of chemical modifications that regulate gene expression, cellular function, and responses to physiological and pathological cues. Yet, unlike the genome, no reference map of RNA modifications is available for any human cell. To generate this resource, the Human RNome Project Consortium analyzed a shared RNA preparation from the well-characterized GM12878 B-cell line using short-read sequencing, long-read direct RNA sequencing, and mass spectrometry, generating more than 7.1 billion sequencing reads spanning approximately 1.2 trillion nucleotides. The resulting maps of the human RNome reveal that RNA modifications are organized according to function, transcript architecture, and cellular identity. Modifications concentrate at functional centers of ribosomal and transfer RNAs, follow the canonical topology of N6-methyladenosine in coding transcripts, and form coordinated hotspots in immune regulatory genes. This first reference human RNome provides a foundation for understanding how RNA chemistry shapes cellular identity, human disease, and the development of RNA-based therapeutics.

59 BASIC BIOLOGICAL SCIENCES

Cloning the promoter for transforming growth factor-beta type III receptor. Basal and conditional expression in fetal rat osteoblasts

Transforming growth factor-beta binds to three high affinity cell surface molecules that directly or indirectly regulate its biological effects. The type III receptor (TRIII) is a proteoglycan that lacks significant intracellular signaling or enzymatic motifs but may facilitate transforming growth factor-beta binding to other receptors, stabilize multimeric receptor complexes, or segregate growth factor from activating receptors. Because various agents or events that regulate osteoblast function rapidly modulate TRIII expression, we cloned the 5' region of the rat TRIII gene to assess possible control elements. DNA fragments from this region directed high reporter gene expression in osteoblasts. Sequencing showed no consensus TATA or CCAAT boxes, whereas several nuclear factors binding sequences within the 3' region of the promoter co-mapped with multiple transcription initiation sites, DNase I footprints, gel mobility shift analysis, or loss of activity by deletion or mutation. An upstream enhancer was evident 5' proximal to nucleotide -979, and a silencer region occurred between nucleotides -2014 and -2194. Glucocorticoid sensitivity mapped between nucleotides -687 and -253, whereas bone morphogenetic protein 2 sensitivity co-mapped within the silencer region. Thus, the TRIII promoter contains cooperative basal elements and dispersed growth factor- and hormone-sensitive regulatory regions that can control TRIII expression by osteoblasts.

Non-NASA Center

Identification of a QTL region for tomato brown rugose fruit virus resistance in Solanum pimpinellifolium

Abstract Tomato (Solanum lycopersicumL.), one of the most widely grown vegetables in the world, has been seriously impacted in the past decade by the emerging tomato brown rugose fruit virus (ToBRFV). ToBRFV is a seed-borne tobamovirus, with ability to overcome the commonly usedTm-2 2 resistance gene in tomato. The objective of this study was to conduct quantitative trait locus (QTL) mapping and identify single-nucleotide polymorphism (SNP) markers associated with ToBRFV resistance in tomato. Two F 2 populations were used for QTL mapping: One derived from a cross betweenS. pimpinellifoliumUSVL333 (PI 390718) × USVL332 (PI 390717) and another from ‘Moneymaker’ × USVL332 (PI 390717), with population sizes of 195 and 79 plants, respectively. The resistance trait was derived from theS. pimpinellifoliumaccession USVL332 (PI 390717). A major QTL for ToBRFV resistance was identified on chromosome 11 (SL4.0ch11), with the peak located at approximately 46.84 Mbp. This QTL spans a 22-kb interval between 46,825,788 bp and 46,847,421 bp, as determined through both genome-wide association study (GWAS) and QTL linkage mapping. Three SNP markers, SL4.0ch11_46825788, SL4.0ch11_46847421, and SL4.0ch11_46850215, demonstrated the most significant association with high LOD values (LOD = 13 in the Blink model) in GWAS analysis. In this genomic region, two disease resistance gene analogs, Solyc11g062150 (TIR-NBS-LRR resistance protein, Toll-Interleukin receptor) and Solyc11g062180 (disease resistance protein, leucine-rich repeat), were identified, which may serve as candidates for ToBRFV resistance. The QTL identified in this study could be valuable for plant breeders in facilitating tomato breeding with ToBRFV resistance.

Agriculture

Long-range correlation properties of coding and noncoding DNA sequences: GenBank analysis

An open question in computational molecular biology is whether long-range correlations are present in both coding and noncoding DNA or only in the latter. To answer this question, we consider all 33301 coding and all 29453 noncoding eukaryotic sequences--each of length larger than 512 base pairs (bp)--in the present release of the GenBank to dtermine whether there is any statistically significant distinction in their long-range correlation properties. Standard fast Fourier transform (FFT) analysis indicates that coding sequences have practically no correlations in the range from 10 bp to 100 bp (spectral exponent beta=0.00 +/- 0.04, where the uncertainty is two standard deviations). In contrast, for noncoding sequences, the average value of the spectral exponent beta is positive (0.16 +/- 0.05) which unambiguously shows the presence of long-range correlations. We also separately analyze the 874 coding and the 1157 noncoding sequences that have more than 4096 bp and find a larger region of power-law behavior. We calculate the probability that these two data sets (coding and noncoding) were drawn from the same distribution and we find that it is less than 10(-10). We obtain independent confirmation of these findings using the method of detrended fluctuation analysis (DFA), which is designed to treat sequences with statistical heterogeneity, such as DNA's known mosaic structure ("patchiness") arising from the nonstationarity of nucleotide concentration. The near-perfect agreement between the two independent analysis methods, FFT and DFA, increases the confidence in the reliability of our conclusion.

Non-NASA Center

Delayed translational silencing of ceruloplasmin transcript in gamma interferon-activated U937 monocytic cells: role of the 3' untranslated region

Ceruloplasmin (Cp) is an acute-phase protein with ferroxidase, amine oxidase, and pro- and antioxidant activities. The primary site of Cp synthesis in human adults is the liver, but it is also synthesized by cells of monocytic origin. We have shown that gamma interferon (IFN-gamma) induces the synthesis of Cp mRNA and protein in monocytic cells. We now report that the induced synthesis of Cp is terminated by a mechanism involving transcript-specific translational repression. Cp protein synthesis in U937 cells ceased after 16 h even in the presence of abundant Cp mRNA. RNA isolated from cells treated with IFN-gamma for 24 h exhibited a high in vitro translation rate, suggesting that the transcript was not defective. Ribosomal association of Cp mRNA was examined by sucrose centrifugation. When Cp synthesis was high, i.e., after 8 h of IFN-gamma treatment, Cp mRNA was primarily associated with polyribosomes. However, after 24 h, when Cp synthesis was low, Cp mRNA was primarily in the nonpolyribosomal fraction. Cytosolic extracts from cells treated with IFN-gamma for 24 h, but not for 8 h, contained a factor which blocked in vitro Cp translation. Inhibitor expression was cell type specific and present in extracts of human cells of myeloid origin, but not in several nonmyeloid cells. The inhibitory factor bound to the 3' untranslated region (3'-UTR) of Cp mRNA, as shown by restoration of in vitro translation by synthetic 3'-UTR added as a "decoy" and detection of a binding complex by RNA gel shift analysis. Deletion mapping of the Cp 3'-UTR indicated an internal 100-nucleotide region of the Cp 3'-UTR that was required for complex formation as well as for silencing of translation. Although transcript-specific translational control is common during development and differentiation and global translational control occurs during responses to cytokines and stress, to our knowledge, this is the first report of translational silencing of a specific transcript following cytokine activation.

NASA Discipline Regulatory Physiology

Three-dimensional tertiary structure of yeast phenylalanine transfer RNA

Results of an analysis and interpretation of a 3-A electron density map of yeast phenylalanine transfer RNA. Some earlier detailed assignments of nucleotide residues to electron density peaks are found to be in error, even though the overall tracing of the backbone conformation of yeast phenylalanine transfer RNA was generally correct. A new, more comprehensive interpretation is made which makes it possible to define the tertiary interactions in the molecule. The new interpretation makes it possible to visualize a number of tertiary interactions which not only explain the structural role of most of the bases which are constant in transfer RNAs, but also makes it possible to understand in a direct and simple fashion the chemical modification data on transfer RNA. In addition, this pattern of tertiary interactions provides a basis for understanding the general three-dimensional folding of all transfer RNA molecules.

Kim, S. H.

An Open-Science Approach to Address Individual Response to Simulated GCR In Genetically Diverse Populations of Mice and Humans

This project addresses the challenge of understanding and predicting individual radiation sensitivity by integrating genetics, demographics and biomarker characteristics across species (mice and humans). We hypothesize that ex vivo DNA repair response to GCR components is a central determinant of cancer risk from space radiation and can serve as a biomarker of radiation risk in combination with genetics. Automated image quantification of 53BP1+ radiation-induced foci (RIF) during the first 4-48 h post-irradiation was performed as a function of dose and LET in non-immortalized primary skin fibroblasts derived from 76 mice across 15 strains (5 inbred reference strains and 10 collaborative-cross strains) exposed to X rays (0.1, 1 and 4 Gy), 350 MeV/n 40Ar and 600 MeV/n 56Fe (1.1 and 3 particles/100sq. μm), as well as in peripheral blood mononuclear cells (PBMCs) from 768 healthy donors (matched ethnicity, 50/50 male/female, 18-70 years old) exposed to gamma rays (0.1 and 1 Gy), 350 MeV/n 28Si, 350 MeV/n 40Ar and 600 MeV/n 56Fe (1.1 and 3 particles/100sq. μm). A genome-wide association study (GWAS) was performed on the mouse strains between DNA damage responses to space radiation and single nucleotide polymorphisms (SNPs). We found SNPs, which were significantly associated to the RIF phenotype, mapped to genes and pathways that are functionally linked to health hazards for deep space exploration (e.g. carcinogenesis, nervous system damage and immune dysfunction). Some of these SNPs were located within protein coding regions, potentially interfering with protein functions and providing promising genetic targets for countermeasures. We also found correlations between both spontaneous and radiation-induced DNA damage and SNPs mapped to pathways associated with cellular metabolism. GWAS is undergoing for the human data. All data have been made available via the NASA Space Biology Open-Science database (genelab.nasa.gov) and we will discuss how various genomic and transcriptomic datasets can be accessed for modeling and integrated using machine learning methods for discovering new radiation biology.

Sylvain V Costes

Scaling features of noncoding DNA

We review evidence supporting the idea that the DNA sequence in genes containing noncoding regions is correlated, and that the correlation is remarkably long range--indeed, base pairs thousands of base pairs distant are correlated. We do not find such a long-range correlation in the coding regions of the gene, and utilize this fact to build a Coding Sequence Finder Algorithm, which uses statistical ideas to locate the coding regions of an unknown DNA sequence. Finally, we describe briefly some recent work adapting to DNA the Zipf approach to analyzing linguistic texts, and the Shannon approach to quantifying the "redundancy" of a linguistic text in terms of a measurable entropy function, and reporting that noncoding regions in eukaryotes display a larger redundancy than coding regions. Specifically, we consider the possibility that this result is solely a consequence of nucleotide concentration differences as first noted by Bonhoeffer and his collaborators. We find that cytosine-guanine (CG) concentration does have a strong "background" effect on redundancy. However, we find that for the purine-pyrimidine binary mapping rule, which is not affected by the difference in CG concentration, the Shannon redundancy for the set of analyzed sequences is larger for noncoding regions compared to coding regions.

Non-NASA Center

Molecular architecture and functional dynamics of the pre-incision complex in nucleotide excision repair

Nucleotide excision repair (NER) is vital for genome integrity. Yet, our understanding of the complex NER protein machinery remains incomplete. Combining cryo-EM and XL-MS data with AlphaFold2 predictions, we build an integrative model of the NER pre-incision complex(PInC). Here TFIIH serves as a molecular ruler, defining the DNA bubble size and precisely positioning the XPG and XPF nucleases for incision. Using simulations and graph theoretical analyses, we unveil PInC’s assembly, global motions, and partitioning into dynamic communities. Remarkably, XPG caps XPD’s DNA-binding groove and bridges both junctions of the DNA bubble, suggesting a novel coordination mechanism of PInC’s dual incision. XPA rigging interlaces XPF/ERCC1 with RPA, XPD, XPB, and 5' ssDNA, exposing XPA’s crucial role in licensing the XPF/ERCC1 incision. Mapping disease mutations onto our models reveals clustering into distinct mechanistic classes, elucidating xeroderma pigmentosum and Cockayne syndrome disease etiology.

60 APPLIED LIFE SCIENCES

In vivo mapping of mutagenesis sensitivity of human enhancers

Distant-acting enhancers are central to human development1. However, our limited understanding of their functional sequence features prevents the interpretation of enhancer mutations in disease2. Here we determined the functional sensitivity to mutagenesis of human developmental enhancers in vivo. Focusing on seven enhancers that are active in the developing brain, heart, limb and face, we created over 1,700 transgenic mice for over 260 mutagenized enhancer alleles. Systematic mutation of 12-base-pair blocks collectively altered each sequence feature in each enhancer at least once. We show that 69% of all blocks are required for normal in vivo activity, with mutations more commonly resulting in loss (60%) than in gain (9%) of function. Using predictive modelling, we annotated critical nucleotides at the base-pair resolution. The vast majority of motifs predicted by these machine learning models (88%) coincided with changes in in vivo function, and the models showed considerable sensitivity, identifying 59% of all functional blocks. Taken together, our results reveal that human enhancers contain a high density of sequence features that are required for their normal in vivo function and provide a rich resource for further exploration of human enhancer logic.

Kosicki, Michael

Identification and mapping of quantitative trait loci for Fusarium head blight resistance in a synthetic hexaploid × hard red spring wheat population

Abstract Fusarium head blight (FHB), caused byFusarium graminearumSchwabe, is one of the most devastating diseases in wheat (Triticum aestivumL.). The synthetic hexaploid wheat line Largo was developed from a cross between the durum wheat [T. turgidumssp.durum(Desf.) Husn.] variety Langdon and theAegilops tauschiiCosson accession PI 268210, and it was previously found to have a moderate level of FHB resistance. This study was conducted to identify quantitative trait loci (QTL) associated with FHB resistance using a population of 188 recombinant inbred lines (RILs) from a cross between Largo and the susceptible wheat line ND495. The RILs were evaluated for Type II resistance in two greenhouse and two field environments. The disease severity and 90K single‐nucleotide polymorphism marker data were used for QTL analysis, which revealed six QTL on chromosomes 1D, 2D, 5B, and 7D. Four QTL (QFhb.rwg‐1D,QFhb.rwg‐5B,QFhb.rwg‐7D.1, andQFhb.rwg‐7D.3) from Largo had minor effects, whereas two QTL (QFhb.rwg‐2DandQFhb.rwg‐7D.2) from ND495 showed large effects on FHB resistance. The result suggested that ND495 may possess suppressor or susceptibility gene(s) suppressing or masking FHB resistance controlled by the resistance QTL. Among these QTL, four coincided with previously reported QTL, includingFhb9, and two (QFhb.rwg‐1DandQFhb.rwg‐7D.1) are likely novel QTL. From the six QTL regions, 10 Kompetitive allele‐specific PCR markers were developed and validated for marker‐assisted selection. The QTL detected from the resistant and susceptible parents enhance our understanding of FHB resistance expression and provide new resources for improving FHB resistance in wheat.

Genetics & Heredity

NuclPred v1

This tool takes a genome assembly as input and predicts per-site nucleosome occupancy as output. Trained on physical maps of nucleosome binding preferences across the fungal kingdom, NuclPred can be applied broadly across fungi (and other eukaryotes). This breadth, combined with its accuracy, means it could have both basic and applied biological implications, for example in understanding eukaryotic gene regulation and genetic engineering. Almost universally across eukaryotes, nucleosomes - each wrapping ~150 base pairs of DNA - serve to package DNA inside the nucleus, with major consequences on DNA access, gene activity and DNA integration. NuclPred was generated using a supervised deep learning approach combining convolutional and recurrent neural networks to take DNA features (nucleotides, GC content and structural information) as input, then use that information to predict the physical attractiveness DNA sequences might have for forming nucleosomes. With this information at hand, researchers can design more efficient CRISPR constructs, explore the interplay between DNA signatures and other regulators impact nucleosome locations, predict expression patterns, etc. This tool will be published as part of a manuscript currently under revision at iScience (draft attached).

Mondo, Stephen

Structural characterization and regulatory element analysis of the heart isoform of cytochrome c oxidase VIa

In order to investigate the mechanism(s) governing the striated muscle-specific expression of cytochrome c oxidase VIaH we have characterized the murine gene and analyzed its transcriptional regulatory elements in skeletal myogenic cell lines. The gene is single copy, spans 689 base pairs (bp), and is comprised of three exons. The 5'-ends of transcripts from the gene are heterogeneous, but the most abundant transcript includes a 5'-untranslated region of 30 nucleotides. When fused to the luciferase reporter gene, the 3.5-kilobase 5'-flanking region of the gene directed the expression of the heterologous protein selectively in differentiated Sol8 cells and transgenic mice, recapitulating the pattern of expression of the endogenous gene. Deletion analysis identified a 300-bp fragment sufficient to direct the myotube-specific expression of luciferase in Sol8 cells. The region lacks an apparent TATA element, and sequence motifs predicted to bind NRF-1, NRF-2, ox-box, or PPAR factors known to regulate other nuclear genes encoding mitochondrial proteins are not evident. Mutational analysis, however, identified two cis-elements necessary for the high level expression of the reporter protein: a MEF2 consensus element at -90 to -81 bp and an E-box element at -147 to -142 bp. Additional E-box motifs at closely located positions were mutated without loss of transcriptional activity. The dependence of transcriptional activation of cytochrome c oxidase VIaH on cis-elements similar to those found in contractile protein genes suggests that the striated muscle-specific expression is coregulated by mechanisms that control the lineage-specific expression of several contractile and cytosolic proteins.

Non-NASA Center