Search NASA⌕ Search

SEARCH · Search NASA

Results for “epigenomics”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

Statistical relationships across epigenomes using large-scale hierarchical clustering

Recent advances in genomics and sequencing platforms have revolutionized our ability to create immense data sets, particularly for studying epigenetic regulation of gene expression. However, the avalanche of epigenomic data is difficult to parse for biological interpretation given nonlinear complex patterns and relationships. This attractive challenge in epigenomic data lends itself to machine learning for discerning infectivity and susceptibility. In this study, we explore over 3000 epigenomes of uninfected individuals and provide a framework to characterize the relationships among epigenetic modifiers, their modifiers, genetic loci, and specific immune cell types across all chromosomes using hierarchical clustering. Hierarchical clustering of epigenomic data revealed consistent epigenetic patterns across chromosomes, demonstrating that variation due to epigenetic modifiers is greater than variation between cell types. Gene Ontology and KEGG pathway analyses indicated significant enrichment of genes involved in chromatin remodeling, mRNA splicing, immune responses, and the regulation of microRNAs and snoRNAs. Epigenetic modifiers frequently formed biologically relevant clusters, including the cohesin complex, RNA Polymerase II transcription factors, and PRC2 complex members. These clustering behaviors remained consistent across all chromosomes, supported by entropy analysis and high Adjusted Rand Index scores, indicating robust cross-chromosomal similarity. Co-occurrence analysis further revealed specific sets of modifiers that consistently appeared together within clusters, reflecting shared biological functions and interactions. Validation using another dataset confirmed the reproducibility of these clustering patterns and modifier co-occurrence relationships, underscoring the reliability and generalizability of the methodology.

97 MATHEMATICS AND COMPUTING↗

Integrative analysis of the 3D genome and epigenome in mouse embryonic tissues

While a rich set of putative cis-regulatory sequences involved in mouse fetal development have been annotated recently on the basis of chromatin accessibility and histone modification patterns, delineating their role in developmentally regulated gene expression continues to be challenging. To fill this gap, here we mapped chromatin contacts between gene promoters and distal sequences across the genome in seven mouse fetal tissues and across six developmental stages of the forebrain. We identified 248,620 long-range chromatin interactions centered at 14,138 protein-coding genes and characterized their tissue-to-tissue variations and developmental dynamics. Integrative analysis of the interactome with previous epigenome and transcriptome datasets from the same tissues revealed a strong correlation between the chromatin contacts and chromatin state at distal enhancers, as well as gene expression patterns at predicted target genes. We predicted target genes of 15,098 candidate enhancers and used them to annotate target genes of homologous candidate enhancers in the human genome that harbor risk variants of human diseases. We present evidence that schizophrenia and other adult disease risk variants are frequently found in fetal enhancers, providing support for the hypothesis of fetal origins of adult diseases.

59 BASIC BIOLOGICAL SCIENCES↗

The genomic and epigenomic evolutionary history of papillary renal cell carcinomas

Intratumor heterogeneity (ITH) and tumor evolution have been well described for clear cell renal cell carcinomas (ccRCC), but they are less studied for other kidney cancer subtypes. Here we investigate ITH and clonal evolution of papillary renal cell carcinoma (pRCC) and rarer kidney cancer subtypes, integrating whole-genome sequencing and DNA methylation data. In 29 tumors, up to 10 samples from the center to the periphery of each tumor, and metastatic samples in 2 cases, enable phylogenetic analysis of spatial features of clonal expansion, which shows congruent patterns of genomic and epigenomic evolution. In contrast to previous studies of ccRCC, in pRCC, driver gene mutations and most arm-level somatic copy number alterations (SCNAs) are clonal. These findings suggest that a single biopsy would be sufficient to identify the important genetic drivers and that targeting large-scale SCNAs may improve pRCC treatment, which is currently poor. While type 1 pRCC displays near absence of structural variants (SVs), the more aggressive type 2 pRCC and the rarer subtypes have numerous SVs, which should be pursued for prognostic significance.

59 BASIC BIOLOGICAL SCIENCES↗

Unified epigenomic, transcriptomic, proteomic, and metabolomic taxonomy of Alzheimer’s disease progression and heterogeneity

Alzheimer’s disease (AD) is a heterogeneous disorder with abnormalities in multiple biological domains. In an advanced machine learning analysis of postmortem brain and in vivo blood multi-omics molecular data ( N = 1863), we integrated epigenomic, transcriptomic, proteomic, and metabolomic profiles into a multilevel biological AD taxonomy. We obtained a personalized multilevel molecular index of AD dementia progression that predicts severity of neuropathologies, and identified three robust molecular-based subtypes that explain much of the pathologic and clinical heterogeneity of AD. These subtypes present distinct patterns of alteration in DNA methylation, RNA, proteins, and metabolites, identifiable in the brain and subsequently in blood. In addition, the genetic variations that predispose to the various AD subtypes in brain predict distinct spatial patterns of alteration in cell types, suggesting a unique influence of each putative AD variant on neuropathological mechanisms. These observations support that an individually tailored multi-omics molecular taxonomy of AD may represent distinct targets for preventive or treatment interventions.

60 APPLIED LIFE SCIENCES↗

Human Liver Epithelial Cells (HuH7) Response to HCoV-229E Infection Epigenomics (ATAC-Seq) (ACS-DP4)

The purpose of this experiment was to evaluate how wild-type Human coronavirus strain 229E (HCoV-299E) infection alters chromatin accessibility in infected cells. Sample data was obtained from mock-infected cells, UV-inactivated virus treated cells, and replication competent HCoV-229E infected immortalized human liver cells (HuH7) at 24 hours post infection. Samples were processed using ATAC-seq methods for reported bar coded libraries. Sample data was acquired using an Illumina Hi-Seq 2500 sequencer system and further processed for ATAC-Seq expression analysis.

59 BASIC BIOLOGICAL SCIENCES↗

Human Liver Epithelium Response to HCoV-229E Infection Epigenomics (ACS-DP4)

The purpose of this experiment was to evaluate how wild-type Human coronavirus strain 229E (HCoV-229E) infection alters chromatin accessibility in infected cells only. Sample data was obtained for mock and infected (standard and UV-inactivated) immortalized human liver cells (HuH-7) and collected 24 hrs. post infection. Samples were processed using assay for transposase-accessible chromatin using high-throughput sequencing (ATAC-Seq) and generated bar coded library samples were evaluated for RNA sequencing (RNA-Seq) expression analysis. Processed ATAC-Seq datasets are openly accessible from the download button and contain secondary processed RNA-Seq results files and supporting metadata materials. Data download includes a sample naming key, infection titer metadata, normalized counts, and relevant computational source code information supporting data transparency and reuse.

59 BASIC BIOLOGICAL SCIENCES↗

Exploring Plant Cis –Regulatory Elements at Single–Cell Resolution: Overcoming Biological and Computational Challenges to Advance Plant Research

Cis-regulatory elements (CREs) are important sequences for gene expression and for plant biological processes such as development, evolution, domestication, and stress response. However, studying CREs in plant genomes has been challenging. The totipotent nature of plant cells, coupled with inability to maintain plant cell types in culture and the inherent technical challenges posed by the cell wall have limited our understanding of how plant cell types acquire and maintain their identities and respond to the environment via CRE usage. Furthermore, advances in single cell epigenomics have revolutionized the field identifying cell-type-specific CREs. These new technologies have the potential to significantly advance our understanding of plant CRE biology, and shed light on how the regulatory genome gives rise to diverse plant phenomena. However, there are significant biological and computational challenges associated with analyzing single cell epigenomic datasets. In this review, we discuss the historical and foundational underpinnings of plant single-cell research, challenges and common pitfalls in analysis of plant single-cell epigenomic data, and highlight biological challenges unique to plants. Additionally, we discuss how the application of single-cell epigenomic data in various contexts stands to transform our understanding of the importance of CREs in plant genomes.

59 BASIC BIOLOGICAL SCIENCES↗

The changing mouse embryo transcriptome at whole tissue and single-cell resolution

During mammalian embryogenesis, differential gene expression gradually builds the identity and complexity of each tissue and organ system 1 . Here we systematically quantified mouse polyA-RNA from day 10.5 of embryonic development to birth, sampling 17 tissues and organs. The resulting developmental transcriptome is globally structured by dynamic cytodifferentiation, body-axis and cell-proliferation gene sets that were further characterized by the transcription factor motif codes of their promoters. We decomposed the tissue-level transcriptome using single-cell RNA-seq (sequencing of RNA reverse transcribed into cDNA) and found that neurogenesis and haematopoiesis dominate at both the gene and cellular levels, jointly accounting for one-third of differential gene expression and more than 40% of identified cell types. By integrating promoter sequence motifs with companion ENCODE epigenomic profiles, we identified a prominent promoter de-repression mechanism in neuronal expression clusters that was attributable to known and novel repressors. Focusing on the developing limb, single-cell RNA data identified 25 candidate cell types that included progenitor and differentiating states with computationally inferred lineage relationships. We extracted cell-type transcription factor networks and complementary sets of candidate enhancer elements by using single-cell RNA-seq to decompose integrative cis -element (IDEAS) models that were derived from whole-tissue epigenome chromatin data. These ENCODE reference data, computed network components and IDEAS chromatin segmentations are companion resources to the matching epigenomic developmental matrix, and are available for researchers to further mine and integrate.

59 BASIC BIOLOGICAL SCIENCES↗

Genome-wide DNA methylation patterns harbour signatures of hatchling sex and past incubation temperature in a species with environmental sex determination

Conservation of thermally sensitive species depends on monitoring organismal and population-level responses to environmental change in real time. Epigenetic processes are increasingly recognized as key integrators of environmental conditions into developmentally plastic responses, and attendant epigenomic data sets hold potential for revealing cryptic phenotypes relevant to conservation efforts. Here, we demonstrate the utility of genome-wide DNA methylation (DNAm) patterns in the face of climate change for a group of especially vulnerable species, those with temperature-dependent sex determination (TSD). Due to their reliance on thermal cues during development to determine sexual fate, contemporary shifts in temperature are predicted to skew offspring sex ratios and ultimately destabilize sensitive populations. Using reduced-representation bisulphite sequencing, we profiled the DNA methylome in blood cells of hatchling American alligators (Alligator mississippiensis), a TSD species lacking reliable markers of sexual dimorphism in early life stages. We identified 120 sex-associated differentially methylated cytosines (DMCs; FDR < 0.1) in hatchlings incubated under a range of temperatures, as well as 707 unique temperature-associated DMCs. We further developed DNAm-based models capable of predicting hatchling sex with 100% accuracy (in 20 training samples and four test samples) and past incubation temperature with a mean absolute error of 1.2°C (in four test samples) based on the methylation status of 20 and 24 loci, respectively. Though largely independent of epigenomic patterning occurring in the embryonic gonad during TSD, DNAm patterns in blood cells may serve as nonlethal markers of hatchling sex and past incubation conditions in conservation applications. These findings also raise intriguing questions regarding tissue-specific epigenomic patterning in the context of developmental plasticity.

59 BASIC BIOLOGICAL SCIENCES↗

MAPLE v.1.0

SAND2025-00659O MAPLE is a software tool that uses epigenomic data to predict gene expression. MAPLE uses a set of epigenomic modifications to determine the effect on gene expression in a specific subset of species. The algorithm can be trained on additional species and epigenomic modifications, enhancing its predictive capabilities. EAGLE employs a hybrid neural network architecture, featuring a convolutional front-end and a multi-head attention layer, to process pre-processed signal data as input. Sandia National Laboratories is a multimission laboratory managed and operated by National Technology & Engineering Solutions of Sandia, LLC, a wholly owned subsidiary of Honeywell International Inc., for the U.S. Department of Energy's National Nuclear Security Administration under contract DE-NA0003525.

Davis IV, Warren↗

From 2D to 4D: a containerized workflow and browser to explore dynamic chromatin architecture

Background Characterizing the physical organization of the genome is essential for understanding long-range gene regulation, chromatin compartmentalization, and epigenetic accessibility. Hi-C experiments generate two-dimensional (2D) genome-wide contact maps of chromatin interactions by capturing the spatial proximity between genomic loci, which reveal interaction frequencies but lack the spatial resolution needed to interpret the three-dimensional (3D) genome structure(s). Emerging evidence suggests that epigenetic regulation is closely linked to 3D genome architecture, and that structural changes over time (4D) drive key biological processes in development, disease, and environmental response. Thus, integrating 3D structure with functional data is critical for a more complete understanding of genome regulation. Previous work, most notably the 4DHiC chromosome modeling framework, has shown that physical multi-dimensional modeling approaches rooted in polymer physics and molecular dynamics can resolve these structures at biologically meaningful resolutions by integrating temporal Hi-C data with physical constraints to uncover dynamic chromosome reorganization. Thus, molecular dynamics simulations, constrained by Hi-C contact matrices, can resolve fine-scale structural changes and reveal functionally significant transitions in chromatin conformation. Results Herein, we present the 4D Genome Browser Workflow (4DGBWorkflow) and the 4D Genome Browser (4DGB). The algorithm is based on the 4DHiC method, and the containerized tool is an end-to-end workflow that can transform, filter, and view 4D epigenomics and chromatin datasets, allowing non-specialists to apply three-dimensional modeling principles to diverse datasets and experimental conditions. The software executes on a laptop running macOS, Linux or Windows. From input Hi-C files (.hic), the 4DGBWorkflow produces 3D reconstructions of chromosomes, integrates the reconstruction with track data (e.g., epigenetic marks, transcriptome profiles), and provides comparative visualization of the results in a single workflow. Conclusions The 4DGBWorkflow and 4D Genome Browser are open-source tools for comparative analysis and visualization of 4D chromosome datasets, including chromatin architecture and epigenomic signals. Automatic integration of Hi-C data with molecular dynamics democratizes the construction of time resolved 3D genome structures, simplifying complex simulations and data integration schemes.

3D Genome Browser↗

Genome-wide fetalization of enhancer architecture in heart disease

Heart disease is associated with re-expression of key transcription factors normally active only during prenatal development of the heart. However, the impact of this reactivation on the regulatory landscape in heart disease is unclear. Here, we use RNA-seq and ChIP-seq targeting a histone modification associated with active transcriptional enhancers to generate genome-wide enhancer maps from left ventricle tissue from up to 26 healthy controls, 18 individuals with idiopathic dilated cardiomyopathy (DCM), and five fetal hearts. Healthy individuals have a highly reproducible epigenomic landscape, consisting of more than 33,000 predicted heart enhancers. In contrast, we observe reproducible disease-associated changes in activity at 6,850 predicted heart enhancers. Combined analysis of adult and fetal samples reveals that the heart disease epigenome and transcriptome both acquire fetal-like characteristics, with 3,400 individual enhancers sharing fetal regulatory properties. We also provide a comprehensive data resource (http://heart.lbl.gov) for the mechanistic exploration of DCM etiology.

60 APPLIED LIFE SCIENCES↗

Improved quality metrics for association and reproducibility in chromatin accessibility data using mutual information

Correlation metrics are widely utilized in genomics analysis and often implemented with little regard to assumptions of normality, homoscedasticity, and independence of values. This is especially true when comparing values between replicated sequencing experiments that probe chromatin accessibility, such as assays for transposase-accessible chromatin via sequencing (ATAC-seq). Such data can possess several regions across the human genome with little to no sequencing depth and are thus non-normal with a large portion of zero values. Despite distributed use in the epigenomics field, few studies have evaluated and benchmarked how correlation and association statistics behave across ATAC-seq experiments with known differences or the effects of removing specific outliers from the data. Here, we developed a computational simulation of ATAC-seq data to elucidate the behavior of correlation statistics and to compare their accuracy under set conditions of reproducibility. Using these simulations, we monitored the behavior of several correlation statistics, including the Pearson’s R and Spearman’s ρ coefficients as well as Kendall’s τ and Top–Down correlation. We also test the behavior of association measures, including the coefficient of determination R 2 , Kendall’s W, and normalized mutual information. Our experiments reveal an insensitivity of most statistics, including Spear man’s ρ, Kendall’s τ, and Kendall’s W, to increasing differences between simulated ATAC-seq replicates. The removal of co-zeros (regions lacking mapped sequenced reads) between simulated experiments greatly improves the estimates of correlation and association. After removing co-zeros, the R 2 coefficient and normalized mutual information display the best performance, having a closer one-to-one relationship with the known portion of shared, enhanced loci between simulated replicates. When comparing values between experimental ATAC-seq data using a random forest model, mutual information best predicts ATAC-seq replicate relationships. Collectively, this study demonstrates how measures of correlation and association can behave in epigenomics experiments. We provide improved strategies for quantifying relationships in these increasingly prevalent and important chromatin accessibility assays.

59 BASIC BIOLOGICAL SCIENCES↗