Search NASASearch

SEARCH · Search NASA

Results for “massive datasets”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 37 records · Page 2

TxDOT Road Elevation Model Dataset

This dataset provides three formats of Road Elevation Model (REM) data: 3D road line/polygon GeoPackage (GPKG), road lidar LAZ and COPC LAZ, and road digital surface model (DSM) GeoTIFF. Data are produced from the ~50TB TxGIO (formerly TNRIS) state lidar collections. This dataset is currently organized by maintenance section in each TxDOT district. Computation is done on GPU computing resources at Oak Ridge National Laboratory (ORNL), through a Strategic Partnership Project with UT Austin and an NSF ACCESS computing allocation award that enables fast massive data movement between TACC Corral and ORNL CADES/OLCF using Globus. In addition to this release from ORNL, a copy of this dataset can also be downloaded at https://web.corral.tacc.utexas.edu/nfiedata/road3d/.

13 HYDRO ENERGY

Increased inflammation as well as decreased endoplasmic reticulum stress and translation differentiate pancreatic islets from donors with pre-symptomatic stage 1 type 1 diabetes and non-diabetic donors

Aims/hypothesis Progression to type 1 diabetes is associated with genetic factors, the presence of autoantibodies and a decline in beta cell insulin secretion in response to glucose. Very little is known regarding the molecular changes that occur in human insulin-secreting beta cells prior to the onset of type 1 diabetes. Herein, we applied an unbiased proteomics approach to identify changes in proteins and potential mechanisms of islet dysfunction in islet-autoantibody-positive organ donors with pre-symptomatic stage 1 type 1 diabetes (HbA1c ≤42 mmol/mol [6.0%]). We aimed to identify pathways in islets that are indicative of beta cell dysfunction. Methods Multiple islet sections were collected through laser microdissection of frozen pancreatic tissues from organ donors positive for single or multiple islet autoantibodies (AAb + , n=5), and age (±2 years)- and sex-matched non-diabetic (ND) control donors (n=5) obtained from the Network for Pancreatic Organ donors with Diabetes (nPOD). Islet sections were subjected to MS-based proteomics and analysed with label-free quantification followed by pathway and functional annotations. Results Analyses resulted in ~4500 proteins identified with low false discovery rate (<1%), with 2165 proteins reliably quantified in every islet sample. We observed large inter-donor variations that presented a challenge for statistical analysis of proteome changes between donor groups. We therefore focused on only the donors with stage 1 type 1 diabetes who were positive for multiple autoantibodies (mAAb + , n=3) and genetic risk compared with their matched ND controls (n=3) for the final statistical analysis. Approximately 10% of the proteins (n=202) were significantly different (unadjusted p<0.025, q<0.15) for mAAb + vs ND donor islets. The significant alterations clustered around major functions for upregulation in the immune response and glycolysis, and downregulation in endoplasmic reticulum (ER) stress response as well as protein translation and synthesis. The observed proteome changes were further supported by several independent published datasets, including a proteomics dataset from in vitro proinflammatory cytokine-treated human islets and single-cell RNA-seq datasets from AAb + individuals. Conclusions/interpretation In situ human islet proteome alterations in stage 1 type 1 diabetes centred around several major functional categories, including an expected increase in immune response genes (elevated antigen presentation/HLA), with decreases in protein synthesis and ER stress response, as well as compensatory metabolic response. The dataset serves as a proteomics resource for future studies on beta cell changes during type 1 diabetes progression and pathogenesis. Data availability The LC-MS raw datasets that support the findings of this study have been deposited in the online repository: MassIVE (https://massive.ucsd.edu/ProteoSAFe/static/massive.jsp) with accession no. MSV000090212.

Autoantibody-positive

In Silico Human Mobility Data Science: Leveraging Massive Simulated Mobility Data (Vision Paper)

Human mobility data science using trajectories or check-ins of individuals has many applications. Recently, we have seen a plethora of research efforts that tackle these applications. However, research progress in this field is limited by a lack of large and representative datasets. The largest and most commonly used dataset of individual human trajectories captures fewer than 200 individuals, while datasets of individual human check-ins capture fewer than 100 check-ins per city per day. Thus, it is not clear if findings from the human mobility data science community would generalize to large populations. Since obtaining massive, representative, and individual-level human mobility data is hard to come by due to privacy considerations, the vision of this work is to embrace the use of data generated by large-scale socially realistic microsimulations. Informed by both real data and leveraging social and behavioral theories, massive spatially explicit microsimulations may allow us to simulate entire megacities at the person level. The simulated worlds, which do not capture any identifiable personal information, allow us to perform “in silico” experiments using the simulated world as a sandbox in which we have perfect information and perfect control without jeopardizing the privacy of any actual individual. In silico experiments have become commonplace in other scientific domains such as chemistry and biology, permitting experiments that foster the understanding of concepts without any harm to individuals. This work describes challenges and opportunities for leveraging massive and realistic simulated alternate worlds for in silico human mobility data science.

97 MATHEMATICS AND COMPUTING

TrioSim: A Lightweight Simulator for Large-Scale DNN Workloads on Multi-GPU Systems

Deep Neural Networks (DNNs) have become increasingly capable of performing tasks ranging from image recognition to content generation. The training and inference of DNNs heavily rely on GPUs, as GPUs' massively parallel architecture delivers extremely high computing capability. With the growing complexity of DNNs and the size of training datasets, training DNNs with a large number of GPUs is becoming a prevalent strategy. Researchers have been exploring how to design software and hardware systems for GPU farms to achieve the best utilization, efficiency, and DNN accuracy during training or inference. However, when designing and deploying such systems, designers usually rely on testing on physical hardware platforms equipped with many GPUs, incurring high costs that are almost prohibitive for system designers to test different configurations and designs, even for highly resourceful companies. While an alternative solution is to test on GPU simulators, they are often too slow for these l

Li, Ying [William & Mary, Williamsburg, VA, USA] (

A Million Person Study Innovation: Evaluating Cognitive Impairment and other Morbidity Outcomes from Chronic Radiation Exposure Through Linkages with the Centers for Medicaid and Medicare Services Assessment and Claims Data

Here, the study of One Million U.S. Radiation Workers and Veterans, the Million Person Study (MPS), examines the health consequences, both cancer and non-cancer, of exposure to ionizing radiation received gradually over time. Recently the MPS has focused on mortality patterns from neurological and behavioral conditions, e.g., Parkinson's disease, Alzheimer's disease, dementia, and motor neuron disease such as amyotrophic lateral sclerosis. A fuller picture of radiation-related late effects comes from studying both mortality and the occurrence (incidence) of conditions not leading to death. Accordingly, the MPS is identifying neurocognitive diagnoses from fee-for-service insurance claims from the Centers for Medicare and Medicaid Services (CMS), among Medicare beneficiaries beginning in 1999 (the earliest date claims data are available). Linkages to date have identified ∼540,000 workers with available health information. Such linkages provide individual information on important co-factor and confounding variables such as smoking, alcohol consumption, blood pressure, obesity, diabetes and many other health and demographic characteristics. The total person-level set of time-dependent variables, outcomes, organ-specific dose measures, co-factors, and demographics will be massive and much too large to be evaluated with standard software. Thus, development of specialized open-source software designed for large datasets (Colossus) is nearly complete. The wealth of information available from CMS claims data, coupled with individual dose reconstructions, will thus greatly enhance the quality and precision of health evaluations for this new field of low-dose radiation and neurocognitive effects.

Dauer, Lawrence T.

Organisation of Diverse Mechanisms of Secondary Ice Production among Basic Convective and Stratiform Cloud-types

This 3-year DoE-funded joint project had the over-arching aim of understanding how ice is initiated in clouds of various types. Focus was given to processes of fragmentation of pre-existing ice, which can occur in positive feedback loops (‘ice multiplication’). A basic question to address was which fragmentation processes prevail in which basic cloud-types. The approach was to use cloud models and field observations, while pioneering our own lab observations of ice initiation to break the deadlock from the past lack of lab observations. Historically, the tendency of the cloud physics community to avoid doing lab observations has allowed a vast gap in knowledge about ice initiation to persist for decades. During the first part of the project, new formulations were created to treat two overlooked types of fragmentation of ice. First, sublimational breakup of ice was treated based on a theoretical formula that we fitted to a pooled dataset of lab observations published previously in the literature. Second, a new mode of fragmentation of freezing raindrops was treated, which involves a supercooled drop being hit by a more massive ice particle. Some of the secondary droplets from the impact freeze. This work was done at Manchester University by Co-I Connolly. Then during the second part, both formulations were implemented in our ‘aerosol-cloud model’ (AC). AC has a hybrid bin/bulk microphysics scheme, and now represents four processes of SIP. The accuracy of AC was evaluated for four cases typifying four basic cloud-types: slightly cold-based stratiform cloud and cold-, warm- and very warm-based convective clouds. We discovered that the warmth of cloud-base, especially in the tropics, promotes SIP processes of raindrop-freezing fragmentation and rime-splintering, and surprisingly, sublimational breakup too. It was found that breakup in ice-ice collisions is ubiquitous. Finally, a portable laboratory chamber was constructed at Lund and deployed in northern Sweden to observe breakup in graupel-snow collisions outdoors. This was seen to be even more prolific than treated in our 2018 formulation. Papers describing results are either published or soon to be published.

54 ENVIRONMENTAL SCIENCES

Energy dataset of Frontier supercomputer for waste heat recovery

The Hewlett Packard Enterprise–Cray EX Frontier is the world’s first and fastest exascale supercomputer, hosted at the Oak Ridge Leadership Computing Facility in Tennessee, United States. Frontier is a significant electricity consumer, drawing 8–30 MW; this massive energy demand produces significant waste heat, requiring extensive cooling measures. Although harnessing this waste heat for campus heating is a sustainability goal at Oak Ridge National Laboratory (ORNL), the 30 °C–38 °C waste heat temperature poses compatibility issues with standard HVAC systems. Heat pump systems, prevalent in residential settings and some industries, can efficiently upgrade low-quality heat to usable energy for buildings. Thus, heat pump technology powered by renewable electricity offers an efficient, cost-effective solution for substantial waste heat recovery. However, a major challenge is the absence of benchmark data on high-performance computing (HPC) heat generation and waste heat profiles. This paper reports power demand and waste heat measurements from an ORNL HPC data centre, aiming to guide future research on optimizing waste heat recovery in large-scale data centres, especially those of HPC calibre.

97 MATHEMATICS AND COMPUTING

An expanded registry of candidate cis -regulatory elements

Mammalian genomes contain millions of regulatory elements that control the complex patterns of gene expression. Previously, the ENCODE consortium mapped biochemical signals across hundreds of cell types and tissues and integrated these data to develop a registry containing 0.9 million human and 300,000 mouse candidate cis-regulatory elements (cCREs) annotated with potential functions. Here we have expanded the registry to include 2.37 million human and 967,000 mouse cCREs, leveraging new ENCODE datasets and enhanced computational methods. This expanded registry covers hundreds of unique cell and tissue types, providing a comprehensive understanding of gene regulation. Functional characterization data from assays such as STARR-seq, massively parallel reporter assay, CRISPR perturbation and transgenic mouse assays have profiled more than 90% of human cCREs, revealing complex regulatory functions. We identified thousands of novel silencer cCREs and demonstrated their dual enhancer and silencer roles in different cellular contexts. Integrating the registry with other ENCODE annotations facilitates genetic variation interpretation and trait-associated gene identification, exemplified by the identification of KLF1 as a novel causal gene for red blood cell traits. This expanded registry is a valuable resource for studying the regulatory genome and its impact on health and disease.

Moore, Jill E. [Univ. of Massachusetts, Worchester

Positive Neutrino Masses with DESI DR2 via Matter Conversion to Dark Energy

The Dark Energy Spectroscopic Instrument (DESI) is a massively parallel spectroscopic survey on the Mayall telescope at Kitt Peak, which has released measurements of baryon acoustic oscillations determined from over 14 million extragalactic targets. We combine DESI Data Release 2 with CMB datasets to search for evidence of matter conversion to dark energy (DE), focusing on a scenario mediated by stellar collapse to cosmologically coupled black holes (CCBHs). In this physical model, which has the same number of free parameters as Λ⁢CDM, DE production is determined by the cosmic star formation rate density (SFRD), allowing for distinct early- and late-time cosmologies. Using two SFRDs to bracket current observations, we find that the CCBH model: accurately recovers the cosmological expansion history, agrees with early-time baryon abundance measured by BBN, reduces tension with the local distance ladder, and relaxes constraints on the summed neutrino mass ∑𝑚 𝜈 . For these SFRDs, we find a peaked positive ∑𝑚 𝜈 < 0.149 eV (95% confidence) and ∑𝑚 𝜈 = 0.106$^{+0.050}_{−0.069}$ eV, respectively, in good agreement with lower limits from neutrino oscillation experiments. A peak in ∑𝑚 𝜈 > 0 results from late-time baryon consumption in the CCBH scenario and is expected to be a general feature of any model that converts sufficient matter to dark energy during and after reionization.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS

An Autonomous MCP Bridge to Rucio: Enhancing Data Management Accessibility for High Energy Physics

The Rucio Data Management System [1] is an important tool used by High Energy Physics experiments, including those at Fermi National Accelerator Laboratory, to store and manage exabyte-scale scientific datasets. Despite its central role in coordinating data across globally distributed storage sites, Rucio's command line interface (CLI) presents a steep learning curve, and makes it difficult for scientists to navigate through. To solve this issue, a containerized Model Context Protocol (MCP) [2] server was built that connects Large Language Models directly to Rucio, allowing AI agents to handle data tasks by using simple, natural language rather than memorized terminal commands. The core engineering focus of this project was moving the server away from slow terminal commands that require text parsing and replacing them with a native Python Client API toolset and a planned REST API framework. Moving to the Python API handles data operations directly in memory, which helps clear up formatting errors, provides the AI with clean, structured JSON data and speeds up tool execution. To prove that the system actually works, a benchmarking pipeline was also built with various questions to test the AI across four different model configurations. The questions included finding data scopes, tracking down specific datasets, and checking replication rules. Through benchmarking, early runs showed that with raw terminal text, the model would get confused and stuck, whereas switching to the Python API to feed the AI clean, structured data yielded massive improvement. By creating an intelligent and autonomous bridge to a storage network, this project shows how AI can be implemented in scientific data management, which ultimately helps scientists at Fermilab spend less time sorting through data and more time focusing on their experiments and analysis.

Akella, Kashyap [William Rainey Harper Coll.]

Moving lens effect: Simulations, forecasts, and foreground mitigation

The peculiar motion of massive objects across the line of sight imprints a dipolar temperature anisotropy pattern on the cosmic microwave background known as the moving lens effect. This effect provides a unique probe of the transverse components of the peculiar velocity field, but has not yet been detected due to its small size. We implement and validate a stacking estimator for the moving lens signal using a galaxy catalog as a tracer of massive haloes combined with reconstructed velocities from the galaxy number density field. Using simulations, we forecast detection prospects for the moving lens signal from current and upcoming microwave background and galaxy surveys. Here, we demonstrate a new foreground mitigation strategy likely sufficient for current datasets, and discuss various sources of systematic error and noise. Upcoming galaxy surveys will provide high-significance statistical detections of the moving lens effect.

Astrophysical & cosmological simulations

Harnessing on-machine metrology data for prints with a surrogate model for laser powder directed energy deposition

In this study, we leverage the massive amount of multi-modal on-machine metrology data generated from Laser Powder Directed Energy Deposition (LP-DED) to construct a comprehensive surrogate model of the 3D printing process. By employing Dynamic Mode Decomposition with Control (DMDc), a data-driven technique, we capture the complex physics inherent in this extensive dataset. This physics-based surrogate model emphasizes thermodynamically significant quantities, enabling us to accurately predict key process outcomes. The model ingests 21 process parameters, including laser power, scan rate, and position, while providing outputs such as melt pool temperature, melt pool size, and other essential observables. Furthermore, it incorporates uncertainty quantification to provide bounds on these predictions, enhancing reliability and confidence in the results. We then deploy the surrogate model on a new, unseen part and monitor the printing process as validation of the method. Our experimental results demonstrate that the predictions align with actual measurements with high accuracy, confirming the effectiveness of our approach. Furthermore, this methodology not only facilitates real-time predictions but also operates at process-relevant speeds, establishing a basis for implementing feedback control in LP-DED.

Digital twins

The stellar origins of 96 Zr excesses in presolar graphites from the Murchison meteorite

Context. Zirconium-96 is a stable isotope that can be synthesized under different neutron-rich nucleosynthetic conditions. Astrophysical models predict its production to occur in various stellar environments: from low-to-intermediate-mass asymptotic giant branch (AGB) stars to massive stars and core-collapse supernovae. Aims. Detections of 96Zr excesses, in combination with other isotopic measurements from presolar grains can provide unique constraints on its stellar origin. Presolar grains are microscopic particles found in primitive Solar System materials, which formed in stellar winds and supernova ejecta. The isotopic composition of each grain can provide us a snapshot of the nucleosynthetic processes that took place during the parent star’s lifetime. Methods. In this study, we measured the stable isotopes of C, N, O, Mo, Zr, and Ru in high-density presolar graphite grains from the Murchison meteorite and found four grains that contain positive isotopic anomalies in 96 Zr carried by their internal subgrains. We analyzed multi-element isotopic datasets from each grain to explore the source of the observed 96 Zr excesses. Results. Comparisons with stellar models indicate that two grains likely condensed in an intermediate-mass AGB star with initial metallicity of Z ≤ 0.014. Their 96 Zr/ 94 Zr ratios also match those predicted for born-again AGB stars undergoing a very late thermal pulse and rapidly accreting white dwarfs. After comparing the relative populations of the aforementioned dust-producing stars, we propose rapidly accreting white dwarfs as a new, and more likely, stellar source for one of the presolar grains. The remaining two grains could have originated in the supernova ejecta of massive stars, due to correlated excesses in the p-nuclides, 92, 94 Mo. Thus, grains with 96 Zr anomalies can have a variety of stellar origins, in agreement with theoretical studies. Conclusions. Our study highlights the importance of multi-element analysis in constraining the types of stars where presolar grains have condensed. These data will help improve our understanding of various nucleosynthesis processes in different stellar phases.

79 ASTRONOMY AND ASTROPHYSICS

SPT clusters with DES and HST weak lensing. II. Cosmological constraints from the abundance of massive halos

We present cosmological constraints from the abundance of galaxy clusters selected via the thermal Sunyaev-Zel’dovich (SZ) effect in South Pole Telescope (SPT) data with a simultaneous mass calibration using weak gravitational lensing data from the Dark Energy Survey (DES) and the Hubble Space Telescope (HST). The cluster sample is constructed from the combined SPT-SZ, SPTpol ECS, and SPTpol 500d surveys, and comprises 1,005 confirmed clusters in the redshift range 0.25–1.78 over a total sky area of 5200 deg 2 . We use DES Year 3 weak-lensing data for 688 clusters with redshifts 𝑧 < 0.95 and HST weak-lensing data for 39 clusters with 0.6 < 𝑧 < 1.7. The weak-lensing measurements enable robust mass measurements of sample clusters and allow us to empirically constrain the SZ observable-mass relation without having to make strong assumptions about, e.g., the hydrodynamical state of the clusters. For a flat Λ⁢ CDM cosmology, and marginalizing over the sum of massive neutrinos, we measure Ω m = 0.286 ± 0.032, 𝜎 8 = 0.817 ± 0.026, and the parameter combination 𝜎 8 ⁢(Ω m /0.3) 0.25 = 0.805 ± 0.016. Our measurement of 𝑆 8 ≡ 𝜎 8 ⁢$\sqrt{Ω_{m}/0.3}$ = 0.795 ± 0.029 and the constraint from Planck CMB anisotropies (2018 TT, TE, EE+lowE) differ by 1.1⁢𝜎. In combination with that Planck dataset, we place a 95% upper limit on the sum of neutrino masses ∑𝑚 𝜈 < 0.18 eV. When additionally allowing the dark energy equation of state parameter 𝑤 to vary, we obtain 𝑤 = −1.45 ± 0.31 from our cluster-based analysis. In combination with Planck data, we measure 𝑤 =−1.3⁢4$^{+0.22}_{−0.15}$, or a 2.2⁢𝜎 difference with a cosmological constant. We use the cluster abundance to measure 𝜎8 in five redshift bins between 0.25 and 1.8, and we find the results to be consistent with structure growth as predicted by the Λ⁢ CDM model fit to Planck primary CMB data.

79 ASTRONOMY AND ASTROPHYSICS

Using Temporal Information from Human Mobility Data to Detect Anchor Points

Spatiotemporal mobility data are available in massive quantities, but large quantities of data typically include fewer variables or data fields. Often, the only available fields are User ID, Longitude, Latitude, Timestamp (ULLT). This raises an important question: how much can we infer about human mobility patterns using only these four fields? With ULLT data, we do not know individuals' socioeconomic status information or when they are visiting their anchor points (AP) or locations (such as homes, places of employment, or schools), and it is a modern challenge to use this data to infer these characteristics. When detecting anchor locations with limited input information, verification and validation (VV) are significant challenges. This paper addresses the problem of identifying individuals' anchor locations using only temporal information from spatiotemporal datasets with limited attributes. Our approach does not explicitly use latitude and longitude during analysis. Locationbased information is only employed in the preprocessing stage to identify periods of movement (trips) and stops (dwelling). Beyond this step, all analysis is based on temporal patterns. In theory, if stops and dwell times could be detected through alternative means, our method could function entirely without location-based input. We demonstrate this methodology on the 2017 National Household Travel Survey (NHTS) data, because it includes a carefully designed and collected time use survey with representative sampling and labeled ground truth. The high-quality survey data allows us to test the accuracy of our methods because NHTS contains intended place labels and agent/user characteristics. We have also applied our validated AP identification algorithm on very large-scale GPS based trajectory data for Patterns-of-Life (PoL) assessment and other applications, but due to space limit that could not be presented here.

McBride, Liz [ORNL] (ORCID:0000000286925869)

Massively parallel reporter assays and mouse transgenic assays provide correlated and complementary information about neuronal enhancer activity

High-throughput massively parallel reporter assays (MPRAs) and phenotype-rich in vivo transgenic mouse assays are two potentially complementary ways to study the impact of noncoding variants associated with psychiatric diseases. Here, we investigate the utility of combining these assays. Specifically, we carry out an MPRA in induced human neurons on over 50,000 sequences derived from fetal neuronal ATAC-seq datasets and enhancers validated in mouse assays. We also test the impact of over 20,000 variants, including synthetic mutations and 167 common variants associated with psychiatric disorders. We find a strong and specific correlation between MPRA and mouse neuronal enhancer activity. Four out of five tested variants with significant MPRA effects affected neuronal enhancer activity in mouse embryos. Mouse assays also reveal pleiotropic variant effects that could not be observed in MPRA. Our work provides a catalog of functional neuronal enhancers and variant effects and highlights the effectiveness of combining MPRAs and mouse transgenic assays.

Kosicki, Michael

The Artificial Scientist: in-Transit Machine Learning of Plasma Simulations

Large-scale simulations or scientific experiments produce petabytes of data per run. This poses massive challenges for I/O and storage when scientific analysis workflows are run manually offline. Unsupervised deep learning-based techniques to extract patterns and non-linear relations from these large amounts of data provide a way to build scientific understanding from raw data, reducing the need for manual pre-selection of analysis steps, but require exascale compute and memory to process the full dataset available. In this paper, we demonstrate a heterogeneous streaming workflow in which plasma simulation data is streamed directly to a Machine Learning (ML) application training a model on the simulation data in-transit, completely circumventing the capacity-constrained filesystem bottleneck. This workflow employs openPMD to provide a high level interface to describe scientific data and also uses ADIOS2, to transfer volumes of data that exceed the capabilities of the filesystem. We employ experience replay to avoid catastrophic forgetting in learning from this non-steady state process in a continual manner and adapt it to improve model convergence while learning in-transit. As a proof-of-concept, we approach the ill-posed inverse problem of predicting particle dynamics from radiation in a particle-incell (PIConGPU) simulation of the Kelvin-Helmholtz instability (KHI). We detail hardware-software co-design challenges as we scale PIConGPU to full Frontier, the Top-1 system as of June 2024 Top500 list.

Kelling, Jeffrey [Helmholtz-Zentrum Dresden Rossen

Chromosome-scale Genome Assembly of the Most Abundant Ectomycorrhizal Fungus Cenococcum Geophilum Reveals Massive TE Expansion and RIP Defense Mechanism

Transposable elements (TEs) play crucial roles in genome evolution and ecological adaptation in fungi, yet their dynamics in ectomycorrhizal species remain poorly understood. Cenococcum geophilum, the most widespread ectomycorrhizal fungus in boreal and temperate forests with its large, repeat-rich genome, represents an ideal system to investigate TE-mediated adaptation to the physical environment and symbiotic lifestyle. However, previous studies have been limited by fragmented genome assemblies that prevented the resolution of repeat-rich regions. We assembled a telomere-to-telomere reference genome of C. geophilum strain 1.58 using PacBio HiFi and Hi-C datasets, resulting in a 178.54 Mbp genome with seven contiguous chromosomes. We identified 14,145 genes and over 78% of the genome consists of transposable elements (TEs). Of these, 94% are affected by repeat-induced point mutations (RIP), a genome defense mechanism that acts during the sexual reproduction phase, indicating cryptic or ancient sexual reproduction in this putatively asexual fungus. Long terminal repeat retrotransposons, LINEs, and DNA transposons dominate, with three TE families (Ty3, Ty1, and Tad1) contributing over 60% of the genome size, indicating recent transposition bursts. Screening of 15 additional C. geophilum strains revealed recent and lineage-specific TE expansions, implying that several TEs escaped the RIP machinery and retained potential activity. Supporting TE activity in the context of symbiosis, we found 56 TEs differentially transcribed between ectomycorrhizal and free-living mycelium tissues. An even higher number (n = 66) of TEs were differentially expressed between stress resistance morphology (i.e. sclerotia) and free-living mycelium. This supports that TEs are differentially regulated as a response to symbiotic and stress-related conditions. Our results demonstrate that the C. geophilum genome expansion was driven by a few lineage-specific TE families in recent history, with high RIP activity attesting to sexual reproduction. We also provide insights how TEs could respond to lifestyle transitions and traits associated with desiccation resistance.

Cenococcum geophilum