Search NASA⌕ Search

SEARCH · Search NASA

Results for “cluster analysis”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 325 records · Page 18

Harnessing the predicted maize pan-interactome for putative gene function prediction and prioritization of candidate genes for important traits

Abstract The recent assembly and annotation of the 26 maize nested association mapping population founder inbreds have enabled large-scale pan-genomic comparative studies. These studies have expanded our understanding of agronomically important traits by integrating pan-transcriptomic data with trait-specific gene candidates from previous association mapping results. In contrast to the availability of pan-transcriptomic data, obtaining reliable protein–protein interaction (PPI) data has remained a challenge due to its high cost and complexity. We generated predicted PPI networks for each of the 26 genomes using the established STRING database. The individual genome-interactomes were then integrated to generate core- and pan-interactomes. We deployed the PPI clustering algorithm ClusterONE to identify numerous PPI clusters that were functionally annotated using gene ontology (GO) functional enrichment, demonstrating a diverse range of enriched GO terms across different clusters. Additional cluster annotations were generated by integrating gene coexpression data and gene description annotations, providing additional useful information. We show that the functionally annotated PPI clusters establish a useful framework for protein function prediction and prioritization of candidate genes of interest. Our study not only provides a comprehensive resource of predicted PPI networks for 26 maize genomes but also offers annotated interactome clusters for predicting protein functions and prioritizing gene candidates. The source code for the Python implementation of the analysis workflow and a standalone web application for accessing the analysis results are available at https://github.com/eporetsky/PanPPI.

Genetics & Heredity↗

Pando

SAND2025-02006O Pando is a distributed data analysis software tool. It is designed to handle large-scale graph analysis problems, often with a specific focus on blockchain/cryptocurrency data. Pando handles scalability by running on a distributed cluster of servers. Users can customize the output using the program’s plugin/extension design methodology. Sandia National Laboratories is a multimission laboratory managed and operated by National Technology & Engineering Solutions of Sandia, LLC, a wholly owned subsidiary of Honeywell International Inc., for the U.S. Department of Energy’s National Nuclear Security Administration under contract DE-NA0003525.

Gabert, Kasimir↗

Dense autoencoders, clustering techniques, and semi-supervised learning for HPGe $γ$-spectra

Classifying high-resolution gamma spectra by their isotopic content is an essential task in nuclear forensics and other applications. Traditional analysis methods are often time-intensive, but machine learning (ML) may help analysts quickly process many spectra. Such methods tend to rely on abundant, well-labeled data for training. Historical gamma data exists in various fields but is not uniformly useful for supervised ML due to inconsistent labeling. Here, to address some of these challenges, we present a method to classify and organize unlabeled data from high-purity germanium detectors using an autoencoding neural network (autoencoder). We trained dense autoencoders to compress gamma data into latent representations that enable efficient data characterization. By clustering the encoded spectra or lower-dimensional mappings of them, we identified and removed portions of over-abundant data categories, resulting in a more balanced dataset and improved autoencoder performance. This encoding and clustering pipeline also enabled the organization of spectra into self-consistent categories. Finally, we found that encoded representations showed potential as inputs for semi-supervised learning of nuclide identification (NID) labels, achieving an average F1 score of 0.85 ± 0.03 when mapping encodings to a set of 65 isotope labels.

Autoencoders↗

A User-Friendly GUI Tool for Automated Microstructural Analysis of Fiber-Reinforced Composites and Porous Structures

Understanding and quantifying microstructural features such as fiber orientation and porosity is critical for predicting the mechanical behavior and performance of fiber-reinforced polymer composites. Traditional manual analysis is time-consuming, subjective, and unsuitable for high-throughput datasets. We present a graphical user interface (GUI) application that automates the analysis of microscopy images to extract key microstructural metrics, including fiber orientation tensors, fiber orientation distribution, porosity and pore size distribution. The app integrates multiple image segmentation techniques including global and local thresholding, clustering, and region-based approaches, offering flexibility for different types of image qualities and features. Users can load microstructural images, select regions of interest and segmentation techniques tailored to their image dataset. It also addresses a critical challenge in fiber orientation analysis: the ambiguities caused by touching, overlapping, or partially cut fibers. It supports autorun examples for standardized workflows, enabling reproducible analysis and facilitating training and benchmarking. This tool significantly reduces manual intervention, enhances consistency, and accelerates data generation for structure–property modeling, process optimization, and digital materials research. The tool is intended for use by materials scientists, engineers, and researchers engaged in composite characterization, quality control, and machine learning-based microstructural studies.

Chawla, Komal [ORNL] (ORCID:0000000190327565)↗

Toward an Understanding of Linear Scaling Relations through Energy Decomposition Analysis

The discovery of linear scaling relations has fundamentally changed the field of heterogeneous catalysis. The scaling relations have been rationalized based on the d-band theory, specifically a separation of sp and d electron contributions to adsorption energies. Within the framework of energy decomposition analysis, a full understanding of such a separation would require one to further break down the adsorption energy into distinct energy components such as electrostatics, polarization, charge transfer, and van der Waals interactions, and to examine the sp and d contributions to each of them. As a step in this direction, we analyzed the interaction energy between CH x (x = 1–4) adsorbates and fcc(100) transition metal surfaces (M = Cu, Ag, Au, Rh, and Pt), with the surfaces represented both as slabs in plane-wave density functional theory (pw-DFT) calculations and as atomic clusters in atomic-orbital basis density functional theory (ao-DFT) calculations. Through an absolutely localized molecular orbital (ALMO) based energy decomposition analysis of the ao-DFT adsorption energy, each of the interaction energy components (electrostatics, polarization, van der Waals, and charge transfer) was found to follow its own scaling relations, with an intricate interplay among these energy components yielding the overall scaling relations for the total adsorption energies. Using the recently introduced ALMO-based polarization and charge-transfer analysis schemes, we further dissected polarization into metal surface and adsorbate contributions, and charge transfer into metal → adsorbate and adsorbate → metal contributions. The contributions from the sp and d electrons of the metal to these terms were further quantified, and the dominant role of the metal d electrons was reaffirmed. These results shed light on how CHx adsorbates interact with metal surfaces and further reveal the physical origin of the scaling relations.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

DESI 2024 V: Full-Shape galaxy clustering from galaxies and quasars

We present the measurements and cosmological implications of the galaxy two-point clustering using over 4.7 million unique galaxy and quasar redshifts in the range 0.1 < z < 2.1 divided into six redshift bins over a ∼ 7,500 square degree footprint, from the first year of observations with the Dark Energy Spectroscopic Instrument (DESI Data Release 1). By fitting the full power spectrum, we extend previous DESI DR1 baryon acoustic oscillation (BAO) measurements to include redshift-space distortions and signals from the matter-radiation equality scale. For the first time, this Full-Shape analysis is blinded at the catalogue-level to avoid confirmation bias and the systematic errors are accounted for at the two-point clustering level, which automatically propagates them into any cosmological parameter. When analyzing the data in terms of compressed model-agnostic variables, we obtain a combined precision of 4.7% on the amplitude of the redshift space distortion (RSD) signal reaching a similar precision with just one year of DESI data than with twenty years of observation from the previous generation survey. We also analyze the data to directly constrain the cosmological parameters within the ΛCDM model using perturbation theory and combine this information with the reconstructed DESI DR1 galaxy BAO. Using a Big Bang Nucleosynthesis Gaussian prior on the baryon density parameter, ω b , and a weak Gaussian prior on the spectral index, n s , we constrain the matter density is Ω m = 0.296±0.010 and the Hubble constant H 0 = (68.63 ± 0.79)[km s -1 Mpc -1 ]. Additionally, we measure the amplitude of clustering σ 8 = 0.841±0.034. The DESI DR1 galaxy clustering results are in agreement with the ΛCDM model based on general relativity with parameters consistent with those from Planck. The cosmological interpretation of these results in combination with DESI DR1 Ly-α forest data and external datasets are presented in the companion paper [1].

79 ASTRONOMY AND ASTROPHYSICS↗

Signature analysis of high-throughput transcriptomics screening data for mechanistic inference and chemical grouping

Abstract High-throughput transcriptomics (HTTr) uses gene expression profiling to characterize the biological activity of chemicals in in vitro cell-based test systems. As an extension of a previous study testing 44 chemicals, HTTr was used to screen an additional 1,751 unique chemicals from the EPA’s ToxCast collection in MCF7 cells using 8 concentrations and an exposure duration of 6 h. We hypothesized that concentration-response modeling of signature scores could be used to identify putative molecular targets and cluster chemicals with similar bioactivity. Clustering and enrichment analyses were conducted based on signature catalog annotations and ToxPrint chemotypes to facilitate molecular target prediction and grouping of chemicals with similar bioactivity profiles. Enrichment analysis based on signature catalog annotation identified known mechanisms of action (MeOAs) associated with well-studied chemicals and generated putative MeOAs for other active chemicals. Chemicals with predicted MeOAs included those targeting estrogen receptor (ER), glucocorticoid receptor (GR), retinoic acid receptor (RAR), the NRF2/KEAP/ARE pathway, AP-1 activation, and others. Using reference chemicals for ER modulation, the study demonstrated that HTTr in MCF7 cells was able to stratify chemicals in terms of agonist potency, distinguish ER agonists from antagonists, and cluster chemicals with similar activities as predicted by the ToxCast ER Pathway model. Uniform manifold approximation and projection (UMAP) embedding of signature-level results identified novel ER modulators with no ToxCast ER Pathway model predictions. Finally, UMAP combined with ToxPrint chemotype enrichment was used to explore the biological activity of structurally related chemicals. The study demonstrates that HTTr can be used to inform chemical risk assessment by determining in vitro points of departure, predicting chemicals’ MeOA and grouping chemicals with similar bioactivity profiles.

Toxicology↗

High-Pressure and Temperature Effects on the Clustering Ability of Monohydroxy Alcohols

This study examined the clustering behavior of monohydroxy alcohols, where hydrogen-bonded clusters of up to a hundred molecules on the nanoscale can form. By performing X-ray diffraction experiments at different temperatures and under high pressure, we investigated how these conditions affect the ability of alcohols to form clusters. The pioneering high-pressure experiment performed on liquid alcohols contributes to the emerging knowledge in this field. Implementation of molecular dynamics simulations yielded excellent agreement with the experimental results, enabling the analysis of theoretical models. Here we show that at the same global density achieved either by alteration of pressure or temperature, the local aggregation of molecules at the nanoscale may significantly differ. Surprisingly, high pressure not only promotes the formation of hydrogen-bonded clusters but also induces the serious reorganization of molecules. This research represents a milestone in understanding association under extreme thermodynamic conditions in other hydrogen bonding systems such as water.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Fiducial-cosmology-dependent systematics for the DESI 2024 full-shape analysis

We assess the impact of the fiducial cosmology choice on cosmological inference from full-shape (FS) fits of the galaxy power spectrum in the DESI 2024 Data Release 1 (DR1). Using a suite of AbacusSummit DR1 mock catalogues based on the Planck 2018 best-fit cosmology, we quantify potential systematic shifts introduced by analysing the data under five secondary cosmologies — featuring variations in matter density, thawing dark energy, higher effective number of neutrino species, reduced clustering amplitude, and the DESI DR1 BAO best-fit w 0 w a CDM cosmology — relative to DESI's baseline Planck 2018 cosmology. We investigate two complementary FS analysis approaches: full-modelling (FM) and ShapeFit (SF), each with distinct sensitivities to the assumed fiducial model. Across all tracers, we find for FM that systematic shifts induced by fiducial cosmology mismatches remain well below the DESI DR1 statistical uncertainties, with maximum deviations of 0.22σ DR1 in ΛCDM scenarios and 0.12σ DR1+SN when including SN Ia mock data in extended w 0 w a CDM fits. For SF, the shifts in the compressed parameters remain below 0.45σ DR1 for all tracers and cosmologies.

dark energy experiments↗

MVP: a modular viromics pipeline to identify, filter, cluster, annotate, and bin viruses from metagenomes

While numerous computational frameworks and workflows are available for recovering prokaryote and eukaryote genomes from metagenome data, only a limited number of pipelines are designed specifically for viromics analysis. With many viromics tools developed in the last few years alone, it can be challenging for scientists with limited bioinformatics experience to easily recover, evaluate quality, annotate genes, dereplicate, assign taxonomy, and calculate relative abundance and coverage of viral genomes using state-of-the-art methods and standards. Here, we describe Modular Viromics Pipeline (MVP) v.1.0, a user-friendly pipeline written in Python and providing a simple framework to perform standard viromics analyses. MVP combines multiple tools to enable viral genome identification, characterization of genome quality, filtering, clustering, taxonomic and functional annotation, genome binning, and comprehensive summaries of results that can be used for downstream ecological analyses. Overall, MVP provides a standardized and reproducible pipeline for both extensive and robust characterization of viruses from large-scale sequencing data including metagenomes, metatranscriptomes, viromes, and isolate genomes. As a typical use case, we show how the entire MVP pipeline can be applied to a set of 20 metagenomes from wetland sediments using only 10 modules executed via command lines, leading to the identification of 11,656 viral contigs and 8,145 viral operational taxonomic units (vOTUs) displaying a clear beta-diversity pattern. Further, acting as a dynamic wrapper, MVP is designed to continuously incorporate updates and integrate new tools, ensuring its ongoing relevance in the rapidly evolving field of viromics. MVP is available at https://gitlab.com/ccoclet/mvp and as versioned packages in PyPi and Conda.

59 BASIC BIOLOGICAL SCIENCES↗

Neural network-based model of galaxy power spectrum: fast full-shape galaxy power spectrum analysis

ABSTRACT We present a neural network-based emulator for the galaxy redshift-space power spectrum that enables several orders of magnitude acceleration in the galaxy clustering parameter inference, while preserving 3$\sigma$ accuracy better than 0.5 per cent up to $k_{\mathrm{max}}$ = 0.25 $\, h\text{Mpc}^{-1}$ within Lambda-cold dark matter ($\Lambda$CDM) and around 0.5 per cent $w_0$–$w_a$CDM. Our surrogate model only emulates the galaxy bias-invariant terms of one-loop perturbation theory predictions, these terms are then combined analytically with galaxy bias terms, counter-terms, and stochastic terms in order to obtain the non-linear redshift-space galaxy power spectrum. This allows us to avoid any galaxy bias prescription in the training of the emulator, which makes it more flexible. Moreover, we include the redshift $z \in [0,1.4]$ in the training which further avoids the need for re-training the emulator. We showcase the performance of the emulator in recovering the cosmological parameters of $\Lambda$CDM by analysing the suite of 25 AbacusSummit simulations that mimic the Dark Energy Spectroscopic Instrument luminous red galaxies at $z=0.5$ and 0.8, together as the emission line galaxies at $z=0.8$. We obtain similar performance in all cases, demonstrating the reliability of the emulator for any galaxy sample at any redshift in $0 \lt z \lt 1.4$. We will make our emulator public at github repository.

Trusov, Svyatoslav (ORCID:0000000224146720)↗

Geometry, spin coupling, and dielectric control of redox potentials in [4Fe–4S] Clusters

Iron–sulfur (Fe–S) clusters are common biological cofactors that facilitate vital redox reactions. Despite extensive research, the molecular basis of redox potential tuning in ferredoxin-like proteins remains an active area of debate. In this study, we combine statistical analysis of over one thousand [4Fe–4S]-containing protein structures from the Protein Data Bank (PDB) with broken-symmetry and extended broken-symmetry density functional theory to examine how cysteine ligand orientations and environmental screening affect redox properties of the clusters. We identified five main ligand configurations, three of which are predominant in natural structures. Among these, the adiabatic electron affinity differs by less than 0.1 V, indicating that, while geometry plays a secondary role, it allows localized fine-tuning of redox properties. In contrast, electrostatic and solvation effects primarily determine the overall potential range.

Computational Chemistry↗

Plant genotype and rhizobia strain combinations strongly influence the transcriptome under heavy metal stress conditions in Medicago truncatula

Heavy metals such as cadmium (Cd) and mercury (Hg) pose significant threats to plant health and food safety as they are absorbed from the environment. Legumes are generally considered sensitive to heavy metals but possess standing genetic variation for accumulation and tolerance to toxic ions. We conducted a transcriptomic analysis on hydroponically and soil grown Medicago truncatula plants to investigate gene expression responses to Cd and Hg exposure in roots, leaves, and nodules. By using plant genotypes with varying metal tolerance or accumulation levels, we observed distinct clustering of gene ontologies, indicating tissue-specific, genotype-specific, and metal-specific gene expression patterns. Considering the symbiotic relationship between legumes and nitrogen-fixing bacteria, we further examined plant phenotypes and transcriptomes of plant genotypes with contrasting Hg accumulation levels and inoculated them with high or low Hg-tolerant Sinorhizobium medicae strains that have presence-absence variation for a mercury reductase (Mer) operon. Host plants inoculated with the Hg-tolerant rhizobia strain possessing a Mer operon exhibited less reduction in nodule number and plant biomass. A smaller reduction in iron (Fe) distribution in nodules after Hg stress was measured using X-ray Fluorescence (XRF) imaging. Dual transcriptome (host plant and bacteria) analysis of nodules revealed a remarkable decrease in the number of differentially expressed genes (DEGs) and clustering of gene ontologies in plants inoculated with the Hg-tolerant rhizobia strain, including symbiosis related genes. This finding suggests that the Hg-tolerant rhizobia strain has the potential to mitigate Hg stress in host plants. Furthermore, we observed genotype by-genotype interactions between the high Hg accumulating plant genotype and the Hg-tolerant rhizobia strain. These findings provide insights into enhancing plant resilience in contaminated environments through optimizing legume-rhizobia interactions for heavy metal tolerance.

59 BASIC BIOLOGICAL SCIENCES↗

Validation of the DESI DR2 measurements of baryon acoustic oscillations from galaxies and quasars

The Dark Energy Spectroscopic Instrument (DESI) Data Release 2 (DR2) galaxy and quasar clustering data represents a significant expansion of data from Data Release 1 (DR1), providing improved statistical precision in baryon acoustic oscillation (BAO) constraints across multiple tracers, including bright galaxies, luminous red galaxies, emission line galaxies, and quasars. In this paper, we validate the BAO analysis of DR2. We present the results of robustness tests on the blinded DR2 data and, after unblinding, consistency checks on the unblinded DR2 data. All results are compared with those obtained from a suite of mock catalogs that replicate the selection and clustering properties of the DR2 sample. We confirm the consistency of DR2 BAO measurements with DR1 while achieving a reduction in statistical uncertainties due to the increased survey volume and completeness. The combined BAO precision, including both statistical and systematic errors, improves from ∼0.52% in DR1 to 0.30% in DR2—a factor of 1.7 gain. We assess the impact of analysis choices, including different data vectors (correlation function vs power spectrum), modeling approaches and systematics treatments, and an assumption of the Gaussian likelihood, finding that our BAO constraints are stable across these variations and assumptions with a few minor refinements to the baseline setup of the DR1 BAO analysis. We summarize a series of pre-unblinding tests that confirmed the readiness of our analysis pipeline, the final systematic errors, and the DR2 BAO analysis baseline. The successful completion of these tests led to the unblinding of the DR2 BAO measurements, ultimately leading to the DESI DR2 cosmological analysis, with their implications for the expansion history of the Universe and the nature of dark energy presented in the DESI key paper (companion paper).

79 ASTRONOMY AND ASTROPHYSICS↗

SIDM Concerto: Compilation and Data Release of Self-interacting Dark Matter Zoom-in Simulations

We present SIDM Concerto: 14 cosmological zoom-in simulations in cold dark matter (CDM) and self-interacting dark matter (SIDM) models based on the Symphony and Milky Way-est suites. SIDM Concerto includes one Large Magellanic Cloud– (LMC-) mass system (host mass ∼10 11 M ⊙ ), two Milky Way (MW) analogs (∼10 12 M ⊙ ), two group-mass hosts (∼10 13 M ⊙ ), and one low-mass cluster (∼10 14 M ⊙ ). Each host contains ≈2 × 10 7 particles and is run in CDM and one or more strong, velocity-dependent SIDM models. Our analysis of SIDM (sub)halo populations over seven subhalo mass decades reveals that (1) the fraction of core-collapsed isolated halos and subhalos peaks at a maximum circular velocity corresponding to the transition of the SIDM cross section from a v −4 to v 0 scaling; (2) SIDM subhalo mass functions are suppressed by ≈50% relative to CDM in LMC, MW, and group-mass hosts but are consistent with CDM in the low-mass cluster host; (3) subhalos’ inner density profile slopes, which are more diverse in SIDM than in CDM, are sensitive to both the amplitude and shape of the SIDM cross section. Our simulations provide a benchmark for testing SIDM predictions with astrophysical observations of field and satellite galaxies, strong lensing systems, and stellar streams. Data products are publicly available at doi:10.5281/zenodo.14933624.

dark matter↗

Unsupervised Learning for Improved Gamma-Ray Spectrometry in Pixelated Cadmium Zinc Telluride (CZT) Detectors

Machine learning has been found to be ubiquitously useful across many industries, presenting an opportunity to improve radiation detection performance using data-driven algorithms. Improved detector resolution can aid in the detection, identification, and quantification of radionuclides. Here, in this work, a novel, data-driven, unsupervised learning approach is developed to improve detector spectral characteristics by learning, and subsequently rejecting, poorly performing regions of the pixelated detector. Feature engineering is used to fit individual characteristic photo peaks to a Doniach lineshape with a linear background model. Then, principal component analysis is used to learn a lower-dimension latent space representation of each photo peak where the pixels are clustered, and subsequently ranked, based on the cluster mean distance to an optimal point. Pixels within the worst cluster(s) are rejected to improve the full-width at half-maximum (FWHM) by 10% to 15% (relative to the bulk detector) at 50% net efficiency when applied to training data obtained from measurements of a 100 μCi 154 Eu source using a H3D M400i pixelated cadmium zinc telluride detector. These results compare well with, but do not outperform, a greedy algorithm that accumulates pixels in order of FWHM from lowest to highest used as a benchmark. In the future, this approach can be extended to include the detector energy and angular response. Finally, the model is applied to newly seen natural and enriched uranium spectra relevant for nuclear safeguards applications.

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND ↗

Dark Energy Survey year 6 results: Magnification modeling and its impact on galaxy clustering and galaxy-galaxy lensing cosmology

Gravitational lensing magnification alters the observed spatial distribution of galaxies and must be accounted for to prevent biases in cosmological probes of the large-scale structure. We investigate its effects on the Dark Energy Survey Year 6 galaxy clustering and galaxy-galaxy lensing analyses using the fiducial lens (position tracer) sample M ag L im++. Magnification bias is parameterized by a coefficient that describes the response of the number of selected objects per unlensed area element to a change in the lensing convergence. We quantify this coefficient using the BALROG synthetic source injection catalog to account for the complexity of the selection function, and compare these results with simplified estimates. The resulting values of the magnification coefficients for each redshift bin are [3.16 ± 0.08, 2.76 ± 0.21, 4.09 ± 0.15, 4.42 ± 0.16, 4.90 ± 0.29, 4.83 ± 0.25]. Relative to Year 3, this analysis provides more precise and accurate magnification bias estimates through a larger BALROG area and reweighting to better match the data properties. Here, the cosmological results are robust when tested against various magnification parameter prior choices and also when adding cross-clustering between lens redshift bins. Neglecting magnification, however, introduces significant systematic shifts: relative to the fiducial analysis with Gaussian priors centered on the BALROG -derived estimates, we observe shifts of 1.37σ in S 8 and -0.84σ in Ω m (with cosmic shear included: -0.61σ in S 8 and -0.71σ in Ω m ), in agreement with findings from simulated data, demonstrating that magnification must be modeled to avoid biases. Freeing the magnification bias in lens bin 2 leads to unphysical negative values, further justifying its exclusion from the fiducial Year 6 analysis.

Cosmological parameters↗

Telomere-to-telomere assemblies of chromosome 10 reveal complex adaptive variation of 3-ketoacyl-CoA-synthases in Populus trichocarpa likely driven by Helitrons

The model woody plant Populus trichocarpa displays an atypical alkene-diverse wax cuticle likely driven by copy number variation (CNV) of 3-ketoacyl-CoA synthases ( KCS ), which has been difficult to confirm with short-read assemblies. Long-read sequencing enables the development of telomere-to-telomere resources to detect cryptic variation, including CNVs, which are currently missed. Integrating this information can improve genomic prediction for breeding and provide insights into the evolutionary basis of important traits. Our analysis of 78 long-read haplotypes from chromosome 10 identified more than twice as many KCS genes as previously reported, and numerous intragenic non-synonymous substitutions. Random Forest predictive models highlighted the importance of Potri.010G079500 in producing very long chain alkenes; however, its absence did not predict previously reported alkene-deficient phenotypes. Instead, alkene levels are best predicted by the combinations of KCS copies. Additionally, amino acid substitutions clustered around ligand and donor binding pockets, suggesting they contribute to differing wax cuticle composition. Finally, each KCS gene and copy was linked to a Helitron transposon. A phylogenetic analysis suggests Helitrons are the evolutionary mechanism for generating KCS tandem arrays. Long-read generated telomere-to-telomere assemblies of P. trichocarpa chromosome 10 revealed large-effect loci critical to genetic studies that are unattainable from short-reads. This new resource produced novel insights into genome structure and function, and a novel mechanism for generating tandem gene duplication. Our results highlight that, given current challenges in annotation and assembly, detailed and focused long-read sequences are key to interpreting complex genomic regions that contain tandem copy number variants.

09 BIOMASS FUELS↗