Search NASA⌕ Search

SEARCH · Search NASA

Results for “validation data sets”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 271 records · Page 15

JGI-Trichoderma v1.0

There is a series of Python and bash scripts to parse genomics datasets used to evaluate the coevolution of gene families and the feature importance of gene families using an SVM classifier. - Cover analysis: takes a list of single-copy genes in a set of genomes, aligns and builds the gene trees to determine if two gene families have a signature of covariation with one another. It parses the files to run phykit cover script described here: https://jlsteenwyk.com/PhyKIT/usage/index.html - SVM-classifier: This Python script is an SVM-based genomic classifier designed for biological data analysis. It combines machine learning with feature selection to identify important genomic markers and classify biological samples. Core Functionality: The script uses Support Vector Machines from scikit-learn to classify genomic data, incorporating SelectKBest for automated feature selection and leave-one-out cross-validation for performance assessment. It operates in multiple modes: feature ranking, optimal combination discovery, and sample prediction. Primary Applications: Genomic sample classification and biomarker discovery Feature importance analysis in high-dimensional biological datasets Prediction of sample categories based on genomic profiles Research applications requiring robust classification of biological data Key Advantages: High-dimensional handling: SVMs excel with genomic data's typical high feature-to-sample ratios Integrated feature selection: Reduces noise and computational overhead while identifying key markers Probability estimation: Provides confidence scores essential for biological interpretation Validation robustness: Leave-one-out cross-validation ensures reliable performance metrics Operational flexibility: Multiple analysis modes support different research phases from exploration to prediction

Stecca Steindorff, Andrei [Lawrence Berkeley Natio↗

Hydrogen Production System Scaling Using a High-Fidelity Simulation-Optimization Framework

Proton exchange membrane (PEM) electrolyzers are widely used for hydrogen production, yet few validated, high-fidelity tools can reliably guide scale-up. Using measured performance from a 50-hour hardware-in-the-loop pilot test, a physics-based, plant-level model of a 1.25 MW PEM electrolyzer and its balance-of-plant (BoP) subsystems is developed and validated. The model couples electrochemistry and thermal/flow submodels and is calibrated against pilot test data via a genetic algorithm (GA) workflow. Validation yields a mean absolute percentage error (APE) of 0.43% for cell voltage and stack power. Two scale-out strategies are then benchmarked under a common 7-day wind-and-photovoltaic (PV) profile: (i) linear duplication of 1.25 MW blocks and (ii) shared-BoP architectures. Sharing BoP between stacks reduces BoP energy by 27% at 10 MW and 34% at 100 MW (vs. linear duplication) and improves system specific energy consumption (SEC) to 52.9 and 52.6 kWh/kg, respectively (from 54.0 kWh/kg with linear duplication). Partial-load studies (25-100% set-point) show that cumulative hydrogen production remains nearly constant down to 50% load because all cases use the same weekly renewable-energy input. Below 50%, the power cap limits how much energy can be used within 168 h, which reduces hydrogen output. The model further indicates that the practical operating optimum lies between 50% and 85% load, where efficiency gains begin to appear without significant loss in hydrogen output. Moreover, the efficiency gains at lower loads are offset by reduced production. The validated framework supports scenario-based engineering trade-off studies for large configurations (10-100 MW) and for operating policies under variable renewables.

08 HYDROGEN↗

Taming nuclear mass models with Gaussian processes

We propose a new set of nuclear mass predictions based on multiple theoretical mass models. By employing Gaussian process regression with the Matérn kernel, we achieved root-mean-square (rms) deviations below 100 keV for the training dataset. The best-performing mass models achieved rms deviations below 150 keV for the new precise mass data from AME2020, whereas the ensemble average showed robust performance across the nuclear chart. Our approach uniquely combines: (1) systematic refinement of eight mass models through their residuals, (2) physics-informed features, including magic numbers, nucleon parity numbers, neutron excess, and nuclear collectivity, and (3) theory-to-theory validation demonstrating robust extrapolation capability. We find that the Matérn kernel provides superior uncertainty quantification compared to the RBF kernel, with a length-scale analysis revealing enhanced inter-nuclei correlations. We provide complete mass predictions for all unknown nuclides in AME2020, offering valuable constraints for nuclear structure studies and astrophysical modeling when used with proper uncertainty propagation.

Gaussian processes↗

Validation of the ERO2.0 code using W7-X and JET experiments and predictions for ITER operation

Abstract The paper provides an overview of recent modelling of global material erosion and deposition in the fusion devices Wendelstein 7-X (W7-X), JET and ITER using the Monte-Carlo code ERO2.0. For validating the modelling tool in a three-dimensional environment, W7-X simulations are performed to describe carbon erosion from the graphite test divertor units, which were equipped in operational phase OP 1.2 and analysed post-mortem. Synthetic spectroscopy of carbon line emission is compared with experimental results from the divertor spectrometer measurement system, showing a good agreement in the e-folding lengths in the radial intensity profiles of carbon. In the case of metallic wall materials, earlier modelling of the Be/W environment in JET and ITER is revisited and extended with an updated set of sputtering and reflection data, as well as including the mixing model for describing the Be/W dynamics in the divertor. Motivated by recent H/D/T isotope experiments in JET, limited and diverted configuration pulses are modelled, showing the expected trend of both Be and W erosion increasing with isotope mass. For the JET diverted configuration pulses, it is shown that Be migrates predominantly to the upper part of the inner divertor where it initially leads to strong W erosion. With longer exposure time, the growth of a Be deposited layer leads to a reduction of W erosion in that region. A similar trend is observed in simulations of the ITER baseline Q = 10 scenario, however with a more symmetric Be migration pattern leading to deposition also on the outer divertor.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

Adaptive Reinforcement Learning (ARL) Control of a Multi-port Resonant Converter in UAV Systems

This study presents an adaptive reinforcement learning (ARL) control framework for a multi-port resonant converter used in hybrid unmanned aerial vehicle (UAV) power systems. The converter integrates high-frequency half-bridge input ports connected to a rectified engine–generator set and a battery energy storage system, along with a semi-bridgeless active rectifier supplying the propulsion load. A deep RL agent is trained to dynamically regulate inter-port phase-shift commands in real time based on flight conditions and load power demand. The ARL controller autonomously identifies phase-shift combinations that maximize conversion efficiency while maintaining stable and coordinated power flow, even under rapidly varying operating scenarios. This data-driven approach eliminates the need for explicit system modeling or extensive manual tuning and enables coordinated control among multiple power ports without inter-port communication. Experimental results validate that the ARL based strategy achieves reliable power sharing and consistently high-efficiency operation across diverse UAV operating conditions.

Asa, Erdem [ORNL] (ORCID:0000000190884812)↗

Optimizing Hydronic Heating for Comfort and Performance in Multifamily Housing

Inefficient control settings in multifamily boilers often lead to substantial energy and cost penalties. To address this, a Fault Detection and Diagnostic (FDD) tool was developed to automate data analysis and identify operational faults such as suboptimal outdoor temperature sensor placement, misconfigured outdoor air reset (OAR) curves, excess boiler cycling, and domestic hot water (DHW) setpoint errors. By comparing pre- and post-implementation periods and applying engineering models, the tool quantifies energy savings and reduces manual analysis time by over 90%. Testing on over 100 monitored sites and a targeted subset of 12 buildings showed an average 11% energy savings from remote optimization; further validation across 19 OAR curve changes confirmed the tool’s accuracy, predicting actual savings within ±5% for most cases. Simple payback can be under three years for many multifamily buildings, though rising hardware, labor, and fuel costs create uncertainties, and decarbonization goals increasingly shift focus to electrification. The FDD tool remains invaluable for optimizing existing boilers, enhancing future electrification measures, and adapting to new technologies by refining building load estimates. In doing so, it supports both near-term efficiency and long-term transitions to low-carbon alternatives, ensuring buildings achieve substantial cost and energy benefits throughout their system lifecycles.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗

Validation of a Custom Ball-on-Ring Apparatus and Consideration of Common Issues for Use in Further Testing (SULI Deliverables)

Ceramic materials are well-known for their high hardness and strength but are limited in their application due to low toughness and sudden failure. As a potential solution, inspiration can be taken from dental enamel nanostructure, where undulating rods cause cracks to branch or deflect, increasing the energy needed to cause total fracture of a ceramic part. Following the dental enamel structure, a novel ceramic which uses 3D printed Yttria stabilized Zirconia rods in an alumina matrix was developed. To test this bio-inspired ceramic material, a proper testing apparatus needed to be created and tested to verify its accuracy. For this project, a bespoke ball on ring testing apparatus was created and tested using both conventionally sintered alumina disks and purchased alumina disks to validate its accuracy. By comparing the Weibull distribution of rupture strengths measured by the tests to literature values, it was shown that the testing frame had a wide distribution of strength which did not align with literature values on the lower end. Through fractography, it was found that some samples fractured from the contact stress induced by the ball indenter, which could not be used to calculate rupture strength. This fracture was often linked to low stress to failure, which was initiated by a flaw on the surface near the indenter which acted as a stress concentrator. Removal of these samples from the data set increased the accuracy of the reported rupture strength values for the ceramic. Considering the equations for the magnitude of contact and flexural strength, along with observations of initiating flaws, several measures can be taken for testing the bio-inspired composite. These measures include proper polishing of both sides of the sample, reducing sample thickness, and potentially using a softer indenter material.

36 - MATERIALS SCIENCE↗

APOLLO: a facility-scale differentiable virtual accelerator for Fermilab

As the design complexity of modern accelerators grows, there is more interest in using advanced simulations that have fast execution time or yield additional insights like gradients. The FAST/IOTA facility has been working on implementing and experimentally validating an end-to-end digital twin that is both fast and gradient-aware, allowing for rapid prototyping of new software and experiments with minimal beam time costs. Our framework integrates physics and ML codes for linac and ring simulation through a set of generic interfaces between surrogate and physics-based sections. To reproduce device inputs and outputs, system state is exposed as a deterministic discrete event simulator. Because Fermilab is undergoing control system transition, both EPICS and ACNET frontends are supported. Recently, we have begun transitioning to a new community lattice standard, PALS, as well as developing standardized infrastructure for data ingest and normalization to prepare for model calibration during FAST proton injector commissioning. We discuss implementation details as well as challenges, and future plans to extend modelling to main complex proton accelerators like PIPII and Booster.

Kuklev, Nikita [Fermilab]↗

Smart culture medium optimization for recombinant protein production: Experimental, modeling, and AI/ML-driven strategies

Recombinant protein production (RPP) is central to biotechnology, where recombinant proteins are used as either end products or catalysts in the synthesis of chemicals, fuels, and materials. Among the major cost drivers, culture medium plays a pivotal role in determining protein yield and quality. This review presents a comprehensive perspective on the critical stages of “smart” culture medium optimization: planning, screening, modeling, optimization, and validation. In the planning stage, we examine the nutritional and energetic roles of medium components, including carbon, nitrogen, amino acids, salts, and trace metals, and their impacts on culture parameters such as pH, oxidative state, and osmolality. We highlight the variability in trace metal content due to water sources, culture vessels, and raw materials, which can substantially influence RPP. The screening stage covers Design of Experiments (DoE) approaches, assessing their theoretical basis, implementation, and limitations. For modeling, we describe methods that integrate experimental data to develop predictive models for smart medium formulation. Model-based optimization strategies can then be employed to select optimal media compositions for a given application. The validation stage aims to evaluate model predictions and provide feedback for model training and refinement. Finally, we survey mechanistic and artificial intelligence/machine learning (AI/ML)-driven models as integrated, transformational tools for predictive modeling of bioprocess conditions, nutrient availability, cellular metabolism, and protein quality, with the goal of optimizing culture media to enhance protein yields while reducing costs and environmental impact. We conclude by addressing the challenges of translating laboratory-scale medium optimization to industrial-scale settings and exploring future AI/ML-driven approaches that may overcome current bottlenecks and accelerate medium design for RPP. Overall, this review provides a unified framework for advancing smart medium design in RPP.

Artificial Intelligence/Machine Learning (AI/ML)↗

Comparing Compressed and Full-Modeling analyses with FOLPS: implications for DESI 2024 and beyond

The Dark Energy Spectroscopic Instrument (DESI) will provide unprecedented information about the large-scale structure of our Universe. In this work, we study the robustness of the theoretical modelling of the power spectrum of F OLPS , a novel effective field theory-based package for evaluating the redshift space power spectrum in the presence of massive neutrinos. We perform this validation by fitting the AbacusSummit high-accuracy N -body simulations for Luminous Red Galaxies, Emission Line Galaxies and Quasar tracers, calibrated to describe DESI observations. We quantify the potential systematic error budget of F OLPS finding that the modelling errors are fully sub-dominant for the DESI statistical precision within the studied range of scales. Additionally, we study two complementary approaches to fit and analyse the power spectrum data, one based on direct Full-Modelling fits and the other on the ShapeFit compression variables, both resulting in very good agreement in precision and accuracy. In each of these approaches, we study a set of potential systematic errors induced by several assumptions, such as the choice of template cosmology, the effect of prior choice in the nuisance parameters of the model, or the range of scales used in the analysis. Furthermore, we show how opening up the parameter space beyond the vanilla ΛCDM model affects the DESI observables. These studies include the addition of massive neutrinos, spatial curvature, and dark energy equation of state. We also examine how relaxing the usual Cosmic Microwave Background and Big Bang Nucleosynthesis priors on the primordial spectral index and the baryonic matter abundance, respectively, impacts the inference on the rest of the parameters of interest. This paper pathways towards performing a robust and reliable analysis of the shape of the power spectrum of DESI galaxy and quasar clustering using F OLPS .

79 ASTRONOMY AND ASTROPHYSICS↗

Temperature-dependent mechanical properties and crystal plasticity parameters for additively manufactured Haynes-214 alloy: Experiments and numerical modeling

Our experimental mechanical testing data demonstrated that the additively manufactured (AM) laser powder bed fusion (L-PBF) Haynes-214 alloy exhibits non-linear mechanical properties as the temperature rises from ambient to 870 °C. Crystal plasticity (CP) simulations provide an effective approach to gaining deeper insights into microstructure-property linkages under thermomechanical loading. This method can reduce the need for costly high-temperature mechanical testing while accounting for the effects of crystallographic texture and grain morphology on the mechanical behavior of AM materials. However, calibrating a CP model is time-consuming because individual simulations are computationally expensive and hundreds (or more) of iterations over parameter sets may be required. To address this issue, we have designed a machine learning-differential evolution (ML-DE) CP framework that can accurately interpolate the tensile properties of AM L-PBF Haynes-214 alloy across a wide temperature range from ambient to 870 °C, with minimal reliance on experimental data. The framework uses electron backscatter diffraction (EBSD) measurements to generate statistically equivalent microstructural volume elements to serve as inputs to the CP modeling framework. Stress–strain curves were generated from 1000 CP simulations, which serve as the training data set for the three ML regression algorithms explored: linear, extra-trees, and multi-layer perceptron. These three regression models were independently evaluated to compare their efficiency and identify the most suitable algorithm for the given problem. Results revealed that the extra-trees ML regressor outperforms the other models in both qualitative and quantitative aspects with an R 2 of 0.98. Subsequently, the differential evolution optimization approach is employed to calibrate the ML-based CP material parameters with experimental results obtained at various temperatures. Finally, temperature-dependent CP material parameters are formulated. The effectiveness and efficiency of the designed framework are validated through comparison with experimental results, demonstrating a high degree of agreement. These calibrated parametric constitutive equations enable further use of the CP model to study the deformation behavior of this alloy under a wide range of thermo-mechanical loading conditions.

36 MATERIALS SCIENCE↗

Geometry-complete diffusion for 3D molecule generation and optimization

Abstract Generative deep learning methods have recently been proposed for generating 3D molecules using equivariant graph neural networks (GNNs) within a denoising diffusion framework. However, such methods are unable to learn important geometric properties of 3D molecules, as they adopt molecule-agnostic and non-geometric GNNs as their 3D graph denoising networks, which notably hinders their ability to generate valid large 3D molecules. In this work, we address these gaps by introducing the Geometry-Complete Diffusion Model (GCDM) for 3D molecule generation, which outperforms existing 3D molecular diffusion models by significant margins across conditional and unconditional settings for the QM9 dataset and the larger GEOM-Drugs dataset, respectively. Importantly, we demonstrate that GCDM’s generative denoising process enables the model to generate a significant proportion of valid and energetically-stable large molecules at the scale of GEOM-Drugs, whereas previous methods fail to do so with the features they learn. Additionally, we show that extensions of GCDM can not only effectively design 3D molecules for specific protein pockets but can be repurposed to consistently optimize the geometry and chemical composition of existing 3D molecules for molecular stability and property specificity, demonstrating new versatility of molecular diffusion models. Code and data are freely available on GitHub .

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Optical constants of magnetron sputtered aluminum in the range 17–1300 eV with improved accuracy and ultrahigh resolution in the L absorption edge region

This work determines a new set of EUV/x-ray optical constants for aluminum (Al), one of the most important materials in science and technology. Absolute photoabsorption (transmittance) measurements in the 17–1300 eV spectral range were performed on freestanding Al films protected by carbon (C) layers, to prevent oxidation. The dispersive portion of the refractive index was obtained via the Kramers–Kronig transformation. Our data provide significant improvements in accuracy compared to previously tabulated values and reveal fine structure in the Al L 1 and L 2,3 regions, with photon energy step sizes as small as 0.02 eV. The implications of this work in the successful realization of EUV/x-ray instruments and in the validation of atomic and molecular physics models are also discussed.

74 ATOMIC AND MOLECULAR PHYSICS↗

Coupling Remote Sensing With a Process Model for the Simulation of Rangeland Carbon Dynamics

Rangelands provide significant environmental benefits through many ecosystem services, which may include soil organic carbon (SOC) sequestration. However, quantifying SOC stocks and monitoring carbon (C) fluxes in rangelands are challenging due to the considerable spatial and temporal variability tied to rangeland C dynamics as well as limited data availability. We developed the Rangeland Carbon Tracking and Management (RCTM) system to track long-term changes in SOC and ecosystem C fluxes by leveraging remote sensing inputs and environmental variable data sets with algorithms representing terrestrial C-cycle processes. Bayesian calibration was conducted using quality-controlled C flux data sets obtained from 61 Ameriflux and NEON flux tower sites from Western and Midwestern US rangelands to parameterize the model according to dominant vegetation classes (perennial and/or annual grass, grass-shrub mixture, and grass-tree mixture). The resulting RCTM system produced higher model accuracy for estimating annual cumulative gross primary productivity (GPP) (R 2 > 0.6, RMSE <390 g C m -2 ) relative to net ecosystem exchange of CO 2 (NEE) (R 2 > 0.4, RMSE <180 g C m -2 ). Model performance in estimating rangeland C fluxes varied by season and vegetation type. The RCTM captured the spatial variability of SOC stocks with R 2 = 0.6 when validated against SOC measurements across 13 NEON sites. Model simulations indicated slightly enhanced SOC stocks for the flux tower sites during the past decade, which is mainly driven by an increase in precipitation. Future efforts to refine the RCTM system will benefit from long-term network-based monitoring of vegetation biomass, C fluxes, and SOC stocks.

54 ENVIRONMENTAL SCIENCES↗

High-accuracy emulators for observables in ΛCDM, N eff, Σ m ν, and w cosmologies

ABSTRACT We use the emulation framework CosmoPower to construct and publicly release neural network emulators of cosmological observables, including the cosmic microwave background (CMB) temperature and polarization power spectra, matter power spectrum, distance-redshift relation, baryon acoustic oscillation (BAO) and redshift-space distortion (RSD) observables, and derived parameters. We train our emulators on Einstein–Boltzmann calculations obtained with high-precision numerical convergence settings, for a wide range of cosmological models including ΛCDM, wCDM, ΛCDM + Neff, and ΛCDM + Σmν. Our CMB emulators are accurate to better than 0.5 per cent out to ℓ = 104, which is sufficient for Stage-IV data analysis, and our P(k) emulators reach the same accuracy level out to $k=50 \, \, \mathrm{Mpc}^{-1}$, which is sufficient for Stage-III data analysis. We release the emulators via an online repository (CosmoPower Organisation), which will be continually updated with additional extended cosmological models. Our emulators accelerate cosmological data analysis by orders of magnitude, enabling cosmological parameter extraction analyses, using current survey data, to be performed on a laptop. We validate our emulators by comparing them to class and camb and by reproducing cosmological parameter constraints derived from Planck TT, TE, EE, and CMB lensing data, as well as from the Atacama Cosmology Telescope Data Release 4 CMB data, Dark Energy Survey Year-1 galaxy lensing and clustering data, and Baryon Oscillation Spectroscopic Survey Data Release 12 BAO and RSD data.

Astronomy & Astrophysics↗

Direct Feed High-Level Waste APPS Model Glass Testing (DFHLW APPS) Matrix

This report summarizes the data collected during the batching and melting of the Direct Feed High-Level Waste APPS Model Glass Matrix (DFHLW APPS) to serve as a quality-assured validation of the Aspen Process Performance Simulation (APPS) formulation method. Of 15 glasses tested, 12 satisfied all target property constraints. Two glasses, APPS-05 and -06, formed nepheline on canister centerline cooling heat-treatment and failed the Product Consistency Test response limits. Glass APPS-07-2 formed unacceptably high concentrations of crystals (primarily Na3Nd(PO4)2) when heat treated at 950 °C. All other glasses were found to be satisfactory. The measured property values were compared to predicted values from a set of current models. In many cases the current models were found to be inadequate for design of DFHLW glasses. These models are being adjusted to correct for mispredictions. Other models, e.g., density, toxicity characteristic leaching procedure, and sulfur solubility, are adequate for formulation of DFHLW glasses.

12 MANAGEMENT OF RADIOACTIVE AND NON-RADIOACTIVE W↗

Production of alternate realizations of DESI fiber assignment for unbiased clustering measurement in data and simulations

A critical requirement of spectroscopic large scale structure analyses is correcting for selection of which galaxies to observe from an isotropic target list. This selection is often limited by the hardware used to perform the survey which will impose angular constraints of simultaneously observable targets, requiring multiple passes to observe all of them. In SDSS this manifested solely as the collision of physical fibers and plugs placed in plates. In DESI, there is the additional constraint of the robotic positioner which controls each fiber being limited to a finite patrol radius. A number of approximate methods have previously been proposed to correct the galaxy clustering statistics for these effects, but these generally fail on small scales. To accurately correct the clustering we need to upweight pairs of galaxies based on the inverse probability that those pairs would be observed (Bianchi & Percival 2017). This paper details an implementation of that method to correct the Dark Energy Spectroscopic Instrument (DESI) survey for incompleteness. To calculate the required probabilities, we need a set of alternate realizations of DESI where we vary the relative priority of otherwise identical targets. These realizations take the form of alternate Merged Target Ledgers (AMTL), the files that link DESI observations and targets. We present the method used to generate these alternate realizations and how they are tracked forward in time using the real observational record and hardware status, propagating the survey as though the alternate orderings had been adopted. We detail the first applications of this method to the DESI One-Percent Survey (SV3) and the DESI year 1 data. We include evaluations of the pipeline outputs, estimation of survey completeness from this and other methods, and validation of the method using mock galaxy catalogs.

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND ↗

Statistical relationships across epigenomes using large-scale hierarchical clustering

Recent advances in genomics and sequencing platforms have revolutionized our ability to create immense data sets, particularly for studying epigenetic regulation of gene expression. However, the avalanche of epigenomic data is difficult to parse for biological interpretation given nonlinear complex patterns and relationships. This attractive challenge in epigenomic data lends itself to machine learning for discerning infectivity and susceptibility. In this study, we explore over 3000 epigenomes of uninfected individuals and provide a framework to characterize the relationships among epigenetic modifiers, their modifiers, genetic loci, and specific immune cell types across all chromosomes using hierarchical clustering. Hierarchical clustering of epigenomic data revealed consistent epigenetic patterns across chromosomes, demonstrating that variation due to epigenetic modifiers is greater than variation between cell types. Gene Ontology and KEGG pathway analyses indicated significant enrichment of genes involved in chromatin remodeling, mRNA splicing, immune responses, and the regulation of microRNAs and snoRNAs. Epigenetic modifiers frequently formed biologically relevant clusters, including the cohesin complex, RNA Polymerase II transcription factors, and PRC2 complex members. These clustering behaviors remained consistent across all chromosomes, supported by entropy analysis and high Adjusted Rand Index scores, indicating robust cross-chromosomal similarity. Co-occurrence analysis further revealed specific sets of modifiers that consistently appeared together within clusters, reflecting shared biological functions and interactions. Validation using another dataset confirmed the reproducibility of these clustering patterns and modifier co-occurrence relationships, underscoring the reliability and generalizability of the methodology.

97 MATHEMATICS AND COMPUTING↗