Search NASASearch

SEARCH · Search NASA

Results for “Factor Analysis, Statistical”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 109 records · Page 6

Data from: 'Abiotic influences on continuous conifer forest structure across a subalpine watershed'

This package archives the core data used for analysis and inference in 'Abiotic influences on continuous conifer forest structure across a subalpine watershed' (Worsham et al., 2025). All data were collected in the East River, Washington Gulch, Slate River, and Coal Creek watersheds of Colorado. In the paper, we quantified the relative influence of climate, topographic, edaphic, and geologic factors on conifer stand structure and composition, and their functional relationships, at the watershed scale. We used waveform LiDAR data to derive spatially continuous stand structure metrics. We fused these with a species-level classification map to estimate tree species abundance. We applied generalized additive and generalized boosted models to evaluate the covariability of structural and compositional metrics with abiotic variables. The package contains the essential products required for reproducing our analysis and the tables and figures reported in the publication. The products comprise four classes: (1) geospatial data, (2) tabular data used for inferential analysis, (3) tabular data describing analytical results and performance statistics, and (4) a data user guide. (1) includes discretized waveform LiDAR data, locations and attributes of individual tree crowns, sampling locations and domain boundaries, a canopy height model, and raster files of estimated forest structural and compositional metrics at 100 m grid scale. (2) includes all response and explanatory variable values applied in inferential models. Response variables include conifer forest stand density, basal area, 95th percentile height, quadratic mean diameter, and others. Explanatory variables include climatic water deficit, actual evapotranspiration, elevation, heat load, soil available water content, and others. (3) includes results of training and testing several individual tree detection (ITD) algorithms, as well as inferential modeling results. (4) is a PDF user guide for this data package, including detailed descriptions and data dictionaries for all files. The data package root contains 17 assets: 8 compressed tape archive (.tar.gz) files, 5 comma-separated values (.csv) files, 3 Geographic Tagged Image File Format (GeoTIFF) (.tif) files, and 1 Portable Document Format (.pdf) file. The compressed .tar.gz archives contain ESRI shapefiles (.shp) .tif, compressed LASer (.laz), and .csv files. The archives must first be decompressed using the widely distributed command-line software utility TAR. All other files, including constituent files within the .tar.gz archives, can be opened in the open-source R statistical computing environment. Alternatively, .csv files may also be read in any simple text editor software or Microsoft Excel. Geospatial files including .shp and .tif files can also be opened in GIS software, such as QGIS (open-source) or ESRI ArcGIS (proprietary). The .pdf Data User Guide can be read with Adobe Acrobat Reader or other compatible readers.

2018 NEON and 2025 CHESS Campaigns

DEPRECATED AI-Batt-OS (Autonomous Identification of Battery Life Models - Open Source) [SWR 21-17]

DEPRECATED. This repository was archived by the owner on Jun 30, 2026. It is now read-only. Open source implementation of some of the methods utilized by AI-Batt, a battery lifetime modeling and analysis toolkit provided by the National Laboratory of the Rockies (NLR). This software demonstrates the use of bi-level optimization and symbolic regression techniques to semi-autonomously identify algebraic models predicting the capacity fade of lithium-ion batteries during calendar aging. Modeling the degradation of batteries is a complex task, due to the difficulty in separating the time-dependent and time-independent factors impacting cell level degradation, across multiple data series with different numbers of measurements and/or data quality. Bi-level optimization enables model parameters to be optimized to either the entire data set or to individual data series, allowing statistical disambiguation of global behaviors (data series independent) and local behaviors (data series dependent). Symbolic regression is used to automatically search for optimal low-dimesional models predicting the variation of locally optimized parameters versus time-independent experimental variables from millions of possible models, resulting in a more accurate and repeatable model identification process than is possible by a manual search. The provided tools also implement cross-validation and bootstrap resampling schemes, empowering statistical model comparison/selection and quantification of model uncertainties. An example script replicates the results from the manuscript "Challenging Practices of Algebraic Battery Life Models through Statistical Validation and Model Identification via Machine-Learning", submitted to ECS. All code is written in MATLAB. Requires the Statistics and Machine Learning Toolbox. Contact Dr. Paul Gasper at Paul.Gasper@nlr.gov for any questions.

Gasper, Paul [National Renewable Energy Lab. (NREL

Extensive analysis of reconstruction algorithms for DESI 2024 baryon acoustic oscillations

Reconstruction of the baryon acoustic oscillation (BAO) signal has been a standard procedure in BAO analyses over the past decade and has helped to improve the BAO parameter precision by a factor of ∼2 on average. The Dark Energy Spectroscopic Instrument (DESI) BAO analysis for the first year (DR1) data uses the “standard” reconstruction framework, in which the displacement field is estimated from the observed density field by solving the linearized continuity equation in redshift space, and galaxy and random positions are shifted in order to partially remove non-linearities. There are several approaches to solving for the displacement field in real survey data, including the multigrid (MG), iterative Fast Fourier Transform (iFFT), and iterative Fast Fourier Transform particle (iFFTP) algorithms. In this work, we analyze these algorithms and compare them with various metrics including two-point statistics and the displacement itself using realistic DESI mocks. We focus on three representative DESI samples, the emission line galaxies (ELG), quasars (QSO), and the bright galaxy sample (BGS), which cover the extreme redshifts and number densities, and potential wide-angle effects. We conclude that the MG and iFFT algorithms agree within 0.4% in post-reconstruction power spectrum on BAO scales with the RecSym convention, which does not remove large-scale redshift space distortions (RSDs), in all three tracers. The RecSym convention appears to be less sensitive to displacement errors than the RecIso convention, which attempts to remove large-scale RSDs. However, iFFTP deviates from the first two; thus, we recommend against using iFFTP without further development. In addition, we provide the optimal settings for reconstruction for five years of DESI observation. The analyses presented in this work pave the way for DESI DR1 analysis as well as future BAO analyses.

79 ASTRONOMY AND ASTROPHYSICS

SUNBIRD : a simulation-based model for full-shape density-split clustering

Combining galaxy clustering information from regions of different environmental densities can help break cosmological parameter degeneracies and access non-Gaussian information from the density field that is not readily captured by the standard two-point correlation function (2PCF) analyses. However, modelling these density-dependent statistics down to the non-linear regime has so far remained challenging. We present a simulation-based model that is able to capture the cosmological dependence of the full shape of the density-split clustering (DSC) statistics down to intra-halo scales. Our models are based on neural-network emulators that are trained on high-fidelity mock galaxy catalogues within an extended-ΛCDM framework, incorporating the effects of redshift-space, Alcock–Paczynski distortions, and models of the halo–galaxy connection. Our models reach sub-percent level accuracy down to $1 \, h^{-1}\text{Mpc}$ and are robust against different choices of galaxy–halo connection modelling. When combined with the galaxy 2PCF, DSC can tighten the constraints on ω cdm , σ 8 , and n s by factors of 2.9, 1.9, and 2.1, respectively, compared to a 2PCF-only analysis. DSC additionally puts strong constraints on environment-based assembly bias parameters.

79 ASTRONOMY AND ASTROPHYSICS

A Framework Using Applied Process Analysis Methods to Assess Water Security in the Vu Gia–Thu Bon River Basin, Vietnam

The Vu Gia–Thu Bon (VG–TB) river basin is facing numerous challenges to water security, particularly in light of the increasing impacts of climate change. These challenges, including salinity intrusion, shifts in rainfall patterns, and reduced water supply in downstream areas, are of great concern. This study comprehensively assessed the current state of water security in the basin using robust statistical analysis methods such as the Process Analysis Method (PAM), SMART principle, and Analytic Hierarchy Process (AHP). This resulted in the development of a comprehensive assessment framework for water security in the VG–TB river basin. This framework identified five key dimensions, with basin development activities (0.32), the ability to meet water needs (0.24), and natural disaster resilience (0.19) being the most crucial and water resource potential being the least crucial (0.11) according to the AHP methodology. The latter also highlighted 15 indicators, four of which are particularly influential, including waste resources (0.54), flood (0.53), water storage capacity (0.45), and basin governance (0.42). Furthermore, 28 variables with high weight factors were identified. This framework aligns with the UN-Water water security definition and addresses the global water sustainability criteria outlined in Sustainable Development Goal 6 (SDG6). It enables the computation of a comprehensive Water Security Index (WSI) for specific regions, providing a strong foundation for decision-making and policy formulation. It aims to enhance water security in the context of climate change and support sustainable basin development, thereby guiding future research and policy decisions in water resource management.

54 ENVIRONMENTAL SCIENCES

A Computational Review of Privacy-Preserving Mechanisms for the Smart Grid

Smart grid technologies have rapidly become one of the largest and most comprehensive sources of data for the modern utility. For the most part, data streams are seen as an essential tool that enable utilities to carry their day-to-day business operations, but they also create the need for efficient and secure data management strategies. In the context of the smart grid, ensuring data privacy is becoming an increasing concern due to a combination of factors that range from shifts in operational paradigms and rapid technology evolution to changes in legislation. Furthermore, researchers have highlighted the risks associated with improperly protected energy records. For example, energy consumption data from homes could be used to infer the behaviors and habits of home occupants through activity recognition or user profiling (Fan, 2017), which may lead to unfair service pricing, targeted advertising, or other personal security violations. Similarly, Electric Vehicles’ (EVs) charging metadata could be used to reveal private information about the owner such as their payment methods, preferred charging stations, and other locational and timing information that could be used to reconstruct the vehicle owner’s behaviors. The privacy of user data, even when used for statistical analysis or machine learning training processes, also needs to be carefully considered, as an individual’s private traits may still be vulnerable if their inclusion/exclusion greatly impacts the result or could be linked to a public dataset through cross-reference. The breach of user privacy also has severe impacts for organizations that store, transmit, or work on the data in the form of diminishing the public’s trust in them while potentially incurring legal consequences (e.g., fines and suspensions under the European Union General Data Protection Regulation, Health Insurance Portability and Accountability Act, etc.). Because of these risks, several privacy-preserving mechanisms are available to help organizations comply with privacy legislations and prevent the unauthorized and malicious use of user data. In light of these concerns, this report focuses on performing a computational review of privacy-preserving mechanisms that have received a significant amount of interest in literature. It specifically focuses on 1) homomorphic encryption, 2) zero-knowledge proofs, 3) differential privacy, and 4) federated learning. It is worth noting that although many of the methods presented in this document rely on cryptographic primitives, their intent is not to provide perfect secrecy, but rather to enable users to maintain privacy, and thus they shall not be compared or equated to other constructs that are aimed to address cybersecurity constructs.

24 POWER TRANSMISSION AND DISTRIBUTION

Bayesian Framework for Bioburden Density Estimation in Planetary Protection

To comply with the international planetary protection policy set forth by the Committee on Space Research and NASA Agency level requirements, spacecraft destined to biologically sensitive planetary bodies have to minimize terrestrial biological contamination. Analysis, testing and inspection are the standard forward verification activities that are used to demonstrate compliance with the biological contamination requirements. For testing of spacecraft surface areas, a swab or wipe sample is collected from surfaces prior to last access and subsequently processed in the lab using NASA Approved Planetary Protection Methods for Culture Based Assays. Raw data resulting from this assay is then statistically treated employing a mathematical paradigm stemming from the 1970’s Viking Lander Project to generate the bioburden density and total microbial bioburden present. This standard approach arbitrarily accounts for error and provides an upper conservative bound as it reports the maximum number of spores estimated to be present on flight hardware surfaces. A bioburden density estimate factors in the following variables: the observed bioburden count, representative volume processed, sampling efficiencies. Notably, to account for error in the approach, a 0 observed count is arbitrarily changed to a count of 1 for each hardware grouping. The data generated by spacecraft bioburden verification campaigns in the past have resulted in <80% of wipes and <90% of swabs containing a bioburden count of 0. As such, having a robust and well documented statistical approach for dealing with the probability of low incident rates is necessary to be able to estimate spacecraft bioburden. Being able to statistically describe the bioburden distribution and associated confidence level is a gamechanger for the development of bioburden allocations during mission design and will allow for tighter management of risk throughout spacecraft build. Thus, Empirical Bayes statistical approach was evaluated to estimate the microbial bioburden on spacecraft to mitigate the aforementioned mathematical concerns and provide a probabilistic bioburden distribution of the flight hardware surface. For application of this approach to performing bioburden calculations, a range of non-informative prior assumptions on hardware surfaces are explored for Bayesian analyses while informative priors using posterior distributions from prior assays are utilized for Empirical Bayes analyses. Several non-informative priors are currently under investigation to assess fitness including use of these priors to serve as a foundation to build off of NASA specification values or a basis of risk to account for unknowns during the integration and testing process. Informative priors under consideration are generated using sampled bioburden values from hardware originating within like processing environments (e.g. vendor cleaning process or similar assembly process), temporal spacecraft status events as a prediction for hardware cleanliness of future samples, and heritage system bioburden actuals to predict allocation for subsequent missions. Informative priors and probabilistic bioburden distributions are then validated using data sets from the Mars Exploration Rover, Mars Science Laboratory, and InSight missions. Using Empirical Bayes approach to generate a probabilistic bioburden distribution as demonstrated through mission use cases provides a valid approach for use in the end-to-end requirements verification process.

97 - MATHEMATICS AND COMPUTING

Numerical Analysis of High-Order Modes in SRF Resonators for Particle Accelerators

Over the past decades, superconducting technology has rapidly evolved towards high accelerating gradients and low surface resistance, making it possible to operate particle accelerators with high average beam currents and large duty factors. However, RF losses due to coherent excitation of the HOM become the limiting factor for these regimes. Unlike the cavity operating mode, which is tuned separately, the HOM parameters can significantly vary from one cavity to another due to finite mechanical tolerances during the manufacturing process. Thus, it is of utmost importance to know the HOM parameter spread in advance in order to predict unexpected cryogenic losses, overheating of beam line components and maintain stable beam dynamics. In this paper, we present a method for generating cavity geometry with an arbitrary spread of mechanical imperfections and numerically evaluating HOM statistics. Knowing the spread of HOM parameters, we calculated the probability of resonant HOM losses in SRF accelerating cavities used in CW beam current machines such as the PIP-II and LCLS-II linacs, as well as for the SRF crab-cavity for the ILC project. Finally, we present experimental results of HOM spectra measurements in hundreds of 1.3 GHz cavities installed in LCLS-II cryomodules. Studying the effects of HOM excitation results in specifications of the SRF cavity and cryomodule and can significantly impact the efficiency and reliability of the machine operation.

Lunin, Andrei [Fermilab] (ORCID:0000000290096792)

Causal Directions Matter: How Environmental Factors Drive Convective Cloud Detrainment Heights

This study investigates how environmental factors influence the level of maximum detrainment (LMD) in deep convective clouds. Through a novel application of the Linear Non‐Gaussian Acyclic Model (LiNGAM), we discover causal structures between environmental variables and LMD, observed at six tropical sites operated by the Atmospheric Radiation Measurement (ARM) user facility. LiNGAM effectively identifies causal directions among variables of interest, revealing robust relationships such as those among the lifting condensation level (LCL), level of free convection (LFC), and convective inhibition (CIN), aligning with prior knowledge. Relative humidity is shown to directly influence LMD; however, this relationship exhibits strong nonlinearity and becomes difficult to detect when the contrast between oceanic and continental environments is excluded from the analysis. This study highlights the importance of establishing causal relationships before performing statistical inference.

54 ENVIRONMENTAL SCIENCES

Biopolymer-Templated Titania Film Formation for Nanostructured Coatings Revealed by Machine Learning-Supported Time-Resolved Analysis

This study presents a machine learning approach to derive the film formation of biopolymer-templated titania nanostructures during spray deposition, in combination with in situ grazing-incidence small-angle X-ray scattering (GISAXS). A neural network trained on synthetic GISAXS data directly predicts domain-size distributions from experimental two-dimensional scattering patterns, capturing the full kinetics of nanostructure evolution with high temporal resolution. The predictions reveal hierarchical size distributions and periodic growth features, consistent with layer-by-layer spray deposition and validated by complementary scanning electron microscopy (SEM) imaging. Quantitative comparison with conventional parametric GISAXS fits shows good qualitative agreement, with systematic differences explained by domain-shape assumptions and resolved by applying a geometric scaling factor. Simulated SEM-like surfaces derived from neural network outputs reproduce the porous, foam-like nanoscale morphology observed experimentally, reinforcing the method’s credibility. This integrated approach enables real-time, nondestructive, statistically averaged monitoring of bulk nanostructure development in functional coatings, offering a scalable methodology to accelerate the characterization and process control of sustainably manufactured nanostructured titania films for energy-related applications such as photocatalysis and photovoltaics.

Heger, JulianEliah

Study of e + e − → π + π − π 0 at s from 2.00 to 3.08 GeV at BESIII

With the data samples taken at center-of-mass energies from 2.00 to 3.08 GeV with the BESIII detector at the BEPCII collider, a partial wave analysis on the e + e − → π + π − π 0 process is performed. The Born cross sections for e + e − → π + π − π 0 and its intermediate processes e + e − → ρ π and ρ ( 1450 ) π are measured as functions of s . The results for e + e − → π + π − π 0 are consistent with previous results measured with the initial state radiation method within one standard deviation, and improve the uncertainty by a factor of ten. By fitting the line shapes of the Born cross sections for the e + e − → ρ π and e + e − → ρ ( 1450 ) π , a structure with mass M = 2119 ± 11 ± 15 MeV / c 2 and width Γ = 69 ± 30 ± 5 MeV is observed with a significance of 5.9 σ , where the first uncertainties are statistical and the second ones are systematic. This structure can be interpreted as an excited ω state. Published by the American Physical Society 2024

Astronomy & Astrophysics

Highly accelerated life testing (HALT): A review from a statistical perspective

Despite its use in one form or another for at least four decades, HALT and related techniques [e.g., highly accelerated-stress screening (HASS) and stress audits (HASA)] are not well understood within the statistical community and remain controversial. This largely reflects a conflict in motivation between engineers, testing under harsh conditions to discover and eliminate failure modes, and statisticians, taking a more cautious approach to develop quantitative estimates of parameters such as mean time between failures (MTBF). Here, this review article will clarify HALT concepts and methods and explain where it fits within the universe of methods that involve the application of accelerating factors to compress the time required to evaluate or enhance product reliability. A major distinction is between methods such as HALT, a high-stress test-analyze-fix-test iterative process directed at improving reliability by discovering and fixing weak points in a design, and quantitative accelerated life testing (QALT), whose goal is the estimation of product life for a fixed design. We discuss methods such as physics of failure that offer some hope of bridging the gap between the qualitative nature of HALT, and purely quantitative statistical methods. We present a variety of engineering applications of HALT including metal fatigue, piping and pressure vessels, structural damage, radiation damage, and rotating machinery. We also discuss potential synergies between HALT and QALT, such as rapid identification, through HALT, of failure modes requiring quantitative analysis. For further study, extensive references to the applicable literature are provided as well as an appendix that describes related methods.

97 MATHEMATICS AND COMPUTING

Intrinsic Kinetics of Polyethylene Terephthalate Pyrolysis via Micropyrolysis and Multivariate Chromatographic Analysis

This study provides an in-depth investigation of the primary decomposition of polyethylene terephthalate (PET) via pyrolysis, employing an experimental-analytic workflow that integrates design of experiments (DoE), micropyrolysis coupled with comprehensive two-dimensional gas chromatography (GC×GC), and multivariate data analysis to verify intrinsic kinetic conditions and elucidate evolving product distributions for mapping key reaction pathways. Peaks that could not be identified using commercial spectral libraries were assigned using Mass Frontier simulations, enabling the identification of divinyl terephthalate, ethyl vinyl terephthalate, and 2-(benzoyloxy)ethyl vinyl terephthalate. A polar×polar (non-orthogonal) column set tailored for the detection of carboxylic acids enhanced the quantification of benzoic acid, 4-vinylbenzoic acid, 4-ethylbenzoic acid, and methylbenzoic acid by up to 6-fold relative to an orthogonal column combination (non-polar×mid-polar). Moreover, pyrolysis variables were systematically evaluated using a Box- Behnken design (BBD), encompassing pyrolysis temperature (500−600 °C), sample weight (50−150 μg), and carrier gas flow rate (100−300 mL min −1 ). Among these, pyrolysis temperature was the only statistically significant factor influencing product yields, ranging from 58.78 to 84.26 wt %. In contrast, neither the sample weight nor the carrier gas flow rate had a significant effect on product yields within the evaluated experimental space. At 600 °C, the major pyrolysis products were benzoic acid (up to 20.20 ± 1.46 wt %) and CO 2 (up to 21.28 ± 1.46 wt %), which can be produced through decarboxylation reactions. These findings underscore the critical importance of selecting appropriate analytical columns for the accurate quantification of heteroatomcontaining products such as carboxylic acids, which may otherwise be underestimated or undetected due to their reactivity with the stationary phase of non-polar and mid-polar columns, as well as other GC components. They also highlight the importance of selecting pyrolysis conditions for investigating the primary decomposition of PET under an isothermal kinetically limited regime.

aromatic compounds

Stacked reverberation mapping of high-redshift quasars in DESI. I. Feasibility analysis

The broad-line region of quasars has long been probed by reverberation mapping techniques that measure time lags between continuum and broad emission-line variations. Stacked reverberation mapping has been proposed as a less observationally expensive alternative to traditional methods. This ensemble approach also reduces biases from small-number statistics. The Dark Energy Spectroscopic Instrument (DESI) is conducting the most extensive spectroscopic survey of quasars to date. We create mock light curves emulating expected DESI quasar observations at redshifts $1.48\lt z\lt 5.2$ and luminosities $44.68 \le \log \lambda L_{1350 \mathring{\rm A}{}} / \mathrm{erg\, s^{-1}} \le 45.99$ to test stacked reverberation mapping feasibility using sparse spectroscopic data paired with well-sampled photometric data. The pipeline, using the lag estimation code JAVELIN (Just Another Vehicle for Estimating Lags In Nuclei), successfully recovers the simulated C IV lags within 1σ of the true values using spectroscopic light curves composed of only a few spectral epochs (2–10) with irregular cadences. We investigate how observational factors, including C IV flux error magnitude, number of stacked quasars, and spectral epoch count, affect performance. This work motivates a pathway for future stacked reverberation mapping projects with large-scale spectroscopic surveys of quasars having $\ge 2$ spectroscopic observations. Our results suggest an economical alternative for constraining and extending the radius–luminosity relation to higher redshifts and luminosities. Subsequently, this relation can be employed more reliably in single-epoch black hole mass measurements and quasar cosmology in these distant regimes.

quasars: general, quasars: supermassive black hole

Effect of adaptive cruise control on fuel consumption in real-world driving conditions

This paper presents a comprehensive analysis of the impact of adaptive cruise control on energy consumption in real-world driving conditions based on a natural experiment: a large-scale observational dataset of driving data from a diverse fleet of vehicles and drivers. The analysis is conducted at two different fidelity levels: (1) a macroscopic trip-level benefit estimate that compares trips with and without cruise control in a counterfactual way using statistical methods, and (2) a situation-based comparison achieved through the segmentation of trips into distinct driving situations such as acceleration, braking, cruising, and other maneuvers. The results of this research show that the effect of cruise control on energy consumption varies across different driving situations and levels of analysis. In a macroscopic trip-level analysis, cruise control engagement is associated with a slight increase in fuel consumption across the fleet. As revealed later by the situation-based analysis, this result can be attributed to the negative impact of cruise control on energy consumption in cruising mode, which is the most common driving situation. However, the situation-based comparison demonstrates that cruise control can provide fuel consumption benefits in situations involving acceleration and braking, particularly when a preceding vehicle is present. The study also emphasizes the importance of controlling for various factors that can influence both fuel consumption and the likelihood of cruise control engagement to properly evaluate its effects.

33 ADVANCED PROPULSION SYSTEMS

Precision Measurement of the Neutron Magnetic Form Factor via the Ratio Method at Jefferson Lab Hall A

Protons and neutrons, collectively known as nucleons, are composed of quarks and gluons. The Sachs electromagnetic form factors encode information about the spatial distributions of charge and magnetization in the nucleon, particularly at low momentum transfer. In particular, the neutron magnetic form factor (GMn) provides crucial information about the distribution of magnetization inside the neutron and helps constrain theoretical models of nucleon structure. Quasi-elastic electron scattering from deuterium was measured up to Q^2=13.5 GeV^2 using the Super BigBite Spectrometer in Hall A at Jefferson Lab. In this work, the neutron magnetic form factor GMn was extracted at Q^2 = 3.0 GeV^2 and Q^2=4.5 GeV^2 using the Ratio Method. These results represent a subset of the full dataset collected in this experiment, which extended to significantly higher Q^2. The extracted GMn values agree with the existing global fit within approximately two standard deviations at Q^2=3.0 and show excellent agreement at Q^2=4.5. The measurements achieved systematic uncertainties of about 2% and statistical uncertainties below 0.5%, among the most precise determinations of GMn at these kinematics. These results demonstrate the robustness of the experimental technique and provide an important validation point for future extractions at higher Q^2, where data remain scarce. In addition, the GRINCH heavy gas Cherenkov detector—a key component of the experimental apparatus—was commissioned and achieved an electron detection efficiency of approximately 97%, supporting reliable particle identification. Together, the analysis presented here advances both our understanding of nucleon structure and the validation of the experimental methods and instrumentation used to access it.

Satnik, Maria [College of William and Mary, Willia

Simulated wildfire burned area over the CONUS during 2001-2020

Wildfires have shown increasing trends in both frequency and severity across the Contiguous United States (CONUS). However, process-based fire models have difficulties in accurately simulating the burned area over the CONUS due to a simplification of the physical process and cannot capture the interplay among fire, ignition, climate, and human activities. The deficiency of burned area simulation deteriorates the description of fire impact on energy balance, water budget, and carbon fluxes in the Earth System Models (ESMs). Alternatively, machine learning (ML) based fire models, which capture statistical relationships between the burned area and environmental factors, have shown promising burned area predictions and corresponding fire impact simulation. We develop a hybrid framework (ML4Fire-XGB) that integrates a pretrained eXtreme Gradient Boosting (XGBoost) wildfire model with the Energy Exascale Earth System Model (E3SM) land model (ELM). A Fortran-C-Python deep learning bridge is adapted to support online communication between ELM and the ML fire model. Specifically, the burned area predicted by the ML-based wildfire model is directly passed to ELM to adjust the carbon pool and vegetation dynamics after disturbance, which are then used as predictors in the ML-based fire model in the next time step. Evaluated against the historical burned area from Global Fire Emissions Database 5 from 2001-2020, the ML4Fire-XGB model outperforms process-based fire models in terms of spatial distribution and seasonal variations. Sensitivity analysis confirms that the ML4Fire-XGB well captures the responses of the burned area to rising temperatures. The ML4Fire-XGB model has proved to be a new tool for studying vegetation-fire interactions, and more importantly, enables seamless exploration of climate-fire feedback, working as an active component in E3SM.

Liu, Ye

Statistical relationships across epigenomes using large-scale hierarchical clustering

Recent advances in genomics and sequencing platforms have revolutionized our ability to create immense data sets, particularly for studying epigenetic regulation of gene expression. However, the avalanche of epigenomic data is difficult to parse for biological interpretation given nonlinear complex patterns and relationships. This attractive challenge in epigenomic data lends itself to machine learning for discerning infectivity and susceptibility. In this study, we explore over 3000 epigenomes of uninfected individuals and provide a framework to characterize the relationships among epigenetic modifiers, their modifiers, genetic loci, and specific immune cell types across all chromosomes using hierarchical clustering. Hierarchical clustering of epigenomic data revealed consistent epigenetic patterns across chromosomes, demonstrating that variation due to epigenetic modifiers is greater than variation between cell types. Gene Ontology and KEGG pathway analyses indicated significant enrichment of genes involved in chromatin remodeling, mRNA splicing, immune responses, and the regulation of microRNAs and snoRNAs. Epigenetic modifiers frequently formed biologically relevant clusters, including the cohesin complex, RNA Polymerase II transcription factors, and PRC2 complex members. These clustering behaviors remained consistent across all chromosomes, supported by entropy analysis and high Adjusted Rand Index scores, indicating robust cross-chromosomal similarity. Co-occurrence analysis further revealed specific sets of modifiers that consistently appeared together within clusters, reflecting shared biological functions and interactions. Validation using another dataset confirmed the reproducibility of these clustering patterns and modifier co-occurrence relationships, underscoring the reliability and generalizability of the methodology.

97 MATHEMATICS AND COMPUTING