Search NASASearch

SEARCH · Search NASA

Results for “Model comparison”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 37 records · Page 2

DECOVALEX-2023: Task C Final Report

The Full-scale Emplacement (FE) heater experiment at the Mont Terri Underground Rock Laboratory (URL) was designed and conducted by Nagra to replicate an emplacement tunnel of Nagra’s reference repository design at 1:1 scale. Alongside testing the technical feasibility of constructing disposal tunnels, emplacing waste containers in the tunnels and then backfilling them, the main goals of the FE experiment are (1) to obtain a better understanding of the coupled effects of induced thermo-hydro-mechanical (THM) processes that may occur and (2) to validate existing coupled THM models (Müller et al., 2017). A key aspect of ensuring safety for repositories located in low-permeability rock involves minimizing any damage to the rock itself, thereby preserving its integrity and promoting a stable environment Amongst a number of processes that could damage the rock is the increase in pore pressure due to thermal loading caused by heat emitted from the waste. To reduce the potential damage of the rock, it is important to analyse the evolution of heat over time due to the heat load of the containers and assess possible consequences by coupled THM models. The aim of Task C of DECOVALEX-2023 was to build 3D numerical models of the FE experiment, focussing in particular on the heating induced pore pressure change in the Opalinus Clay. Data from a large number of sensors were available from the FE experiment for model comparison. These sensors measured temperature and relative humidity in the bentonite around the heaters, and temperature, pressure and displacement/strain in the surrounding Opalinus clay. Data were available from the start of excavation (April 2012) up to August 2020 for most sensors (more than 5 years from the start of heating in December 2014). To fulfil the overall aim of the task, the work was broken down into a number of steps, starting with simpler models to build confidence in each team’s approach and then moving to more complex models that better represent the FE experiment. Step 0 consisted of 2D benchmark models, gradually increasing the number of processes that are represented from thermal (T) only models in Step 0a, to coupled thermal hydraulic (TH) models in Step 0b with a representation of changing porosity, to coupled thermo-hydro-mechanical (THM) models in Step 0c, where porosity changes are calculated by the mechanical model. A detailed specification of processes, parameters, initial and boundary conditions was provided for this step, with the ambition that all teams would work towards close agreement in their model results, thus building confidence in the model implementations. vi It was not straightforward to achieve agreement between the teams, so additional steps (Step 0b2, 0b3, 0c2, 0c3) were added along with derivation of some analytical solutions against which the models could be compared. The reasons for the differences between teams were investigated and found to be caused primarily by different conceptual model assumptions (including temperature dependence of the thermal expansion of water), different model formulations (including porosity evolution) and differences in modelled domain sizes, boundary conditions and grid discretisation. This demonstrates that comparisons between multiple modelling teams and/or comparison with analytical results and experimental data are highly beneficial in providing an indication of uncertainty in model predictions. At the conclusion of Step 0, almost all teams had achieved a close agreement in model results and those that had not achieved an agreement knew the reason for this. Step 1 moved from 2D models to 3D models of the FE experiment without adding technical features like shotcrete or EDZ, and only considering the heating phase. Initially the 3D model was tightly specified to continue to build confidence in the model implementations (Step 1a). The results of Step 1a were compared to the data from the FE-experiment without the teams seeing the data. The teams were then provided with a sub-set of the data from the FE-experiment and invited to consider how best to use the large dataset for model comparison (Step 1b). Teams were then asked to use the data provided to calibrate their models, only changing material property values rather than adding features or processes to their models (Step 1c). In Step 1, teams were asked to only model the heating phase of the experiment, so pressure in the Opalinus Clay was reported as change in pressure since the initial conditions were specified rather than modelled. The change from 2D to 3D models was accompanied by an increase in the dispersion of results between the teams. Some of this was resolved during the task, but some remained and is potentially due to model discretisation. Calibration of parameters was useful in improving the fit of the models to the data but the remaining differences indicated that the models were missing features or processes. In Step 2, the teams were asked to update their models with additional features and processes as well as calibrating parameters to try and improve the fit of the models to the data. Teams were encouraged to represent ventilation of the open FE tunnel prior to backfilling with heaters and bentonite and in Step 2, the absolute pressure in the Opalinus Clay was compared between the teams. Teams took different approaches, but there was consideration of adding shotcrete and an EDZ into the model, representing stress change during excavation and different approaches to modelling ventilation of the FE tunnel. Overall, the documented results showed a very good agreement for temperature. The results for porewater pressure evolution showed a significant improvement for most teams compared to Step 1c with a good agreement to the measurements for several teams whereas some teams overpredicted the pressure increase and others overpredicted the drainage effect especially for the sensors close to the heater. Step 3 was an opportunity for teams to use the models developed in Step 1 and Step 2 to make predictions about the temperature and pressure changes that will be expected at the FE experiment over the next few years in light of the planned changes in thermal output of the heaters.

12 MANAGEMENT OF RADIOACTIVE AND NON-RADIOACTIVE W

Policy implications of net-zero emissions: A multi-model analysis of United States emissions and energy system impacts

Many countries, subnational jurisdictions, and companies are setting net-zero emissions goals; however, questions remain about strategies to reach these targets, policy measures, technology gaps, and economic impacts. Here, we investigate the potential policy implications of reaching economy-wide net-zero CO 2 emissions across the United States by 2050 using results from a multi-model comparison with 14 energy-economic models. Model results suggest that achieving net-zero CO 2 targets depends on policies that accelerate deployment of zero- and low-emitting technologies that have seen rapid cost reductions in recent years (including wind, solar, battery storage, and electric vehicles) as well as relatively nascent options (including carbon capture and storage, advanced biofuels, low-carbon hydrogen, advanced nuclear, and long-duration energy storage). While net-zero policies are likely to lower fossil fuel consumption, including considerable coal and petroleum reductions, achieving net-zero emissions does not necessarily mean phasing out all fossil fuels. Model results indicate that the Inflation Reduction Act’s energy and climate provisions amplify near-term decarbonization but that net-zero policies have larger impacts on long-run outcomes. Stringent climate policy can have large fiscal impacts on tax revenue and government spending—revenues from carbon pricing and subsidies for carbon removal range from 0.1 % to 3.7 % of GDP in 2050 across models. Each dollar per metric ton carbon price leads to a 0.06 % to 0.31 % reduction in economy-wide CO 2 emissions relative to a reference scenario with current policies. Spending on energy across the economy decreases relative to today for many models under reference and net-zero policies, especially as a share of GDP, due primarily to end-use electrification and energy efficiency.

54 ENVIRONMENTAL SCIENCES

Modeling the formation of Sedan Crater using the FLAG and HOSS codes

Numerical modeling of explosion crater formation requires accounting for complex physical processes. Numerical validation of explosion cratering is an important step in modeling and requires experimental data for comparison. Models using discrete elements and continuum models have both benefits and drawbacks to their approaches. In this work, we consider both an arbitrary Lagrangian–Eulerian (ALE) hydrocode and a finite discrete element method (FDEM) approach to modeling the formation of the Sedan crater, the largest human-made crater in the United States. The Sedan crater formed from an underground nuclear detonation in the Nevada desert as part of Project Plowshare. Our models show that the continuum approach of the hydrocode matched well compared to early test time prior to the mound rupture and subsequent fireball venting, when most of the alluvium exhibited fluid behavior. Our FDEM approach matched the final crater dimensions well, after material had settled back into the crater, when material strength and solid mechanics play key roles. Our work shows how leveraging the benefits of multiple numerical approaches can lead to better understanding of complex physical problems, especially problems with limited experimental data. By using a continuum approach to early-time hydrodynamics and an FDEM approach to later-time solid mechanics, we can better understand the different physical regimes of explosion crater formation.

36 MATERIALS SCIENCE

Improving neutrino-nuclei interaction models: Recommendations and case studies on Peelle’s Pertinent Puzzle

Improving the modeling of neutrino-nuclei interactions using data-driven methods is crucial for high-precision neutrino oscillation experiments. This paper investigates Peelle’s Pertinent Puzzle (PPP) in the context of neutrino measurements, a longstanding challenge to fitting theoretical models to experimental data. Inconsistencies in data-model comparisons hinder efforts to enhance the accuracy and reliability of model predictions. We analyze various sources contributing to these inconsistencies and propose strategies to address them, supported by practical case studies. We advocate for incorporating model fitting exercises as a standard practice in cross section publications to enhance the robustness of results. We use a common analysis framework to explore PPP-related challenges with MicroBooNE and T2K data in an unified manner. Our findings offer valuable insights for improving the accuracy and reliability of neutrino-nuclei interaction models, particularly by systematically tuning models using data.

73 NUCLEAR PHYSICS AND RADIATION PHYSICS

Assessment of Climate Change Impacts on Renewable Energy Resources in Western North America

We examine a 25 km resolution climate model dataset to evaluate how regional climate change impacts solar and wind energy under a high-emission scenario. Our study considers the Western Electricity Coordinating Council (WECC) region, which covers the western United States and southwestern Canada, focusing specifically on locations with existing solar and wind infrastructure. First, we conduct a historical model comparison of solar and wind energy capacity factors to highlight model uncertainties across the study area. Using future climate projections, we then assess the seasonal patterns of solar and wind capacity factors for three timeframes: historical, mid-century, and end of century. Additionally, we estimate the frequency of solar and wind resource droughts during these periods for the entire WECC and its five operational subregions, finding that certain subregions are more susceptible to energy droughts due to limited renewable resources. Finally, we present day-ahead capacity factor forecasts to support energy storage planning and provide estimates of offshore wind energy capacity within the WECC. Our results indicate that offshore wind capacity factors are nearly twice as high as onshore values, with less seasonal variation, which suggests that offshore wind could offer a more consistent renewable energy supply in the future.

climate change

Persistent Sampling: Enhancing the Efficiency of Sequential Monte Carlo

Sequential Monte Carlo (SMC) samplers are powerful tools for Bayesian inference but suffer from high computational costs due to their reliance on large particle ensembles for accurate estimates. We introduce persistent sampling (PS), an extension of SMC that systematically retains and reuses particles from all prior iterations to construct a growing, weighted ensemble. By leveraging multiple importance sampling and resampling from a mixture of historical distributions, PS mitigates the need for excessively large particle counts, directly addressing key limitations of SMC such as particle impoverishment and mode collapse. Crucially, PS achieves this without additional likelihood evaluations-weights for persistent particles are computed using cached likelihood values. This framework not only yields more accurate posterior approximations but also produces marginal likelihood estimates with significantly lower variance, enhancing reliability in model comparison. Furthermore, the persistent ensemble enables efficient adaptation of transition kernels by leveraging a larger, decorrelated particle pool. Experiments on high-dimensional Gaussian mixtures, hierarchical models, and non-convex targets demonstrate that PS consistently outperforms standard SMC and related variants, including recycled and waste-free SMC, achieving substantial reductions in mean squared error for posterior expectations and evidence estimates, all at reduced computational cost. PS thus establishes itself as a robust, scalable, and efficient alternative for complex Bayesian inference tasks.

Karamanis, Minas

The relative importance of wind and hydroclimate drivers in modulating the interannual variability of dust emissions in Earth system models

Windblown dust emissions are controlled by near-surface wind speed and sediment erodibility, the latter modulated by hydroclimate and land-use conditions. Accurate representations of these drivers are critical for reproducing historical dust variability and projecting future dust changes in Earth system models (ESMs). This study examines the discrepancies among 21 ESMs in the relative importance of wind speed versus five hydroclimate drivers in explaining the historical (1980–2014) variability of dust emissions from global drylands. In hyperarid areas, models show poor agreement in the simulated dust variability, with only 9 % out of 210 inter-model comparisons exhibiting significant positive correlations. In contrast, arid and semiarid areas exhibit a dual pattern driven by a “double-edged sword” effect of land surface memory: models with coherent hydroclimate variability show better agreement, whereas those with divergent hydroclimate representations show larger disagreement. While the ESMs capture the dominant role of wind speed in hyperarid areas, they diverge markedly in the relative contributions of wind and hydroclimate drivers in arid and semiarid areas. Replacing the Zender et al. (2003) dust scheme with the Kok et al. (2014) scheme in CESM and E3SM generally strengthens hydroclimate influences while reducing wind speed contributions to simulated dust variability. MERRA-2 reanalysis produces stronger wind influences than most ESMs across all dryland regions. These results underscore the need for improved near-surface wind simulations in hyperarid areas and more realistic land surface and hydroclimate representations in arid and semiarid areas to reduce uncertainties in global dust emission simulations.

Li, Xinzhu [Michigan Technological University, Hou

FY-25 Progress on Computational Modeling of the Water Based NSTF

This report summarizes the system level modeling using RELAP5-3D of the Natural Convection Shutdown Heat Removal Test Facility (NSTF) completed in FY25. This year’s work focuses on a new tank configuration where the inlet of the tank was lowered in elevation by 45”. The stability boundaries of the NSTF are thoroughly studied and stability maps are constructed based on the stability and the oscillation patterns of the system. Five distinct operational modes are identified, namely single-phase liquid, uniform double peak oscillations, uniform sinusoidal oscillations, stable two-phase flow, and non-uniform oscillations. Next, the riser inlet throttling case of experimental test Run-104 is simulated with the RELAP5 model where good agreement is obtained between the model and the experimental data. The simulation also highlights the effects of backflow of water from the tank to the upper region of the chimney. Additionally, the decay heat removal test of Run-99 is simulated with the RELAP5 model. Comparison is carried out between this run and a similar run with the mid-tank inlet of Run-74 performed in FY22. With the lower tank inlet, the RELAP5 model is able to predict the experimental data more accurately than the previous mid tank inlet configuration. The discrepancy in model prediction accuracy highlights the non-symmetrical spatial effects in the tank that would otherwise be more easily captured with higher fidelity models. Lastly, two exploratory studies are conducted to investigate the behaviors of the NSTF when 1) heating is provided to the downcomer and 2) a bypass channel is added between the horizontal chimney section to the downcomer.

42 ENGINEERING

High-Burnup LOCA Burst Susceptibility BISON Analysis in PWRs and BWRs

Accurately assessing high-burnup fuel behavior during loss-of-coolant accidents (LOCAs) is essential for understanding fuel fragmentation, relocation, and dispersal (FFRD) risks across the US light-water reactor fleet. This work updates previous Nuclear Energy Advanced Modeling and Simulation (NEAMS) Program multiphysics LOCA analyses for a pressurized water reactor (PWR) and a boiling water reactor (BWR) by incorporating recent model and material property advancements in the BISON fuel performance code, including a high-burnup structure (HBS) model, revised cladding burst criteria, and updated thermal–mechanical correlations. This update was needed to support ongoing industry initiatives and upcoming regulatory changes. Full-core, rod-resolved operating histories generated using Virtual Environment for Reactor Analysis (VERA) and system-level LOCA conditions obtained from TRACE were applied to statistically representative rod samples in BISON to evaluate burst behavior and FFRD susceptibility. These calculations used two cladding burst correlations and three fuel pulverization models so that the predictions of these models could be compared. The updated PWR simulations show markedly improved numerical stability as the number of crashed simulations decreased by 95% compared to the previous study, and hence higher confidence in results. The updated PWR simulations predicted cladding bursts exclusively among once-burned, high-power rods, with two different cladding burst models identifying the same burst-susceptible population. Resulting FFRD susceptibility estimates are significantly reduced compared with earlier studies, driven by cooler predicted fuel and plenum temperatures, lower hoop strains, and reduced fission gas release in the updated models. In contrast, none of the BWR rods were predicted to burst under either burst criterion, reaffirming minimal BWR FFRD susceptibility even with updated HBS and material models. Comparisons between the PWR and BWR end-of-cycle predictions are made. Comparison with prior work highlights significant shifts in PWR fuel performance metrics and confirmation of earlier BWR conclusions. Overall, the updated results underscore the importance of having high-resolution detailed modeling capability and continuously integrating evolving material models and physics into high-resolution multiphysics simulations. The unified assessment presented here strengthens confidence in predicting high-burnup LOCA behavior by improving agreement between different cladding burst correlations. These results also provide an improved foundation for future BISON model development, FFRD susceptibility calculations.

11 NUCLEAR FUEL CYCLE AND FUEL MATERIALS

An international study on THM modelling of the full-scale heater experiment at Mont Terri laboratory

We present results from an international model comparison study of the Full-Scale Emplacement (FE) experiment in Opalinus Clay at the Mont Terri Laboratory, Switzerland. Based on a provided parameter set the teams decided which parameters they adopted for their models, whether they considered the excavation and the ventilation phase in addition to the heating phase and if they included technical features like the shotcrete or the EDZ. The teams were able to reproduce the measured parameters temperature, relative humidity and pore pressure. The modelled results for temperature agree very closely between the teams especially in the sensors in Opalinus Clay. All teams were able to reproduce the redistribution of water in the bentonite backfill due to heating. The evolution of the relative humidity showed similar trends with differences in the intensity of the dry out effect. To model the pore pressure evolution is more complex because it comprises the full interaction of the coupled THM processes. The spread between the pore pressure modelled by the teams was larger, with some teams overestimating the pressure increase due to heating and some teams overestimating the extent of drainage. The agreement of modelled results with measurements improves with larger distance to the heater. We conclude that the EDZ and the shotcrete potentially influence the behaviour of the rock causing higher differences closer to the heater. Further research is needed to better implement those influences into the models. Based on the calibrated models, the future evolution of temperature, relative humidity and pore pressure was predicted over the next 10 years following a change of the heat power applied in 2023 and 2024. Again, the predicted temperatures agree very closely between the teams. Most teams do not expect an increase in relative humidity during the next 10 years after the initial dry-out.

58 GEOSCIENCES

Contrasting Parametric Sensitivities in Two Global Vegetation Models Using Parameter Perturbation Ensembles

Uncertainty in land model projections remains high and the roles of parametric and structural uncertainty are difficult to disentangle. To compare parametric sensitivity across model structures we present two parameter perturbation ensembles using the Community Land Model (CLM) operating in satellite phenology mode. The ensembles contrast two vegetation modules: (a) the default CLM vegetation module and (b) the Functionally Assembled Terrestrial Ecosystem Simulator (CLM-FATES). We perturbed over 300 parameters and quantified their effects on biophysical fluxes globally and across biomes. Most parameters have minimal impact on biophysical fluxes, with only a few substantially influencing results. While both models exhibit similar parameter sensitivity for some fluxes, CLM-FATES shows larger spread in gross primary productivity (GPP), driven by strong sensitivity to carboxylation rate. CLM-FATES also shows a weaker GPP response to soil hydrology parameters and exhibits higher water use efficiency (WUE). Cross-model comparisons reveal similar sensitivities for some parameters (e.g., leaf dimension) but divergent responses to others (e.g., stomatal intercept), highlighting underlying structural differences. Differences in WUE and sensitivity to hydrology and stomatal conductance parameters underscore how model structure fundamentally alters parametric sensitivity. The data sets generated from these ensembles can be used to identify influential parameters and guide future calibration efforts.

Foster, A. C. [NSF National Center for Atmospheric

Data from: "Towards CONUS-Wide ML-Augmented Conceptually-Interpretable Modeling of Catchment-Scale Precipitation-Storage-Runoff Dynamics"

This data package was generated to support the manuscript “Towards CONUS-Wide Machine Learning-Augmented Conceptually Interpretable Modeling of Catchment-Scale Precipitation-Storage-Runoff Dynamics.” It provides input files, model outputs, plotting data, scripts, notebooks, and documentation used to develop, evaluate, and reproduce Mass-Conserving Perceptron (MCP)-based hydrologic modeling experiments across 513 selected Catchment Attributes and Meteorology for Large-sample Studies in the United States (CAMELS-US) basins. The files are organized by modeling component and analysis purpose, including rainfall–runoff experiments, snow module experiments, coupled hydrologic-snow experiments, Long Short-Term Memory (LSTM) benchmark results, model skill metrics, initialization and epoch records, cell-state normalization files, Akaike Information Criterion (AIC)-based model comparison files, and data used to generate manuscript figures. Tabular files can be opened using standard spreadsheet software or Python/R data-analysis tools. Python scripts, Jupyter notebooks, and selected MATLAB scripts are included for model execution, postprocessing, plotting, and statistical analysis. Quality assurance and quality control were conducted through the source-data selection and modeling workflow. Meteorological forcing, streamflow, and static catchment attributes were derived from the CAMELS-US dataset, and snow water equivalent data were derived from the University of Arizona (UA) Snow Water Equivalent dataset. Selected basins and time periods were screened during the associated research workflow to avoid missing observations or poor-quality cases. Static geospatial features were processed primarily using Quantum Geographic Information System (QGIS) and Geospatial Data Abstraction Library (GDAL) workflows. Additional details are provided in the associated manuscript and documentation.

ESS-DIVE CSV File Formatting Guidelines Reporting

chatHPC: Empowering HPC users with large language models

The ever-growing number of pre-trained large language models (LLMs) across scientific domains presents a challenge for application developers. While these models offer vast potential, fine-tuning them with custom data, aligning them for specific tasks, and evaluating their performance remain crucial steps for effective utilization. However, applying these techniques to models with tens of billions of parameters can take days or even weeks on modern workstations, making the cumulative cost of model comparison and evaluation a significant barrier to LLM-based application development. To address this challenge, we introduce an end-to-end pipeline specifically designed for building conversational and programmable AI agents on high performance computing (HPC) platforms. Our comprehensive pipeline encompasses: model pre-training, fine-tuning, web and API service deployment, along with crucial evaluations for lexical coherence, semantic accuracy, hallucination detection, and privacy considerations. Here, we demonstrate our pipeline through the development of chatHPC, a chatbot for HPC question answering and script generation. Leveraging our scalable pipeline, we achieve end-to-end LLM alignment in under an hour on the Frontier supercomputer. We propose a novel self-improved, self-instruction method for instruction set generation, investigate scaling and fine-tuning strategies, and conduct a systematic evaluation of model performance. The established practices within chatHPC will serve as a valuable guidance for future LLM-based application development on HPC platforms.

97 MATHEMATICS AND COMPUTING

Assessing High Burnup U-19Pu-10Zr Fuel Performance against Historical and Modeled Behavior

Advancing the deployment of sodium-cooled fast reactors (SFRs) requires thorough testing of metallic fuel pins under accident conditions to establish safe operational limits of high burnup fuel. To conduct transient testing, a comprehensive understanding of steady-state fuel behavior obtained through both experimental characterization and accurate predictive capabilities is needed. This study comparatively assesses the steady-state irradiation performance of two high burnup U-19Pu-10Zr fuel pins, DP-36 and DP-40, irradiated under prototypic fast reactor conditions in preparation for planned safety testing at the Transient Reactor Test Facility. Since DP-40 was designated for use in the test and DP-36 serves as its sibling pin, non-destructive, engineering-scale post-irradiation examinations (PIE) were conducted on both pins while destructive examinations were performed exclusively on DP-36. The results were then assessed against historical performance data from similar fuel pins irradiated in the Experimental Breeder Reactor-II. Additionally, the steady-state irradiation of each pin was modeled using the BISON fuel performance code to assess the accuracy of current modeling capabilities in predicting the baseline irradiation behavior. Non-destructive examinations included neutron radiography to measure fuel column elongation, gamma scanning to verify pin integrity and fission product migration, and profilometry to assess dimensional changes. Benchmarking against existing PIE data revealed consistent patterns in axial fuel column growth and cladding diametral strain, though both pins exhibited longer low-density “fluff” structures, which can have implications for core reactivity and source term calculations. Destructive examinations on DP-36 included fission gas release analysis and sectioning for optical microscopy, which showed more complex constituent redistribution patterns than the traditionally accepted 3-ring model. The axial evolution of fractional areas and porosities of each of the redistributed zones were quantified and presented. Modeling comparisons showed agreement in fractional fission gas release but consistently overestimated axial and radial swelling and disagreed with measured axial porosity patterns. These conservative overpredictions suggested that the pins would appear closer to failure or operational limits at the start of transient tests, potentially leading to higher strain accumulation during the transient. While conservative estimates provide safety margins, they can negatively impact fuel economics. A review of the swelling models identified areas for improvement in the gaseous swelling, solid swelling, and fuel hot-pressing models when applied to ternary fuel. The results of this study highlight the critical importance of conducting pre-test characterization on both test and sibling pins to accurately capture steady-state fuel behavior, providing a precise baseline for post-test evaluations and essential inputs for transient modeling of the planned experiments. The analysis also revealed significant data gaps that require further investigation to enhance the understanding and prediction of fuel swelling and pore dynamics. Collecting comprehensive data across different irradiation conditions, burnup levels, and fuel compositions are essential for refining existing models and developing mechanistic models for both binary and ternary metallic fuels, ultimately improving the integration of modeling and experimental approaches in accident testing.

11 - NUCLEAR FUEL CYCLE AND FUEL MATERIALS

Localized Interactions in Neutrino Simulations

The Deep Underground Neutrino Experiment (DUNE) requires precise modeling of neutrino--nucleus interactions to achieve its neutrino-oscillation measurement goals. GENIE, the Monte Carlo event generator used by DUNE, exhibits a known discrepancy with MicroBooNE measurements of transverse kinematic imbalance (TKI): the data display a larger high-TKI tail while maintaining a peak similar to that predicted by the baseline GENIE model. Previous variations of final-state interaction (FSI) strength affected both the peak and the tail and therefore did not resolve the discrepancy. This work investigates whether inconsistencies between local and global treatments of intranuclear physics contribute to the observed mismodeling. The GENIE FSI routines were modified to use the particle position when generating nucleons, thereby introducing a local-density treatment in the hA and hN intranuclear models for their 2018 and 2025 implementations. Meson-exchange-current (MEC) localization was also tested for the hN 2018 model. Comparisons of the intranuclear scattering-center momentum and its radial dependence confirm that the localization was implemented as intended. Localization produces modest changes in the proton momentum spectrum, primarily at low momentum, but only a small change in the TKI distribution between approximately \SI{0.25}{\giga\electronvolt} and \SI{0.40}{\giga\electronvolt}. These changes are insufficient to account for the discrepancy with MicroBooNE data. Although a localized treatment improves the internal consistency of the GENIE model, the origin of the TKI discrepancy remains unresolved.

Bulla, Braden [Unlisted, US, IL]

A Benchmarking Framework for Evaluating Large Language Model Capabilities in Nuclear Reactor Safety Applications

Large language models (LLMs) are increasingly capable of answering technical questions, synthesizing domain knowledge, and supporting engineering workflows. For nuclear science and engineering, these capabilities require careful, domain-specific evaluation before they can be credibly incorporated into safety-related activities, regulatory review, or technical decision support. This paper presents preliminary results from benchmarking framework for evaluating LLM capabilities in nuclear contexts. The framework is organized into three evaluation categories: nuclear fundamentals, general dual-use knowledge, and plant specific knowledge. These categories are intended to distinguish general nuclear engineering competence from broader technical reasoning and more context-dependent nuclear knowledge. Initial evaluations focus on nuclear fundamentals using questions representative of the knowledge expected of a nuclear professional engineer. Results indicate that contemporary frontier models perform at a high level and substantially exceed the performance of older model generations, with some models approaching saturation of the current benchmark. These findings suggest both the rapid improvement of LLM capabilities in specialized technical domains and the need for more discriminating evaluation methods. The paper presents the benchmark structure, preliminary model-comparison results, and ongoing work. This work supports development of verifiable, responsible, and safety-conscious methods for assessing AI systems in nuclear engineering applications.

21 - SPECIFIC NUCLEAR REACTORS AND ASSOCIATED PLAN

Diagnosing the representation of surface and layered soil moisture in Earth system models

Surface soil moisture (mrsos) and vertically integrated soil moisture (mrsol) over the top 10 cm should, by definition, be physically consistent in Earth System Models (ESMs). However, an evaluation of nine CMIP6 models reveals substantial inconsistencies: in some models, mrsos and integrated mrsol agree globally; in others, they align only in specific regions; and in a few, they diverge across all grid cells. These discrepancies arise from a combination of factors, including metadata errors, inconsistent variable definitions, or diagnostic sequencing within the model. We demonstrate how such issues can lead to significant biases, even when both variables are present and seemingly well-defined. As model complexity increases and multi-model comparisons become more common, assumptions about variable equivalence may lead to flawed conclusions. This study highlights the need for routine consistency checks, improved metadata standards, and community-wide practices that ensure reliability of derived variables across ESM outputs, particularly in preparation for CMIP7.

Earth system models

Comparison of structurally diverse simulation models for prediction of epidemic outcomes caused by a long-distance dispersed pathogen

Long-distance dispersal (LDD) pathogens pose substantial challenges for epidemic control due to their ability to generate new infection foci at great distances. While various modeling approaches have been developed to understand and manage such outbreaks, little work has compared how models of different structures behave under shared conditions. Here, in this study, we compare four structurally distinct epidemiological models — EPIMUL, GEMF, PoPS, and Warwick — each adapted to simulate the spread of wheat stripe rust (WSR), a wind-dispersed LDD pathogen, under identical epidemiological parameters and dispersal kernel. Using data from a controlled field experiment, we evaluate the ability of each model to replicate disease prevalence under nine intervention scenarios that vary in timing and culling area. While the models differ substantially in design — ranging from spatial grid-based to network-based and raster-based frameworks — the shared dispersal kernel allowed for close alignment in their predictions. All models accurately captured general epidemic trends, particularly the strong effect of early intervention on disease suppression. We qualitatively compared their behavioral responses across scenarios and also evaluated an ensemble prediction by averaging across model outputs. Our findings highlight how integrating shared epidemiological components into distinct modeling frameworks can improve consistency and accuracy, while reinforcing the importance of early culling in managing LDD pathogen outbreaks.

Dispersal kernel