Practical VVUQ–The Interplay of Model Form, Uncertainty Quantification and Error in Computational Predictions from a Validation Assessment Perspective [Slides]
Abstract not provided.
SEARCH · Search NASA
Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.
Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.
Abstract not provided.
Distributed temperature sensing (DTS) using fiber optic sensors (FOS) offers a promising method for temperature measurements in advanced reactors, such as sodium fast reactors and molten salt cooled reactors. To support the calibration and validation of DTS measurements, Argonne National Laboratory developed the Validation, Optical Calibration, and Learning (VOCAL) software package. This report describes the integration of a local large language model (LLM) with a retrieval-augmented generation (RAG) system into the VOCAL interface to serve as an interactive user assistant. The LLM framework enhances the VOCAL platform’s accessibility to users by explaining interface components, clarifying inputs and outputs, and answering user queries dynamically in real-time. The accuracy of the LLM assistant performance was evaluated with 20 queries regarding the interface and its parameters using experimental data from the Thermal Hydraulic Experimental Test Article (THETA) facility. Results demonstrate that the LLM achieved a 95% accuracy rate, with a BERTScore of 0.8816 and SBERT value of 0.7417. Furthermore, validation of the RAG system within the LLM framework showed optimal accuracy with k-values between 1 and 2 using the k-refinement convergence test. The prompt perturbation analysis demonstrated good initial consistency for the RAG system, exhibiting the highest accuracy under punctuation variations and the greatest sensitivity under query reordering. Notably, the model’s errors were limited to data retrieval failures rather than factual hallucinations, reinforcing its baseline reliability. The integration of LLM provides a highly accurate, userfriendly enhancement to the VOCAL platform without disrupting its core computational capabilities for FOS calibration and validation.
With fault-tolerant quantum computing (FTQC) on the horizon, it is critical to understand sources of logical errors in plausible hardware implementations of quantum error-correcting codes. Detailed error modeling of computational instructions on particular FTQC architectures will enable the better prediction of error propagation in FT-encoded quantum circuits while revealing where greater attention is needed in hardware design. In this work, we consider logical error rates for the surface code implemented on a hypothetical grid-based trapped-ion quantum charge-coupled device architecture. Specifically, we construct logical channels for the idling surface code and examine its diamond error under a mixed coherent and stochastic circuit-level noise model inspired by trapped ions. We include the coherent dephasing noise that is known to accumulate during physical qubit idling and transport in these systems, determining idling and transport durations using the time-resolved output of an open-source trapped-ion surface code compiler. To estimate expectation values of logical Pauli observables following hardware circuits containing non-Clifford sources of noise, we utilize a Monte Carlo technique to sample from an underlying quasiprobability distribution of Clifford circuits that we independently simulate in a phase-sensitive fashion. We verify error suppression up to code distance 𝑑 = 11 at coherent dephasing rates near and below those of current-generation trapped-ion quantum computers and find that logical error rates align with those of analogous fully stochastic simulations in this regime. Exploring higher dephasing rates at 𝑑 = 3−5, we find evidence for growing coherent rotations about all three logical Pauli axes, increased diagonal logical error process matrix elements relative to those of stochastic simulations, and a reduced dephasing rate threshold. Overall, our work paves a way toward realistic hardware emulation of small fault-tolerant quantum processes, e.g., members of an FTQC instruction set.
Critical heat flux (CHF) marks the transition from nucleate to film boiling, where heat transfer to the working fluid can rapidly deteriorate. Accurate CHF prediction is essential for efficiency, safety, and preventing equipment damage, particularly in nuclear reactors. Although widely used, empirical correlations frequently exhibit discrepancies when compared to experimental data, limiting their reliability in diverse operational conditions. Traditional machine learning (ML) approaches have demonstrated potential for CHF prediction but often suffer from limited interpretability, data scarcity, and insufficient knowledge of physical principles. Hybrid model approaches, which combine data-driven ML with base models, mitigate these concerns by incorporating prior knowledge of the domain. This study integrates an externally trained purely data-driven ML model and two hybrid models (using the Biasi and Bowring CHF correlations) within the CTF subchannel code via a custom Fortran framework. Performance was evaluated using two validation cases: a subset of the Nuclear Regulatory Commission (NRC) CHF database and the Bennett dryout experiments. In both cases, the hybrid models demonstrated significantly lower error metrics compared to conventional empirical correlations, with the best models often reducing relative error by about 5 percentage points. The pure ML model achieved comparable accuracy, outperforming the hybrid Biasi model in the NRC test case (3.3% versus 5.5% relative error) but exhibiting slightly higher error against the hybrid Bowring model in the Bennett test case (7.7% versus 6.1%). Trend analysis of error parity indicated that ML-based models reduced the tendency for CHF overprediction, improving overall accuracy. These results demonstrate that ML-based CHF models can be effectively integrated into subchannel codes and could potentially increase performance compared to conventional methods.
In preparation for the first cosmological measurements from the full shape of the Lyman-α (Lyα) forest from DESI, we must carefully model all relevant systematics that might bias our analysis. It was shown in Youles et al. (2022) that random quasar redshift errors produce a smoothing effect on the mean quasar continuum in the Lyα forest region. This, in turn, gives rise to spurious features in the Lyα autocorrelation and its cross-correlation with quasars. Using synthetic data sets based on the DESI survey, we confirm that the impact on BAO measurements is small, but that a bias is introduced to parameters which depend on the full shape of our correlations. We combine a model of this contamination in the cross-correlation (Youles et al. 2022) with a new model we introduce here for the auto-correlation. These are parametrised by 3 parameters, which, when included in a joint fit to both correlation functions, successfully eliminate any impact of redshift errors on our full-shape constraints. We also present a strategy for removing this contamination from real data, by removing ∼0.3% of correlating pairs.
When a new, better-formulated physical parameterization is introduced into a global atmospheric model, aspects of the global model solutions are sometimes degraded. Then, in order to use the new global model to address science questions, there is an incentive to restore its accuracy. Oftentimes this restoration is achieved by tuning of model parameter values. Unfortunately, the retuning process is expensive because characterizing the parameter dependence requires numerous time-consuming global simulations. To reduce the cost of tuning, this manuscript introduces a “poor man's” model tuner, “QuadTune”. QuadTune carves the globe into regions and approximates the model parameter dependence through the use of an uncorrelated quadratic emulator (i.e., response surface). The simplicity of the emulator reduces the required number of global model simulations and aids explainability of tuner behavior. Tuning removes parametric error but leaves behind model structural error. Structural error manifests itself as regional residual biases, such as stubborn biases and tuning trade-offs. To visualize these residual biases, QuadTune's software includes a set of diagnostic plots. This paper illustrates the use of the plots for characterizing residual biases with an example tuning problem.
Here, we apply Bayesian techniques to compare a simple, empirical model for jet quenching in heavy-ion collisions to centrality-dependent jet R AA measured by ATLAS for Pb + Pb collisions at $\sqrt{s_{NN}}$ = 5.02 TeV. We find that the R AA values for central collisions are adequately described with a model for the mean p T -dependent jet energy loss using only two parameters. This model is extended by incorporating two-dimensional initial geometry information from TRENTo and compared to centrality-dependent R AA values. We find that the results are sensitive to the value of the jet-quenching formation time, τ ƒ , and that the optimal value of τ ƒ varies with the assumed path-length dependence of the energy loss. We construct a covariance error matrix for the data from the p T -dependent contributions to the ATLAS systematic errors and perform Bayesian calibrations for several different assumptions for the systematic error correlations. We show that the most-probable functions and $χ^2_d$ values are sensitive to assumptions made when fitting to correlated errors. This work demonstrates the utility of a simple model that can quickly demonstrate the constraining power of jet-quenching observables with corresponding uncertainties and guide future studies using more sophisticated models.
Many physical processes in science and engineering are naturally represented by operators between infinite-dimensional function spaces. The problem of operator learning, in this context, seeks to extract these physical processes from empirical data, which is challenging due to the infinite or high dimensionality of data. An integral component in addressing this challenge is model reduction, which reduces both the data dimensionality and problem size. In this paper, we utilize low-dimensional nonlinear structures in model reduction by investigating Auto-Encoder-based Neural Network (AENet). AENet first learns the latent variables of the input data and then learns the transformation from these latent variables to corresponding output data. Our numerical experiments validate the ability of AENet to accurately learn the solution operator of nonlinear partial differential equations. Furthermore, we establish a mathematical and statistical estimation theory that analyzes the generalization error of AENet. Finally, our theoretical framework shows that the sample complexity of training AENet is intricately tied to the intrinsic dimension of the modeled process, while also demonstrating the robustness of AENet to noise.
Abstract Bayesian regression determines model parameters by minimizing the expected loss, an upper bound to the true generalization error. However, this loss ignores model form error, or misspecification, meaning parameter uncertainties are significantly underestimated and vanish in the large data limit. As misspecification is the main source of uncertainty for surrogate models of low-noise calculations, such as those arising in atomistic simulation, predictive uncertainties are systematically underestimated. We analyze the true generalization error of misspecified, near-deterministic surrogate models, a regime of broad relevance in science and engineering. We show that posterior parameter distributions must cover every training point to avoid a divergence in the generalization error and design a compatible ansatz which incurs minimal overhead for linear models. The approach is demonstrated on model problems before application to thousand-dimensional datasets in atomistic machine learning. Our efficient misspecification-aware scheme gives accurate prediction and bounding of test errors in terms of parameter uncertainties, allowing this important source of uncertainty to be incorporated in multi-scale computational workflows.
This paper provides a comprehensive overview of how fitting of baryon acoustic oscillations (BAO) is carried out within the upcoming Dark Energy Spectroscopic Instrument’s (DESI) 2024 results using its DR1 data set, and the associated systematic error budget from theory and modelling of the BAO. We derive new results showing how non-linearities in the clustering of galaxies can cause potential biases in measurements of the isotropic (α iso ) and anisotropic (α ap ) BAO distance scales, and how these can be effectively removed with an appropriate choice of reconstruction algorithm. We then demonstrate how theory leads to a clear choice for how to model the BAO and develop, implement, and validate a new model for the remaining smooth-broad-band (i.e. without BAO) component of the galaxy clustering. Finally, we explore the impact of all remaining modelling choices on the BAO constraints from DESI using a suite of high-precision simulations, arriving at a set of best practices for DESI BAO fits, and an associated theory and modelling systematic error. Overall, our results demonstrate the remarkable robustness of the BAO to all our modelling choices and motivate a combined theory and modelling systematic error contribution to the post-reconstruction DESI BAO measurements of no more than 0.1 per cent (0.2 per cent) for its isotropic (anisotropic) distance measurements. We expect the theory and best practices laid out to here to be applicable to other BAO experiments in the era of DESI and beyond.
Lithium-ion batteries with silicon anodes promise high energy density but are limited by calendar lifetime. Reducing the long iteration time to obtain experimental results requires predicting calendar lifetime early in a cell's life. In this study, we demonstrate that lightweight machine learning models with feature engineering can provide calendar lifetime estimates from early electrochemical signals. After 1 month of electrochemical aging, the best models achieve 10% error in calendar-life prediction and can separate "bad" from "good" lifetime cells with a mean F1 score of 0.857. As battery systems exhibit inherent variability, four methods for uncertainty quantification are compared, and confidence intervals are demonstrated with an uncertainty of +-3.6 months in lifetime prediction. A feature importance analysis indicates that early patterns in voltage decay are the strongest indicators of calendar lifetime. Finally, this modeling approach has high error when generalizing to new electrode chemistries or testing conditions but with appropriately low confidence.
The Dark Energy Spectroscopic Instrument (DESI) will provide unprecedented information about the large-scale structure of our Universe. In this work, we study the robustness of the theoretical modelling of the power spectrum of F OLPS , a novel effective field theory-based package for evaluating the redshift space power spectrum in the presence of massive neutrinos. We perform this validation by fitting the AbacusSummit high-accuracy N -body simulations for Luminous Red Galaxies, Emission Line Galaxies and Quasar tracers, calibrated to describe DESI observations. We quantify the potential systematic error budget of F OLPS finding that the modelling errors are fully sub-dominant for the DESI statistical precision within the studied range of scales. Additionally, we study two complementary approaches to fit and analyse the power spectrum data, one based on direct Full-Modelling fits and the other on the ShapeFit compression variables, both resulting in very good agreement in precision and accuracy. In each of these approaches, we study a set of potential systematic errors induced by several assumptions, such as the choice of template cosmology, the effect of prior choice in the nuisance parameters of the model, or the range of scales used in the analysis. Furthermore, we show how opening up the parameter space beyond the vanilla ΛCDM model affects the DESI observables. These studies include the addition of massive neutrinos, spatial curvature, and dark energy equation of state. We also examine how relaxing the usual Cosmic Microwave Background and Big Bang Nucleosynthesis priors on the primordial spectral index and the baryonic matter abundance, respectively, impacts the inference on the rest of the parameters of interest. This paper pathways towards performing a robust and reliable analysis of the shape of the power spectrum of DESI galaxy and quasar clustering using F OLPS .
In the burgeoning field of quantum computing, the precise design and optimization of quantum pulses are essential for enhancing qubit operation fidelity. This study focuses on refining the pulse engineering techniques for superconducting qubits, employing a detailed analysis of square and Gaussian pulse envelopes under various approximation schemes. We evaluated the effects of coherent errors induced by naive pulse designs. Furthermore, we identified the sources of these errors in the Hamiltonian model’s approximation level. We mitigated these errors through adjustments to the external driving frequency and pulse durations, thus implementing a pulse scheme with stroboscopic error reduction. Our results demonstrate that these refined pulse strategies improve performance and reduce coherent errors. Moreover, the techniques developed herein are applicable across different quantum architectures, such as ion-trap, atomic, and photonic systems.
Traffic crashes significantly contribute to global fatalities, particularly in urban areas, highlighting the need to evaluate the relationship between urban environments and traffic safety. This study extends former spatial modeling frameworks by drawing paths between global models, including spatial lag (SLM), and spatial error (SEM), and local models, including geographically weighted regression (GWR), multi-scale geographically weighted regression (MGWR), and multi-scale geographically weighted regression with spatially lagged dependent variable (MGWRL). Utilizing the proposed framework, this study analyzes severe traffic crashes in relation to urban built environments using various spatial regression models within Leon County, Florida. According to the results, SLM outperforms OLS, SEM, and GWR models. Local models with lagged dependent variables outperform both the global and generic versions of the local models in all performance measures, whereas MGWR and MGWRL outperform GWR and GWRL. Local models performed better than global models, showing spatial non-stationarity; so, the relationship between the dependent and independent variables varies over space. The better performance of models with lagged dependent variables signifies that the spatial distribution of severe crashes is correlated. Finally, the better performance of multi-scale local models than classical local models indicates varying influences of independent variables with different bandwidths. According to the MGWRL model, census block groups close to the urban area with higher population, higher education level, and lower car ownership rates have lower crash rates. On the contrary, motor vehicle percentage for commuting is found to have a negative association with severe crash rate, which suggests the locality of the mentioned associations.
Distributed training increases the number of batches processed per iteration either by scaling-out (adding more nodes) or scaling-up (increasing the batch-size). However, the largest configuration does not necessarily yield the best performance. Horizontal scaling introduces additional communication overhead, while vertical scaling is constrained by computation cost and device memory limits. Thus, simply increasing the batch-size leads to diminishing returns: training time and cost decrease initially but eventually plateaus, creating a knee-point in the time/cost vs. batch-size pareto curve. The optimal batch-size therefore depends on the underlying model, data and available compute resources. Large batches also suffer from worse model quality due to the well-known “generalization gap”. In this paper, we present Tula, an online service that automatically optimizes time, cost, and convergence quality for large-batch training of convolutional models. It combines parallel-systems modeling with statistical performance prediction to identify the optimal batchsize. Tula predicts training time and cost within 7.5−14% error across multiple models, and achieves up to 20× overall speedup and improves test accuracy by ≈9% on average over standard large-batch training on various vision tasks, thus successfully mitigating the generalization gap and accelerating training at the same time.
Machine-learned interatomic potentials (ML-IAPs) continue to gain popularity as accurate, computationally efficient replacements for traditional, physics-based interatomic potentials and expensive ab initio methods. Uncertainty quantification (UQ) of ML-IAPs is a growing area of research as UQ is critical in many applications of IAPs, such as developing curated datasets, active learning-based data augmentation, self-improving models, and estimating the uncertainty of molecular dynamics simulations. In this paper, we construct and benchmark a series of different neural network potentials (NNPs) with varying network architectures to determine the performance of these models with respect to both the mean and uncertainty calibration error. Each NNP method is specifically designed to predict either epistemic or aleatoric uncertainty with particular focus on the differences in behavior between the epistemic and aleatoric uncertainty estimates. We benchmark these methods using multiple datasets common in the ML-IAP literature. The results show that the aleatoric uncertainty from single-shot model architectures is a competitive alternative to ensemble-based epistemic uncertainty predictions in regions of sufficient data-density. However, in regions where the representative data is sparse, aleatoric uncertainty models tend to overpredict and epistemic methods tend to underpredict the actual model error. We conclude that the type of UQ is crucial when discussing performance of probabilistic model results as different methods have different performance characteristics depending on the regime in which they are evaluated. Therefore, the type of UQ method should be carefully evaluated against both the data characteristics and requirements for the intended application.
In this study, we investigated strategies to address trust issues arising from errors in large language models (LLMs). The study examined the impact of confidence scores, system capability explanations, and user feedback on trust restoration post-error. 68 participants viewed the responses of an LLM to 20 general trivia questions, with an error introduced on the third trial. Each participant was presented with one mitigation strategy. Participants rated their overall trust in the model and the reliability of the answer. Results showed an immediate drop in trust after the error; however, there were no differences across the three strategies in trust recovery. All conditions had a logarithmic trend in trust recovery following error. Differences in overall trust were predicted by perceived reliability of the answer, suggesting that participants were evaluating results critically and using that to inform their trust in the model. Qualitative data supported this finding; participants expressed lasting distrust despite the LLM’s later accuracy. Results showcase the need to prioritize accuracy in LLM deployment, because early errors may irrevocably damage user trust calibration and later adoption.
Bias correction is a crucial step in using Earth system model outputs for assessments, as it adjusts systematic errors by comparing the model to observations. However, standard methods – ranging from mean-based linear scaling to distribution-based quantile mapping typically treat bias correction as a single-scale process, overlooking the fact that biases can manifest differently across daily, seasonal, and annual timescales. In this study, we propose a novel, timescale-aware bias-correction approach built on Empirical Mode Decomposition. By decomposing the meteorological signal into multiple oscillatory components and aggregating them to represent distinct timescales, we apply targeted corrections to each component, thereby preserving both short- and long-term structure in the data. Experimental illustrations show that the timescale-aware EMDBC framework matches the performance of conventional quantile-delta mapping (QDM) at the native daily scale and achieves progressively larger bias reductions at bi-weekly, seasonal, and annual scales. As a result, the proposed approach offers a more robust path to accurate and reliable Earth system projections, strengthening their utility for resilience and adaptation planning.