Search NASASearch

SEARCH · Search NASA

Results for “Bayesian statistical modeling”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 55 records · Page 3

Data Assimilation for Robust UQ Within Agent-Based Simulation on HPC Systems

Agent-based simulation provides a powerful tool for in silico system modeling. However, these simulations do not provide built-in methods for uncertainty quantification (UQ). Within these types of models a typical approach to UQ is to run multiple realizations of the model then compute aggregate statistics. This approach is limited due to the compute time required for a solution. When faced with an emerging biothreat, public health decisions need to be made quickly and solutions for integrating near real-time data with analytic tools are needed. We propose an integrated Bayesian UQ framework for agent-based models based on sequential Monte Carlo sampling. Given streaming or static data about the evolution of an emerging pathogen this Bayesian framework provides a distribution over the parameters governing the spread of a disease through a population. These estimates of the spread of a disease may be provided to public health agencies seeking to abate the spread. By coupling agent-based simulations with Bayesian modeling in a data assimilation, our proposed framework provides a powerful tool for modeling dynamical systems in silico. We propose a method which reduces model error and provides a range of realistic possible outcomes. Moreover, our method addresses two primary limitations of ABMs: the lack of UQ and an inability to assimilate data. Our proposed framework combines the flexibility of an agent-based model with UQ provided by the Bayesian paradigm in a workflow which scales well to HPC systems. We provide algorithmic details and results on a simulated outbreak with both static and streaming data.

Spannaus, Adam [ORNL] (ORCID:0000000225213657)

Hierarchical Bayesian Inverse Problems: A High-Dimensional Statistics Viewpoint

This paper analyzes hierarchical Bayesian inverse problems using techniques from highdimensional statistics. Furthermore, our analysis leverages a property of hierarchical Bayesian regularizers that we call approximate decomposability to obtain non-asymptotic bounds on the reconstruction error attained by maximum a posteriori estimators. The new theory explains how hierarchical Bayesian models that exploit sparsity, group sparsity, and sparse representations of the unknown parameter can achieve accurate reconstructions in high-dimensional settings.

MAP estimation

Neural network emulation of flow in heavy-ion collisions at intermediate energies

Applications of new techniques in machine learning are speeding up progress in research in various fields. In this work, we construct and evaluate a deep neural network (DNN) to be used within a Bayesian statistical framework as a faster and more reliable alternative to the Gaussian process (GP) emulator of an isospin-dependent Boltzmann-Uehling-Uhlenbeck (IBUU) transport model simulator of heavy-ion reactions at intermediate beam energies. We found strong evidence of the DNN being able to emulate the IBUU simulator's prediction on the strengths of protons' directed and elliptical flow very efficiently even with small training datasets and with accuracy about ten times higher than the GP. Here, limitations of our present work and future improvements are also discussed.

73 NUCLEAR PHYSICS AND RADIATION PHYSICS

Semi-Analytical Hierarchical Bayesian Inference of Nonlinear Model Structure in Stochastic Dynamics: Applied to Compartmental Models of Infectious Diseases

A Bayesian computational framework for parsimonious inference in stochastic nonlinear dynamical systems is presented. This framework enables the concurrent estimation of system states, time-varying parameters, time-invariant parameters, and the optimal sparsity structure of the model parameters. Because differential equation-based models are often simplified mechanistic or phenomenological representations, robust inference from noisy measurement data requires explicit treatment of model error and uncertainty. Model error and time-varying parameters can be represented as random processes, enabling inference while making minimal assumptions about the underlying sources of discrepancy and variability. Adopting stochastic differential equation representations affords the model significant flexibility, but can also render it susceptible to overfitting during statistical inversion, where the inferred model may track noise rather than the underlying signal. To alleviate the effects of overfitting and to enable the discovery of the optimal sparse representation of the time-invariant parameters, a Bayesian sparse learning algorithm is embedded within the framework. This sparse learning framework adopts an approximate hierarchical Bayesian setting defined by a series of semi-analytical expressions. The model structure inference framework is validated using a stochastic compartmental model for tracking and forecasting active cases of an infectious disease. Compartmental models describe population-level infectious disease dynamics through interactions among population fractions grouped by disease state. Mathematically, such models consist of a system of coupled ordinary differential equations. This example adopts an expressive compartmental model that includes multiple possible interactions between disease states, motivated by early uncertainty surrounding COVID-19 reinfection dynamics and their implications for long-term epidemic forecasting. The sparse learning exercise permits the inference of a priori unknown epidemiological dynamics from simulated public health data, discovering the nested compartmental model that optimizes the trade-off between average data-fit and model complexity. It is shown that inducing sparsity among the model parameters eliminates redundant interactions between compartments, equivalently revealing the optimal coupling structure between differential equations.

97 MATHEMATICS AND COMPUTING

Calibrating Bayesian generative machine learning for Bayesiamplification

Recently, combinations of generative and Bayesian deep learning have been introduced in particle physics for both fast detector simulation and inference tasks. These neural networks aim to quantify the uncertainty on the generated distribution originating from limited training statistics. The interpretation of a distribution-wide uncertainty however remains ill-defined. We show a clear scheme for quantifying the calibration of Bayesian generative machine learning models. For a Continuous Normalizing Flow applied to a low-dimensional toy example, we evaluate the calibration of Bayesian uncertainties from either a mean-field Gaussian weight posterior, or Monte Carlo sampling network weights, to gauge their behaviour on unsteady distribution edges. Well calibrated uncertainties can then be used to roughly estimate the number of uncorrelated truth samples that are equivalent to the generated sample and clearly indicate data amplification for smooth features of the distribution.

97 MATHEMATICS AND COMPUTING

The decay of HIV under anti-retroviral therapy is biphasic even in humanized mice with just T cells

HIV-1 plasma viral load decays in a biphasic manner during antiretroviral therapy (ART). It was hypothesized that this is due to infection of different cell types, namely CD4+ T cells and macrophages. We studied this possibility directly by modeling the decay of HIV-1 in humanized mice. We utilized previously published data from humanized T-cell only mice (TOM) and myeloid-only mice (MOM) infected with HIV-1 and treated with a potent ART regimen. Viral load decay dynamics were modeled using either a single or a biexponential decay fitted using nonlinear mixed effects techniques. Fits were compared using the corrected Bayesian information criterion (BICc). In TOM, the biphasic model was significantly better than a single-phase decay model (ΔBICc ≈ 16) despite additional parameters. In MOM, the biphasic decay was statistically better, but there was substantial uncertainty because the virus goes below detection very fast. The first-phase half-life was consistent between groups (1.2 days in MOM and 1.3 days in TOM) and similar to the half-life estimated in human infection. The second-phase decay in these mice was minimal likely due to low initial viral loads. Additional analyses with mice containing both CD4+ T cells and macrophages or X4-tropic virus-infected MOM mice confirmed the biphasic pattern, demonstrating the robustness of this result. The biphasic decline in HIV-1 occurs, even with only CD4+ T cells, refuting the hypothesis that distinct cell populations (CD4+ T cells and macrophages) drive each decay phase. These findings support an alternative model in which the observed dynamics arise from intrinsic properties of the viral infection lifecycle rather than from cellular compartmentalization.

59 BASIC BIOLOGICAL SCIENCES

Improvement and generalization of ABCD method with Bayesian inference

To find New Physics or to refine our knowledge of the Standard Model at the LHC is an enterprise that involves many factors, such as the capabilities and the performance of the accelerator and detectors, the use and exploitation of the available information, the design of search strategies and observables, as well as the proposal of new models. We focus on the use of the information and pour our effort in re-thinking the usual data-driven ABCD method to improve it and to generalize it using Bayesian Machine Learning techniques and tools. We propose that a dataset consisting of a signal and many backgrounds is well described through a mixture model. Signal, backgrounds and their relative fractions in the sample can be well extracted by exploiting the prior knowledge and the dependence between the different observables at the event-by-event level with Bayesian tools. We show how, in contrast to the ABCD method, one can take advantage of understanding some properties of the different backgrounds and of having more than two independent observables to measure in each event. In addition, instead of regions defined through hard cuts, the Bayesian framework uses the information of continuous distribution to obtain soft-assignments of the events which are statistically more robust. To compare both methods we use a toy problem inspired by pp\to hh\to b\bar b b \bar b p p → h h → b b ‾ b b ‾ , selecting a reduced and simplified number of processes and analysing the flavor of the four jets and the invariant mass of the jet-pairs, modeled with simplified distributions. Taking advantage of all this information, and starting from a combination of biased and agnostic priors, leads us to a very good posterior once we use the Bayesian framework to exploit the data and the mutual information of the observables at the event-by-event level. We show how, in this simplified model, the Bayesian framework outperforms the ABCD method sensitivity in obtaining the signal fraction in scenarios with 1% and 0.5% true signal fractions in the dataset. We also show that the method is robust against the absence of signal. We discuss potential prospects for taking this Bayesian data-driven paradigm into more realistic scenarios.

Alvarez, Ezequiel

Role of the likelihood for elastic scattering uncertainty quantification

In the last decade, uncertainty quantification (UQ) for optical model potentials (OMPs) has become a focal point for nuclear reaction theory, and several competing approaches for OMP UQ have recently been developed. Here, we clarify recent efforts to compare frequentist and Bayesian approaches in the context of OMP UQ [G. B. King et al., Phys. Rev. Lett. 122, 232502 (2019)]. We replicate a portion of that OMP UQ study but use independent statistical tools. Specifically, we compare two methods for OMP parameter inference from elastic scattering data: the Levenberg-Marquardt algorithm for χ 2 minimization on one hand and Markov chain Monte Carlo (MCMC) sampling on the other. Separately, we assess the common practice of using a renormalized likelihood (χ 2 /N), N being the number of data points, instead of the canonical weighted-least-squares likelihood (χ 2 ), as a way of accounting for unknown data correlations. Here, we show that for a generic linear model and for a five-parameter OMP analysis, frequentist and uniform-prior Bayesian approaches recover the same optimum and uncertainty estimates—not systematically larger uncertainties for the Bayesian approach, as was concluded in G. B. King et al., Phys. Rev. Lett. 122, 232502 (2019). Further, we show that if an additional, near-degenerate parameter is introduced into the same OMP analysis such that the parameter posterior becomes non-Gaussian, then covariance-based estimates of uncertainty become unreliable. Finally, we show that regardless of optimization approach, if χ 2 /N is used for the likelihood, the resulting parametric uncertainties increase by $\sqrt{N}$, and that this is responsible for the conclusions drawn in the revisited study. Based on our replication results, we find that a fortuitous cancellation of unreported errors and the renormalization factor can lead to improvement in empirical coverages, as was the case in the original comparative study. We emphasize that developing and applying a realistic likelihood function is an essential task in a UQ analysis, and that several recent UQ studies that employed a renormalized likelihood (i.e., including a 1/N factor) may have yielded unrealistically large uncertainties for elastic-scattering observables. If the parameter posterior deviates from multivariate-normal, a sampling-based approach like MCMC has a clear advantage over methods that assume the Laplace approximation holds. We note that empirical coverage can serve as an important internal check for the analyst whose model or data may have additional, unaccounted-for uncertainties.

73 NUCLEAR PHYSICS AND RADIATION PHYSICS

Assessing correlated truncation errors in modern nucleon-nucleon potentials

We test the BUQEYE model of correlated effective field theory (EFT) truncation errors on Reinert, Krebs, and Epelbaum's semilocal momentum-space implementation of the chiral EFT (𝜒⁢EFT ) expansion of the nucleon-nucleon (NN) potential. This Bayesian model hypothesizes that dimensionless coefficient functions extracted from the order-by-order corrections to NN observables can be treated as draws from a Gaussian process (GP). We combine a variety of graphical and statistical diagnostics to assess when predicted observables have a 𝜒⁢EFT convergence pattern consistent with the hypothesized GP statistical model. Our conclusions are that, first, the BUQEYE model is generally applicable to the potential investigated here, which enables statistically principled estimates of the impact of higher EFT orders on observables. Second, parameters defining the extracted coefficients such as the expansion parameter 𝑄 must be well chosen for the coefficients to exhibit a regular convergence pattern—a property we exploit to obtain posterior distributions for such quantities. Third, the assumption of GP stationarity across lab energy and scattering angle is not generally met; this necessitates adjustments in future work. We provide a workflow and interpretive guide for our analysis framework, and show what can be inferred about probability distributions for 𝑄, the EFT breakdown scale Λ 𝑏 , the scale associated with soft physics in the 𝜒⁢EFT potential 𝑚 eff , and the GP hyperparameters. All our results can be reproduced using a publicly available Jupyter notebook, which can be straightforwardly modified to analyze other 𝜒⁢EFT NN potentials.

Bayesian methods

Neutrinoless Double-Beta Decay Investigations of 82Se Using Three Shell Model Hamiltonians

Neutrinoless double-beta decay is considered one of the most promising processes that would help clarify some of the symmetry-breaking problems in our understanding of the observable universe. Recent studies of neutrinoless double-beta decay matrix elements have employed statistical approaches based on modified shell model effective Hamiltonians for 48Ca (Phys. Rev. C 106, 054302 (2022)) and 136Xe (Phys. Rev. C 107, 045501 (2023)). The analyses rely on inducing perturbations in the starting effective Hamiltonians to observe the behavior of a wide range of observables, besides the 0νββ) NME, that are compared with experimental data. Following a Bayesian model averaging approach, the range of probable values for the neutrinoless double-beta decay matrix elements is presented. In this paper, we present a similar study for 82Se, which is described in the same model space as 76Ge that is under experimental observation. Due to its faster calculation time compared to 76Ge, 82Se can be used as an appropriate substitute in our complex statistical study. Using the calculations performed for the statistical analysis of the neutrinoless double-beta decay matrix elements, we also search for the correlations between the observables that we can compare to experimental data.

Science & Technology - Other Topics

Dark Energy Survey Year 3 results: optimized $w$CDM simulation-based inference with weak lensing map-level hybrid statistics

We present cosmological constraints from the Dark Energy Survey Year 3 (DES Y3) weak lensing data using hierarchical hybrid statistics within a Bayesian simulation-based inference framework that is based on the Gower Street simulations. To maximize the precision of the inference, we have developed a new, information-theory based, data compression of the weak lensing maps to just seven highly informative summary statistics. The hybrid scheme exploits the high information content of the power spectrum, compressing both the power spectrum and neural-based summaries that are designed to extract further information. Our simulation-based approach enables principled forward modelling of all major sources of systematic uncertainty and survey properties into realistic mock observations, including the survey mask, photometric redshift uncertainties, intrinsic galaxy alignments, multiplicative shear calibration bias, source galaxy clustering, non-Gaussian shape noise, and non-linear structure formation. The summary statistics are then used in a Bayesian simulation-based inference pipeline. The inference is validated through coverage tests and checks for robustness against baryonic feedback. Assuming a $w$CDM cosmology, our analysis yields $S_8 = 0.808 \pm 0.017$, $Ω_{\rm m} = 0.325 \pm 0.024$, and $w < -0.766$ (marginalized posterior 68 per cent credible intervals). This rigorous combination of information theory, physics- and neural network-based extreme data compression, and principled Bayesian analysis improves the figure of merit for $(Ω_{\rm m}, S_8, w)$ by 60 per cent over the previous state-of-the-art, and by almost a factor of 3 over two-point analyses of the same data. They are the most precise joint constraints on $(Ω_{\rm m}, S_8, w)$ from weak gravitational lensing data alone of any survey to date. We intend to apply this analysis to the more recent DES Y6 data.

Williamson, J. [University Coll. London]

Systematic Bayesian evaluation of resonance parameters in 19 Ne for the 15 O ⁡(𝛼, 𝛾)⁢ 19 Ne and 18 F ⁡(𝑝, 𝛼)⁢ 15 O reactions

Here, we present a comprehensive evaluation of the nuclear structure properties of 19 Ne using a novel and rigorous Bayesian statistical framework. Precise characterization of 19 Ne resonance parameters is critical for accurately determining reaction rates of the astrophysically significant 15 O ⁡(𝛼, 𝛾)⁢ 19 Ne and 18 F ⁡(𝑝, 𝛼)⁢ 15 O reactions, which govern breakout from the hot CNO cycle in x-ray bursts and influence 𝛾-ray emission in novae, respectively. By reconstructing likelihood functions from published experimental data—including asymmetric uncertainties and upper or lower limits—we derive posterior distributions for resonance energies, decay widths, and branching ratios. Our Bayesian approach systematically incorporates previously reported discrepancies among measurements, providing a statistically robust and consistent treatment of these uncertainties. The evaluated resonance parameters and associated uncertainties provide crucial input for stellar nucleosynthesis modeling, contributing to a refined understanding of explosive astrophysical phenomena.

Kim, Sohyun H. [Sungkyunkwan Univ., Suwon (Republi

Statistical inference of anomalous thermal transport with uncertainty quantification for interpretive 2D SOL models

The critical task of inferring anomalous cross-field transport coefficients is addressed in simulations of boundary plasmas with fluid models. A workflow for parameter inference in the UEDGE fluid code is developed using Bayesian optimization with parallelized sampling and integrated uncertainty quantification. In this workflow, transport coefficients are inferred by maximizing their posterior probability distribution, which is generally multidimensional and non-Gaussian. Uncertainty quantification is integrated throughout the optimization within the Bayesian framework that combines diagnostic uncertainties and model limitations. As a concrete example, we infer the anomalous electron thermal diffusivity $\chi_\perp$ from an interpretive 2D model describing electron heat transport in the conduction-limited region with radiative power loss. The workflow is first benchmarked against synthetic data and then tested on H-, L-, and I-mode discharges to match their midplane temperature and divertor heat flux profiles. We demonstrate that the workflow efficiently infers diffusivity and its associated uncertainty, generating 2D profiles that match 1D measurements. Future efforts will focus on incorporating more complicated fluid models and analyzing transport coefficients inferred from a large database of experimental results.

Bayesian optimization

Taylor approximation variance reduction for approximation errors in PDE-constrained Bayesian inverse problems

In numerous applications, surrogate models are used as a replacement for accurate parameter-to-observable mappings when solving large-scale inverse problems governed by partial differential equations (PDEs). The surrogate model may be a computationally cheaper alternative to the accurate parameter-to-observable mappings and/or may ignore additional unknowns or sources of uncertainty. The Bayesian approximation error (BAE) approach provides a means to account for the induced uncertainties and approximation errors, i.e. the errors between the accurate parameter-to-observable mapping and the surrogate. The statistics of these errors are, however, in general unknown a priori, and are thus calculated using Monte Carlo sampling. Although the sampling is typically carried out offline, i.e. before considering the data, the process can still represent a computational bottleneck. In this work, we develop a scalable computational approach for reducing the costs associated with the sampling stage of the BAE approach. Specifically, we consider the Taylor expansion of the accurate and surrogate forward models with respect to the uncertain parameter fields either as a control variate for variance reduction or as a means to directly and efficiently approximate the mean and covariance of the approximation errors. We propose efficient methods for evaluating the expressions for the mean and covariance of the Taylor approximations based on linear(-ized) PDE solves. Furthermore, the proposed approach is independent of the dimension of the uncertain parameter, depending instead on the intrinsic dimension of the data, ensuring scalability to high-dimensional problems. The potential benefits of the proposed approach are demonstrated for two high-dimensional inverse problems governed by PDE examples, namely for the estimation of a distributed Robin boundary coefficient in a linear diffusion problem, and for a coefficient estimation problem governed by a nonlinear diffusion problem.

Bayesian approximation error

Bayesian Physics Informed Spatio-Temporal Network for Streamflow Data Imputation

Reliable reconstruction of incomplete streamflow records is critical for improving hydrological forecasting, flood preparedness, and water resource management. However, large observational gaps and uncertainties in governing physical parameters limit the accuracy of traditional statistical and machinelearning imputation frameworks. To address these challenges, we develop a Bayesian Physics-Informed Spatio-Temporal Network (BPI-STNet) that jointly captures spatial and temporal dependencies while enforcing hydrologic consistency through embedded physical constraints. The framework integrates a GraphSAGE-LSTM architecture to model spatial connectivity across gauges and temporal flow dynamics, coupled with a Bayesian update mechanism to estimate uncertain parameters in a simplified water-balance framework. Unlike conventional physics-informed networks that rely on sampling-based posterior estimation, BPI-STNet derives an analytic solution to the inverse problem, allowing closed-form Bayesian updates of uncertain parameters Λ={α,β,k} using Gaussian priors and likelihoods. Applied to daily observations from the Susquehanna River Basin (1980-2022), BPI-STNet achieves substantial improvements over a purely data-driven RGNN baseline, which reduced RMSE by 23 % and MAE by 9 %, and achieving an average NSE values up to 0.96. The results demonstrate that coupling Bayesian inference with physics-informed learning yields physically consistent, uncertainty-aware reconstructions that preserve the temporal persistence and statistical distribution of observed flows. The proposed framework establishes a generalizable paradigm for data-sparse hydrologic systems where both data fidelity and physical interpretability are essential.

Krishnan Kutty Ambika, Anukesh [ORNL] (ORCID:00000

sparse_bias

This is a python package used to fit an unknown function from data that potential contains systematic biases related to metadata. The model fits the unknown function and uses a Bayesian horseshoe prior model to impose sparsity on the bias terms. This code has been generalized from research code developed for AIACHNE into a package that should have more general application in a wider class of statistical models.

Walton, Noah

Inverse prediction of PuO2 processing conditions using Bayesian seemingly unrelated regression with functional data

Over the past decade, a variety of innovative methodologies have been developed to better characterize the relationships between processing conditions and the physical, morphological, and chemical features of special nuclear material (SNM). Different processing conditions generate SNM products with different features, which are known as “signatures” because they are indicative of the processing conditions used to produce the material. These signatures can potentially allow a forensic analyst to determine which processes were used to produce the SNM and make inferences about where the material originated. This article investigates a statistical technique for relating processing conditions to the morphological features of PuO 2 particles. We develop a Bayesian implementation of seemingly unrelated regression (SUR) to inverse-predict unknown PuO 2 processing conditions from known PuO 2 features. Model results from simulated data demonstrate the usefulness of the technique. Applied to empirical data from a bench-scale experiment specifically designed with inverse prediction in mind, our model successfully predicts nitric acid concentration, while results for Pu concentration and precipitation temperature were equivalent to a simple mean model. Our technique compliments other recent methodologies developed for forensic analysis of nuclear material and can be generalized across the field of chemometrics for application to other materials.

11 NUCLEAR FUEL CYCLE AND FUEL MATERIALS

A novel framework for increasing research transparency: Exploring the connection between diversity and innovation

A split sample/dual method research protocol is demonstrated to increase transparency while reducing the probability of false discovery. We apply the protocol to examine whether diversity in ownership teams increases or decreases the likelihood of a firm reporting a novel innovation using data from the 2018 United States Census Bureau’s Annual Business Survey. Transparency is increased in three ways: 1) all specification testing and identifying potentially productive models is done in an exploratory subsample that 2) preserves the validity of hypothesis test statistics fromde novoestimation in the holdout confirmatory sample with 3) all findings publicly documented in an earlier registered report and in this journal publication. Bayesian estimation procedures that leverage information from the exploratory stage included in the confirmatory stage estimation replace traditional frequentist null hypothesis significance testing. In addition to increasing statistical power by using information from the full sample, Bayesian methods directly estimate a probability distribution for the magnitude of an effect, allowing much richer inference. Estimated magnitudes of diversity along academic discipline, race, ethnicity, and foreign-born status dimensions are positively associated with innovation. A maximally diverse ownership team on these dimensions would be roughly six times more likely to report new-to-market innovation than a homophilic team.

Science & Technology - Other Topics