Search NASA⌕ Search

SEARCH · Search NASA

Results for “Bayesian sampling”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 109 records · Page 6

Non-smooth Bayesian optimization in tuning scientific applications

Tuning algorithmic parameters to optimize the performance of large, complicated computational codes is an important problem involving finding the optima and identifying regimes defined by non-smooth boundaries in black-box functions. Within the Bayesian optimization framework, the Gaussian process surrogate model produces smooth mean functions, but functions in the tuning problem are often non-smooth, which is exacerbated by the fact that we usually have limited sequential samples from the black-box function. Here, motivated by these issues encountered in tuning, we propose a novel Gaussian process model called a clustered Gaussian process (cGP), where the components are dynamically updated by clustering. In our studies, the performance of cGP can be better than stationary GPs in nearly 90% of the experiments and better than non-stationary GPs in nearly 70% of the repeated experiments while requiring less computational cost. cGP provides a novel approach for dynamic GP, computes more efficiently than recursive partitioning, and discovers non-smoothness regimes. We provide extensive experiments including high-performance computing (HPC) and industrial simulation functions to show the effectiveness of our methods.

97 MATHEMATICS AND COMPUTING↗

Role of the likelihood for elastic scattering uncertainty quantification

In the last decade, uncertainty quantification (UQ) for optical model potentials (OMPs) has become a focal point for nuclear reaction theory, and several competing approaches for OMP UQ have recently been developed. Here, we clarify recent efforts to compare frequentist and Bayesian approaches in the context of OMP UQ [G. B. King et al., Phys. Rev. Lett. 122, 232502 (2019)]. We replicate a portion of that OMP UQ study but use independent statistical tools. Specifically, we compare two methods for OMP parameter inference from elastic scattering data: the Levenberg-Marquardt algorithm for χ 2 minimization on one hand and Markov chain Monte Carlo (MCMC) sampling on the other. Separately, we assess the common practice of using a renormalized likelihood (χ 2 /N), N being the number of data points, instead of the canonical weighted-least-squares likelihood (χ 2 ), as a way of accounting for unknown data correlations. Here, we show that for a generic linear model and for a five-parameter OMP analysis, frequentist and uniform-prior Bayesian approaches recover the same optimum and uncertainty estimates—not systematically larger uncertainties for the Bayesian approach, as was concluded in G. B. King et al., Phys. Rev. Lett. 122, 232502 (2019). Further, we show that if an additional, near-degenerate parameter is introduced into the same OMP analysis such that the parameter posterior becomes non-Gaussian, then covariance-based estimates of uncertainty become unreliable. Finally, we show that regardless of optimization approach, if χ 2 /N is used for the likelihood, the resulting parametric uncertainties increase by $\sqrt{N}$, and that this is responsible for the conclusions drawn in the revisited study. Based on our replication results, we find that a fortuitous cancellation of unreported errors and the renormalization factor can lead to improvement in empirical coverages, as was the case in the original comparative study. We emphasize that developing and applying a realistic likelihood function is an essential task in a UQ analysis, and that several recent UQ studies that employed a renormalized likelihood (i.e., including a 1/N factor) may have yielded unrealistically large uncertainties for elastic-scattering observables. If the parameter posterior deviates from multivariate-normal, a sampling-based approach like MCMC has a clear advantage over methods that assume the Laplace approximation holds. We note that empirical coverage can serve as an important internal check for the analyst whose model or data may have additional, unaccounted-for uncertainties.

73 NUCLEAR PHYSICS AND RADIATION PHYSICS↗

NOvA 2024 official data release (26.61E20 neutrino + 12.5E20 antineutrino)

This data release corresponds to the Bayesian 2024 analysis of NOvA $\nu_e$ appearance and $\nu_{\mu}$ disappearance data, corresponding to analysis described in https://arxiv.org/abs/2509.04361. File `NOvA_2024_data_histograms.root` contains data histograms for all the NOvA data samples. Exposure: * neutrino-enhanced beam: 26.61E20 protons on target * antineutrino-enhanced beam: 12.5E20 protons on target External constraints: * ss2th12=0.851, dm21=7.53e-5 are fixed at 2019 PDG values, with negligible effect on NOvA predictions. * ss2th13 & dm32: * RCDB1D: 1D constraint from Daya Bay for ss2th13, 0.0851+/-0.0024 * RCDB2D: Correlated 2D constraint from Daya Bay on ss2th13 & dm32, available in their official 2023 data release: https://journals.aps.org/prl/abstract/10.1103/PhysRevLett.130.161802#supplemental An additional file containing the predictions (`NOvA_2024_prediction_with_systs_histograms.root`) for all channels’ signal and background components computed at the NOvA best-fit oscillation parameters and systematic pull terms. Numu samples include the no-oscillation case as well. Please note that any fits performed with these histograms are not expected to exactly reproduce the official NOvA results as parameterizations of the numerous systematic uncertainties considered in the official fits are not included in this release. The zip file also contains the 2D credible interval contours, with details in a README.md

Sztuc, Artur [University Coll. London] (ORCID:0000↗

Uncertainty Quantification for Neutron Shield Using Convolutional Neural Networks

Uncertainty quantification from radiation transport calculations was conducted using a Bayesian inference approach. A surrogate model, using a convolutional neural network, was employed to emulate the neutron fluence, which was simulated with a Monte Carlo radiation transport model. This allowed for a computationally cheap approach to evaluate input parameters and to sample their corresponding posterior probability distributions. Experimental data from the literature were employed to perform uncertainty quantification studies for concrete shields. As a result, the method is a nonintrusive approach that enables studies with multiple input parameters and can be applied to any radiation transport model.

Bayesian inference↗

Machine Learning-Guided Optimization of SABRE Hyperpolarization for α-Ketoglutarate in Acetone–Water

Signal amplification by reversible exchange (SABRE) is a hyperpolarization method that polarizes target nuclei of metabolites quickly and efficiently. Recent SABRE advances, including Ace-SABRE, yield biocompatible, aqueous solutions of hyperpolarized markers for metabolic monitoring. Building on recent advancements, expanding the substrate scope of Ace-SABRE is desirable. However, SABRE polarization is sensitive to many different parameters; therefore, traditional optimization approaches are experimentally time-consuming. In this proof-of-concept application of machine learning (ML), Bayesian optimization (BO) is used for four important input parameters to model the complex SABRE dynamics while saving experimental time. The presented ML model also provides chemical insights that enable predictions of sample compositions for increased polarization levels. In this paper, we transition from an original average free polarization of p = ∼0.90% to a maximum observed free polarization of p = ∼6.6% for 1- 13 C alpha-ketoglutarate (AKG) with 13 C at natural abundance, utilizing both direct outputs as well as chemical insights revealed by the ML model.

Catalysts↗

Rapid Bayesian High Entropy Alloy Designs Fabricated via Wire Arc Additive Manufacturing

Purpose: This project seeks to demonstrate a new high-throughput (rapid) alloy design technique applied to creating new high entropy alloys (HEAs) for extreme environments. High entropy alloys shift the design paradigm from being focused on a single principal element (e.g. nickel-based alloys) to target alloys that include high atomic fractions (X >10%) of multiple elements. These HEA materials can exhibit sluggish diffusion and enhanced corrosion resistance, ideal for potential applications in advanced ultra supercritical (A-USC) steam cycles for power generation. Scope: The addition of multiple elements in high atomic fractions creates an enormous design space that cannot easily be investigated by traditional material design strategies such as designed of experiments (DOE). This project utilizes a Bayesian machine learning algorithm that has been modified to work with calculation of phase diagrams (CALPHAD) software. This Bayesian algorithm reduces manual inputs and increase the likelihood of achieving an optimal solution. Compositional inputs to this algorithm will be assessed using existing material property models for high temperature strength and corrosion resistance. The target for alloy performance will be a 15% (~100 ⁰C) increase in allowable service temperature beyond heat-resistant stainless steels while maintaining or improving alloy cost and corrosion resistance. Haynes 230 was selected as a baseline, which is 57 wt% Ni with 22 wt% Cr 14 wt% W, and 2 wt% Mo as solid solution strengtheners. In addition to rapid design via Bayesian machine learning, the alloys were rapidly fabricated using a multi-wire arc additive manufacturing (mWAAM) technique which allows for precise control of alloy composition and assessing of alloy design “windows” to study composition effects. Build speeds for wire-arc additive processes are among the highest for additive technologies enabling rapid and reliable sample fabrication when compared to conventional methods such as arc button melting. The mWAAM samples will be rapidly characterized via instrumented indentation for room temperature modulus and strength and for elevated temperature strength via hot hardness tests. After being screened with hardness testing, potential alloys will be further evaluated with conventional microscopy techniques including scanning electron microscopy (SEM) and transmission electron microscopy (TEM) to assess agreement with modeling results. The most promising compositions will also be evaluated by printing full sized tensile specimens for mechanical behavior tests at elevated temperatures. Results: Bayesian machine learning of a single performance function was initially used to optimize five performance metrics: 1) single phase stability, 2) yield strength, 3) creep resistance (low diffusion coefficient), 4) freezing range (weldability), and 5) material cost. The single performance function was suboptimal as assumptions had to be made about the results while formulating the optimization. A goal-oriented Bayesian optimization strategy (Hanaoka, 2021) was implemented with CALPHAD for use with the five metrics above. This multi-objective Bayesian optimization (MOBO) enabled the design of NiCrCoFe alloys with V and W additions. A base composition of NiCoCr was selected as Ni provides a stable FCC matrix, Cr aids corrosion/oxidation resistance, and Co is a solid-solutions strengthener that also improves creep by increasing the activation energy. Fe helps reduce diffusion coefficients and cost. Finally, V and W were selected for their reasonable solubility and high atomic misfit to aid in solid solution strengthening. Cracking of the mWAAM specimens was an early issue, and the Easton solidification cracking model (Easton et al., 2014a) was selected for addition to the MOBO function. High performing alloys fabricated by mWAAM included Ni 28 Cr 25 Co 26 Fe 15 V 8 and Ni 62 Cr 18 Co 1 Fe 3 W 15 . It was observed that even after adapting the mWAAM process for W, the W did not fully dissolve. To fully evaluate the Ni 62 Cr 18 Co 1 Fe 3 W 15 composition, a cored wire (80-20 NiCr sheath/powder core) was manufactured and printed via WAAM, and HIP’ing was utilized to homogenize and densify the printed alloy. The V and W alloys produced met metrics 1 (solid solution), 4 (solidification cracking), and 5 (cost). However, an unmodeled mechanism of thermal stress cracking was identified in the WAAM produced materials, perhaps exacerbated by the lack of grain boundary strengthening elements (B, C). Conclusions & Recommendations: A high-throughput (rapid) alloy design technique was applied to designing and manufacturing new high entropy alloys (HEAs) for extreme environments utilizing MOBO and mWAAM. The developed process was rapid and effective in addressing the mechanisms included in the model. The lack of grain boundary strengthening element additions (e.g., B, C) was a simplification that likely produced thermal stress cracking that turned into a large part of the investigation. Additions on the order of 0.005 wt% B and 0.05 wt% C likely would have minimized thermal stress grain boundary cracking. Overall, the high throughput design strategy is promising for rapid design of metrics-driven alloys for advanced ultra supercritical (A-USC) steam cycles for power generation. The MOBO and mWAAM process could be commercialized to accelerate metrics-driven alloy design. In addition, the cored-wire process utilized for scale-up is a promising high-volume process for WAAM alloy development and scale-up.

36 MATERIALS SCIENCE↗

Bayesian Framework for Bioburden Density Estimation in Planetary Protection

To comply with the international planetary protection policy set forth by the Committee on Space Research and NASA Agency level requirements, spacecraft destined to biologically sensitive planetary bodies have to minimize terrestrial biological contamination. Analysis, testing and inspection are the standard forward verification activities that are used to demonstrate compliance with the biological contamination requirements. For testing of spacecraft surface areas, a swab or wipe sample is collected from surfaces prior to last access and subsequently processed in the lab using NASA Approved Planetary Protection Methods for Culture Based Assays. Raw data resulting from this assay is then statistically treated employing a mathematical paradigm stemming from the 1970’s Viking Lander Project to generate the bioburden density and total microbial bioburden present. This standard approach arbitrarily accounts for error and provides an upper conservative bound as it reports the maximum number of spores estimated to be present on flight hardware surfaces. A bioburden density estimate factors in the following variables: the observed bioburden count, representative volume processed, sampling efficiencies. Notably, to account for error in the approach, a 0 observed count is arbitrarily changed to a count of 1 for each hardware grouping. The data generated by spacecraft bioburden verification campaigns in the past have resulted in <80% of wipes and <90% of swabs containing a bioburden count of 0. As such, having a robust and well documented statistical approach for dealing with the probability of low incident rates is necessary to be able to estimate spacecraft bioburden. Being able to statistically describe the bioburden distribution and associated confidence level is a gamechanger for the development of bioburden allocations during mission design and will allow for tighter management of risk throughout spacecraft build. Thus, Empirical Bayes statistical approach was evaluated to estimate the microbial bioburden on spacecraft to mitigate the aforementioned mathematical concerns and provide a probabilistic bioburden distribution of the flight hardware surface. For application of this approach to performing bioburden calculations, a range of non-informative prior assumptions on hardware surfaces are explored for Bayesian analyses while informative priors using posterior distributions from prior assays are utilized for Empirical Bayes analyses. Several non-informative priors are currently under investigation to assess fitness including use of these priors to serve as a foundation to build off of NASA specification values or a basis of risk to account for unknowns during the integration and testing process. Informative priors under consideration are generated using sampled bioburden values from hardware originating within like processing environments (e.g. vendor cleaning process or similar assembly process), temporal spacecraft status events as a prediction for hardware cleanliness of future samples, and heritage system bioburden actuals to predict allocation for subsequent missions. Informative priors and probabilistic bioburden distributions are then validated using data sets from the Mars Exploration Rover, Mars Science Laboratory, and InSight missions. Using Empirical Bayes approach to generate a probabilistic bioburden distribution as demonstrated through mission use cases provides a valid approach for use in the end-to-end requirements verification process.

97 - MATHEMATICS AND COMPUTING↗

Sequential Kalman tuning of the t -preconditioned Crank-Nicolson algorithm: efficient, adaptive and gradient-free inference for Bayesian inverse problems

Ensemble Kalman Inversion (EKI) has been proposed as an efficient method for the approximate solution of Bayesian inverse problems with expensive forward models. However, when applied to the Bayesian inverse problem EKI is only exact in the regime of Gaussian target measures and linear forward models. Here, in this work we propose embedding EKI and Flow Annealed Kalman Inversion, its normalizing flow (NF) preconditioned variant, within a Bayesian annealing scheme as part of an adaptive implementation of the t-preconditioned Crank-Nicolson (tpCN) sampler. The tpCN sampler differs from standard pCN in that its proposal is reversible with respect to the multivariate t-distribution. The more flexible tail behaviour allows for better adaptation to sampling from non-Gaussian targets. Within our Sequential Kalman Tuning (SKT) adaptation scheme, EKI is used to initialize and precondition the tpCN sampler for each annealed target. The subsequent tpCN iterations ensure particles are correctly distributed according to each annealed target, avoiding the accumulation of errors that would otherwise impact EKI. We demonstrate the performance of SKT for tpCN on three challenging numerical benchmarks, showing significant improvements in the rate of convergence compared to adaptation within standard SMC with importance weighted resampling at each temperature level, and compared to similar adaptive implementations of standard pCN. The SKT scheme applied to tpCN offers an efficient, practical solution for solving the Bayesian inverse problem when gradients of the forward model are not available. Code implementing the SKT schemes for tpCN is available at https://github.com/RichardGrumitt/KalmanMC.

97 MATHEMATICS AND COMPUTING↗

Bayesian analysis of (3 +1)⁢D relativistic nuclear dynamics with the RHIC beam energy scan data

This work presents a Bayesian inference study for relativistic heavy-ion collisions in the beam energy scan program at the BNL Relativistic Heavy-Ion Collider. The theoretical model simulates event-by-event (3+1)-dimensional [(3+1)⁢D] collision dynamics using hydrodynamics and hadronic transport theory. We analyze the model's 20-dimensional posterior distributions obtained using three model emulators with different accuracy and demonstrate the essential role of training an accurate model emulator in the Bayesian analysis. Our analysis provides robust constraints on the quark-gluon plasma's transport properties and various aspects of (3+1)⁢D relativistic nuclear dynamics. By running full model simulations with 100 parameter sets sampled from the posterior distribution, we make predictions for p T -differential observables and estimate their systematic theory uncertainty. Here, a sensitivity analysis is performed to elucidate how individual experimental observables respond to different model parameters, providing useful physics insights into the phenomenological model for heavy-ion collisions.

73 NUCLEAR PHYSICS AND RADIATION PHYSICS↗

Uncovering multiscale structure-property correlations via active learning in scanning tunneling microscopy

Atomic arrangements and local sub-structures fundamentally influence emergent material functionalities. These structures are conventionally probed using spatially resolved studies and the property correlations are deciphered by a researcher based on sequential explorations, thereby limiting the efficiency and scope. Here we demonstrate a multi-scale Bayesian deep-learning based framework that automatically correlates material structure with its electronic properties using scanning tunneling microscopy (STM) measurements in real-time. Its predictions are used to autonomously direct exploration toward regions of the sample that optimize a given material property. This method is deployed on a low-temperature ultra-high vacuum STM to understand the structure-property relationship in a europium-based semimetal, EuZn 2 As 2 , a promising candidate relevant to magnetism-driven topological phenomena. The framework employs a sparse-sampling approach to efficiently construct the scalar-property space using minimal measurements, about 1–10% of the data required in standard hyperspectral methods. Moreover, we formulate the problem hierarchically across length scales, implementing autonomous workflow to locate mesoscopic and atomic structures that correspond to a target material property. This framework offers the choice to design scalar-property from the spectroscopic data to steer sample exploration. Our findings reveal correlations of the electronic properties unique to surface terminations, local defect density, and point defects.

75 CONDENSED MATTER PHYSICS, SUPERCONDUCTIVITY AND↗

Bayesian chain graph models to characterize microbe-environment dynamics

Microbiome data require statistical models that can simultaneously decode microbes' reaction to the environment and interactions among microbes. While a multiresponse linear regression model seems like a straight-forward solution, we argue that treating it as a graphical model is problematic given that the regression coefficient matrix does not encode the conditional dependence structure between response and predictor nodes. This observation is especially important in biological settings when we have prior knowledge on the edges from specific experimental interventions that can only be properly encoded under a conditional dependence model. Here, we propose a chain graph model with two sets of nodes (predictors and responses) whose solution yields a graph with edges that indeed represent conditional dependence, thus agreeing with the experimenter's intuition on the average behavior of nodes under treatment. The solution to our model is sparse via the Bayesian linear regression (LASSO). In addition, we propose an adaptive extension so that different shrinkages can be applied to different edges to incorporate edge-specific prior knowledge. Our model is computationally inexpensive through an efficient Gibbs sampling algorithm and can account for binary, counting, and compositional responses via an appropriate hierarchical structure. We test the performance of our model in a variety of simulated datasets, thereby showing superior performance to state-of-the-art approaches. We further apply our model to human gut and soil microbial compositional datasets, and we highlight that CG-LASSO can estimate biologically meaningful network structures in the data.

compositional data↗

Uranium particle age dating, aggregation, and model age best estimators

We present important aspects of uranium particle age dating by Large-Geometry Secondary Ion Mass Spectrometry (LG-SIMS) that can introduce bias and increase model age uncertainties, especially for small, young, and/or low-enriched particles. This metrology is important for applications related to International Nuclear Safeguards. We explore influential factors related to model age estimation, including the effects of evolving surface chemistry on inter-element measurements of particles (e.g., Th and U), detector background, and aggregation methods using simulated and actual particle samples. We introduce a new model age estimator, called “mid68”, that supplements 95% confidence intervals, providing a “best estimate” and uncertainty about the most likely age. The mid68 estimator can be calculated using the Feldman and Cousins method or Bayesian methods and provides a value with a symmetric uncertainty that can be used for calculations and approximate aggregation of processed model age values when the raw data and correction factors are not available. For particles yielding low 230 Th counts amidst nonzero detector background, their underlying model age probability distributions are asymmetric, so the mid68 estimator provides additional robust information regarding the underlying model age likelihood. This study provides a comprehensive and timely examination of critical aspects of uranium particle age dating as more laboratories establish particle chronometry capabilities.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Selection function of clusters in Dark Energy Survey year 3 data from cross-matching with South Pole Telescope detections

Context. Galaxy clusters selected based on overdensities of galaxies in photometric surveys provide the largest cluster samples. However, modeling the selection function of such samples is complicated by noncluster members projected along the line of sight (projection effects) and the potential detection of unvirialized objects (contamination). Aims. We empirically constrained the magnitude of these effects by cross-matching galaxy clusters selected in the Dark Energy Survey data with the redMaPPer algorithm with significant detections in three South Pole Telescope surveys (SZ, pol-ECS, pol-500d). Methods. For matched clusters, we augmented the redMaPPer catalog with the SPT detection significance. For unmatched objects we used the SPT detection threshold as an upper limit on the SZe signature. Using a Bayesian population model applied to the collected multiwavelength data, we explored various physically motivated models to describe the relationship between observed richness and halo mass. Results. Our analysis reveals a clear preference for models with an additional skewed scatter component associated with projection effects over a purely log-normal scatter model. We rule out significant contamination by unvirialized objects at the high-richness end of the sample. While dedicated simulations offer a well-fitting calibration of projection effects, our findings suggest the presence of redshift-dependent trends that these simulations may not have captured. Our findings highlight that modeling the selection function of optically detected clusters remains a complicated challenge that requires a combination of simulation and data-driven approaches.

79 ASTRONOMY AND ASTROPHYSICS↗

Karhunen–Loève deep learning method for surrogate modeling and approximate Bayesian parameter estimation

We evaluate the performance of the Karhunen-Loève Deep Neural Network (KL-DNN) framework for surrogate modeling and approximate Bayesian parameter estimation in partial differential equation models. In the surrogate model, the Karhunen-Loève (KL) expansions are used for the dimensionality reduction of the number of unknown parameters and variables, and a deep neural network is employed to relate the reduced space of parameters to that of the state variables. The KL-DNN surrogate model is used to formulate a maximum-a-posteriori-like least-squares problem, which is randomized to draw samples of the posterior distribution of the parameters. We test the proposed framework for a hypothetical unconfined aquifer via comparison with the forward MODFLOW and inverse PEST++ iterative ensemble smoother (IES) solutions as well as the state-of-the-art Fourier neural operator (FNO) and deep operator networks (DeepONets) operator learning surrogate models. Our results show that the KL-DNN surrogate model outperforms FNO and DeepONet for forward predictions. For solving inverse problems, the randomized algorithm provides the same or more accurate Bayesian predictions of the parameters than IES as evidenced by the higher log-predictive probability of both the estimated parameter field and the forecast hydraulic head. The posterior mean obtained from the randomized algorithm is closer to the reference parameter field than that obtained with FNO as the maximum a posteriori estimate.

Approximate Bayesian inference↗

Analysis of differential scanning calorimetry data for aged plutonium

Differential scanning calorimetry data for samples of a 52 year old plutonium alloy with 3.3 at. % Ga that were heated beyond the melting point is analyzed using transition state theory to find activation energies for the δ to ε and ε to liquid phase transitions. A Bayesian statistical method involving a Gaussian process model is used to find mean values and confidence intervals for the activation energies. The activation energy for the δ to ε phase transition increases by 3.3 ± 3.8% per decade, relative to the case when all age related plutonium lattice point defects have been removed through annealing. The corresponding increase in activation energy for the ε to liquid transition is shown to be 7.1 ± 1.8% per decade. It is postulated that the change in activation energy with age for both phase transitions is caused, in part, by the accumulation of the same type of lattice point defects associated with the observed increase in elastic bulk modulus over time.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Designing an Optimal Sensor Network via Minimizing Information Loss

Optimal experimental design is a classic topic in statistics, with many well-studied problems, applications, and solutions. The design problem we study is the placement of sensors to monitor spatiotemporal processes, explicitly accounting for the temporal dimension in our modeling and optimization. We observe that recent advancements in computational sciences often yield large datasets based on physics-based simulations, which are rarely leveraged in experimental design. We introduce a novel model-based sensor placement criterion, along with a highly-efficient optimization algorithm, which integrates physics-based simulations and Bayesian experimental design principles to identify sensor networks that “minimize information loss” from simulated data. Our technique relies on sparse variational inference and (separable) Gauss-Markov priors, and thus may adapt many techniques from Bayesian experimental design. We validate our method through a case study monitoring air temperature in Phoenix, Arizona, using state-of-the-art physics-based simulations. Our results show our framework to be superior to random or quasi-random sampling, particularly with a limited number of sensors. We conclude by discussing practical considerations and implications of our framework, including more complex modeling tools and real-world deployments.

54 ENVIRONMENTAL SCIENCES↗

ELG×LRG Distribution through Dark Matter Halo Dynamics

We investigate the clustering and halo occupation distribution (HOD) of DESI Y1 emission-line (ELGs) and luminous red (LRGs) galaxies at 0.8 < z < 1.1, including their cross-correlation (ELG×LRG), using the A BACUS S UMMIT suite and a new Halo Occupation Model (H OME ) for galaxy multitracers. This integrates intrahalo dynamics, halo exclusion, and quenching, bridging insights from hydrodynamical, HOD, abundance-matching, and semianalytic studies. Leveraging full phase-space information from the Uchuu N-body simulation, and sampling satellites from dark-matter particle positions via physically motivated prescriptions, Home reproduces the anisotropic clustering down to s = 200 h −1 kpc with unprecedented accuracy. Model parameters are inferred solely from two-point statistics using a two-level Bayesian framework, yielding high-fidelity ELG, LRG, and cross-reference catalogs. We find that satellite ELGs behave as incoherent flows within their parent halos, dominating the clustering below 4 h −1 Mpc. The HOD from the best-fit Home has the following properties: (i) 90.50% (85.91%) of ELGs (LRGs) are central galaxies without satellites, residing in halos of M vir ∼ 6.6 × 10 11 (1.2 × 10 13 ) h −1 M ⊙ ; (ii) the ELG×LRG cross-correlation is governed by central-central pairs and shaped by halo exclusion on 2–5 h −1 Mpc scales; (iii) 9.50% (14.09%) of ELGs (LRGs) are satellites, of which 1.09% (3.52%) inhabit halos with a central galaxy of the same species in a maximally conformal configuration, 7.02% (0.005%) orbit complementary hosts in a minimally conformal state, and 0.58% (10.57%) are orphans. The high sensitivity of Home precisely captures the dynamics of satellites in different host environments, opening a promising avenue for understanding systematics and the dynamical nature of dark matter, potentially distinguishing gravity models.

Favole, Ginevra [Universidad de La Laguna (Spain);↗

Taylor approximation variance reduction for approximation errors in PDE-constrained Bayesian inverse problems

In numerous applications, surrogate models are used as a replacement for accurate parameter-to-observable mappings when solving large-scale inverse problems governed by partial differential equations (PDEs). The surrogate model may be a computationally cheaper alternative to the accurate parameter-to-observable mappings and/or may ignore additional unknowns or sources of uncertainty. The Bayesian approximation error (BAE) approach provides a means to account for the induced uncertainties and approximation errors, i.e. the errors between the accurate parameter-to-observable mapping and the surrogate. The statistics of these errors are, however, in general unknown a priori, and are thus calculated using Monte Carlo sampling. Although the sampling is typically carried out offline, i.e. before considering the data, the process can still represent a computational bottleneck. In this work, we develop a scalable computational approach for reducing the costs associated with the sampling stage of the BAE approach. Specifically, we consider the Taylor expansion of the accurate and surrogate forward models with respect to the uncertain parameter fields either as a control variate for variance reduction or as a means to directly and efficiently approximate the mean and covariance of the approximation errors. We propose efficient methods for evaluating the expressions for the mean and covariance of the Taylor approximations based on linear(-ized) PDE solves. Furthermore, the proposed approach is independent of the dimension of the uncertain parameter, depending instead on the intrinsic dimension of the data, ensuring scalability to high-dimensional problems. The potential benefits of the proposed approach are demonstrated for two high-dimensional inverse problems governed by PDE examples, namely for the estimation of a distributed Robin boundary coefficient in a linear diffusion problem, and for a coefficient estimation problem governed by a nonlinear diffusion problem.

Bayesian approximation error↗