Search NASA⌕ Search

SEARCH · Search NASA

Results for “Bayesian model calibration”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 109 records · Page 6

Chiral interactions and superfluidity in the calcium isotopic chain

We perform ab initio calculations of three-point mass differences in the odd- and even-mass 39−49Ca isotopes to probe nuclear superfluidity via empirical neutron pairing gaps. We also quantify the sensitivity of those gaps to the parameters of the interaction at mean-field level. Recent studies employing accurate chiral nuclear interactions have found these gaps to be too small. We show that experimental values can be reproduced at mean-field level by substantially increasing the attraction of the singlet 𝑆 -wave two-nucleon contact interaction, but doing so induces an unphysical bound state of the dineutron. The sensitivity of these predictions to the full calibration of the nuclear interaction is then studied by performing Bayesian posterior sampling in a delta-full chiral effective field theory at third chiral order. We find that pairing gaps remain largely unaffected, leaving the explanation of nuclear superfluidity as a future task for improved many-body modeling and refined interactions at higher chiral orders.

Hagen, Gaute [ORNL] (ORCID:0000000160191687)↗

The SRG/eROSITA All-Sky Survey: Dark Energy Survey year 3 weak gravitational lensing by eRASS1 selected galaxy clusters

Context. Number counts of galaxy clusters across redshift are a powerful cosmological probe if a precise and accurate reconstruction of the underlying mass distribution is performed – a challenge called mass calibration. With the advent of wide and deep photometric surveys, weak gravitational lensing (WL) by clusters has become the method of choice for this measurement. Aims. We measured and validated the WL signature in the shape of galaxies observed in the first three years of the Dark Energy Survey (DES Y3) caused by galaxy clusters and groups selected in the first all-sky survey performed by SRG (Spectrum Roentgen Gamma)/eROSITA (eRASS1). These data were then used to determine the scaling between the X-ray photon count rate of the clusters and their halo mass and redshift. Methods. We empirically determined the degree of cluster member contamination in our background source sample. The individual cluster shear profiles were then analyzed with a Bayesian population model that self-consistently accounts for the lens sample selection and contamination and includes marginalization over a host of instrumental and astrophysical systematics. To quantify the accuracy of the mass extraction of that model, we performed mass measurements on mock cluster catalogs with realistic synthetic shear profiles. This allowed us to establish that hydrodynamical modeling uncertainties at low lens redshifts (z < 0.6) are the dominant systematic limitation. At high lens redshift, the uncertainties of the sources’ photometric redshift calibration dominate. Results. With regard to the X-ray count rate to halo mass relation, we determined its amplitude, its mass trend, the redshift evolution of the mass trend, the deviation from self-similar redshift evolution, and the intrinsic scatter around this relation. Conclusions. The mass calibration analysis performed here sets the stage for a joint analysis with the number counts of eRASS1 clusters to constrain a host of cosmological parameters. We demonstrate that WL mass calibration of galaxy clusters can be performed successfully with source galaxies whose calibration was performed primarily for cosmic shear experiments, opening the way for the cluster cosmological exploitation of future optical and NIR surveys like Euclid and LSST.

79 ASTRONOMY AND ASTROPHYSICS↗

pop-cosmos : redshifts and physical properties of KiDS-1000 galaxies

ABSTRACT Principled Bayesian inference of galaxy properties has not previously been performed for wide-area weak-lensing surveys with millions of sources. We address this gap by applying the pop-cosmos generative model to perform spectral energy distribution (SED) fitting for 4 million KiDS (Kilo-Degree Survey)-1000 galaxies. Calibrated on deep COSMOS2020 photometric data, pop-cosmos specifies a physically motivated prior over the galaxy population up to $z \simeq 6$ in stellar population synthesis (SPS) parameter space. Using the Speculator SPS emulator with GPU (graphics processing unit)-accelerated Markov Chain Monte Carlo sampling, we perform full posterior inference at 8.2 GPU seconds per galaxy, obtaining joint constraints on galaxy redshifts and physical properties. We validate photometric redshifts against $\sim \!185\,\!000$ KiDS galaxies cross-matched to Dark Energy Spectroscopic Instrument Data Release 1 spectroscopic samples, achieving low bias ($2\times 10^{-3}$), scatter ($\sigma _{\mathrm{MAD}}=0.03$), and outlier fraction (3.2 per cent) for the Bright Galaxy Survey, with comparable performance (bias $3\times 10^{-2}$, $\sigma _{\mathrm{MAD}}=0.05$, 1.0 per cent outliers) for luminous red galaxies (LRGs). Within the LRG sample, we identify massive, dusty, star-forming contaminants at $z \simeq 0.4$ satisfying standard colour selections for quenched populations. We infer trends in stellar mass, star formation, metallicity, and dust across five tomographic redshift bins consistent with established scaling relations. Using specific star formation rate constraints, we identify $\sim$7 per cent of KiDS-1000 galaxies as quenched, versus 37 per cent implied by conservative colour cuts. This enables the construction of weak-lensing samples defined by physical properties while mitigating intrinsic alignment systematics and preserving statistical power. Our analysis validates pop-cosmos out of sample, establishing it as a scalable approach for galaxy evolution and cosmological analyses with photometric surveys.

Halder, Anik [Institute of Astronomy and Kavli Ins↗

Searches for New Physics With Muon Conversion at Fermilab and Triboson Production at the LHC

We report on several efforts to search for physics beyond the standard model of particle physics at broad energy scales. The Mu2e experiment at Fermilab will search for charged lepton flavor violation via the muon to electron conversion process, which is suppressed in the Standard Model. Mu2e will be operated at a low energy, yet can probe New Physics at very high mass scales (O(1e3 - 1e4 ) TeV). At high energies, the CMS experiment at the CERN LHC continues to deliver an impressive suite of Standard Model measurements and limits on a variety of New Physics signatures. Mu2e is under construction and slated to collect its first physics data in the coming years. This thesis describes work done during the construction phase of Mu2e and focuses on two critical areas: magnetic field modeling and statistical analysis. We describe a novel method for field modeling which we validate using a simulated dataset representing the expected magnetic field in the Detector Solenoid. This method blends a standard least-squares fitting technique that utilizes physically motivated analytical model functions with a novel physics informed network that is constructed to obey Maxwell’s equations. We show the technique can model the field with an accuracy of 10−7 despite the presence of injected noise in the pseudo-measurements at the 10−5 level. We then present preliminary results of the calibration of 3D Hall probes at the sub-10−4 level. These probes will be used to directly measure the Mu2e Detector Solenoid magnetic field on a sparse grid; these measurements serve as the input to the field model fitting. Finally, we describe the first implementation of both an unbinned shape analysis and a Bayesian interpretation applied to Mu2e pseudo-data. Up to 20% tighter limits can be set by the shape analysis compared to a standard cut & count analysis. The AlCap experiment collected data at PSI in 2015 to measure several important quantities related to nuclear muon capture on an aluminum target, which is a significant background process for Mu2e. The neutron emission from muon capture can introduce background hits in the Mu2e detectors and can increase radiation damage in various elements of the apparatus. We present measurements of the neutron group fluence and mean neutron multiplicity for muon capture on aluminum nuclei. Finally, we discuss an analysis of triboson production at CMS using an Effective Field Theory framework. Standard Model triboson production, which was first observed at CMS in 2020, has a relatively small cross section and provides direct access to both anomalous triple gauge couplings and quartic gauge couplings. These couplings, interpreted in the Standard Model Effective Field Theory, are studied in the present work. We target the boosted regime where the background rate is low and yields are enhanced when dimension-6 and dimension-8 Wilson coefficients are non-zero. We do not observe an excess in the data and therefore set bounds on the Wilson coefficients. For dimension-6 coefficients the tightest observed (expected) bounds are set on cW /Λ2 where Λ is the mass scale of new physics; the bounds are [−0.13, 0.12] TeV−2 ([−0.12, 0.12] TeV−2 ) at 95% CL. The tightest bounds in dimension-8 are set on fT,0 / Λ4 ; the observed (expected) bounds at 95% CL are [−0.63, 0.69] TeV−4 ([−0.54, 0.62] TeV−4 ). Additional results are presented which include scenarios where multiple Wilson coefficients are non-zero, the application of signal model clipping to address unitarity violation in Effective Field Theories, and a novel template fit developed for easier reinterpretation of our results.

Kampa, Cole Erik [Northwestern U. (main)] (ORCID:0↗

Confidence-weighted integration of human and machine judgments for superior decision-making

Large language models (LLMs) can surpass humans in certain forecasting tasks. What role does this leave for humans in the overall decision process? One possibility is that humans, despite performing worse than LLMs, can still add value when teamed with them. A human and machine team can surpass each individual teammate when team members’ confidence is well calibrated and team members diverge in which tasks they find difficult (i.e., calibration and diversity are needed). We simplified and extended a Bayesian approach to combining judgments using a logistic regression framework that integrates confidence-weighted judgments for any number of team members. Using this straightforward method, we demonstrated its effectiveness in both image classification and neuroscience forecasting tasks. Combining human judgments with one or more machines consistently improved overall team performance. Our hope is that this simple and effective strategy for integrating the judgments of humans and machines will lead to productive collaborations.

97 MATHEMATICS AND COMPUTING↗

Emulating ab initio computations of infinite nucleonic matter

We construct efficient emulators for the computation of the infinite nuclear matter equation of state. These emulators are based on the subspace-projected coupled-cluster method for which we here develop a new algorithm called small-batch voting to eliminate spurious states that might appear when emulating quantum many-body methods based on a non-Hermitian Hamiltonian. The efficiency and accuracy of these emulators facilitate a rigorous statistical analysis within which we explore nuclear matter predictions for > 10 6 different parametrizations of a chiral interaction model with explicit Δ -isobars at next-to-next-to leading order. Constrained by nucleon-nucleon scattering phase shifts and bound-state observables of light nuclei up to He 4 , we use history matching to identify nonimplausible domains for the low-energy coupling constants of the chiral interaction. Within these domains we perform a Bayesian analysis using sampling and importance resampling with different likelihood calibrations and study correlations between interaction parameters, calibration observables in light nuclei, and nuclear matter saturation properties. Published by the American Physical Society 2024

73 NUCLEAR PHYSICS AND RADIATION PHYSICS↗

Nuclear-matter saturation and symmetry energy within Δ -full chiral effective field theory

Nuclear saturation and the symmetry energy are key properties of low-energy nuclear physics that depend on fine details of the nuclear interaction. The equation of state around saturation is also an important anchor for extrapolations to higher densities and studies of neutron stars. Here we develop a unified statistical framework that uses realistic nuclear forces to link the theoretical modeling of finite nuclei and infinite nuclear matter. We construct fast and accurate emulators for nuclear-matter observables and employ an iterative history-matching approach to explore and reduce the enormous parameter domain of Δ -full chiral interactions. We perform rigorous uncertainty quantification and find that model calibration including O 16 observables gives saturation predictions that are more precise than those that only use few-body data. Published by the American Physical Society 2024

73 NUCLEAR PHYSICS AND RADIATION PHYSICS↗

Detecting outbreaks using a spatial latent field

In this paper, we present a method for estimating the infection-rate of a disease as a spatial-temporal field. Our data comprises time-series case-counts of symptomatic patients in various areal units of a region. We extend an epidemiological model, originally designed for a single areal unit, to accommodate multiple units. The field estimation is framed within a Bayesian context, utilizing a parameterized Gaussian random field as a spatial prior. We apply an adaptive Markov chain Monte Carlo method to sample the posterior distribution of the model parameters condition on COVID-19 case-count data from three adjacent counties in New Mexico, USA. Our results suggest that the correlation between epidemiological dynamics in neighboring regions helps regularize estimations in areas with high variance (i.e., poor quality) data. Using the calibrated epidemic model, we forecast the infection-rate over each areal unit and develop a simple anomaly detector to signal new epidemic waves. Our findings show that anomaly detector based on estimated infection-rates outperforms a conventional algorithm that relies solely on case-counts.

Safta, Cosmin [Sandia National Laboratories (SNL-C↗

Measurement of $\nu_\mu$ CC Interactions With Two-Proton Final State in MINERvA

This dissertation presents a measurement of charged–current (CC) muon–neutrino interactions with exactly two protons and no pions in the final state (CC~$2p\,0\pi$), using data collected by the MINERvA detector in the NuMI medium–energy beam at Fermilab. Such two–proton topologies are a sensitive probe of nuclear dynamics in the few–GeV regime, including multi–nucleon correlations (npnh, notably $2p2h$) and intranuclear final–state interactions (FSI) such as pion absorption and nucleon rescattering. A precise experimental characterization of these processes is essential both for neutrino–interaction theory and for reducing systematic uncertainties in oscillation experiments that rely on accurate modeling of neutrino–nucleus interactions. Events are selected by requiring a $\nu_\mu$ CC interaction with a reconstructed $\mu^-$ and two proton tracks originating from a common vertex in MINERvA’s finely segmented scintillator tracker, with no reconstructed mesons. Muon charge and momentum are constrained by matching to the MINOS Near Detector, while proton identification exploits energy–loss profiles and stopping–proton features. Backgrounds from pion–producing channels that enter the signal region through FSI or reconstruction effects are constrained with data–driven sidebands (Michel–electron and isolated–cluster “blob” samples) and tuned via a simultaneous fit across signal and sideband regions. To correct detector resolution and acceptance effects, the analysis employs iterative Bayesian unfolding with extensive validation: statistical pseudo–experiments, and robustness checks against generator systematic “universes” and additional strong shape warps. Single–differential cross sections are reported for three observables tailored to the two–proton final state: the opening–angle cosine $\cos\!\left(\theta_{pp}\right)$, the leading–proton momentum, and the subleading–proton momentum. Systematic uncertainties include contributions from neutrino flux, interaction modeling (e.g., npnh and resonance parameters, pion FSI), and detector response (calibration, reconstruction efficiencies). The resulting distributions provide targeted constraints on the interplay of multi–nucleon dynamics and FSI that shape CC~$2p\,0\pi$ final states on hydrocarbon. Comparisons to modern GENIE–based simulations highlight kinematic regions where model components require refinement. These measurements thus inform generator tuning and improve the reliability of neutrino–energy reconstruction strategies for current and future long–baseline oscillation programs.

Syrotenko, Vladyslav S. [Tufts U.]↗

Selection function of clusters in Dark Energy Survey year 3 data from cross-matching with South Pole Telescope detections

Context. Galaxy clusters selected based on overdensities of galaxies in photometric surveys provide the largest cluster samples. However, modeling the selection function of such samples is complicated by noncluster members projected along the line of sight (projection effects) and the potential detection of unvirialized objects (contamination). Aims. We empirically constrained the magnitude of these effects by cross-matching galaxy clusters selected in the Dark Energy Survey data with the redMaPPer algorithm with significant detections in three South Pole Telescope surveys (SZ, pol-ECS, pol-500d). Methods. For matched clusters, we augmented the redMaPPer catalog with the SPT detection significance. For unmatched objects we used the SPT detection threshold as an upper limit on the SZe signature. Using a Bayesian population model applied to the collected multiwavelength data, we explored various physically motivated models to describe the relationship between observed richness and halo mass. Results. Our analysis reveals a clear preference for models with an additional skewed scatter component associated with projection effects over a purely log-normal scatter model. We rule out significant contamination by unvirialized objects at the high-richness end of the sample. While dedicated simulations offer a well-fitting calibration of projection effects, our findings suggest the presence of redshift-dependent trends that these simulations may not have captured. Our findings highlight that modeling the selection function of optically detected clusters remains a complicated challenge that requires a combination of simulation and data-driven approaches.

79 ASTRONOMY AND ASTROPHYSICS↗

Q-Cluster: Quantum Error Mitigation Through Noise-Aware Unsupervised Learning

Quantum error mitigation (QEM) is critical in reducing the impact of noise in the pre-fault-tolerant era, and is expected to complement error correction in fault-tolerant quantum computing (FTQC). In this work, we propose a novel QEM approach, Q-Cluster, that uses unsupervised learning (clustering) to reshape the measured bit-string distribution. Our approach starts with a simplified bit-flip noise model. It first performs clustering on noisy measurement results, i.e., bit-strings, based on the Hamming distance. The centroid of each cluster is calculated using a qubit-wise majority vote. Next, the noisy distribution is adjusted with the clustering outcomes and the bitflip error rates using Bayesian inference. Our simulation results show that Q-Cluster can mitigate high noise rates (up to 40% per qubit) with the simple bit-flip noise model. However, real quantum computers do not fit such a simple noise model. To address the problem, we (a) apply Pauli twirling to tailor the complex noise channels to Pauli errors, and (b) employ a machine learning model, ExtraTrees regressor, to estimate an effective bit-flip error rate using a feature vector consisting of machine calibration data (gate & measurement error rates), circuit features (number of qubits, numbers of different types of gates, etc.) and the shape of the noisy distribution (entropy). Our experimental results show that our proposed Q-Cluster scheme improves the fidelity by a factor of 1.46x, on average, compared to the unmitigated output distribution, for a set of low-entropy benchmarks on five different IBM quantum machines. Our approach outperforms the state-of-art QEM approaches RZNE [28], M3 [24], Hammer [35], and QBEEP [33] by 1.26x,1.29x,1.47x, and 2.65 x, respectively.

42 ENGINEERING↗

Universal reduced basis for the calibration of covariant energy density functionals

The reduced basis method is used to construct a “universal” basis of Dirac orbitals that may be applicable throughout the nuclear chart to calibrate covariant energy density functionals. Relative to the successful development of a reduced basis emulator for the nonrelativistic Schrödinger equation, the Dirac equation adds an extra layer of complexity due to the existence of negative energy states, which complicates building an efficient reduced basis. However, once this problem is mitigated, the resulting reduced basis is able to accurately and efficiently reproduce the high-fidelity model at a fraction of the computational cost. We are confident that the resulting reduced basis will serve as a foundational element in developing rapid and accurate emulators. In turn, these emulators will play a critical role in the Bayesian optimization of covariant energy density functionals.

Bayesian methods↗

Nuclear Data Adjustment for Nonlinear Applications in the OECD/NEA WPNCS SG14 Benchmark—A Bayesian Inverse UQ-Based Approach for Data Assimilation

The Organisation for Economic Co-operation and Development Working Party on Nuclear Criticality Safety has proposed a benchmark exercise to assess the performance of current nuclear data adjustment techniques applied to nonlinear applications and experiments with low correlation to applications. This work introduces Bayesian inverse uncertainty quantification (IUQ) employing scientific machine learning surrogate models as a method for nuclear data adjustments in this benchmark, and compares IUQ to the more traditional methods of generalized linear least squares (GLLS) and Monte Carlo Bayes (MOCABA). Posterior predictions from IUQ showed agreement with GLLS and MOCABA for linear applications. Here, when comparing GLLS, MOCABA, and IUQ posterior predictions to computed model responses using adjusted parameters, we observe that the GLLS predictions failed to replicate the computed response distributions for nonlinear applications, while MOCABA showed near agreement, and IUQ used the computed model responses directly. We also discuss observations on why experiments with low correlation to applications can be informative to nuclear data adjustments and identify some properties useful in selecting experiments for inclusion in nuclear data adjustment. Performance in this benchmark indicates potential for Bayesian IUQ in nuclear data adjustments.

Bayesian calibration↗

Conformal Hierarchical Simulation-Based Inference with Local Validity

Trustworthy and interpretable uncertainty quantification is a long-standing challenge in artificial intelligence. Simulation-based inference (SBI) comprises a broad swath of approaches for estimating latent parameters with uncertainties. Although flexible neural density estimators in SBI can be remark- ably expressive capturing highly structured, high-dimensional posteriors their credible regions can be badly mis-calibrated and are often only accompanied by heuristic coverage checks. We present the first SBI framework that delivers finite-sample local valid coverage guarantees that hold in the neighborhood of each observation. Our framework can couple any off-the-shelf hierarchical SBI engine with a confor- mal Bayesian post-processing step that operates on the posterior predictive density. A kernel-weighted conformity score adapts the conformal quantile to the local geometry of the data, yielding prediction sets that are simultaneously (i) marginally calibrated, (ii) locally valid, and (iii) hierarchical, handling global and observation-specific parameters in a single pass. Through experiments on synthetic data and benchmarks from neuroscience and physics, we show that our approach attains 1 − α coverage, where prior SBI methods under- or over-cover. Our approach also maintains a competitive, credible set size with minimal computational overhead. Finally, our approach can be used to make predictions on real data and give valid credible regions modulo weight-initialization-based model mis-specification.

Trivedi, Shubhendu [Fermilab]↗

The DESI-Lensing Mock Challenge: large-scale cosmological analysis of 3x2-pt statistics

The current generation of large galaxy surveys will test the cosmological model by combining multiple types of observational probes. Realising the statistical promise of these new datasets requires rigorous attention to all aspects of analysis including cosmological measurements, modelling, covariance and parameter likelihood. In this paper we present the results of an end-to-end simulation study designed to test the analysis pipeline for the combination of the Dark Energy Spectroscopic Instrument (DESI) Year 1 galaxy redshift dataset and separate weak gravitational lensing information from the Kilo-Degree Survey, Dark Energy Survey and Hyper-Suprime-Cam Survey. Our analysis employs the 3x2-pt correlation functions including cosmic shear and galaxy-galaxy lensing, together with the projected correlation function of the spectroscopic DESI lenses. We build realistic simulations of these datasets including galaxy halo occupation distributions, photometric redshift errors, weights, multiplicative shear calibration biases and magnification. We calculate the analytical covariance of these correlation functions including the Gaussian, noise and super-sample contributions, and show that our covariance determination agrees with estimates based on the ensemble of simulations. We use a Bayesian inference platform to demonstrate that we can recover the fiducial cosmological parameters of the simulation within the statistical error margin of the experiment, investigating the sensitivity to scale cuts. This study is the first in a sequence of papers in which we present and validate the large-scale 3x2-pt cosmological analysis of DESI-Y1.

79 ASTRONOMY AND ASTROPHYSICS↗

The Dark Energy Survey supernova program: a reanalysis of cosmology results and evidence for evolving dark energy with an updated Type Ia supernova calibration

We present improved cosmological constraints from a re-analysis of the Dark Energy Survey (DES) 5-year sample of Type Ia supernovae (DES-SN5YR). This re-analysis includes an improved photometric cross-calibration, recent white dwarf observations to cross-calibrate between DES and low-redshift surveys, retraining the salt3 light-curve model and fixing a numerical approximation in the host-galaxy colour law. Our fully recalibrated sample, which we call DES-Dovekie, comprises ~1600 likely Type Ia SNe from DES and ~200 low-redshift SNe from other surveys. With DES-Dovekie, we obtain Ω m = 0.330 ± 0.015 in flat Lambda-cold dark matter (⁠ΛCDM) which changes Ω m by –0.022 compared to DES-SN5YR. Combining DES-Dovekie with cosmic microwave background data from Planck, Atacama Cosmology Telescope, and South Pole Telescope and the DESI DR2 measurements in a flat CDM cosmology, we find ω 0 = –0.803 ± 0.054 and ω a = –0.72 ± 0.21⁠. Our results hold a significance of 3.2σ, reduced from 4.2σ for DES-SN5YR, to reject the null hypothesis that the data are compatible with the cosmological constant. This significance is equivalent to a Bayesian model preference odds of approximately 5:1 in favour of the flat ω 0 ω a CDM model. Using generally accepted thresholds for model preference, our updated data exhibits only a weak preference for evolving dark energy.

dark energy↗

Evaluating Probabilistic Deep Learning Methods for Uncertainty Quantification of Precipitation Bias Correction

Climate models often exhibit biases in their precipitation predictions, particularly underestimating high-intensity events and overestimating low precipitation. Deep learning approaches offer promising solutions, but their epistemic uncertainty associated with a deep learning–based bias correction method has not previously been quantified for reliable downstream climate impact studies. While methods for capturing the epistemic uncertainty in deep learning frameworks exist, there is currently no consensus on the best method. In this work, we compare three uncertainty quantification (UQ) methods—Deep Ensembles (DEns), Monte Carlo Dropout (MCD), and Flipout—by assessing the reliability of their uncertainty estimates using standard measures such as sharpness and calibration. These UQ methods are applied to an existing deep learning precipitation bias correction model known as UFNet: a coupled U-Net and fully connected neural network. The methods utilized to assess the models’ uncertainties are 1) calibration, which ensures that the expected probabilities of the model align with reality and 2) sharpness, which is a measure of the precision of the model’s probabilistic predictions. Of the three UQ methods evaluated, the DEns and MCD methods demonstrated the best-calibrated performance (expected calibration error of 0.36 and 0.35, respectively), compared to Flipout (0.58). In contrast, Flipout had the sharpest predictions and the highest metric performance in bias correcting precipitation—especially for higher-order moments such as kurtosis with a spatial correlation of 72% compared to 32% and 55% spatial correlation for DEns and MCD, respectively. Of the three UQ methods, MCD was found to be the most suitable method for UQ purposes based on its calibration, sharpness, and computational requirements.

Bayesian methods↗

Bayesian Optimization of Catalysis with In-Context Learning

Large language models (LLMs) can perform accurate classification with zero or few examples through in-context learning (ICL), allowing the model to observe query-relevant examples at inference time and eliminating the need for additional weight updates to generalize beyond its original training data. We extend this capability to regression with uncertainty estimation using frozen LLMs (e.g., GPT-4o, Gemini), enabling Bayesian optimization (BO) in natural language without explicit model training or feature engineering. We apply this to materials discovery by representing materials as synthesis and testing procedures for use in natural language prompts. This Bayesian, design-first approach prioritizes optimization toward target material properties before detailed characterization, in contrast to conventional experimental workflows that often emphasize characterization of suboptimal materials. On benchmarks like aqueous solubility and oxidative coupling of methane (OCM), BO-ICL matches or outperforms Gaussian processes. In live experiments on the reverse water–gas shift (RWGS) reaction, BO-ICL identifies multimetallic catalysts that approach equilibrium CO yield within 6 and 10 iterations from a pool of 3,700 and 360,000 candidates, respectively. Our method redefines materials representation and accelerates discovery, with broad applications across catalysis, materials science, and AI.

Calibration↗