Search NASA⌕ Search

SEARCH · Search NASA

Results for “Bayesian sampling”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 55 records · Page 3

Mars Sample Return: Risk Management & Sample Safety Assessment

Returning samples from Mars has long been a major planetary science objective due to the high scientific value and transformative potential of the resulting data. An exciting dimension of this objective is the potential for the detection of ancient microbiological life, and the possibility of improving our understanding of the evolution of habitable environments on Mars and the development of life on Earth. To ensure that returned samples meet stringent planetary protection requirements and do not expose Earth to potential biohazards, the joint NASA/ESA Sample Receiving Project (SRP) assembled the Sample Safety Assessment Protocol Tiger Team (SSAP-TT). Members were recruited with the specific goal of creating a multi-disciplinary and internationally distributed team of experts in their respective fields across the federal government, academia, and private industry. This team was chartered with reassessing previous sample safety assessment strategies, defining what constitutes a biological hazard, developing a protocol to test for potential biohazards, and establishing a statistical framework to determine if samples are “safe” for release. The team developed a three-step protocol, supported by a Bayesian statistical framework, to assess whether returned samples contain potential biohazards that could present a risk to Earth’s biosphere. Initial conclusions indicated that an effective and comprehensive safety assessment protocol is feasible using modern techniques and does not require an excessive amount of sample consumption or traditional microbiological detection methodology. Herein, we will present an overview of the MSR SRP, the proposed safety assessment protocol, and how aspects of this novel approach can be applied to biological assessment in healthcare product manufacturing practices.

Alvin L Smith↗

Hierarchical Gaussian Random Field Sampling for Multilevel Markov Chain Monte Carlo: Coupling Stochastic Partial Differential Equation and the Karhunen–Loève Decomposition

This work introduces structure preserving hierarchical decompositions for sampling Gaussian random fields (GRFs) within the context of multilevel Bayesian inference in high-dimensional space. Existing scalable hierarchical sampling methods, such as those based on stochastic partial differential equations (SPDEs), often reduce the dimensionality of the sample space at the cost of accuracy of inference. Other approaches, such that those based on Karhunen-Loève (KL) expansions, offer sample space dimensionality reduction but sacrifice GRF representation accuracy and ergodicity of the Markov chain Monte Carlo (MCMC) sampler and are computationally expensive for high-dimensional problems. The proposed method integrates the dimensionality reduction capabilities of KL expansions with the scalability of SPDE-based sampling, thereby providing a robust, unified framework for high-dimensional uncertainty quantification (UQ) that is scalable and accurate, preserves ergodicity, and offers dimensionality reduction of the sample space. The hierarchy in our multilevel algorithm is derived from the geometric multigrid hierarchy. By constructing a hierarchical decomposition that maintains the covariance structure across the levels in the hierarchy, the approach enables efficient coarse-to-fine sampling while ensuring that all samples are drawn from the desired distribution. The effectiveness of the proposed method is demonstrated on a benchmark subsurface flow problem, demonstrating its effectiveness in improving computational efficiency and statistical accuracy. Furthermore, our proposed technique is more efficient and accurate and displays better convergence properties than existing methods for high-dimensional Bayesian inference problems.

Gaussian random fields↗

Ensemble variational Fokker-Planck methods for data assimilation

Particle flow filters solve Bayesian inference problems by smoothly transforming a set of particles into samples from the posterior distribution. Particles move in state space under the flow of an McKean-Vlasov-Itˆo process. This work introduces the Variational Fokker-Planck (VFP) framework for data assimilation, a general approach that includes previously known particle flow filters as special cases. The McKean-Vlasov-Itˆo process that transforms particles is defined via an optimal drift that depends on the selected diffusion term. It is established that the underlying probability density - sampled by the ensemble of particles - converges to the Bayesian posterior probability density. For a finite number of particles the optimal drift contains a regularization term that nudges particles toward becoming independent random variables. Based on this analysis, we derive computationally-feasible approximate regularization approaches that penalize the mutual information between pairs of particles, and avoid particle collapse. Moreover, the diffusion plays a role akin to a particle rejuvenation approach that aims to alleviate particle collapse. The VFP framework is very flexible. Different assumptions on prior and intermediate probability distributions can be used to implement the optimal drift, and localization and covariance shrinkage can be applied to alleviate the curse of dimensionality. A robust implicit-explicit method is discussed for the efficient integration of stiff McKean- Vlasov-Itˆo processes. Here, the effectiveness of the VFP framework is demonstrated on three progressively more challenging test problems, namely the Lorenz ’63, Lorenz ’96 and the quasi-geostrophic equations.

97 MATHEMATICS AND COMPUTING↗

Resolving Mixtures of Soot Characterized by SP-AMS Spectra Using a Latent Dirichlet Allocation Model

Soot produced by detonation or combustion events exhibits different chemical properties depending on the fuel, device construction, and environmental conditions in which the event occurs. These properties can be useful for defining relevant signatures for probabilistically identifying the different types of events that occurred, based on the soot that is produced from these events. However, it is rare to observe samples of soot from a detonation or combustion that are not contaminated by outside particles. In this paper, we present a method for resolving mixtures of soot to determine the contributions of sources that may be present in samples of recovered soot. We use Latent Dirichlet Allocation to describe the generative process for a sample of recovered soot, and use Variational Bayesian Inference to learn about the parameters associated with the generative model. We demonstrate the utility of this method by considering real samples of mixtures of soot under various frameworks to show that the model is able to identify the different components present in a sample of soot as well as their mixing proportions.

54 ENVIRONMENTAL SCIENCES↗

Probabilistic Calibration of Expensive Models using Efficiently Trained Surrogates

Calibration of computational models in the presence of uncertainty is often cast as a Bayesian inference problem and solved via sampling methods, e.g., Markov chain Monte Carlo. When the computational model is expensive, this task becomes intractable due to the large number of samples required to accurately estimate the posterior distribution of the calibration parameters. A popular solution to this problem is to use machine learning to develop a faster-to-evaluate, lower-fidelity substitute for the original model to serve as a surrogate while solving the inference problem. Although considered an offline cost, generating training data to construct this surrogate model can still be an expensive task in practice. An active learning algorithm is presented that focuses training on improving surrogate accuracy specifically in and around the bulk of the posterior distribution, as this is where the model is exercised during calibration. Candidate samples are drawn from families of distributions related to an approximation of the posterior. The sample maximizing predictive variance is then selected for evaluation by the original computational model, yielding a label for the training point. Iterating this approach increases efficiency relative to space filling designs (e.g., Latin hypercube sampling) by avoiding low probability points. Practical considerations are discussed, including the benefits of using a sequential Monte Carlo sampling approach, convergence heuristics, and the importance of both exploration and exploitation given that the true posterior is unknown a priori.

uncertainty quantification↗

Dimensional Reduction for Sampled Priors and Application to Photometric Redshift Distributions

A typical Bayesian inference on the values of some parameters of interest q from some data D involves running a Markov Chain (MC) to sample from the posterior $p$($q$,$n$|$D$) $\propto$ $\mathcal{L}$($D$|$q$,$n$)$p$(q)$p$($n$), where n are some nuisance parameters with a separable prior. In some cases, the nuisance parameters are high-dimensional, and their prior p(n) is itself defined only by a set of samples that have been drawn from some other MC. The MC for the posterior will typically require evaluation of p(n) at arbitrary values of n, i.e., one needs to provide a density estimator over the full n space from the provided samples. But the high dimensionality of n hinders both the density estimation and the efficiency of the MC for the posterior. We describe a solution to this problem: a linear compression of the n space into a much lower-dimensional space u, which projects away directions in n space that cannot appreciably alter $\mathcal{L}$. The algorithm for doing so is a slight modification to principal components analysis, and is less restrictive on p(n) than other proposed solutions to this issue. We demonstrate this “mode projection” technique using the analysis of 2-point correlation functions of weak lensing fields and galaxy density in the Dark Energy Survey, where n is a binned representation of the redshift distribution n(z) of the galaxies.

79 ASTRONOMY AND ASTROPHYSICS↗

Consequential improvement acquisition function for efficient multi-fidelity Bayesian optimization

Abstract Surrogate-based Bayesian optimization has been widely applied in design optimization to increase sampling efficiency. However, the cost for each evaluation of the objective function can still be very high when physical experiments or large-scale simulations are involved. Multi-fidelity Bayesian optimization is the new approach to further improve the sampling efficiency by reducing the number of expensive samples at the highest fidelity level and supplementing them with less expensive ones at low-fidelity levels. In this paper, a new consequential improvement (CI) acquisition function is proposed to allow for the simultaneous selection of the solution and the fidelity level in problems with a known hierarchy of fidelity levels. The new CI acquisition function incorporates the consequential effectiveness of objective improvement with the considerations of cost, accuracy, and validity differences between high- and low-fidelity samples in engineering practice. The new method of multi-fidelity Bayesian optimization based on the CI is demonstrated with several analytical and simulation-based design examples. In the simulation-based design optimization example, the results show that the CI acquisition function has a decisive advantage in the sampling efficiency over the other methods of multi-fidelity Bayesian optimization with simultaneous selection. The results indicate that the proposed method is particularly advantageous in solving high-dimensional problems and when large cost ratios between high- and low-fidelity evaluations exist and high-fidelity validation is mandatory. Furthermore, the method robustly avoids the prevalent issue of over sampling at low-fidelity levels.

Aydogdu, Ibrahim [Georgia Institute of Technology,↗

The SRG/eROSITA All-Sky Survey: Dark Energy Survey year 3 weak gravitational lensing by eRASS1 selected galaxy clusters

Context. Number counts of galaxy clusters across redshift are a powerful cosmological probe if a precise and accurate reconstruction of the underlying mass distribution is performed – a challenge called mass calibration. With the advent of wide and deep photometric surveys, weak gravitational lensing (WL) by clusters has become the method of choice for this measurement. Aims. We measured and validated the WL signature in the shape of galaxies observed in the first three years of the Dark Energy Survey (DES Y3) caused by galaxy clusters and groups selected in the first all-sky survey performed by SRG (Spectrum Roentgen Gamma)/eROSITA (eRASS1). These data were then used to determine the scaling between the X-ray photon count rate of the clusters and their halo mass and redshift. Methods. We empirically determined the degree of cluster member contamination in our background source sample. The individual cluster shear profiles were then analyzed with a Bayesian population model that self-consistently accounts for the lens sample selection and contamination and includes marginalization over a host of instrumental and astrophysical systematics. To quantify the accuracy of the mass extraction of that model, we performed mass measurements on mock cluster catalogs with realistic synthetic shear profiles. This allowed us to establish that hydrodynamical modeling uncertainties at low lens redshifts (z < 0.6) are the dominant systematic limitation. At high lens redshift, the uncertainties of the sources’ photometric redshift calibration dominate. Results. With regard to the X-ray count rate to halo mass relation, we determined its amplitude, its mass trend, the redshift evolution of the mass trend, the deviation from self-similar redshift evolution, and the intrinsic scatter around this relation. Conclusions. The mass calibration analysis performed here sets the stage for a joint analysis with the number counts of eRASS1 clusters to constrain a host of cosmological parameters. We demonstrate that WL mass calibration of galaxy clusters can be performed successfully with source galaxies whose calibration was performed primarily for cosmic shear experiments, opening the way for the cluster cosmological exploitation of future optical and NIR surveys like Euclid and LSST.

79 ASTRONOMY AND ASTROPHYSICS↗

Active Learning‐Driven Inkless Additive Nanomanufacturing for Printed Electronics

Inkless additive nanomanufacturing for printed electronics promises broad material and substrate versatility, yet the high-dimensional print parameter space makes tuning print parameters time-intensive. We present a Bayesian optimization study that constructs a digital twin from printed-silver data to benchmark surrogate models, acquisition functions, and batch sizes head-to-head to achieve user-specified target resistance. Tested surrogate models included Gaussian process, random forest, and Bayesian neural network surrogates with expected improvement and confidence bound acquisition functions. In total, we evaluate 48 unique model configurations alongside a random sampling baseline for comparison. For printed silver, the Bayesian neural network with a batch size of one achieved the lowest average cumulative regret, approximately four times more efficient on average than random sampling. To balance performance and substrate space, a random forest model with expected improvement and a batch size of four was chosen as the model for validation testing. Applying this chosen configuration to copper with an additional print parameter, the model achieved a resistance within 0.15 Ω of a 1 Ω target in fewer than 30 printed lines across five validation sets. Altogether, the workflow yields a tuned and validated model that efficiently guides experiments toward the target while simultaneously learning the parameter space.

Bevel, Colton [Auburn University, AL (United State↗

Probabilistic Damage Characterization Using the Computationally-Efficient Bayesian Approach

This work presents a computationally-ecient approach for damage determination that quanti es uncertainty in the provided diagnosis. Given strain sensor data that are polluted with measurement errors, Bayesian inference is used to estimate the location, size, and orientation of damage. This approach uses Bayes' Theorem to combine any prior knowledge an analyst may have about the nature of the damage with information provided implicitly by the strain sensor data to form a posterior probability distribution over possible damage states. The unknown damage parameters are then estimated based on samples drawn numerically from this distribution using a Markov Chain Monte Carlo (MCMC) sampling algorithm. Several modi cations are made to the traditional Bayesian inference approach to provide signi cant computational speedup. First, an ecient surrogate model is constructed using sparse grid interpolation to replace a costly nite element model that must otherwise be evaluated for each sample drawn with MCMC. Next, the standard Bayesian posterior distribution is modi ed using a weighted likelihood formulation, which is shown to improve the convergence of the sampling process. Finally, a robust MCMC algorithm, Delayed Rejection Adaptive Metropolis (DRAM), is adopted to sample the probability distribution more eciently. Numerical examples demonstrate that the proposed framework e ectively provides damage estimates with uncertainty quanti cation and can yield orders of magnitude speedup over standard Bayesian approaches.

Warner, James E.↗

Probing Non-Standard Interactions in NOvA with Bayesian MCMC Methods

We present a search for non-standard neutrino interactions (NSI) using NOvA's joint $\nu_e$ appearance and $\nu_\mu$ disappearance samples in both neutrino and antineutrino beam modes. A Bayesian Markov Chain Monte Carlo approach is used to map the posterior distribution over NSI parameters $\varepsilon_{e\mu}$, $\varepsilon_{e\tau}$, and $\varepsilon_{\mu\tau}$, simultaneously with standard oscillation parameters. We investigate the effect of prior choice for the complex NSI parameters and present projected sensitivity to off-diagonal NSI.

Huang, Xiaoyan [Mississippi U.]↗

Reduced-dimension Bayesian optimization for model calibration of transient vapor compression cycles

Development and calibration of first-principles dynamic models of vapor compression cycles (VCCs) is of critical importance for applications that include control design and fault detection and diagnostics. Nevertheless, the inherent complexity of models that are represented by large systems of differential–algebraic equations leads to significant challenges for model calibration processes that utilize classical gradient-based methods. Bayesian optimization (BO) is a sample-efficient and gradient-free approach using a probabilistic surrogate model and optimal search over a feasible parameter space. Despite the benefits of BO in reducing computational costs, challenges remain in dealing with a high-dimensional calibration task resulting from a large set of parameters that have significant impacts on system behavior and need to be calibrated simultaneously. This paper presents a reduced-dimension BO framework for calibrating transient VCCs models where the calibration space is projected to a low-dimensional subspace for accelerating convergence of the solution algorithm and consequently reducing the number of transient simulations. The proposed approach was demonstrated via two case studies associated with different VCC applications where 10 parameters were calibrated in each case using laboratory measurements. The reduced-dimension BO framework only required 1 / 8 th of the iterations associated with a standard BO method that deals with high-dimensional calibration parameters for converged solutions and yielded comparable accuracy. Furthermore, both calibrated models revealed significant accuracy improvements compared to uncalibrated models.

Ma, Jiacheng↗

Heterogeneity in Short Gamma-Ray Bursts

We analyze the Swift/BAT sample of short gamma-ray bursts, using an objective Bayesian Block procedure to extract temporal descriptors of the bursts' initial pulse complexes (IPCs). The sample comprises 12 and 41 bursts with and without extended emission (EE) components, respectively. IPCs of non-EE bursts are dominated by single pulse structures, while EE bursts tend to have two or more pulse structures. The medians of characteristic timescales - durations, pulse structure widths, and peak intervals - for EE bursts are factors of approx 2-3 longer than for non-EE bursts. A trend previously reported by Hakkila and colleagues unifying long and short bursts - the anti-correlation of pulse intensity and width - continues in the two short burst groups, with non-EE bursts extending to more intense, narrower pulses. In addition we find that preceding and succeeding pulse intensities are anti-correlated with pulse interval. We also examine the short burst X-ray afterglows as observed by the Swift/XRT. The median flux of the initial XRT detections for EE bursts (approx 6 X 10(exp -10) erg / sq cm/ s) is approx > 20 x brighter than for non-EE bursts, and the median X-ray afterglow duration for EE bursts (approx 60,000 s) is approx 30 x longer than for non-EE bursts. The tendency for EE bursts toward longer prompt-emission timescales and higher initial X-ray afterglow fluxes implies larger energy injections powering the afterglows. The longer-lasting X-ray afterglows of EE bursts may suggest that a significant fraction explode into more dense environments than non-EE bursts, or that the sometimes-dominant EE component efficiently p()wers the afterglow. Combined, these results favor different progenitors for EE and non-EE short bursts.

Norris, Jay P.↗

Efficient Neural Network Approaches for Conditional Optimal Transport with Applications in Bayesian Inference

In this work, we present two neural network approaches that approximate the solutions of static and dynamic conditional optimal transport (COT) problems. Both approaches enable conditional sampling and conditional density estimation, which are core tasks in Bayesian inference—particularly in the simulation-based (“likelihood-free”) setting. Our methods represent the target conditional distribution as a transformation of a tractable reference distribution. Obtaining such a transformation, chosen here to be an approximation of the COT map, is computationally challenging even in moderate dimensions. To improve scalability, our numerical algorithms use neural networks to parameterize candidate maps and further exploit the structure of the COT problem. Our static approach approximates the map as the gradient of a partially input convex neural network. It uses a novel numerical implementation to increase computational efficiency compared to state-of-the-art alternatives. Our dynamic approach approximates the conditional optimal transport via the flow map of a regularized neural ODE; compared to the static approach, it is slower to train but offers more modeling choices and can lead to faster sampling. We demonstrate both algorithms numerically, comparing them with competing state-of-the-art approaches, using benchmark datasets and simulation-based Bayesian inverse problems.

97 MATHEMATICS AND COMPUTING↗

A Computationally-Efficient Inverse Approach to Probabilistic Strain-Based Damage Diagnosis

This work presents a computationally-efficient inverse approach to probabilistic damage diagnosis. Given strain data at a limited number of measurement locations, Bayesian inference and Markov Chain Monte Carlo (MCMC) sampling are used to estimate probability distributions of the unknown location, size, and orientation of damage. Substantial computational speedup is obtained by replacing a three-dimensional finite element (FE) model with an efficient surrogate model. The approach is experimentally validated on cracked test specimens where full field strains are determined using digital image correlation (DIC). Access to full field DIC data allows for testing of different hypothetical sensor arrangements, facilitating the study of strain-based diagnosis effectiveness as the distance between damage and measurement locations increases. The ability of the framework to effectively perform both probabilistic damage localization and characterization in cracked plates is demonstrated and the impact of measurement location on uncertainty in the predictions is shown. Furthermore, the analysis time to produce these predictions is orders of magnitude less than a baseline Bayesian approach with the FE method by utilizing surrogate modeling and effective numerical sampling approaches.

Warner, James E.↗

Autonomous organic synthesis for redox flow batteries via flexible batch Bayesian optimization

Traditional trial-and-error methods for materials discovery are inefficient to meet the urgent demands posed by the rapid progression of climate change. This urgency has driven the increasing interest in integrating robotics and machine learning into materials research to accelerate experimental learning. However, idealized decision-making frameworks to achieve maximum sampling efficiency are not always compatible with high-throughput experimental workflows inside a laboratory. For multi-step chemical processes, differences in hardware capacities can complicate the digital framework by introducing constraints on the maximum number of samples in each step of the experiment, hence causing varying batch sizes in variable selection within the same batch. Therefore, designing flexible sampling algorithms is necessary to accommodate the multi-step synthesis with practical constraints unique to each high-throughput workflow. In this work, we designed and employed three strategies on a high-throughput robotic platform to optimize the sulfonation reaction of redox-active molecules used in flow batteries. Our strategies adapt to the multi-step experimental workflow, where their formulation and heating steps are separate, causing varying batch size requirements. By strategically sampling using clustering and mixed-variable batch Bayesian optimization, we were able to iteratively identify optimal conditions that maximize the yields. Our work presents a flexible approach that allows tailoring the machine learning decision-making to suit the practical constraints in individual high-throughput experimental platforms, followed by performing resource-efficient yield optimization using available open-source Python libraries.

Tamura, Clara [Univ. of Washington, Seattle, WA (U↗

Data Assimilation for Robust UQ Within Agent-Based Simulation on HPC Systems

Agent-based simulation provides a powerful tool for in silico system modeling. However, these simulations do not provide built-in methods for uncertainty quantification (UQ). Within these types of models a typical approach to UQ is to run multiple realizations of the model then compute aggregate statistics. This approach is limited due to the compute time required for a solution. When faced with an emerging biothreat, public health decisions need to be made quickly and solutions for integrating near real-time data with analytic tools are needed. We propose an integrated Bayesian UQ framework for agent-based models based on sequential Monte Carlo sampling. Given streaming or static data about the evolution of an emerging pathogen this Bayesian framework provides a distribution over the parameters governing the spread of a disease through a population. These estimates of the spread of a disease may be provided to public health agencies seeking to abate the spread. By coupling agent-based simulations with Bayesian modeling in a data assimilation, our proposed framework provides a powerful tool for modeling dynamical systems in silico. We propose a method which reduces model error and provides a range of realistic possible outcomes. Moreover, our method addresses two primary limitations of ABMs: the lack of UQ and an inability to assimilate data. Our proposed framework combines the flexibility of an agent-based model with UQ provided by the Bayesian paradigm in a workflow which scales well to HPC systems. We provide algorithmic details and results on a simulated outbreak with both static and streaming data.

Spannaus, Adam [ORNL] (ORCID:0000000225213657)↗