Search NASA⌕ Search

SEARCH · Search NASA

Results for “Bayesian sampling”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 181 records · Page 10

Bayesian chain graph models to characterize microbe-environment dynamics

Microbiome data require statistical models that can simultaneously decode microbes' reaction to the environment and interactions among microbes. While a multiresponse linear regression model seems like a straight-forward solution, we argue that treating it as a graphical model is problematic given that the regression coefficient matrix does not encode the conditional dependence structure between response and predictor nodes. This observation is especially important in biological settings when we have prior knowledge on the edges from specific experimental interventions that can only be properly encoded under a conditional dependence model. Here, we propose a chain graph model with two sets of nodes (predictors and responses) whose solution yields a graph with edges that indeed represent conditional dependence, thus agreeing with the experimenter's intuition on the average behavior of nodes under treatment. The solution to our model is sparse via the Bayesian linear regression (LASSO). In addition, we propose an adaptive extension so that different shrinkages can be applied to different edges to incorporate edge-specific prior knowledge. Our model is computationally inexpensive through an efficient Gibbs sampling algorithm and can account for binary, counting, and compositional responses via an appropriate hierarchical structure. We test the performance of our model in a variety of simulated datasets, thereby showing superior performance to state-of-the-art approaches. We further apply our model to human gut and soil microbial compositional datasets, and we highlight that CG-LASSO can estimate biologically meaningful network structures in the data.

compositional data↗

A generalized concept for cost-effective structural design

A generalized concept for cost-effective structural design is introduced. It is assumed that decisions affecting the cost effectiveness of aerospace structures fall into three basic categories: design, verification, and operation. Within these basic categories, certain decisions concerning items such as design configuration, safety factors, testing methods, and operational constraints are to be made. All or some of the variables affecting these decisions may be treated probabilistically. Bayesian statistical decision theory is used as the tool for determining the cost optimum decisions. A special case of the general problem is derived herein, and some very useful parametric curves are developed and applied to several sample structures.

Thomas, J. M.↗

Effects of Dose Error and Sample Size on Sonic Boom Dose-Response Curves

NASA will soon be collecting noise-annoyance community survey data as the X-59 aircraft flies supersonically over several communities in the USA. Sparse measurements of the X-59 sonic thumps will be used together with physics-based simulations to estimate noise doses at survey participant locations. These dose estimates have associated error that affects the accuracy of modeled dose-response curves, which can result in misestimation of annoyance. The precision in dose-response curves is also a consideration in selecting the number of survey participants. To enable pretest studies of dose error and precision, simulated dose-response data were generated based on NASA’s Quiet Supersonic Flights 2018 test. The data included various degrees of dose error and sample size. Frequentist multilevel logistic regression models were fit to the true and perturbed dose-response data. Simple proportional relationships were identified between the model parameters and the perturbation standard deviation. The summary dose-response curves illustrate the impact on accuracy if dose error is not accounted for in the model. The precision in the dose-response curves is also shown as the number of participants and degree of participation is varied. Finally, sampling variability is illustrated by showing the dose-response curves for several replicates with random draws of participants and errors.

X-59↗

Evaluation of Bayesian Sequential Proportion Estimation Using Analyst Labels

The author has identified the following significant results. A total of ten Large Area Crop Inventory Experiment Phase 3 blind sites and analyst-interpreter labels were used in a study to compare proportional estimates obtained by the Bayes sequential procedure with estimates obtained from simple random sampling and from Procedure 1. The analyst error rate using the Bayes technique was shown to be no greater than that for the simple random sampling. Also, the segment proportion estimates produced using this technique had smaller bias and mean squared errors than the estimates produced using either simple random sampling or Procedure 1.

Lennington, R. K.↗

Uranium particle age dating, aggregation, and model age best estimators

We present important aspects of uranium particle age dating by Large-Geometry Secondary Ion Mass Spectrometry (LG-SIMS) that can introduce bias and increase model age uncertainties, especially for small, young, and/or low-enriched particles. This metrology is important for applications related to International Nuclear Safeguards. We explore influential factors related to model age estimation, including the effects of evolving surface chemistry on inter-element measurements of particles (e.g., Th and U), detector background, and aggregation methods using simulated and actual particle samples. We introduce a new model age estimator, called “mid68”, that supplements 95% confidence intervals, providing a “best estimate” and uncertainty about the most likely age. The mid68 estimator can be calculated using the Feldman and Cousins method or Bayesian methods and provides a value with a symmetric uncertainty that can be used for calculations and approximate aggregation of processed model age values when the raw data and correction factors are not available. For particles yielding low 230 Th counts amidst nonzero detector background, their underlying model age probability distributions are asymmetric, so the mid68 estimator provides additional robust information regarding the underlying model age likelihood. This study provides a comprehensive and timely examination of critical aspects of uranium particle age dating as more laboratories establish particle chronometry capabilities.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Bayesian Statistical Models for Community Annoyance Survey Data

This paper demonstrates the use of two Bayesian statistical models to analyze single-event sonic boom exposure and human annoyance data from community response surveys. Each model is fit to data from a NASA pilot study.Unlike many community noise surveys, this study used a panel sample to collect multiple observations per participant instead of a single observation. Thus, a multilevel (also known as hierarchical or mixed-effects) model is used to account for the within-subject correlation in the panel sample data. This paper describes a multilevel logistic regression model and a multilevel ordinal regression model. The paper also proposes a method for calculating a summary dose-response curve from the multilevel models that represents the population. The two models’ summary dose-response curves are visually similar. However, their estimates differ when calculating the noise dose at a fixed percent highly annoyed.

Musical instruments↗

Selection function of clusters in Dark Energy Survey year 3 data from cross-matching with South Pole Telescope detections

Context. Galaxy clusters selected based on overdensities of galaxies in photometric surveys provide the largest cluster samples. However, modeling the selection function of such samples is complicated by noncluster members projected along the line of sight (projection effects) and the potential detection of unvirialized objects (contamination). Aims. We empirically constrained the magnitude of these effects by cross-matching galaxy clusters selected in the Dark Energy Survey data with the redMaPPer algorithm with significant detections in three South Pole Telescope surveys (SZ, pol-ECS, pol-500d). Methods. For matched clusters, we augmented the redMaPPer catalog with the SPT detection significance. For unmatched objects we used the SPT detection threshold as an upper limit on the SZe signature. Using a Bayesian population model applied to the collected multiwavelength data, we explored various physically motivated models to describe the relationship between observed richness and halo mass. Results. Our analysis reveals a clear preference for models with an additional skewed scatter component associated with projection effects over a purely log-normal scatter model. We rule out significant contamination by unvirialized objects at the high-richness end of the sample. While dedicated simulations offer a well-fitting calibration of projection effects, our findings suggest the presence of redshift-dependent trends that these simulations may not have captured. Our findings highlight that modeling the selection function of optically detected clusters remains a complicated challenge that requires a combination of simulation and data-driven approaches.

79 ASTRONOMY AND ASTROPHYSICS↗

Karhunen–Loève deep learning method for surrogate modeling and approximate Bayesian parameter estimation

We evaluate the performance of the Karhunen-Loève Deep Neural Network (KL-DNN) framework for surrogate modeling and approximate Bayesian parameter estimation in partial differential equation models. In the surrogate model, the Karhunen-Loève (KL) expansions are used for the dimensionality reduction of the number of unknown parameters and variables, and a deep neural network is employed to relate the reduced space of parameters to that of the state variables. The KL-DNN surrogate model is used to formulate a maximum-a-posteriori-like least-squares problem, which is randomized to draw samples of the posterior distribution of the parameters. We test the proposed framework for a hypothetical unconfined aquifer via comparison with the forward MODFLOW and inverse PEST++ iterative ensemble smoother (IES) solutions as well as the state-of-the-art Fourier neural operator (FNO) and deep operator networks (DeepONets) operator learning surrogate models. Our results show that the KL-DNN surrogate model outperforms FNO and DeepONet for forward predictions. For solving inverse problems, the randomized algorithm provides the same or more accurate Bayesian predictions of the parameters than IES as evidenced by the higher log-predictive probability of both the estimated parameter field and the forecast hydraulic head. The posterior mean obtained from the randomized algorithm is closer to the reference parameter field than that obtained with FNO as the maximum a posteriori estimate.

Approximate Bayesian inference↗

Analysis of differential scanning calorimetry data for aged plutonium

Differential scanning calorimetry data for samples of a 52 year old plutonium alloy with 3.3 at. % Ga that were heated beyond the melting point is analyzed using transition state theory to find activation energies for the δ to ε and ε to liquid phase transitions. A Bayesian statistical method involving a Gaussian process model is used to find mean values and confidence intervals for the activation energies. The activation energy for the δ to ε phase transition increases by 3.3 ± 3.8% per decade, relative to the case when all age related plutonium lattice point defects have been removed through annealing. The corresponding increase in activation energy for the ε to liquid transition is shown to be 7.1 ± 1.8% per decade. It is postulated that the change in activation energy with age for both phase transitions is caused, in part, by the accumulation of the same type of lattice point defects associated with the observed increase in elastic bulk modulus over time.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Designing an Optimal Sensor Network via Minimizing Information Loss

Optimal experimental design is a classic topic in statistics, with many well-studied problems, applications, and solutions. The design problem we study is the placement of sensors to monitor spatiotemporal processes, explicitly accounting for the temporal dimension in our modeling and optimization. We observe that recent advancements in computational sciences often yield large datasets based on physics-based simulations, which are rarely leveraged in experimental design. We introduce a novel model-based sensor placement criterion, along with a highly-efficient optimization algorithm, which integrates physics-based simulations and Bayesian experimental design principles to identify sensor networks that “minimize information loss” from simulated data. Our technique relies on sparse variational inference and (separable) Gauss-Markov priors, and thus may adapt many techniques from Bayesian experimental design. We validate our method through a case study monitoring air temperature in Phoenix, Arizona, using state-of-the-art physics-based simulations. Our results show our framework to be superior to random or quasi-random sampling, particularly with a limited number of sensors. We conclude by discussing practical considerations and implications of our framework, including more complex modeling tools and real-world deployments.

54 ENVIRONMENTAL SCIENCES↗

ELG×LRG Distribution through Dark Matter Halo Dynamics

We investigate the clustering and halo occupation distribution (HOD) of DESI Y1 emission-line (ELGs) and luminous red (LRGs) galaxies at 0.8 < z < 1.1, including their cross-correlation (ELG×LRG), using the A BACUS S UMMIT suite and a new Halo Occupation Model (H OME ) for galaxy multitracers. This integrates intrahalo dynamics, halo exclusion, and quenching, bridging insights from hydrodynamical, HOD, abundance-matching, and semianalytic studies. Leveraging full phase-space information from the Uchuu N-body simulation, and sampling satellites from dark-matter particle positions via physically motivated prescriptions, Home reproduces the anisotropic clustering down to s = 200 h −1 kpc with unprecedented accuracy. Model parameters are inferred solely from two-point statistics using a two-level Bayesian framework, yielding high-fidelity ELG, LRG, and cross-reference catalogs. We find that satellite ELGs behave as incoherent flows within their parent halos, dominating the clustering below 4 h −1 Mpc. The HOD from the best-fit Home has the following properties: (i) 90.50% (85.91%) of ELGs (LRGs) are central galaxies without satellites, residing in halos of M vir ∼ 6.6 × 10 11 (1.2 × 10 13 ) h −1 M ⊙ ; (ii) the ELG×LRG cross-correlation is governed by central-central pairs and shaped by halo exclusion on 2–5 h −1 Mpc scales; (iii) 9.50% (14.09%) of ELGs (LRGs) are satellites, of which 1.09% (3.52%) inhabit halos with a central galaxy of the same species in a maximally conformal configuration, 7.02% (0.005%) orbit complementary hosts in a minimally conformal state, and 0.58% (10.57%) are orphans. The high sensitivity of Home precisely captures the dynamics of satellites in different host environments, opening a promising avenue for understanding systematics and the dynamical nature of dark matter, potentially distinguishing gravity models.

Favole, Ginevra [Universidad de La Laguna (Spain);↗

Taylor approximation variance reduction for approximation errors in PDE-constrained Bayesian inverse problems

In numerous applications, surrogate models are used as a replacement for accurate parameter-to-observable mappings when solving large-scale inverse problems governed by partial differential equations (PDEs). The surrogate model may be a computationally cheaper alternative to the accurate parameter-to-observable mappings and/or may ignore additional unknowns or sources of uncertainty. The Bayesian approximation error (BAE) approach provides a means to account for the induced uncertainties and approximation errors, i.e. the errors between the accurate parameter-to-observable mapping and the surrogate. The statistics of these errors are, however, in general unknown a priori, and are thus calculated using Monte Carlo sampling. Although the sampling is typically carried out offline, i.e. before considering the data, the process can still represent a computational bottleneck. In this work, we develop a scalable computational approach for reducing the costs associated with the sampling stage of the BAE approach. Specifically, we consider the Taylor expansion of the accurate and surrogate forward models with respect to the uncertain parameter fields either as a control variate for variance reduction or as a means to directly and efficiently approximate the mean and covariance of the approximation errors. We propose efficient methods for evaluating the expressions for the mean and covariance of the Taylor approximations based on linear(-ized) PDE solves. Furthermore, the proposed approach is independent of the dimension of the uncertain parameter, depending instead on the intrinsic dimension of the data, ensuring scalability to high-dimensional problems. The potential benefits of the proposed approach are demonstrated for two high-dimensional inverse problems governed by PDE examples, namely for the estimation of a distributed Robin boundary coefficient in a linear diffusion problem, and for a coefficient estimation problem governed by a nonlinear diffusion problem.

Bayesian approximation error↗

New giant planet beyond the snow line for an extended MOA exoplanet microlens sample

Characterizing a planet detected by microlensing is hard if the planetary signal is weak or the lens-source relative trajectory is far from caustics. However, statistical analyses of planet demography must include those planets to accurately determine occurrence rates. As part of a systematic modelling effort in the context of a >10-yr retrospective analysis of MOA’s survey observations to build an extended MOA statistical sample, we analyse the light curve of the planetary microlensing event MOA-2014-BLG-472. This event provides weak constraints on the physical parameters of the lens, as a result of a planetary anomaly occurring at low magnification in the light curve. We use a Bayesian analysis to estimate the properties of the planet, based on a refined Galactic model and the assumption that all Milky Way’s stars have an equal planet-hosting probability. We find that a lens consisting of a 1.9(+2.2,−1.2)M(J) giant planet orbiting a 0.31(+0.36,−0.19)Mꙩ host at a projected separation of 0.75±0.24au is consistent with the observations and is most likely, based on the Galactic priors. The lens most probably lies in the Galactic bulge, at 7.2(+0.6,−1.7)kpc from Earth. The accurate measurement of the measured planet-to-host star mass ratio will be included in the next statistical analysis of cold planet demography detected by microlensing.

Clément Ranc↗

Statistical data analysis of x-ray spectroscopy data enabled by neural network accelerated Bayesian inference

Bayesian inference applied to x-ray spectroscopy data analysis enables uncertainty quantification necessary to rigorously test theoretical models. However, when comparing to data, detailed atomic physics and radiation transfer calculations of x-ray emission from non-uniform plasma conditions are typically too slow to be performed in line with statistical sampling methods, such as Markov Chain Monte Carlo sampling. Furthermore, differences in transition energies and x-ray opacities often make direct comparisons between simulated and measured spectra unreliable. Here, we present a spectral decomposition method that allows for corrections to line positions and bound–bound opacities to best fit experimental data, with the goal of providing quantitative feedback to improve the underlying theoretical models and guide future experiments. In this work, we use a neural network (NN) surrogate model to replace spectral calculations of isobaric hot-spots created in Kr-doped implosions at the National Ignition Facility. The NN was trained on calculations of x-ray spectra using an isobaric hot-spot model post-processed with Cretin, a multi-species atomic kinetics and radiation code. The speedup provided by the NN model to generate x-ray emission spectra enables statistical analysis of parameterized models with sufficient detail to accurately represent the physical system and extract the plasma parameters of interest.

47 OTHER INSTRUMENTATION↗

A copula-based rank histogram ensemble filter

Serial ensemble filters implement triangular probability transport maps to reduce high-dimensional inference problems to sequences of state-by-state univariate inference problems. The univariate inference problems are solved by sampling posterior probability densities obtained by combining constructed prior densities with observational likelihoods according to Bayes' rule. Many serial filters in the literature focus on representing the marginal posterior densities of each state. However, rigorously capturing the conditional dependencies between the different univariate inferences is crucial to correctly sampling multidimensional posteriors. This work proposes a new serial ensemble filter, called the copula rank histogram filter (CoRHF), that seeks to capture the conditional dependency structure between variables via empirical copula estimates; these estimates are used to rigorously implement the triangular (state-by-state univariate) Bayesian inference. The success of the CoRHF is demonstrated on two-dimensional examples and the Lorenz'63 problem. A practical extension to the high-dimensional setting is developed by localizing the empirical copula estimation, and is demonstrated on the Lorenz'96 problem.

97 MATHEMATICS AND COMPUTING↗

Investigating Kinetic Mechanisms of Soot Formation in Plasma Pyrolysis of Methane via Active Learning (Final Technical Report)

Plasma pyrolysis of methane is an effective route for zero-carbon hydrogen production. Yet, soot generated from pyrolysis of hydrocarbons is detrimental to the climate and human health. There is ample experimental and theoretical evidence that suggests polycyclic aromatic hydrocarbons (PAHs) are the molecular precursors to soot particles. The reaction pathways of PAH formation are intricately dependent on a multitude of process parameters, whose kinetic mechanisms are not well-understood in plasma pyrolysis. This project aims to leverage advances in the kinetic modeling of soot formation in combustion, as well as in surrogate modeling and active learning, to systematically investigate the effects of process parameter on the kinetics of PAH formation in plasma pyrolysis of methane. To this end, we propose to use the PAH formation kinetics model developed by the PPPL/PU group based on the well-established ABF and HACA mechanisms, coupled with low-temperature plasma models. We will develop an active learning (AL) framework based on Bayesian optimization to systematically and data-efficiently explore the complex and multivariable parameter space of plasma pyrolysis in order to quantify the effects of plasma and feed parameters on the ABF and HACA kinetic pathways. AL is the branch of machine learning concerned with systematically querying samples from a system (experimental or computational) to train a data-driven model that maps design parameters to a performance criterion. We will use the data generated via AL to perform global sensitivity analysis, combined with uncertainty quantification, to elucidate the impact of different reaction pathways on minimizing formation of soot precursors. This study will result in an improved understanding of kinetics of PAH formation in plasma pyrolysis and can pave the way for more advanced mechanistic studies (e.g., soot nucleation mechanisms). Additionally, the findings will be useful for establishing practical strategies for increasing the pyrolysis efficiency and producing high-grade carbon for synthesis of nanomaterials.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

A Practical Philosophy of Complex Climate Modelling

We give an overview of the practice of developing and using complex climate models, as seen from experiences in a major climate modelling center and through participation in the Coupled Model Intercomparison Project (CMIP).We discuss the construction and calibration of models; their evaluation, especially through use of out-of-sample tests; and their exploitation in multi-model ensembles to identify biases and make predictions. We stress that adequacy or utility of climate models is best assessed via their skill against more naive predictions. The framework we use for making inferences about reality using simulations is naturally Bayesian (in an informal sense), and has many points of contact with more familiar examples of scientific epistemology. While the use of complex simulations in science is a development that changes much in how science is done in practice, we argue that the concepts being applied fit very much into traditional practices of the scientific method, albeit those more often associated with laboratory work.

complex simulation↗

The Rise and Fall of Star Formation Histories of Blue Galaxies at Redshifts 0.2 < z < 1.4

Popular cosmological scenarios predict that galaxies form hierarchically from the merger of many progenitor, each with their own unique star formation history (SFH). We use the approach recently developed by Pacifici et al. to constrain the SFHs of 4517 blue (presumably star-forming) galaxies with spectroscopic redshifts in the range O.2 < z < 1:4 from the All-Wavelength Extended Groth Strip International Survey (AEGIS). This consists in the Bayesian analysis of the observed galaxy spectral ' energy distributions with a comprehensive library of synthetic spectra assembled using state-of-the-art models of star formation and chemical enrichment histories, stellar population synthesis, nebular emission and attenuation by dust. We constrain the SFH of each galaxy in our sample by comparing the observed fluxes in the B, R,l and K(sub s) bands and rest-frame optical emission-line luminosities with those of one million model spectral energy distributions. We explore the dependence of the resulting SFH on galaxy stellar mass and redshift. We find that the average SFHs of high-mass galaxies rise and fall in a roughly symmetric bell-shaped manner, while those of low-mass galaxies rise progressively in time, consistent with the typically stronger activity of star formation in low-mass compared to high-mass galaxies. For galaxies of all masses, the star formation activity rises more rapidly at high than at low redshift. These findings imply that the standard approximation of exponentially declining SFHs wIdely used to interpret observed galaxy spectral energy distributions is not appropriate to constrain the physical parameters of star-forming galaxies at intermediate redshifts.

Pacifici, Camilla↗