Search NASA⌕ Search

SEARCH · Search NASA

Results for “Bayesian nonparametric models”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

Exploration with Scalable Gaussian Process Reinforcement Learning

Exploration is a challenging problem in reinforcement learning (RL), especially in environments with sparse rewards. Quantifying and utilizing the parametric uncertainty has been shown to be paramount for successful exploration [Osband et al., 2018]. Bayesian, or approximately Bayesian, methods present a principled means of estimating the parametric uncertainty in RL problems. Gaussian processes, nonparametric Bayesian models, are often impractical due to poor scalability and computational bottlenecks. We introduce a scalable Gaussian process RL (GPRL) method which directly induces sparsity in the covariance matrix to facilitate faster computation. This is a departure from previous GPRL methods which instead rely on data reduction and subsampling. We compare various covariance-based exploration techniques (Thompson sampling, upper confidence bound, and probabilistic maximum variance) which leverage our scalable GP framework in sparse reward environments. Finally, we show favorable comparison against the bootstrapped deep Q-Network.

97 MATHEMATICS AND COMPUTING↗

The Attraction Indian Buffet Distribution

We propose the attraction Indian buffet distribution (AIBD), a distribution for binary feature matrices influenced by pairwise similarity information. Binary feature matrices are used in Bayesian models to uncover latent variables (i.e., features) that explain observed data. The Indian buffet process (IBP) is a popular exchangeable prior distribution for latent feature matrices. In the presence of additional information, however, the exchangeability assumption is not reasonable or desirable. The AIBD can incorporate pairwise similarity information, yet it preserves many properties of the IBP, including the distribution of the total number of features. Thus, much of the interpretation and intuition that one has for the IBP directly carries over to the AIBD. A temperature parameter controls the degree to which the similarity information affects feature-sharing between observations. Unlike other nonexchangeable distributions for feature allocations, the probability mass function of the AIBD has a tractable normalizing constant, making posterior inference on hyperparameters straight-forward using standard MCMC methods. A novel posterior sampling algorithm is proposed for the IBP and the AIBD. We demonstrate the feasibility of the AIBD as a prior distribution in feature allocation models and compare the performance of competing methods in simulations and an application.

97 MATHEMATICS AND COMPUTING↗

A Bayesian nonparametric analysis for zero-inflated multivariate count data with application to microbiome study

High-throughput sequencing technology has enabled researchers to profile microbial communities from a variety of environments, but analysis of multivariate taxon count data remains challenging. Here, we develop a Bayesian nonparametric (BNP) regression model with zero inflation to analyse multivariate count data from microbiome studies. A BNP approach flexibly models microbial associations with covariates, such as environmental factors and clinical characteristics. The model produces estimates for probability distributions which relate microbial diversity and differential abundance to covariates, and facilitates community comparisons beyond those provided by simple statistical tests. We compare the model to simpler models and popular alternatives in simulation studies, showing, in addition to these additional community-level insights, it yields superior parameter estimates and model fit in various settings. The model's utility is demonstrated by applying it to a chronic wound microbiome data set and a Human Microbiome Project data set, where it is used to compare microbial communities present in different environments.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Studies in Astronomical Time Series Analysis. VI. Bayesian Block Representations

This paper addresses the problem of detecting and characterizing local variability in time series and other forms of sequential data. The goal is to identify and characterize statistically significant variations, at the same time suppressing the inevitable corrupting observational errors. We present a simple nonparametric modeling technique and an algorithm implementing it-an improved and generalized version of Bayesian Blocks [Scargle 1998]-that finds the optimal segmentation of the data in the observation interval. The structure of the algorithm allows it to be used in either a real-time trigger mode, or a retrospective mode. Maximum likelihood or marginal posterior functions to measure model fitness are presented for events, binned counts, and measurements at arbitrary times with known error distributions. Problems addressed include those connected with data gaps, variable exposure, extension to piece- wise linear and piecewise exponential representations, multivariate time series data, analysis of variance, data on the circle, other data modes, and dispersed data. Simulations provide evidence that the detection efficiency for weak signals is close to a theoretical asymptotic limit derived by [Arias-Castro, Donoho and Huo 2003]. In the spirit of Reproducible Research [Donoho et al. (2008)] all of the code and data necessary to reproduce all of the figures in this paper are included as auxiliary material.

signal detection↗

Uncertainty Estimates of Psychoacoustic Thresholds Obtained from Group Tests

Adaptive psychoacoustic test methods, in which the next signal level depends on the response to the previous signal, are the most efficient for determining psychoacoustic thresholds of individual subjects. In many tests conducted in the NASA psychoacoustic labs, the goal is to determine thresholds representative of the general population. To do this economically, non-adaptive testing methods are used in which three or four subjects are tested at the same time with predetermined signal levels. This approach requires us to identify techniques for assessing the uncertainty in resulting group-average psychoacoustic thresholds. In this presentation we examine the Delta Method of frequentist statistics, the Generalized Linear Model (GLM), the Nonparametric Bootstrap, a frequentist method, and Markov Chain Monte Carlo Posterior Estimation and a Bayesian approach. Each technique is exercised on a manufactured, theoretical dataset and then on datasets from two psychoacoustics facilities at NASA. The Delta Method is the simplest to implement and accurate for the cases studied. The GLM is found to be the least robust, and the Bootstrap takes the longest to calculate. The Bayesian Posterior Estimate is the most versatile technique examined because it allows the inclusion of prior information.

Rathsam, Jonathan↗

Direct nonparametric multimessenger constraints on the equation of state of cold dense nuclear matter

We utilize the now substantial amount of astrophysical observations of neutron stars (NSs), along with perturbative quantum chromodynamics (pQCD) calculations at high density, to directly constrain the NS equation of state (EOS). To this end, we construct nonparametric EOS priors by using Gaussian processes trained on 75 EOSs, which include models with either hadrons, hyperons, or quarks at high densities. We create a prior using the full EOS sample (model agnostic), and one prior for each EOS family to test model discrimination. We introduce a novel inference approach, which allows the simultaneous sampling of intrinsic and extrinsic parameters of binary NS mergers, as well as a nonparametric equation of state. We showcase this method in a Bayesian updating scheme by first performing a complete analysis of the binary NS merger event GW170817 with minimal assumptions, and sequentially adding information from x-ray and radio NS observations, along with pQCD calculations. Besides providing standard constraints, such as the pressure at twice nuclear saturation density 𝑝⁡(2⁢𝜌 sat ) = 4.3$^{+0.6}_{−0.6}$ × 10 34 dyne/cm 2 , at 95% confidence level, for the model agnostic prior, our methodology shows how the choice of EOS families used in conditioning changes the inferred astrophysical properties of the EOS, namely tidal deformability and maximum supported NS mass. We find hyperonic priors predicting higher tidal deformabilities for a 1.4⁢𝑀 ⊙ NS, and hadronic priors being preferred by the considered astrophysical data.

73 NUCLEAR PHYSICS AND RADIATION PHYSICS↗

Nonparametric Inference for the Reproductive Rate in Generalized Compartmental Models

We develop a tractable nonparametric model for the time-varying reproductive rate of infectious diseases that combines the structure of a deterministic compartmental model and a stochastic model for incidence data. We use Bayesian inference to estimate, with uncertainty, the reproductive rate of the Coronavirus 2019 outbreak in the U.S. states of California, Florida, Michigan, New Mexico, New York, and Texas from January 2020 to March 2022. Employing the inferred reproductive rates, we estimate the posterior distribution of the time-varying reproduction numbers for each state. Compering the time-varying reproduction numbers across the states, we identify some epidemic waves, potentially driven from changes in human behavior and virus mutations.

97 MATHEMATICS AND COMPUTING↗

Differentiable Preisach Modeling for Characterization and Optimization of Particle Accelerator Systems with Hysteresis

Future improvements in particle accelerator performance are predicated on increasingly accurate online modeling of accelerators. Hysteresis effects in magnetic, mechanical, and material components of accelerators are often neglected in online accelerator models used to inform control algorithms, even though reproducibility errors from systems exhibiting hysteresis are not negligible in high precision accelerators. Here, we combine the classical Preisach model of hysteresis with machine learning techniques to efficiently create nonparametric, high-fidelity models of arbitrary systems exhibiting hysteresis. We experimentally demonstrate how these methods can be used in situ, where a hysteresis model of an accelerator magnet is combined with a Bayesian statistical model of the beam response, allowing characterization of magnetic hysteresis solely from beam-based measurements. Finally, we explore how using these joint hysteresis-Bayesian statistical models allows us to overcome optimization performance limitations that arise when hysteresis effects are ignored.

43 PARTICLE ACCELERATORS↗

Statistical Uncertainty in Paleoclimate Proxy Reconstructions

A quantitative analysis of any environment older than the instrumental record relies on proxies. Uncertainties associated with proxy reconstructions are often underestimated, which can lead to artificial conflict between different proxies, and between data and models. In this paper, using ordinary least squares linear regression as a common example, we describe a simple, robust and generalizable method for quantifying uncertainty in proxy reconstructions. We highlight the primary controls on the magnitude of uncertainty, and compare this simple estimate to equivalent estimates from Bayesian, nonparametric and fiducial statistical frameworks. We discuss when it may be possible to reduce uncertainties, and conclude that the unexplained variance in the calibration must always feature in the uncertainty in the reconstruction. This directs future research toward explaining as much of the variance in the calibration data as possible. We also advocate for a “data-forward” approach, that clearly decouples the presentation of proxy data from plausible environmental inferences.

58 GEOSCIENCES↗

Bayesian projection pursuit regression

In projection pursuit regression (PPR), a univariate response variable is approximated by the sum of $M$ “ridge functions,” which are flexible functions of one-dimensional projections of a multivariate input variable. Traditionally, optimization routines are used to choose the projection directions and ridge functions via a sequential algorithm, and $M$ is typically chosen via cross-validation. Here, we introduce a novel Bayesian version of PPR, which has the benefit of accurate uncertainty quantification. To infer appropriate projection directions and ridge functions, we apply novel adaptations of methods used for the single ridge function case ($M$=1), called the Bayesian Single Index Model; and use a Reversible Jump Markov chain Monte Carlo algorithm to infer the number of ridge functions $M$. We evaluate the predictive ability of our model in 20 simulated scenarios and for 23 real datasets, in a bake-off against an array of state-of-the-art regression methods. Finally, we generalize this methodology and demonstrate the ability to accurately model multivariate response variables. Its effective performance indicates that Bayesian Projection Pursuit Regression is a valuable addition to the existing regression toolbox.

97 MATHEMATICS AND COMPUTING↗

Confidence Intervals for Laboratory Sonic Boom Annoyance Tests

Commercial supersonic flight is currently forbidden over land because sonic booms have historically caused unacceptable annoyance levels in overflown communities. NASA is providing data and expertise to noise regulators as they consider relaxing the ban for future quiet supersonic aircraft. One deliverable NASA will provide is a predictive model for indoor annoyance to aid in setting an acceptable quiet sonic boom threshold. A laboratory study was conducted to determine how indoor vibrations caused by sonic booms affect annoyance judgments. The test method required finding the point of subjective equality (PSE) between sonic boom signals that cause vibrations and signals not causing vibrations played at various amplitudes. This presentation focuses on a few statistical techniques for estimating the interval around the PSE. The techniques examined are the Delta Method, Parametric and Nonparametric Bootstrapping, and Bayesian Posterior Estimation.

Rathsam, Jonathan↗

Physics-Informed Gaussian Process Inference of Liquid Structure from Scattering Data

We present a nonparametric Bayesian framework to infer radial distribution functions from experimental scattering measurements with uncertainty quantification using nonstationary Gaussian processes. The Gaussian process prior mean and kernel functions are designed to mitigate well-known numerical challenges with the Fourier transform, including discrete measurement binning and detector windowing, while encoding fundamental yet minimal physical knowledge of the liquid structure. We demonstrate uncertainty propagation of the Gaussian process posterior to unmeasured quantities of interest. Experimental radial distribution functions of liquid argon and water with uncertainty quantification are provided as both a proof of principle for the method and a benchmark for molecular models.

Chemical structure↗

Shape Optimization by Bayesian-Validated Computer-Simulation Surrogates

A nonparametric-validated, surrogate approach to optimization has been applied to the computational optimization of eddy-promoter heat exchangers and to the experimental optimization of a multielement airfoil. In addition to the baseline surrogate framework, a surrogate-Pareto framework has been applied to the two-criteria, eddy-promoter design problem. The Pareto analysis improves the predictability of the surrogate results, preserves generality, and provides a means to rapidly determine design trade-offs. Significant contributions have been made in the geometric description used for the eddy-promoter inclusions as well as to the surrogate framework itself. A level-set based, geometric description has been developed to define the shape of the eddy-promoter inclusions. The level-set technique allows for topology changes (from single-body,eddy-promoter configurations to two-body configurations) without requiring any additional logic. The continuity of the output responses for input variations that cross the boundary between topologies has been demonstrated. Input-output continuity is required for the straightforward application of surrogate techniques in which simplified, interpolative models are fitted through a construction set of data. The surrogate framework developed previously has been extended in a number of ways. First, the formulation for a general, two-output, two-performance metric problem is presented. Surrogates are constructed and validated for the outputs. The performance metrics can be functions of both outputs, as well as explicitly of the inputs, and serve to characterize the design preferences. By segregating the outputs and the performance metrics, an additional level of flexibility is provided to the designer. The validated outputs can be used in future design studies and the error estimates provided by the output validation step still apply, and require no additional appeals to the expensive analysis. Second, a candidate-based a posteriori error analysis capability has been developed which provides probabilistic error estimates on the true performance for a design randomly selected near the surrogate-predicted optimal design.

Patera, Anthony T.↗

Multiclass Classification Using Bayesian Multivariate Adaptive Regression Splines

We present a new Bayesian model for the problem of multiclass classification. In this model, the probabilities of class membership of a given observation are determined by the mean of a latent Gaussian distribution. The mean functions of this latent distribution consist of combinations of highly flexible basis functions of the inputs: multivariate adaptive regression splines (MARS), first developed for multiple regression. We use reversible jump Markov chain Monte Carlo to make inference on the classification model, including the number of basis functions. We compare the probabilistic classification performance of our proposed approach to existing methods on simulated and benchmark data, and compare uncertainty estimates on simulated data. Our proposed method compares favorably with existing Bayesian and frequentist multiclass classification methods in out-of-sample probabilistic classification, and uncertainty estimation of these probabilistic classifications. We examine the fit of the proposed method to a data set of hurricane storm surge levels near Delaware Bay, US, and conclude that sea level rise is a key contributor to damage delivered by storm surge.

97 MATHEMATICS AND COMPUTING↗

Detailed examination of astrophysical constraints on the symmetry energy and the neutron skin of 208 Pb with minimal modeling assumptions

We report the symmetry energy and its density dependence are pivotal for many nuclear physics and astrophysics applications, as they determine properties ranging from the neutron-skin thickness of nuclei to the crust thickness and the radius of neutron stars. Recently, PREX-II reported a value of 0.283 ± 0.071 fm for the neutron-skin thickness of 208 Pb, $R^{^{208}Pb}_{skin}$, implying a symmetry-energy slope parameter L of 106 ± 37 MeV, larger than most ranges obtained from microscopic calculations and other nuclear experiments. We use a nonparametric equation of state representation based on Gaussian processes to constrain the symmetry energy S 0 , L, and $R^{^{208}Pb}_{skin}$ directly from observations of neutron stars with minimal modeling assumptions. The resulting astrophysical constraints from heavy pulsar masses, LIGO/Virgo, and NICER favor smaller values of the neutron skin and L, as well as negative symmetry incompressibilities. Combining astrophysical data with chiral effective field theory (χEFT) and PREX-II constraints yields S 0 = 33.0$^{+2.0}_{-1.8}$ MeV, L = 53$^{+14}_{-15}$ MeV, and $R^{^{208}Pb}_{skin}$ = 0.17 $^{+0.04}_{-0.04}$ fm. We also examine the consistency of several individual χEFT calculations with astrophysical observations and terrestrial experiments. We find that there is only mild tension between χ EFT, astrophysical data, and PREX-II's $R^{^{208}Pb}_{skin}$ measurement (p value = 12.3%) and that there is excellent agreement between χEFT, astrophysical data, and other nuclear experiments.

74 ATOMIC AND MOLECULAR PHYSICS↗