Search NASA⌕ Search

SEARCH · Search NASA

Results for “Bayesian statistical modeling”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 181 records · Page 10

Online Multi-Modal Learning and Adaptive Information Trajectory Planning for Autonomous Exploration

In robotic information gathering missions, scientists are typically interested in understanding variables which require proxy measurements from specialized sensor suites to estimate. However, energy and time constraints limit how often these sensors can be used in a mission. Robots are also equipped with cheaper to use navigation sensors such as cameras. In this paper, we explore a challenging planning problem in which a robot is required to learn about a scientific variable of interest in an initially unknown environment by planning informative paths and deciding when and where to use its sensors. To tackle this we present two innovations: a Bayesian generative model framework to automatically learn correlations between expensive science sensors and cheaper to use navigation sensors online, and a sampling based approach to plan for multiple sensors while handling long horizons and budget constraints. Our approach does not grow in complexity with data and is anytime making it highly applicable to field robotics. We tested our approach extensively in simulation and validated it with real data collected during the 2014 Mojave Volatiles Prospector Mission. Our planning algorithm performs statistically significantly better than myopic approaches and at least as well as a coverage-based algorithm in an initially unknown environment while having added advantages of being able to exploit prior knowledge and handle other intricacies of the real world without further algorithmic modifications.

learning↗

The DESI-Lensing Mock Challenge: large-scale cosmological analysis of 3x2-pt statistics

The current generation of large galaxy surveys will test the cosmological model by combining multiple types of observational probes. Realising the statistical promise of these new datasets requires rigorous attention to all aspects of analysis including cosmological measurements, modelling, covariance and parameter likelihood. In this paper we present the results of an end-to-end simulation study designed to test the analysis pipeline for the combination of the Dark Energy Spectroscopic Instrument (DESI) Year 1 galaxy redshift dataset and separate weak gravitational lensing information from the Kilo-Degree Survey, Dark Energy Survey and Hyper-Suprime-Cam Survey. Our analysis employs the 3x2-pt correlation functions including cosmic shear and galaxy-galaxy lensing, together with the projected correlation function of the spectroscopic DESI lenses. We build realistic simulations of these datasets including galaxy halo occupation distributions, photometric redshift errors, weights, multiplicative shear calibration biases and magnification. We calculate the analytical covariance of these correlation functions including the Gaussian, noise and super-sample contributions, and show that our covariance determination agrees with estimates based on the ensemble of simulations. We use a Bayesian inference platform to demonstrate that we can recover the fiducial cosmological parameters of the simulation within the statistical error margin of the experiment, investigating the sensitivity to scale cuts. This study is the first in a sequence of papers in which we present and validate the large-scale 3x2-pt cosmological analysis of DESI-Y1.

79 ASTRONOMY AND ASTROPHYSICS↗

Ripening of Rh Nanoparticle Catalysts in Reverse Water–Gas Shift via a Data-Driven Model Combining Physics, Theory, and Experiment

Degradation via sintering is an ongoing challenge that impedes the broad commercial success of supported metallic nanoparticle catalysts. To mitigate degradation via informed catalyst design and process operations, here we aim to disambiguate the underlying mechanisms of sintering by combining theory and experiment in a quantitative framework. While mechanistic sintering models exist, they only model a single sintering pathway, even though multiple sintering mechanisms can occur simultaneously or dominate at different stages of the process. Data-driven machine learning models have emerged as a means to represent complex processes through data regression. However, machine learning models have very large data needs and lack mechanistic insights due to their black-box encoding. To develop an interpretive model of catalyst degradation via sintering, we constructed a hybrid model combining mechanistic “physics-based” models and data-driven methods to obtain both reliable predictions and mechanistic insights regarding experimentally observed sintering phenomena. Focusing on nanoparticle sintering in the Rh–TiO 2 catalyst for the reverse water–gas shift (RWGS) reaction, the hybrid model couples a mechanistic term for Ostwald ripening with energy values calculated via density functional theory (DFT) with a parametric, data-driven discrepancy function term for unmodeled mechanisms. The hybrid model is trained using Bayesian inference with data collected from small-angle X-ray scattering (SAXS) in situ experiments wherein average nanoparticle diameter versus time was measured at three relevant operating temperatures. The calibrated hybrid model results show that an Ostwald ripening-only model parameterized with fixed DFT energies does not fully capture the time and temperature dependence of the SAXS-observed sintering kinetics, and that an additional functional contribution, or DFT energy calibration, is required to reconcile simulation and experiment. Analysis of the hybrid-model error confirms that the hybrid model outperforms both the purely mechanistic and purely data-driven alternatives in terms of expected predictive accuracy for time-evolving average particle sizes. Furthermore, the results support the hypothesis that the Ostwald ripening mechanism is less important for explaining the sintering phenomena as operating temperature increases under an assumed fixed DFT parameterization. This could be explained in one of two ways: either latent, unmodeled sintering mechanisms dominate at higher temperatures, or the DFT uncertainty increases with temperature. The proposed modeling approach directly links theory to experiments and simulations via a statistical hybrid modeling framework and can be extended to other catalytic systems to improve predictive models and mechanistic understanding.

Bayesian hybrid modeling↗

Designing an Optimal Sensor Network via Minimizing Information Loss

Optimal experimental design is a classic topic in statistics, with many well-studied problems, applications, and solutions. The design problem we study is the placement of sensors to monitor spatiotemporal processes, explicitly accounting for the temporal dimension in our modeling and optimization. We observe that recent advancements in computational sciences often yield large datasets based on physics-based simulations, which are rarely leveraged in experimental design. We introduce a novel model-based sensor placement criterion, along with a highly-efficient optimization algorithm, which integrates physics-based simulations and Bayesian experimental design principles to identify sensor networks that “minimize information loss” from simulated data. Our technique relies on sparse variational inference and (separable) Gauss-Markov priors, and thus may adapt many techniques from Bayesian experimental design. We validate our method through a case study monitoring air temperature in Phoenix, Arizona, using state-of-the-art physics-based simulations. Our results show our framework to be superior to random or quasi-random sampling, particularly with a limited number of sensors. We conclude by discussing practical considerations and implications of our framework, including more complex modeling tools and real-world deployments.

54 ENVIRONMENTAL SCIENCES↗

Profiles of Gamma-Ray Bursts and Their Component Pulses

One physically informative regularity of their otherwise heterogeneous ensemble, is that many Gamma-Ray Bursts consist of well defined pulses. To objectively quantify the temporal structure of BATSE bursts, we have developed an automatic modeling procedure that separates overlapping pulses and determines the energy-dependence of the pulse-shape parameters. No binning of photon arrival times is needed, so when applied to time-tagged events (TTE) the procedure captures variability information down to the shortest time scales present in the raw data. Maximizing the Bayesian likelihood function Pr(data/model) yields estimates of the model parameters, including the number of pulses present, and allows intercomparison of models of different forms. As with any nonlinear optimization, good initial guesses are crucial to avoid convergence to undesirable local minima. We find excellent initial pulse decompositions by wavelet-denoising a cumulative distribution of the raw photon arrival data; differentiation then gives a time profile mostly free of the systematic effects of degraded resolution (as in ordinary Fourier smoothing) and binning. We present statistical information on pulse rise-time, decay-time, peakedness, and amplitudes, plus their energy dependences - both within a single burst and for a large ensemble of bursts.

Scargle, Jeff D.↗

Confidence Intervals for Laboratory Sonic Boom Annoyance Tests

Commercial supersonic flight is currently forbidden over land because sonic booms have historically caused unacceptable annoyance levels in overflown communities. NASA is providing data and expertise to noise regulators as they consider relaxing the ban for future quiet supersonic aircraft. One deliverable NASA will provide is a predictive model for indoor annoyance to aid in setting an acceptable quiet sonic boom threshold. A laboratory study was conducted to determine how indoor vibrations caused by sonic booms affect annoyance judgments. The test method required finding the point of subjective equality (PSE) between sonic boom signals that cause vibrations and signals not causing vibrations played at various amplitudes. This presentation focuses on a few statistical techniques for estimating the interval around the PSE. The techniques examined are the Delta Method, Parametric and Nonparametric Bootstrapping, and Bayesian Posterior Estimation.

Rathsam, Jonathan↗

A flexible class of priors for orthonormal matrices with basis function-specific structure

Statistical modeling of high-dimensional matrix-valued data motivates the use of a low-rank representation that simultaneously summarizes key characteristics of the data and enables dimension reduction. Low-rank representations commonly factor the original data into the product of orthonormal basis functions and weights, where each basis function represents an independent feature of the data. However, the basis functions in these factorizations are typically computed using algorithmic methods that cannot quantify uncertainty or account for basis function correlation structure a priori. While there exist Bayesian methods that allow for a common correlation structure across basis functions, empirical examples motivate the need for basis function-specific dependence structure. We propose a prior distribution for orthonormal matrices that can explicitly model basis function-specific structure. The prior is used within a general probabilistic model for singular value decomposition to conduct posterior inference on the basis functions while accounting for measurement error and fixed effects. We discuss how the prior specification can be used for various scenarios and demonstrate favorable model properties through synthetic data examples. Finally, we apply our method to two-meter air temperature data from the Pacific Northwest, enhancing our understanding of the Earth system’s internal variability.

97 MATHEMATICS AND COMPUTING↗

Dynamical dark energy in light of the DESI DR2 baryonic acoustic oscillations measurements

Understanding whether cosmic acceleration arises from a cosmological constant or a dynamical component is a central goal of cosmology, and the Dark Energy Spectroscopic Instrument (DESI) enables stringent tests with high-precision distance measurements. Here we analyse measurements of baryon acoustic oscillations in DESI Data Release 1 and Data Release 2 and consider type Ia supernovae and a distance prior for the cosmic microwave background. With the larger statistical power and wider redshift coverage of Data Release 2, the preference for dynamical dark energy does not diminish relative to Data Release 1. Using both a shape-function reconstruction and non-parametric approaches with a Horndeski-motivated correlation prior, we find that the equation of state for dark energy w(z) varies with redshift. Baryon acoustic oscillation data alone yield modest constraints, but in combination with independent supernova compilations and the prior for the cosmic microwave background, they strengthen the evidence for dynamics. A Bayesian comparison of models shows moderate support for departures from Λ cold dark matter (ΛCDM) when several degrees of freedom in w(z) are allowed, corresponding to ~3σ tension with ΛCDM (and higher for some datasets). Despite methodological differences, our results are consistent with companion DESI papers, underscoring the complementarity of the approaches. Possible systematics remain under study; forthcoming DESI, Euclid and next-generation cosmic microwave background data will provide decisive tests.

cosmology↗

Dynamical dark energy in light of the DESI DR2 baryonic acoustic oscillations measurements

Understanding whether cosmic acceleration arises from a cosmological constant or a dynamical component is a central goal of cosmology, and the Dark Energy Spectroscopic Instrument (DESI) enables stringent tests with high-precision distance measurements. Here we analyse measurements of baryon acoustic oscillations in DESI Data Release 1 and Data Release 2 and consider type Ia supernovae and a distance prior for the cosmic microwave background. With the larger statistical power and wider redshift coverage of Data Release 2, the preference for dynamical dark energy does not diminish relative to Data Release 1. Using both a shape-function reconstruction and non-parametric approaches with a Horndeski-motivated correlation prior, we find that the equation of state for dark energy w(z) varies with redshift. Baryon acoustic oscillation data alone yield modest constraints, but in combination with independent supernova compilations and the prior for the cosmic microwave background, they strengthen the evidence for dynamics. A Bayesian comparison of models shows moderate support for departures from Λ cold dark matter (ΛCDM) when several degrees of freedom in w(z) are allowed, corresponding to ~3σ tension with ΛCDM (and higher for some datasets). Despite methodological differences, our results are consistent with companion DESI papers, underscoring the complementarity of the approaches. Possible systematics remain under study; forthcoming DESI, Euclid and next-generation cosmic microwave background data will provide decisive tests.

79 ASTRONOMY AND ASTROPHYSICS↗

Optimal Estimation Framework for Ocean Color Atmospheric Correction and Pixel-level Uncertainty Quantification

Ocean color remote sensing requires compensation for atmospheric scattering and absorption (aerosol, Rayleigh, and trace gases), referred to as atmospheric correction (AC). AC allows inference of parameters such as spectrally resolved remote sensing reflectance ( R rs )(λ) ; sr 1 ) at the ocean surface from the top-of-atmosphere reflectance. Often, the uncertainty of this process is not fully explored. Bayesian inference techniques provide a simultaneous AC and uncertainty assessment via a full posterior distribution of the relevant variables, given the prior distribution of those variables and the radiative transfer (RT) likelihood function. Given uncertainties in the algorithm inputs, the Bayesian framework enables better constraints on the AC process by using the complete spectral information compared to traditional approaches that use only a subset of bands for AC. This paper investigates a Bayesian inference research method (Optimal Estimation, OE) for ocean color AC by simultaneously retrieving atmospheric and ocean properties using all visible and near-infrared spectral bands. The OE algorithm analytically approximates the posterior distribution of parameters based on normality assumptions and provides a potentially viable operational algorithm with a reduced computational expense. We developed a Neural Network (NN) RT forward model look-up-table-based emulator to increase algorithm efficiency further and thus speed up the likelihood computations. We then applied the OE algorithm to synthetic data and observations from the MODerate resolution Imaging Spectroradiometer (MODIS) on NASA’s Aqua spacecraft. We compared the R rs )(λ) retrieval and its uncertainty estimates from the OE method with in-situ validation data from the SeaWiFS Bio-optical Archive and Storage System (SeaBASS) and Aerosol Robotic Network Ocean Color (AERONET-OC) datasets. The OE algorithm improved R rs )(λ) estimates relative to the NASA standard operational algorithm by improving all statistical metrics at 443, 555, and 667 nm. Unphysical negative R rs )(λ) , which often appear in complex water conditions, was reduced by a factor of 3. The OE-derived pixel-level R rs )(λ) uncertainty estimates were also assessed relative to in-situ data and were shown to have skill.

Atmospheric correction↗

Deriving cloud droplet number concentration from surface-based remote sensors with an emphasis on lidar measurements

Abstract. Given the importance of constraining cloud droplet number concentrations (Nd) in low-level clouds, we explore two methods for retrieving Nd from surface-based remote sensing that emphasize the information content in lidar measurements. Because Nd is the zeroth moment of the droplet size distribution (DSD), and all remote sensing approaches respond to DSD moments that are at least 2 orders of magnitude greater than the zeroth moment, deriving Nd from remote sensing measurements has significant uncertainty. At minimum, such algorithms require the extrapolation of information from two other measurements that respond to different moments of the DSD. Lidar, for instance, is sensitive to the second moment (cross-sectional area) of the DSD, while other measures from microwave sensors respond to higher-order moments. We develop methods using a simple lidar forward model that demonstrates that the depth to the maximum in lidar-attenuated backscatter (Rmax⁡) is strongly sensitive to Nd when some measure of the liquid water content vertical profile is given or assumed. Knowledge of Rmax⁡ to within 5 m can constrain Nd to within several tens of percent. However, operational lidar networks provide vertical resolutions of > 15 m, making a direct calculation of Nd from Rmax⁡ very uncertain. Therefore, we develop a Bayesian optimal estimation algorithm that brings additional information to the inversion such as lidar-derived extinction and radar reflectivity near the cloud top. This statistical approach provides reasonable characterizations of Nd and effective radius (re) to within approximately a factor of 2 and 30 %, respectively. By comparing surface-derived cloud properties with MODIS satellite and aircraft data collected during the MARCUS and CAPRICORN II campaigns, we demonstrate the utility of the methodology.

54 ENVIRONMENTAL SCIENCES↗

Identifying microbial drivers in biological phenotypes with a Bayesian network regression model

Abstract In Bayesian Network Regression models, networks are considered the predictors of continuous responses. These models have been successfully used in brain research to identify regions in the brain that are associated with specific human traits, yet their potential to elucidate microbial drivers in biological phenotypes for microbiome research remains unknown. In particular, microbial networks are challenging due to their high dimension and high sparsity compared to brain networks. Furthermore, unlike in brain connectome research, in microbiome research, it is usually expected that the presence of microbes has an effect on the response (main effects), not just the interactions. Here, we develop the first thorough investigation of whether Bayesian Network Regression models are suitable for microbial datasets on a variety of synthetic and real data under diverse biological scenarios. We test whether the Bayesian Network Regression model that accounts only for interaction effects (edges in the network) is able to identify key drivers (microbes) in phenotypic variability. We show that this model is indeed able to identify influential nodes and edges in the microbial networks that drive changes in the phenotype for most biological settings, but we also identify scenarios where this method performs poorly which allows us to provide practical advice for domain scientists aiming to apply these tools to their datasets. BNR models provide a framework for microbiome researchers to identify connections between microbes and measured phenotypes. We allow the use of this statistical model by providing an easy‐to‐use implementation which is publicly available Julia package at https://github.com/solislemuslab/BayesianNetworkRegression.jl .

59 BASIC BIOLOGICAL SCIENCES↗

Bayesian event categorization matrix approach for explosion monitoring

Current efforts to correctly categorize natural events from suspected explosion sources with data that is collected by ground- or space-based sensors presents historical challenges that remain unaddressed by the Event Categorization Matrix (ECM) model. Smaller historical events (lower yield explosions) may have data available from fewer measurement techniques than are available today, and therefore, a historical event record can lack a complete set of discriminants. The covariance structures can also differ between such observations of event (source-type) categories. Both obstacles are problematic for the classic ECM model. Our work addresses this gap and presents a Bayesian update to the previous ECM model, termed the Bayesian Event Categorization Matrix model, which can be trained on partial observations and does not rely on a pooled covariance structure. We further augment the ECM model with Bayesian Decision Theory so that false negative or false positive rates of an event categorization can be reduced in an intuitive manner. To demonstrate improved categorization rates for the Bayesian Event Categorization Matrix model, we compare an array of Bayesian and classic models with multiple performance metrics using Monte Carlo experiments. We use both synthetic and real data. Our Bayesian models show consistent gains in overall accuracy and lower false negative rates relative to the classic ECM model. Here, we propose future avenues to improve Bayesian Event Categorization Matrix models’ decision making and predictive capability.

58 GEOSCIENCES↗

The Analysis of the Contribution of Human Factors to the In-Flight Loss of Control Accidents

In-flight loss of control (LOC) is currently the leading cause of fatal accidents based on various commercial aircraft accident statistics. As the Next Generation Air Transportation System (NextGen) emerges, new contributing factors leading to LOC are anticipated. The NASA Aviation Safety Program (AvSP), along with other aviation agencies and communities are actively developing safety products to mitigate the LOC risk. This paper discusses the approach used to construct a generic integrated LOC accident framework (LOCAF) model based on a detailed review of LOC accidents over the past two decades. The LOCAF model is comprised of causal factors from the domain of human factors, aircraft system component failures, and atmospheric environment. The multiple interdependent causal factors are expressed in an Object-Oriented Bayesian belief network. In addition to predicting the likelihood of LOC accident occurrence, the system-level integrated LOCAF model is able to evaluate the impact of new safety technology products developed in AvSP. This provides valuable information to decision makers in strategizing NASA's aviation safety technology portfolio. The focus of this paper is on the analysis of human causal factors in the model, including the contributions from flight crew and maintenance workers. The Human Factors Analysis and Classification System (HFACS) taxonomy was used to develop human related causal factors. The preliminary results from the baseline LOCAF model are also presented.

Ancel, Ersin↗

The Error Distribution of BATSE GRB Location

We develop empirical probability models for BATSE GRB location errors by a Bayesian analysis of the separations between BATSE GRB locations and locations obtained with the InterPlanetary Network (IPN). Models are compared and their parameters estimated using 394 GRBs with single IPN annuli and 20 GRBs with intersecting IPN annuli. Most of the analysis is for the 4B (rev) BATSE catalog; earlier catalogs are also analyzed. The simplest model that provides a good representation of the error distribution has 78% of the locations in a 'core' term with a systematic error of 1.85 degrees and the remainder in an extended tail with a systematic error of 5.36 degrees, implying a 68% confidence region for bursts with negligible statistical errors of 2.3 degrees. There is some evidence for a more complicated model in which the error distribution depends on the BATSE datatype that was used to obtain the location. Bright bursts are typically located using the CONT datatype, and according to the more complicated model, the 68% confidence region for CONT-located bursts with negligible statistical errors is 2.0 degrees.

Briggs, Michael S.↗

The Error Distribution of BATSE Gamma-Ray Burst Locations

Empirical probability models for BATSE gamma-ray burst (GRB) location errors are developed via a Bayesian analysis of the separations between BATSE GRB locations and locations obtained with the Interplanetary Network (IPN). Models are compared and their parameters estimated using 392 GRBs with single IPN annuli and 19 GRBs with intersecting IPN annuli. Most of the analysis is for the 4Br BATSE catalog; earlier catalogs are also analyzed. The simplest model that provides a good representation of the error distribution has 78% of the probability in a "core" term with a systematic error of 1.85 deg and the remainder in an extended tail with a systematic error of 5.1 deg, which implies a 68% confidence radius for bursts with negligible statistical uncertainties of 2.2 deg. There is evidence for a more complicated model in which the error distribution depends on the BATSE data type that was used to obtain the location. Bright bursts are typically located using the CONT data type, and according to the more complicated model, the 68% confidence radius for CONT-located bursts with negligible statistical uncertainties is 2.0 deg.

Briggs, Michael S.↗

Continuous Habitable Zones: Pairing a GCM and Bayesian Framework to Predict Habitable Zone Evolution

In the near-future, new space telescopes like JWST, LUVOIR, and HabEx will begin attempting to explore the properties of atmospheres of potentially habitable planets. This will require a significant amount of time and resources for even a single planet, which makes it essential to prioritize observations by those most-likely to have detectable life. Here we present a statistical method to estimate the probabilities that specific exoplanets have been continuously in the habitable zone of their host stars for more than 2 billion years, the approximate time it took life on Earth to significantly increase the oxygen content of the atmosphere. We introduce the use of statistics of an ensemble of 3D planetary general circulation models to estimate these probabilities, replacing prior 1D model estimates.

habitable planets↗

Pairing a GCM and Bayesian Framework to Predict Habitable Zone Evolution

In the near-future, new space telescopes like JWST, LUVOIR, and HabEx will begin attempting to explore the properties of atmospheres of potentially habitable planets. This will require a significant amount of time and resources for even a single planet, which makes it essential to prioritize observations by those most-likely to have detectable life. Here we present a statistical method to estimate the probabilities that specific exoplanets have been continuously in the habitable zone of their host stars for more than 2 billion years, the approximate time it took life on Earth to significantly increase the oxygen content of the atmosphere. We introduce the use of statistics of an ensemble of 3D planetary general circulation models to estimate these probabilities, replacing prior 1D model estimates.

habitable planets↗