Search NASA⌕ Search

SEARCH · Search NASA

Results for “Bayesian Statistics”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 163 records · Page 9

Hybrid Gibbs Sampling and MCMC for CMB Analysis at Small Angular Scales

A) Gibbs Sampling has now been validated as an efficient, statistically exact, and practically useful method for "low-L" (as demonstrated on WMAP temperature polarization data). B) We are extending Gibbs sampling to directly propagate uncertainties in both foreground and instrument models to total uncertainty in cosmological parameters for the entire range of angular scales relevant for Planck. C) Made possible by inclusion of foreground model parameters in Gibbs sampling and hybrid MCMC and Gibbs sampling for the low signal to noise (high-L) regime. D) Future items to be included in the Bayesian framework include: 1) Integration with Hybrid Likelihood (or posterior) code for cosmological parameters; 2) Include other uncertainties in instrumental systematics? (I.e. beam uncertainties, noise estimation, calibration errors, other).

Gibbs sampling↗

Confidence set inference with a prior quadratic bound

In the uniqueness part of a geophysical inverse problem, the observer wants to predict all likely values of P unknown numerical properties z = (z sub 1,...,z sub p) of the earth from measurement of D other numerical properties y(0)=(y sub 1(0),...,y sub D(0)) knowledge of the statistical distribution of the random errors in y(0). The data space Y containing y(0) is D-dimensional, so when the model space X is infinite-dimensional the linear uniqueness problem usually is insoluble without prior information about the correct earth model x. If that information is a quadratic bound on x (e.g., energy or dissipation rate), Bayesian inference (BI) and stochastic inversion (SI) inject spurious structure into x, implied by neither the data nor the quadratic bound. Confidence set inference (CSI) provides an alternative inversion technique free of this objection. CSI is illustrated in the problem of estimating the geomagnetic field B at the core-mantle boundary (CMB) from components of B measured on or above the earth's surface. Neither the heat flow nor the energy bound is strong enough to permit estimation of B(r) at single points on the CMB, but the heat flow bound permits estimation of uniform averages of B(r) over discs on the CMB, and both bounds permit weighted disc-averages with continous weighting kernels. Both bounds also permit estimation of low-degree Gauss coefficients at the CMB. The heat flow bound resolves them up to degree 8 if the crustal field at satellite altitudes must be treated as a systematic error, but can resolve to degree 11 under the most favorable statistical treatment of the crust. These two limits produce circles of confusion on the CMB with diameters of 25 deg and 19 deg respectively.

Backus, George E.↗

Continuous Habitable Zones: Pairing a GCM and Bayesian Framework to Predict Habitable Zone Evolution

In the near-future, new space telescopes like JWST, LUVOIR, and HabEx will begin attempting to explore the properties of atmospheres of potentially habitable planets. This will require a significant amount of time and resources for even a single planet, which makes it essential to prioritize observations by those most-likely to have detectable life. Here we present a statistical method to estimate the probabilities that specific exoplanets have been continuously in the habitable zone of their host stars for more than 2 billion years, the approximate time it took life on Earth to significantly increase the oxygen content of the atmosphere. We introduce the use of statistics of an ensemble of 3D planetary general circulation models to estimate these probabilities, replacing prior 1D model estimates.

habitable planets↗

Pairing a GCM and Bayesian Framework to Predict Habitable Zone Evolution

In the near-future, new space telescopes like JWST, LUVOIR, and HabEx will begin attempting to explore the properties of atmospheres of potentially habitable planets. This will require a significant amount of time and resources for even a single planet, which makes it essential to prioritize observations by those most-likely to have detectable life. Here we present a statistical method to estimate the probabilities that specific exoplanets have been continuously in the habitable zone of their host stars for more than 2 billion years, the approximate time it took life on Earth to significantly increase the oxygen content of the atmosphere. We introduce the use of statistics of an ensemble of 3D planetary general circulation models to estimate these probabilities, replacing prior 1D model estimates.

habitable planets↗

Structural model optimization using statistical evaluation

The results of research in applying statistical methods to the problem of structural dynamic system identification are presented. The study is in three parts: a review of previous approaches by other researchers, a development of various linear estimators which might find application, and the design and development of a computer program which uses a Bayesian estimator. The method is tried on two models and is successful where the predicted stiffness matrix is a proper model, e.g., a bending beam is represented by a bending model. Difficulties are encountered when the model concept varies. There is also evidence that nonlinearity must be handled properly to speed the convergence.

Collins, J. D.↗

A Ground Flash Fraction Retrieval Algorithm for GLM

A Bayesian inversion method is introduced for retrieving the fraction of ground flashes in a set of N lightning observed by a satellite lightning imager (such as the Geostationary Lightning Mapper, GLM). An exponential model is applied as a physically reasonable constraint to describe the measured lightning optical parameter distributions. Population statistics (i.e., the mean and variance) are invoked to add additional constraints to the retrieval process. The Maximum A Posteriori (MAP) solution is employed. The approach is tested by performing simulated retrievals, and retrieval error statistics are provided. The approach is feasible for N greater than 2000, and retrieval errors decrease as N is increased.

Koshak, William J.↗

Putting Priors in Mixture Density Mercer Kernels

This paper presents a new methodology for automatic knowledge driven data mining based on the theory of Mercer Kernels, which are highly nonlinear symmetric positive definite mappings from the original image space to a very high, possibly infinite dimensional feature space. We describe a new method called Mixture Density Mercer Kernels to learn kernel function directly from data, rather than using predefined kernels. These data adaptive kernels can en- code prior knowledge in the kernel using a Bayesian formulation, thus allowing for physical information to be encoded in the model. We compare the results with existing algorithms on data from the Sloan Digital Sky Survey (SDSS). The code for these experiments has been generated with the AUTOBAYES tool, which automatically generates efficient and documented C/C++ code from abstract statistical model specifications. The core of the system is a schema library which contains template for learning and knowledge discovery algorithms like different versions of EM, or numeric optimization methods like conjugate gradient methods. The template instantiation is supported by symbolic- algebraic computations, which allows AUTOBAYES to find closed-form solutions and, where possible, to integrate them into the code. The results show that the Mixture Density Mercer-Kernel described here outperforms tree-based classification in distinguishing high-redshift galaxies from low- redshift galaxies by approximately 16% on test data, bagged trees by approximately 7%, and bagged trees built on a much larger sample of data by approximately 2%.

Srivastava, Ashok N.↗

Confidence set inference with a prior quadratic bound

In the uniqueness part of a geophysical inverse problem, the observer wants to predict all likely values of P unknown numerical properties z=(z sub 1,...,z sub p) of the earth from measurement of D other numerical properties y (sup 0) = (y (sub 1) (sup 0), ..., y (sub D (sup 0)), using full or partial knowledge of the statistical distribution of the random errors in y (sup 0). The data space Y containing y(sup 0) is D-dimensional, so when the model space X is infinite-dimensional the linear uniqueness problem usually is insoluble without prior information about the correct earth model x. If that information is a quadratic bound on x, Bayesian inference (BI) and stochastic inversion (SI) inject spurious structure into x, implied by neither the data nor the quadratic bound. Confidence set inference (CSI) provides an alternative inversion technique free of this objection. Confidence set inference is illustrated in the problem of estimating the geomagnetic field B at the core-mantle boundary (CMB) from components of B measured on or above the earth's surface.

Backus, George E.↗

Accounting for Epistemic Uncertainty in Mission Supportability Assessment: A Necessary Step in Understanding Risk and Logistics Requirements

Future crewed missions to Mars present a maintenance logistics challenge that is unprecedented in human spaceflight. Mission endurance – defined as the time between resupply opportunities – will be significantly longer than previous missions, and therefore logistics planning horizons are longer and the impact of uncertainty is magnified. Maintenance logistics forecasting typically assumes that component failure rates are deterministically known and uses them to represent aleatory uncertainty, or uncertainty that is inherent to the process being examined. However, failure rates cannot be directly measured; rather, they are estimated based on similarity to other components or statistical analysis of observed failures. As a result, epistemic uncertainty – that is, uncertainty in knowledge of the process – exists in failure rate estimates that must be accounted for. Analyses that neglect epistemic uncertainty tend to significantly underestimate risk. Epistemic uncertainty can be reduced via operational experience; for example, the International Space Station (ISS) failure rate estimates are refined using a Bayesian update process. However, design changes may re-introduce epistemic uncertainty. Thus, there is a tradeoff between changing a design to reduce failure rates and operating a fixed design to reduce uncertainty. This paper examines the impact of epistemic uncertainty on maintenance logistics requirements for future Mars missions, using data from the ISS Environmental Control and Life Support System (ECLS) as a baseline for a case study. Sensitivity analyses are performed to investigate the impact of variations in failure rate estimates and epistemic uncertainty on spares mass. The results of these analyses and their implications for future system design and mission planning are discussed.

Owens, Andrew↗

Mind the Gap: Addressing Data Gaps and Assessing Noise Mismodeling in LISA

Due to the sheer complexity of the Laser Interferometer Space Antenna (LISA) space mission, data gaps arising from instrumental irregularities and/or scheduled maintenance are unavoidable. Focusing on merger-dominated massive black hole binary signals, we test the appropriateness of the Whittle-likelihood on gapped data in a variety of cases. From first principles, we derive the likelihood valid for gapped data in both the time and frequency domains. Cheap-to-evaluate proxies to p-p plots are derived based on a Fisher-based formalism, and verified through Bayesian techniques. Our tools allow to predict the altered variance in the parameter estimates that arises from noise mismodeling, as well as the information loss represented by the broadening of the posteriors. The result of noise mismodeling with gaps is sensitive to the characteristics of the noise model, with strong low-frequency (red) noise and strong high-frequency (blue) noise giving statistically significant fluctuations in recovered parameters. We demonstrate that the introduction of a tapering window reduces statistical inconsistency errors, at the cost of less precise parameter estimates. We also show that the assumption of independence between inter-gap segments appears to be a fair approximation even if the data set is inherently coherent. However, if one instead assumes fictitious correlations in the data stream, when the data segments are actually independent, then the resultant parameter recoveries could be inconsistent with the true parameters. The theoretical and numerical practices that are presented in this work could readily be incorporated into global-fit pipelines operating on gapped data.

LISA↗

MOA-2020-BLG-135Lb: A New Neptune-class Planet for the Extended MOA-II Exoplanet Microlens Statistical Analysis

We report the light-curve analysis for the event MOA-2020-BLG-135, which leads to the discovery of a new Neptune-class planet, MOA-2020-BLG-135Lb. With a derived mass ratio of q=1.52 +0.39 -0.31 x10 -4 and separation s ≈ 1, the planet lies exactly at the break and likely peak of the exoplanet mass-ratio function derived by the Microlensing Observations in Astrophysics (MOA) Collaboration. We estimate the properties of the lens system based on a Galactic model and considering two different Bayesian priors: one assuming that all stars have an equal planet-hosting probability and the other that planets are more likely to orbit more-massive stars. With a uniform host mass prior, we predict that the lens system is likely to be a planet of mass m planet = 11.3 +19.2 -6.9 M ⨁ and a host star of mass M host =0.23 +0.39 -0.14 M ⨀ , located at a distance 𝐃 𝐋 =7.9 +1.0 -1.0 kpc. With a prior that holds that planet occurrence scales in proportion to the host-star mass, the estimated lens system properties are m planet =25 +22 -15 M ⨁ , M M host =0.53 +0.42 -0,32 M ⨀ , and D L =8.3 +0.9 -1.0 .This planet qualifies for inclusion in the extended MOA-II exoplanet microlens sample.

Gravitational microlensing↗

Summary and Annotated Bibliography of Measurement Error Corrections with Potential Application in Future Quesst Mission Community Noise Studies

This document is motivated by likely needs of the Quesst mission community response tests, which will culminate in data collection and estimation of dose-response regression relationships for consideration by domestic and international aviation regulators. Furthermore, basic research questions evaluating interactions between rates of community annoyance, dose levels, and indicators of the presence of rattle, vibration, and startle hinge on hypothesis testing in the context of regression models. For a variety of reasons, noise doses may be known only imprecisely and may not reflect the actual level experienced by responding subjects. These differences between true dose and estimated dose, be they systematic or random, constitute covariate measurement error. Available statistics literature speaks to the impacts of measurement error on regression models, both in terms of bias in estimated coefficients and predicted values, and in terms of the loss of statistical power for hypothesis testing. Given the particulars of a categorical annoyance response variable and a continuous noise dose predictor variable subject to measurement error during testing, the emphasis of this report is on findings and methods pertinent to generalized linear (and mixed) models likely to be employed during the Quesst mission community tests. We reach the following conclusions: 1. Of four reviewed methods, structural Bayesian measurement error models and simulation extrapolation (SIMEX) may be the most readily applicable to Quesst mission community noise study objectives. 2. If warranted, a linear measurement model can help model systematic sources of measurement error that the classical measurement error does not. 3. For its ready implementation and small additional input requirements, simulation extrapolation may be ideally suited for addressing secondary research questions involving interactions between annoyance, noise dose, and other factors through hypothesis testing. 4. For their flexibility and ability to propagate uncertainty, structural Bayesian hierarchical models have great appeal for mission purposes; some care may be needed in developing appropriate probability models describing actual noise exposure during testing. An annotated bibliography logs additional papers and resources that may be of value to analysts in other projects and disciplines.

Dose-Response Model↗

Chrono-Validation of Near-Real-Time Landslide Susceptibility Models via Plugin Statistical Simulations

The idea behind any validation scheme in landslide susceptibility studies is to test whether a model calibrated on a certain data can predict an unknown dataset of the same nature (landslide presences/absences and covariates). Almost the entirety of landslide susceptibility studies are validated by subsetting a single dataset into a training and test sets. This dataset usually corresponds either to event-specific or to historical inventories. Very rarely, a multi-temporal inventory is available and, in the few cases where this condition is met, the validation practices involve training a model on a specific landslide inventory, deriving a single predictive equation and validating it on a subsequent landslide inventory. This commonly leads landslide predictive studies, even those with a strong statistical rigor, to neglect the uncertainty estimation in their modeling scheme. In statistics, validation can also be performed via statistical simulations. This means that after fitting a given model, one can generate any number of predictive functions and test their predictive skills on any type and number of unknown datasets. In this work, we take a similar direction and we apply it to model and validate three separate co-seismic inventories, including an uncertainty estimation phase. We mapped these inventories within the same area in Indonesia, for three earthquakes occurred in 2012, 2017 and 2018. Specifically, we build three event-specific Bayesian Generalize Additive Models of the binomial family. From each model we then simulate 1000 predictive realizations over the remaining two inventories, by using a plug-in scheme where all the morphometric covariates are kept fixed and only the ground motion is replaced according to the prediction target. By doing so, we introduce a new analytical tool for near-real-time landslide predictive purposes, which is able to produce a probabilistic model which stands in between the definitions of susceptibility and hazard. In fact, our model is able to accurately estimate “where” and “when” - although not “how frequently” - landslide have occurred by featuring the multitemporal information of the trigger. In our findings, the simulations are quite similar to the fitted models; and the nine combinations we analyse produce excellent performance. This result confirms the assumption that “the past is the key to the future”, as we show that the relative contribution of each variable and their interactions in each probabilistic model remains practically the same across temporal replicates. This information is not trivial because it supports the routines implemented in global near-real-time applications.

Temporal validation↗

Optimal Estimation Framework for Ocean Color Atmospheric Correction and Pixel-level Uncertainty Quantification

Ocean color remote sensing requires compensation for atmospheric scattering and absorption (aerosol, Rayleigh, and trace gases), referred to as atmospheric correction (AC). AC allows inference of parameters such as spectrally resolved remote sensing reflectance ( R rs )(λ) ; sr 1 ) at the ocean surface from the top-of-atmosphere reflectance. Often, the uncertainty of this process is not fully explored. Bayesian inference techniques provide a simultaneous AC and uncertainty assessment via a full posterior distribution of the relevant variables, given the prior distribution of those variables and the radiative transfer (RT) likelihood function. Given uncertainties in the algorithm inputs, the Bayesian framework enables better constraints on the AC process by using the complete spectral information compared to traditional approaches that use only a subset of bands for AC. This paper investigates a Bayesian inference research method (Optimal Estimation, OE) for ocean color AC by simultaneously retrieving atmospheric and ocean properties using all visible and near-infrared spectral bands. The OE algorithm analytically approximates the posterior distribution of parameters based on normality assumptions and provides a potentially viable operational algorithm with a reduced computational expense. We developed a Neural Network (NN) RT forward model look-up-table-based emulator to increase algorithm efficiency further and thus speed up the likelihood computations. We then applied the OE algorithm to synthetic data and observations from the MODerate resolution Imaging Spectroradiometer (MODIS) on NASA’s Aqua spacecraft. We compared the R rs )(λ) retrieval and its uncertainty estimates from the OE method with in-situ validation data from the SeaWiFS Bio-optical Archive and Storage System (SeaBASS) and Aerosol Robotic Network Ocean Color (AERONET-OC) datasets. The OE algorithm improved R rs )(λ) estimates relative to the NASA standard operational algorithm by improving all statistical metrics at 443, 555, and 667 nm. Unphysical negative R rs )(λ) , which often appear in complex water conditions, was reduced by a factor of 3. The OE-derived pixel-level R rs )(λ) uncertainty estimates were also assessed relative to in-situ data and were shown to have skill.

Atmospheric correction↗

Compound estimation procedures in reliability

At NASA, components and subsystems of components in the Space Shuttle and Space Station generally go through a number of redesign stages. While data on failures for various design stages are sometimes available, the classical procedures for evaluating reliability only utilize the failure data on the present design stage of the component or subsystem. Often, few or no failures have been recorded on the present design stage. Previously, Bayesian estimators for the reliability of a single component, conditioned on the failure data for the present design, were developed. These new estimators permit NASA to evaluate the reliability, even when few or no failures have been recorded. Point estimates for the latter evaluation were not possible with the classical procedures. Since different design stages of a component (or subsystem) generally have a good deal in common, the development of new statistical procedures for evaluating the reliability, which consider the entire failure record for all design stages, has great intuitive appeal. A typical subsystem consists of a number of different components and each component has evolved through a number of redesign stages. The present investigations considered compound estimation procedures and related models. Such models permit the statistical consideration of all design stages of each component and thus incorporate all the available failure data to obtain estimates for the reliability of the present version of the component (or subsystem). A number of models were considered to estimate the reliability of a component conditioned on its total failure history from two design stages. It was determined that reliability estimators for the present design stage, conditioned on the complete failure history for two design stages have lower risk than the corresponding estimators conditioned only on the most recent design failure data. Several models were explored and preliminary models involving bivariate Poisson distribution and the Consael Process (a bivariate Poisson process) were developed. Possible short comings of the models are noted. An example is given to illustrate the procedures. These investigations are ongoing with the aim of developing estimators that extend to components (and subsystems) with three or more design stages.

Barnes, Ron↗

Spatiotemporal Associations Between Social Vulnerability, Environmental Measurements, and COVID-19 in the Conterminous United States

This study summarizes the results from fitting a Bayesian hierarchical spatiotemporal model to coronavirus disease 2019 (COVID-19) cases and deaths at the county level in the United States for the year 2020. Two models were created, one for cases and one for deaths, utilizing a scaled Besag, York, Mollié model with Type I spatial-temporal interaction. Each model accounts for 16 social vulnerability and 7 environmental variables as fixed effects. The spatial pattern between COVID-19 cases and deaths is significantly different in many ways. The spatiotemporal trend of the pandemic in the United States illustrates a shift out of many of the major metropolitan areas into the United States Southeast and Southwest during the summer months and into the upper Midwest beginning in autumn. Analysis of the major social vulnerability predictors of COVID-19 infection and death found that counties with higher percentages of those not having a high school diploma, having non-White status and being Age 65 and over to be significant. Among the environmental variables, above ground level temperature had the strongest effect on relative risk to both cases and deaths. Hot and cold spots, areas of statistically significant high and low COVID-19 cases and deaths respectively, derived from the convolutional spatial effect show that areas with a high probability of above average relative risk have significantly higher Social Vulnerability Index composite scores. The same analysis utilizing the spatiotemporal interaction term exemplifies a more complex relationship between social vulnerability, environmental measurements, COVID-19 cases, and COVID-19 deaths.

spatial epidemiology↗

Estimating the Properties of Hard X-Ray Solar Flares by Constraining Model Parameters

We wish to better constrain the properties of solar flares by exploring how parameterized models of solar flares interact with uncertainty estimation methods. We compare four different methods of calculating uncertainty estimates in fitting parameterized models to Ramaty High Energy Solar Spectroscopic Imager X-ray spectra, considering only statistical sources of error. Three of the four methods are based on estimating the scale-size of the minimum in a hypersurface formed by the weighted sum of the squares of the differences between the model fit and the data as a function of the fit parameters, and are implemented as commonly practiced. The fourth method is also based on the difference between the data and the model, but instead uses Bayesian data analysis and Markov chain Monte Carlo (MCMC) techniques to calculate an uncertainty estimate. Two flare spectra are modeled: one from the Geostationary Operational Environmental Satellite X1.3 class flare of 2005 January 19, and the other from the X4.8 flare of 2002 July 23.We find that the four methods give approximately the same uncertainty estimates for the 2005 January 19 spectral fit parameters, but lead to very different uncertainty estimates for the 2002 July 23 spectral fit. This is because each method implements different analyses of the hypersurface, yielding method-dependent results that can differ greatly depending on the shape of the hypersurface. The hypersurface arising from the 2005 January 19 analysis is consistent with a normal distribution; therefore, the assumptions behind the three non- Bayesian uncertainty estimation methods are satisfied and similar estimates are found. The 2002 July 23 analysis shows that the hypersurface is not consistent with a normal distribution, indicating that the assumptions behind the three non-Bayesian uncertainty estimation methods are not satisfied, leading to differing estimates of the uncertainty. We find that the shape of the hypersurface is crucial in understanding the output from each uncertainty estimation technique, and that a crucial factor determining the shape of hypersurface is the location of the low-energy cutoff relative to energies where the thermal emission dominates. The Bayesian/MCMC approach also allows us to provide detailed information on probable values of the low-energy cutoff, Ec, a crucial parameter in defining the energy content of the flare-accelerated electrons. We show that for the 2002 July 23 flare data, there is a 95% probability that Ec lies below approximately 40 keV, and a 68% probability that it lies in the range 7-36 keV. Further, the low-energy cutoff is more likely to be in the range 25-35 keV than in any other 10 keV wide energy range. The low-energy cutoff for the 2005 January 19 flare is more tightly constrained to 107 +/- 4 keV with 68% probability.

X-rays↗

Differences in the Optical Characteristics of Continental US Ground and Cloud Flashes as Observed from Space

Continental US lightning flashes observed by the Optical Transient Detector (OTD) are categorized according to flash type (ground or cloud flash) using US National Lightning Detection Network (TM) (NLDN) data. The statistics of the ground and cloud flash optical parameters (e.g., radiance, area, duration, number of optical groups, and number of optical events) are inter-compared. On average, the ground flash cloud-top emissions are more radiant, illuminate a larger area, are longer lasting, and have more optical groups and optical events than those cloud-top emissions associated with cloud flashes. Given these differences, it is suggested that the methods of Bayesian Inference could be used to help discriminate between ground and cloud flashes. The ability to discriminate flash type on-orbit is highly desired since such information would help researchers and operational decision makers better assess the intensification, evolutionary state, and severe weather potential of thunderstorms. This work supports risk reduction activities presently underway for the future launch of the GOES-R Geostationary Lightning Mapper (GLM).

Koshak, William↗