Search NASASearch

SEARCH · Search NASA

Results for “Bayesian statistical modeling”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 109 records · Page 6

Bayesian Analysis of the Cosmic Microwave Background

There is a wealth of cosmological information encoded in the spatial power spectrum of temperature anisotropies of the cosmic microwave background! Experiments designed to map the microwave sky are returning a flood of data (time streams of instrument response as a beam is swept over the sky) at several different frequencies (from 30 to 900 GHz), all with different resolutions and noise properties. The resulting analysis challenge is to estimate, and quantify our uncertainty in, the spatial power spectrum of the cosmic microwave background given the complexities of "missing data", foreground emission, and complicated instrumental noise. Bayesian formulation of this problem allows consistent treatment of many complexities including complicated instrumental noise and foregrounds, and can be numerically implemented with Gibbs sampling. Gibbs sampling has now been validated as an efficient, statistically exact, and practically useful method for low-resolution (as demonstrated on WMAP 1 and 3 year temperature and polarization data). Continuing development for Planck - the goal is to exploit the unique capabilities of Gibbs sampling to directly propagate uncertainties in both foreground and instrument models to total uncertainty in cosmological parameters.

methods - statistical

Online Multi-Modal Learning and Adaptive Information Trajectory Planning for Autonomous Exploration

In robotic information gathering missions, scientists are typically interested in understanding variables which require proxy measurements from specialized sensor suites to estimate. However, energy and time constraints limit how often these sensors can be used in a mission. Robots are also equipped with cheaper to use navigation sensors such as cameras. In this paper, we explore a challenging planning problem in which a robot is required to learn about a scientific variable of interest in an initially unknown environment by planning informative paths and deciding when and where to use its sensors. To tackle this we present two innovations: a Bayesian generative model framework to automatically learn correlations between expensive science sensors and cheaper to use navigation sensors online, and a sampling based approach to plan for multiple sensors while handling long horizons and budget constraints. Our approach does not grow in complexity with data and is anytime making it highly applicable to field robotics. We tested our approach extensively in simulation and validated it with real data collected during the 2014 Mojave Volatiles Prospector Mission. Our planning algorithm performs statistically significantly better than myopic approaches and at least as well as a coverage-based algorithm in an initially unknown environment while having added advantages of being able to exploit prior knowledge and handle other intricacies of the real world without further algorithmic modifications.

learning

Profiles of Gamma-Ray Bursts and Their Component Pulses

One physically informative regularity of their otherwise heterogeneous ensemble, is that many Gamma-Ray Bursts consist of well defined pulses. To objectively quantify the temporal structure of BATSE bursts, we have developed an automatic modeling procedure that separates overlapping pulses and determines the energy-dependence of the pulse-shape parameters. No binning of photon arrival times is needed, so when applied to time-tagged events (TTE) the procedure captures variability information down to the shortest time scales present in the raw data. Maximizing the Bayesian likelihood function Pr(data/model) yields estimates of the model parameters, including the number of pulses present, and allows intercomparison of models of different forms. As with any nonlinear optimization, good initial guesses are crucial to avoid convergence to undesirable local minima. We find excellent initial pulse decompositions by wavelet-denoising a cumulative distribution of the raw photon arrival data; differentiation then gives a time profile mostly free of the systematic effects of degraded resolution (as in ordinary Fourier smoothing) and binning. We present statistical information on pulse rise-time, decay-time, peakedness, and amplitudes, plus their energy dependences - both within a single burst and for a large ensemble of bursts.

Scargle, Jeff D.

Confidence Intervals for Laboratory Sonic Boom Annoyance Tests

Commercial supersonic flight is currently forbidden over land because sonic booms have historically caused unacceptable annoyance levels in overflown communities. NASA is providing data and expertise to noise regulators as they consider relaxing the ban for future quiet supersonic aircraft. One deliverable NASA will provide is a predictive model for indoor annoyance to aid in setting an acceptable quiet sonic boom threshold. A laboratory study was conducted to determine how indoor vibrations caused by sonic booms affect annoyance judgments. The test method required finding the point of subjective equality (PSE) between sonic boom signals that cause vibrations and signals not causing vibrations played at various amplitudes. This presentation focuses on a few statistical techniques for estimating the interval around the PSE. The techniques examined are the Delta Method, Parametric and Nonparametric Bootstrapping, and Bayesian Posterior Estimation.

Rathsam, Jonathan

Optimal Estimation Framework for Ocean Color Atmospheric Correction and Pixel-level Uncertainty Quantification

Ocean color remote sensing requires compensation for atmospheric scattering and absorption (aerosol, Rayleigh, and trace gases), referred to as atmospheric correction (AC). AC allows inference of parameters such as spectrally resolved remote sensing reflectance ( R rs )(λ) ; sr 1 ) at the ocean surface from the top-of-atmosphere reflectance. Often, the uncertainty of this process is not fully explored. Bayesian inference techniques provide a simultaneous AC and uncertainty assessment via a full posterior distribution of the relevant variables, given the prior distribution of those variables and the radiative transfer (RT) likelihood function. Given uncertainties in the algorithm inputs, the Bayesian framework enables better constraints on the AC process by using the complete spectral information compared to traditional approaches that use only a subset of bands for AC. This paper investigates a Bayesian inference research method (Optimal Estimation, OE) for ocean color AC by simultaneously retrieving atmospheric and ocean properties using all visible and near-infrared spectral bands. The OE algorithm analytically approximates the posterior distribution of parameters based on normality assumptions and provides a potentially viable operational algorithm with a reduced computational expense. We developed a Neural Network (NN) RT forward model look-up-table-based emulator to increase algorithm efficiency further and thus speed up the likelihood computations. We then applied the OE algorithm to synthetic data and observations from the MODerate resolution Imaging Spectroradiometer (MODIS) on NASA’s Aqua spacecraft. We compared the R rs )(λ) retrieval and its uncertainty estimates from the OE method with in-situ validation data from the SeaWiFS Bio-optical Archive and Storage System (SeaBASS) and Aerosol Robotic Network Ocean Color (AERONET-OC) datasets. The OE algorithm improved R rs )(λ) estimates relative to the NASA standard operational algorithm by improving all statistical metrics at 443, 555, and 667 nm. Unphysical negative R rs )(λ) , which often appear in complex water conditions, was reduced by a factor of 3. The OE-derived pixel-level R rs )(λ) uncertainty estimates were also assessed relative to in-situ data and were shown to have skill.

Atmospheric correction

The Analysis of the Contribution of Human Factors to the In-Flight Loss of Control Accidents

In-flight loss of control (LOC) is currently the leading cause of fatal accidents based on various commercial aircraft accident statistics. As the Next Generation Air Transportation System (NextGen) emerges, new contributing factors leading to LOC are anticipated. The NASA Aviation Safety Program (AvSP), along with other aviation agencies and communities are actively developing safety products to mitigate the LOC risk. This paper discusses the approach used to construct a generic integrated LOC accident framework (LOCAF) model based on a detailed review of LOC accidents over the past two decades. The LOCAF model is comprised of causal factors from the domain of human factors, aircraft system component failures, and atmospheric environment. The multiple interdependent causal factors are expressed in an Object-Oriented Bayesian belief network. In addition to predicting the likelihood of LOC accident occurrence, the system-level integrated LOCAF model is able to evaluate the impact of new safety technology products developed in AvSP. This provides valuable information to decision makers in strategizing NASA's aviation safety technology portfolio. The focus of this paper is on the analysis of human causal factors in the model, including the contributions from flight crew and maintenance workers. The Human Factors Analysis and Classification System (HFACS) taxonomy was used to develop human related causal factors. The preliminary results from the baseline LOCAF model are also presented.

Ancel, Ersin

The Error Distribution of BATSE GRB Location

We develop empirical probability models for BATSE GRB location errors by a Bayesian analysis of the separations between BATSE GRB locations and locations obtained with the InterPlanetary Network (IPN). Models are compared and their parameters estimated using 394 GRBs with single IPN annuli and 20 GRBs with intersecting IPN annuli. Most of the analysis is for the 4B (rev) BATSE catalog; earlier catalogs are also analyzed. The simplest model that provides a good representation of the error distribution has 78% of the locations in a 'core' term with a systematic error of 1.85 degrees and the remainder in an extended tail with a systematic error of 5.36 degrees, implying a 68% confidence region for bursts with negligible statistical errors of 2.3 degrees. There is some evidence for a more complicated model in which the error distribution depends on the BATSE datatype that was used to obtain the location. Bright bursts are typically located using the CONT datatype, and according to the more complicated model, the 68% confidence region for CONT-located bursts with negligible statistical errors is 2.0 degrees.

Briggs, Michael S.

The Error Distribution of BATSE Gamma-Ray Burst Locations

Empirical probability models for BATSE gamma-ray burst (GRB) location errors are developed via a Bayesian analysis of the separations between BATSE GRB locations and locations obtained with the Interplanetary Network (IPN). Models are compared and their parameters estimated using 392 GRBs with single IPN annuli and 19 GRBs with intersecting IPN annuli. Most of the analysis is for the 4Br BATSE catalog; earlier catalogs are also analyzed. The simplest model that provides a good representation of the error distribution has 78% of the probability in a "core" term with a systematic error of 1.85 deg and the remainder in an extended tail with a systematic error of 5.1 deg, which implies a 68% confidence radius for bursts with negligible statistical uncertainties of 2.2 deg. There is evidence for a more complicated model in which the error distribution depends on the BATSE data type that was used to obtain the location. Bright bursts are typically located using the CONT data type, and according to the more complicated model, the 68% confidence radius for CONT-located bursts with negligible statistical uncertainties is 2.0 deg.

Briggs, Michael S.

Continuous Habitable Zones: Pairing a GCM and Bayesian Framework to Predict Habitable Zone Evolution

In the near-future, new space telescopes like JWST, LUVOIR, and HabEx will begin attempting to explore the properties of atmospheres of potentially habitable planets. This will require a significant amount of time and resources for even a single planet, which makes it essential to prioritize observations by those most-likely to have detectable life. Here we present a statistical method to estimate the probabilities that specific exoplanets have been continuously in the habitable zone of their host stars for more than 2 billion years, the approximate time it took life on Earth to significantly increase the oxygen content of the atmosphere. We introduce the use of statistics of an ensemble of 3D planetary general circulation models to estimate these probabilities, replacing prior 1D model estimates.

habitable planets

Pairing a GCM and Bayesian Framework to Predict Habitable Zone Evolution

In the near-future, new space telescopes like JWST, LUVOIR, and HabEx will begin attempting to explore the properties of atmospheres of potentially habitable planets. This will require a significant amount of time and resources for even a single planet, which makes it essential to prioritize observations by those most-likely to have detectable life. Here we present a statistical method to estimate the probabilities that specific exoplanets have been continuously in the habitable zone of their host stars for more than 2 billion years, the approximate time it took life on Earth to significantly increase the oxygen content of the atmosphere. We introduce the use of statistics of an ensemble of 3D planetary general circulation models to estimate these probabilities, replacing prior 1D model estimates.

habitable planets

Derivation of Failure Rates and Probability of Failures for the International Space Station Probabilistic Risk Assessment Study

National Aeronautics and Space Administration s (NASA) International Space Station (ISS) Program uses Probabilistic Risk Assessment (PRA) as part of its Continuous Risk Management Process. It is used as a decision and management support tool to not only quantify risk for specific conditions, but more importantly comparing different operational and management options to determine the lowest risk option and provide rationale for management decisions. This paper presents the derivation of the probability distributions used to quantify the failure rates and the probability of failures of the basic events employed in the PRA model of the ISS. The paper will show how a Bayesian approach was used with different sources of data including the actual ISS on orbit failures to enhance the confidence in results of the PRA. As time progresses and more meaningful data is gathered from on orbit failures, an increasingly accurate failure rate probability distribution for the basic events of the ISS PRA model can be obtained. The ISS PRA has been developed by mapping the ISS critical systems such as propulsion, thermal control, or power generation into event sequences diagrams and fault trees. The lowest level of indenture of the fault trees was the orbital replacement units (ORU). The ORU level was chosen consistently with the level of statistically meaningful data that could be obtained from the aerospace industry and from the experts in the field. For example, data was gathered for the solenoid valves present in the propulsion system of the ISS. However valves themselves are composed of parts and the individual failure of these parts was not accounted for in the PRA model. In other words the failure of a spring within a valve was considered a failure of the valve itself.

Vitali, Roberto

Assimilation of Microwave Observations in the Rainbands of Tropical Cyclones

We propose a novel Bayesian Monte Carlo Integration (BMCI) technique to retrieve the profiles of temperature, water vapor, and cloud liquid/ice water content from microwave cloudy measurements in the presence of tropical cyclones (TC). These retrievals then can either be directly used by meteorologists to analyze the structure of TCs or be assimilated into numerical models to provide accurate initial conditions for the NWP (Numerical Weather Prediction) models. The BMCI technique is applied to the data from the Advanced Technology Microwave Sounder (ATMS) onboard Suomi National Polar-orbiting Partnership (NPP) and Global Precipitation Measurement (GPM) Microwave Imager (GMI). The retrieved profiles are then assimilated into Hurricane WRF (Weather Research and Forecasting) using the GSI (Gridpoint Statistical Interpolation) data assimilation system.

Moradi, Isaac

The Gemini Planet-Finding Campaign: The Frequency of Giant Planets Around Debris Disk Stars

We have completed a high-contrast direct imaging survey for giant planets around 57 debris disk stars as part of the Gemini NICI Planet-Finding Campaign. We achieved median H-band contrasts of 12.4 mag at 0farcs5 and 14.1 mag at 1'' separation. Follow-up observations of the 66 candidates with projected separation <500 AU show that all of them are background objects. To establish statistical constraints on the underlying giant planet population based on our imaging data, we have developed a new Bayesian formalism that incorporates (1) non-detections, (2) single-epoch candidates, (3) astrometric and (4) photometric information, and (5) the possibility of multiple planets per star to constrain the planet population. Our formalism allows us to include in our analysis the previously known β Pictoris and the HR 8799 planets. Our results show at 95% confidence that <13% of debris disk stars have a ≥5 M Jup planet beyond 80 AU, and <21% of debris disk stars have a ≥3 M Jup planet outside of 40 AU, based on hot-start evolutionary models. We model the population of directly imaged planets as d(sq.)N/dMda ∝ m(sub α) a(sup β), where m is planet mass and a is orbital semi-major axis (with a maximum value of a(sub max)). We find that β < -0.8 and/or α > 1.7. Likewise, we find that β < -0.8 and/or a(sub max) < 200 AU. For the case where the planet frequency rises sharply with mass (α > 1.7), this occurs because all the planets detected to date have masses above 5 M(sub Jup), but planets of lower mass could easily have been detected by our search. If we ignore the β Pic and HR 8799 planets (should they belong to a rare and distinct group), we find that <20% of debris disk stars have a ≥3 M(sub Jup) planet beyond 10 AU, and β < -0.8 and/or α < -1.5. Likewise, β < -0.8 and/or a(sub max) < 125 AU. Our Bayesian constraints are not strong enough to reveal any dependence of the planet frequency on stellar host mass. Studies of transition disks have suggested that about 20% of stars are undergoing planet formation; our non-detections at large separations show that planets with orbital separation >40 AU and planet masses >3 M(sub Jup) do not carve the central holes in these disks.

Bayesian formalism

Application of a data-mining method based on Bayesian networks to lesion-deficit analysis

Although lesion-deficit analysis (LDA) has provided extensive information about structure-function associations in the human brain, LDA has suffered from the difficulties inherent to the analysis of spatial data, i.e., there are many more variables than subjects, and data may be difficult to model using standard distributions, such as the normal distribution. We herein describe a Bayesian method for LDA; this method is based on data-mining techniques that employ Bayesian networks to represent structure-function associations. These methods are computationally tractable, and can represent complex, nonlinear structure-function associations. When applied to the evaluation of data obtained from a study of the psychiatric sequelae of traumatic brain injury in children, this method generates a Bayesian network that demonstrates complex, nonlinear associations among lesions in the left caudate, right globus pallidus, right side of the corpus callosum, right caudate, and left thalamus, and subsequent development of attention-deficit hyperactivity disorder, confirming and extending our previous statistical analysis of these data. Furthermore, analysis of simulated data indicates that methods based on Bayesian networks may be more sensitive and specific for detecting associations among categorical variables than methods based on chi-square and Fisher exact statistics.

NASA Discipline Neuroscience

Comparison of Likelihood Methods for Generalized Linear Mixed Models with Application to Quiet Supersonic Flights 2018 Data

Repeated measurement will be a feature of the survey data collected during the Quesst missionX-59 community response tests (CRT). Since each participant will report his or her categorical level of annoyance in response to multiple events, the responses from any single individual may be correlated with one another. Several models within the class of generalized linear mixed models (GLMM) are pertinent to the analysis of correlated categorical outcomes; the random intercept logistic regression model is one example. Both Bayesian and frequentist methods for fitting these models are available, with frequentist methods relying on some form of approximation (of either an integral or the integrand) that appears in the marginal likelihood function. Given several anticipated similarities of the X-59 CRT data to data collected during a past risk reduction, Quiet Supersonic Flights 2018 (QSF18), this short note is intended to create awareness. It documents an instance in which a reported population average dose-response relationship derived from QSF18 single event data was distorted by the integral approximation applied in likelihood-based methods. We review some of the available literature on the topic, compare the outputs of several different computational approaches implemented in available statistical software, and present simple corrective actions that may be useful during the Quesst mission.

dose-response model

Compound estimation procedures in reliability

At NASA, components and subsystems of components in the Space Shuttle and Space Station generally go through a number of redesign stages. While data on failures for various design stages are sometimes available, the classical procedures for evaluating reliability only utilize the failure data on the present design stage of the component or subsystem. Often, few or no failures have been recorded on the present design stage. Previously, Bayesian estimators for the reliability of a single component, conditioned on the failure data for the present design, were developed. These new estimators permit NASA to evaluate the reliability, even when few or no failures have been recorded. Point estimates for the latter evaluation were not possible with the classical procedures. Since different design stages of a component (or subsystem) generally have a good deal in common, the development of new statistical procedures for evaluating the reliability, which consider the entire failure record for all design stages, has great intuitive appeal. A typical subsystem consists of a number of different components and each component has evolved through a number of redesign stages. The present investigations considered compound estimation procedures and related models. Such models permit the statistical consideration of all design stages of each component and thus incorporate all the available failure data to obtain estimates for the reliability of the present version of the component (or subsystem). A number of models were considered to estimate the reliability of a component conditioned on its total failure history from two design stages. It was determined that reliability estimators for the present design stage, conditioned on the complete failure history for two design stages have lower risk than the corresponding estimators conditioned only on the most recent design failure data. Several models were explored and preliminary models involving bivariate Poisson distribution and the Consael Process (a bivariate Poisson process) were developed. Possible short comings of the models are noted. An example is given to illustrate the procedures. These investigations are ongoing with the aim of developing estimators that extend to components (and subsystems) with three or more design stages.

Barnes, Ron

Statistical Symbolic Execution with Informed Sampling

Symbolic execution techniques have been proposed recently for the probabilistic analysis of programs. These techniques seek to quantify the likelihood of reaching program events of interest, e.g., assert violations. They have many promising applications but have scalability issues due to high computational demand. To address this challenge, we propose a statistical symbolic execution technique that performs Monte Carlo sampling of the symbolic program paths and uses the obtained information for Bayesian estimation and hypothesis testing with respect to the probability of reaching the target events. To speed up the convergence of the statistical analysis, we propose Informed Sampling, an iterative symbolic execution that first explores the paths that have high statistical significance, prunes them from the state space and guides the execution towards less likely paths. The technique combines Bayesian estimation with a partial exact analysis for the pruned paths leading to provably improved convergence of the statistical analysis. We have implemented statistical symbolic execution with in- formed sampling in the Symbolic PathFinder tool. We show experimentally that the informed sampling obtains more precise results and converges faster than a purely statistical analysis and may also be more efficient than an exact symbolic analysis. When the latter does not terminate symbolic execution with informed sampling can give meaningful results under the same time and memory limits.

Reliability

Topics in inference and decision-making with partial knowledge

Two essential elements needed in the process of inference and decision-making are prior probabilities and likelihood functions. When both of these components are known accurately and precisely, the Bayesian approach provides a consistent and coherent solution to the problems of inference and decision-making. In many situations, however, either one or both of the above components may not be known, or at least may not be known precisely. This problem of partial knowledge about prior probabilities and likelihood functions is addressed. There are at least two ways to cope with this lack of precise knowledge: robust methods, and interval-valued methods. First, ways of modeling imprecision and indeterminacies in prior probabilities and likelihood functions are examined; then how imprecision in the above components carries over to the posterior probabilities is examined. Finally, the problem of decision making with imprecise posterior probabilities and the consequences of such actions are addressed. Application areas where the above problems may occur are in statistical pattern recognition problems, for example, the problem of classification of high-dimensional multispectral remote sensing image data.

Safavian, S. Rasoul