Search NASA⌕ Search

SEARCH · Search NASA

Results for “Bayesian Statistics”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 127 records · Page 7

Why Do We Find Ourselves Around a Yellow Star Instead of a Red Star?

M-dwarf stars are more abundant than G-dwarf stars, so our position as observers on a planet orbiting a G-dwarf raises questions about the suitability of other stellar types for supporting life. If we consider ourselves as typical, in the anthropic sense that our environment is probably a typical one for conscious observers, then we are led to the conclusion that planets orbiting in the habitable zone of G-dwarf stars should be the best place for conscious life to develop. But such a conclusion neglects the possibility that K-dwarfs or M-dwarfs could provide more numerous sites for life to develop, both now and in the future. In this paper we analyse this problem through Bayesian inference to demonstrate that our occurrence around a G-dwarf might be a slight statistical anomaly, but only the sort of chance event that we expect to occur regularly. Even if M-dwarfs provide more numerous habitable planets today and in the future, we still expect mid G- to early K-dwarfs stars to be the most likely place for observers like ourselves. This suggests that observers with similar cognitive capabilities as us are most likely to be found at the present time and place, rather than in the future or around much smaller stars.

anthropic principle↗

A Guide to the Literature on Learning Graphical Models

This literature review discusses different methods under the general rubric of learning Bayesian networks from data, and more generally, learning probabilistic graphical models. Because many problems in artificial intelligence, statistics and neural networks can be represented as a probabilistic graphical model, this area provides a unifying perspective on learning. This paper organizes the research in this area along methodological lines of increasing complexity.

Buntine, Wray L.↗

The statistical analysis of circadian phase and amplitude in constant-routine core-temperature data

Accurate estimation of the phases and amplitude of the endogenous circadian pacemaker from constant-routine core-temperature series is crucial for making inferences about the properties of the human biological clock from data collected under this protocol. This paper presents a set of statistical methods based on a harmonic-regression-plus-correlated-noise model for estimating the phases and the amplitude of the endogenous circadian pacemaker from constant-routine core-temperature data. The methods include a Bayesian Monte Carlo procedure for computing the uncertainty in these circadian functions. We illustrate the techniques with a detailed study of a single subject's core-temperature series and describe their relationship to other statistical methods for circadian data analysis. In our laboratory, these methods have been successfully used to analyze more than 300 constant routines and provide a highly reliable means of extracting phase and amplitude information from core-temperature data.

NASA Discipline Regulatory Physiology↗

Incorporating Biological Knowledge into Evaluation of Casual Regulatory Hypothesis

Biological data can be scarce and costly to obtain. The small number of samples available typically limits statistical power and makes reliable inference of causal relations extremely difficult. However, we argue that statistical power can be increased substantially by incorporating prior knowledge and data from diverse sources. We present a Bayesian framework that combines information from different sources and we show empirically that this lets one make correct causal inferences with small sample sizes that otherwise would be impossible.

Chrisman, Lonnie↗

JSC Safety and Mission Assurance Data Analysis Overview

These slides describe the data analysis methods that are used to determine inputs for probabilistic risk models supporting the Space Shuttle Program. Other applications can follow a similar path probably using different data sources. Statistical approaches are different and not addressed here. Topics included here: 1) Prior Distribution; 2) Likelihood Data; 3) Bayesian Updating; and 4) Uncertainty and Error. Note: This is a high-level discussion and is not intended to be a tutorial.

Roelant, Henk↗

The Gemini Planet-Finding Campaign: The Frequency of Giant Planets Around Debris Disk Stars

We have completed a high-contrast direct imaging survey for giant planets around 57 debris disk stars as part of the Gemini NICI Planet-Finding Campaign. We achieved median H-band contrasts of 12.4 mag at 0farcs5 and 14.1 mag at 1'' separation. Follow-up observations of the 66 candidates with projected separation <500 AU show that all of them are background objects. To establish statistical constraints on the underlying giant planet population based on our imaging data, we have developed a new Bayesian formalism that incorporates (1) non-detections, (2) single-epoch candidates, (3) astrometric and (4) photometric information, and (5) the possibility of multiple planets per star to constrain the planet population. Our formalism allows us to include in our analysis the previously known β Pictoris and the HR 8799 planets. Our results show at 95% confidence that <13% of debris disk stars have a ≥5 M Jup planet beyond 80 AU, and <21% of debris disk stars have a ≥3 M Jup planet outside of 40 AU, based on hot-start evolutionary models. We model the population of directly imaged planets as d(sq.)N/dMda ∝ m(sub α) a(sup β), where m is planet mass and a is orbital semi-major axis (with a maximum value of a(sub max)). We find that β < -0.8 and/or α > 1.7. Likewise, we find that β < -0.8 and/or a(sub max) < 200 AU. For the case where the planet frequency rises sharply with mass (α > 1.7), this occurs because all the planets detected to date have masses above 5 M(sub Jup), but planets of lower mass could easily have been detected by our search. If we ignore the β Pic and HR 8799 planets (should they belong to a rare and distinct group), we find that <20% of debris disk stars have a ≥3 M(sub Jup) planet beyond 10 AU, and β < -0.8 and/or α < -1.5. Likewise, β < -0.8 and/or a(sub max) < 125 AU. Our Bayesian constraints are not strong enough to reveal any dependence of the planet frequency on stellar host mass. Studies of transition disks have suggested that about 20% of stars are undergoing planet formation; our non-detections at large separations show that planets with orbital separation >40 AU and planet masses >3 M(sub Jup) do not carve the central holes in these disks.

Bayesian formalism↗

Uncertainty Management for Diagnostics and Prognostics of Batteries using Bayesian Techniques

Uncertainty management has always been the key hurdle faced by diagnostics and prognostics algorithms. A Bayesian treatment of this problem provides an elegant and theoretically sound approach to the modern Condition- Based Maintenance (CBM)/Prognostic Health Management (PHM) paradigm. The application of the Bayesian techniques to regression and classification in the form of Relevance Vector Machine (RVM), and to state estimation as in Particle Filters (PF), provides a powerful tool to integrate the diagnosis and prognosis of battery health. The RVM, which is a Bayesian treatment of the Support Vector Machine (SVM), is used for model identification, while the PF framework uses the learnt model, statistical estimates of noise and anticipated operational conditions to provide estimates of remaining useful life (RUL) in the form of a probability density function (PDF). This type of prognostics generates a significant value addition to the management of any operation involving electrical systems.

Saha, Bhaskar↗

Uncertainty Estimates of Psychoacoustic Thresholds Obtained from Group Tests

Adaptive psychoacoustic test methods, in which the next signal level depends on the response to the previous signal, are the most efficient for determining psychoacoustic thresholds of individual subjects. In many tests conducted in the NASA psychoacoustic labs, the goal is to determine thresholds representative of the general population. To do this economically, non-adaptive testing methods are used in which three or four subjects are tested at the same time with predetermined signal levels. This approach requires us to identify techniques for assessing the uncertainty in resulting group-average psychoacoustic thresholds. In this presentation we examine the Delta Method of frequentist statistics, the Generalized Linear Model (GLM), the Nonparametric Bootstrap, a frequentist method, and Markov Chain Monte Carlo Posterior Estimation and a Bayesian approach. Each technique is exercised on a manufactured, theoretical dataset and then on datasets from two psychoacoustics facilities at NASA. The Delta Method is the simplest to implement and accurate for the cases studied. The GLM is found to be the least robust, and the Bootstrap takes the longest to calculate. The Bayesian Posterior Estimate is the most versatile technique examined because it allows the inclusion of prior information.

Rathsam, Jonathan↗

Confidence Intervals for Laboratory Sonic Boom Annoyance Tests

Commercial supersonic flight is currently forbidden over land because sonic booms have historically caused unacceptable annoyance levels in overflown communities. NASA is providing data and expertise to noise regulators as they consider relaxing the ban for future quiet supersonic aircraft. One deliverable NASA will provide is a predictive model for indoor annoyance to aid in setting an acceptable quiet sonic boom threshold. A laboratory study was conducted to determine how indoor vibrations caused by sonic booms affect annoyance judgments. The test method required finding the point of subjective equality (PSE) between sonic boom signals that cause vibrations and signals not causing vibrations played at various amplitudes. This presentation focuses on a few statistical techniques for estimating the interval around the PSE. The techniques examined are the Delta Method, Parametric and Nonparametric Bootstrapping, and Bayesian Posterior Estimation.

Rathsam, Jonathan↗

Rule groupings in expert systems using nearest neighbour decision rules, and convex hulls

Expert System shells are lacking in many areas of software engineering. Large rule based systems are not semantically comprehensible, difficult to debug, and impossible to modify or validate. Partitioning a set of rules found in CLIPS (C Language Integrated Production System) into groups of rules which reflect the underlying semantic subdomains of the problem, will address adequately the concerns stated above. Techniques are introduced to structure a CLIPS rule base into groups of rules that inherently have common semantic information. The concepts involved are imported from the field of A.I., Pattern Recognition, and Statistical Inference. Techniques focus on the areas of feature selection, classification, and a criteria of how 'good' the classification technique is, based on Bayesian Decision Theory. A variety of distance metrics are discussed for measuring the 'closeness' of CLIPS rules and various Nearest Neighbor classification algorithms are described based on the above metric.

Anastasiadis, Stergios↗

A Conceptual Approach to Assimilating Remote Sensing Data to Improve Soil Moisture Profile Estimates in a Surface Flux/Hydrology Model: Overview - Part 1

Knowledge of the amount of water in the soil is of great importance to many earth science disciplines. Soil moisture is a key variable in controlling the exchange of water and energy between the land surface and the atmosphere. Thus, soil moisture information is valuable in a wide range of applications including weather and climate, runoff potential and flood control, early warning of droughts, irrigation, crop yield forecasting, soil erosion, reservoir management, geotechnical engineering, and water quality. Despite the importance of soil moisture information, widespread and continuous measurements of soil moisture are not possible today. Although many earth surface conditions can be measured from satellites, we still cannot adequately measure soil moisture from space. Research in soil moisture remote sensing began in the mid 1970s shortly after the surge in satellite development. Recent advances in remote sensing have shown that soil moisture can be measured, at least qualitatively, by several methods. Quantitative measurements of moisture in the soil surface layer have been most successful using both passive and active microwave remote sensing, although complications arise from surface roughness and vegetation type and density. Early attempts to measure soil moisture from space-borne microwave instruments were hindered by what is now considered sub-optimal wavelengths (shorter than 5 cm) and the coarse spatial resolution of the measurements. L-band frequencies between 1 and 3 GHz (10-30 cm) have been deemed optimal for detection of soil moisture in the upper few centimeters of soil. The Electronically Steered Thinned Array Radiometer (ESTAR), an aircraft-based instrument operating a 1,4 GHz, has shown great promise for soil moisture determination. Initiatives are underway to develop a similar instrument for space. Existing space-borne synthetic aperture radars (SARS) operating at C- and L-band have also shown some potential to detect surface wetness. The advantage of radar is its much higher resolution than passive microwave systems, but it is currently hampered by surface roughness effects and the lack of a good algorithm based on a single frequency and single polarization. In addition, its repeat frequency is generally low (about 40 days). In the meantime, two new radiometers offer some hope for remote sensing of soil moisture from space. The Tropical Rainfall Measuring Mission (TRMM) Microwave Imager (TMI), launched in November 1997, possesses a 10.65 GHz channel and the Advanced Microwave Scanning Radiometer (AMSR) on both the ADEOS-11 and Earth Observing System AM-1 platforms to be launched in 1999 possesses a 6.9 GHz channel. Aside from issues about interference from vegetation, the coarse resolution of these data will provide considerable challenges pertaining to their application. The resolution of TMI is about 45 km and that of AMSR is about 70 km. These resolutions are grossly inconsistent with the scale of soil moisture processes and the spatial variability of factors that control soil moisture. Scale disparities such as these are forcing us to rethink how we assimilate data of various scales in hydrologic models. Of particular interest is how to assimilate soil moisture data by reconciling the scale disparity between what we can expect from present and future remote sensing measurements of soil moisture and modeling soil moisture processes. It is because of this disparity between the resolution of space-based sensors and the scale of data needed for capturing the spatial variability of soil moisture and related properties that remote sensing of soil moisture has not met with more widespread success. Within a single footprint of current sensors at the wavelengths optimal for this application, in most cases there is enormous heterogeneity in soil moisture created by differences in landcover, soils and topography, as well as variability in antecedent precipitation. It is difficult to interpret the meaning of 'mean' soil moisture under such conditions and even more difficult to apply such a value. Because of the non-linear relationships between near-surface soil moisture and other variables of interest, such as surface energy fluxes and runoff, mean soil moisture has little applicability at such large scales. It is for these reasons that the use of remote sensing in conjunction with a hydrologic model appears to be of benefit in capturing the complete spatial and temporal structure of soil moisture. This paper is Part I of a four-part series describing a method for intermittently assimilating remotely-sensed soil moisture information to improve performance of a distributed land surface hydrology model. The method, summarized in section II, involves the following components, each of which is detailed in the indicated section of the paper or subsequent papers in this series: Forward radiative transfer model methods (section II and Part IV); Use of a Kalman filter to assimilate remotely-sensed soil moisture estimates with the model profile (section II and Part IV); Application of a soil hydrology model to capture the continuous evolution of the soil moisture profile within and below the root zone (section III); Statistical aggregation techniques (section IV and Part II); Disaggregation techniques using a neural network approach (section IV and Part III); and Maximum likelihood and Bayesian algorithms for inversely solving for the soil moisture profile in the upper few cm (Part IV).

Crosson, William L.↗

Influence of Coronal Abundance Variations

The PI of this project was Jeff Scargle of NASA/Ames. Co-I's were Alma Connors of Eureka Scientific/Wellesley, and myself. Part of the work was subcontracted to Eureka Scientific via SAO, with Vinay Kashyap as PI. This project was originally assigned grant number NCC2-1206, and was later changed to NCC2-1350 for administrative reasons. The goal of the project was to obtain, derive, and develop statistical and data analysis tools that would be of use in the analyses of high-resolution, high-sensitivity data that are becoming available with new instruments. This is envisioned as a cross-disciplinary effort with a number of "collaborators" including some at SA0 (Aneta Siemiginowska, Peter Freeman) and at the Harvard Statistics department (David van Dyk, Rostislav Protassov, Xiao-li Meng, Epaminondas Sourlas, et al). We have developed a new tool to reliably measure the metallicities of thermal plasma. It is unfeasible to obtain high-resolution grating spectra for most stars, and one must make the best possible determination based on lower-resolution, CCD-type spectra. It has been noticed that most analyses of such spectra have resulted in measured metallicities that were significantly lower than when compared with analyses of high- resolution grating data where available (see, e.g., Brickhouse et al., 2000, ApJ 530,387). Such results have led to the proposal of the existence of so-called Metal Abundance Deficient, or "MAD" stars (e.g., Drake, J.J., 1996, Cool Stars 9, ASP Conf.Ser. 109, 203). We however find that much of these analyses may be systematically underestimating the metallicities, and using a newly developed method to correctly treat the low-counts regime at the high-energy tail of the stellar spectra (van Dyk et al. 2001, ApJ 548,224), have found that the metallicities of these stars are generally comparable to their photospheric values. The results were reported at the AAS (Sourlas, Yu, van Dyk, Kashyap, and Drake, 2000, BAAS 196, v32, #54.02), and at the conference on Statistical Challenges in Modem Astronomy (Sourlas, van Dyk, Kashyap, Drake, and Pease, 2003, SCMA 111, Eds. E.D.Feigelson, G.J.Babu, New York:Springer, p489-490). We also described the limitations of one of the most egregiously misused and misapplied statistical tests in astrophysical literature, the F-test for verifying model components (Protassov, van Dyk, Connors, Kashyap, and Siemiginowska, 2002, ApJ, 571,545). Indeed, a search through the ApJ archives turned up 170 papers in the 5 previous years that used the F-test explicitly in some form or the other, and with the vast majority of them not using it correctly! Indeed, looking at just 4 issues of the ApJ in 2001, we found 13 instances of its use, of which nine were demonstrably incorrect. Clearly, it is difficult to understate the importance of this issue. We also worked on speeding up Bayes Blocks and Sparse Bayes Blocks algorithms to make them more tractable for large searches. We also supported staistics students and postdocs in both explicit physics- model-based (spectra with tens of thousands of atomic lines) and "model-free" -- i.e. non-parametric or semi-parametric -- algorithms. Work on using more of the latter is just beginning; while using multi-scale methods for Poisson imaging has come to hition. In fact, "An Image Restoration Technique with Error Estimates", by D. Esch, A. Connors, M. Karovska, and D. van Dyk, was published by ApJ (Esch et a1.2004, ApJ, 610, 1213). The code has been delivered to M. Karovska for CXC; and is available for beta-testing upon request. The other large project we worked on was on the self-consistent modeling of logN-logs curves in the Poisson limit. logN-logs curves are a fundamental tool in the study of source populations, luminosity functions, and cosmological parameters. However, their determination is hampered by statistical effects such as the Eddington bias, incompleteness due to detection efficiency, faint source flux fluctuations, etc. We have develed a new and powerful method using the full Poisson machinery that allows us to model the logN-logs distribution of X-ray sources in a self-consistent manner. Because we properly account for all the above statistical effects, our modeling is valid over the full range of the data, and not just for strong sources, as is normally done. Using a Bayesian approach and modeling the fluxes with known functional forms such as simple or broken power-laws, and conditioning the expected photon counts on the fluxes, the background contamination, effective area, detector vignetting, and detection probability, we can delve deeply into the low counts regime and extend the usefulness of medium sensitivity surveys such as ChAMP by orders of magnitude. The built-in flexibility of the algorithm also allows a simultaneous analysis of multiple datasets. We have applied this analysis to a set a Chandra observations (Sourlas, Kashyap, Zezas, van Dyk, 2004, HEAD #8, #16.32)

Scargle, Jeffrey D.↗

Computational Inference of Vibratory System with Incomplete Modal Information Using Parallel, Interactive and Adaptive Markov Chains

Inverse analysis of vibratory system is an important subject in fault identification, model updating, and robust design and control. It is challenging subject because 1) the problem is oftentimes underdetermined while the measurements are limited and/or incomplete; 2) many combinations of parameters may yield results that are similar with respect to actual response measurements; and 3) uncertainties inevitably exist. The aim of this research is to leverage upon computational intelligence through statistical inference to facilitate an enhanced, probabilistic framework using incomplete modal response measurement. This new framework is built upon efficient inverse identification through optimization, whereas Bayesian inference is employed to account for the effect of uncertainties. To overcome the computational cost barrier, we adopt Markov chain Monte Carlo (MCMC) to characterize the target function/distribution. Instead of using single Markov chain in conventional Bayesian approach, we develop a new sampling theory with multiple parallel, interactive and adaptive Markov chains and incorporate it into Bayesian inference. This can harness the collective power of these Markov chains to realize the concurrent search of multiple local optima. The number of required Markov chains and their respective initial model parameters are automatically determined via Monte Carlo simulation-based sample pre-screening followed by K-means clustering analysis. These enhancements can effectively address the aforementioned challenges in finite element inverse analysis. The validity of this framework is systematically demonstrated through case studies.

K Zhou↗

Bayesian Framework For Bioburden Density Calculations To Perform Planetary Protection Probabilistic Risk Assessment

The planetary protection discipline aims to minimize the microbial contamination on spacecraft to prevent the inadvertent contamination of other planetary bodies, known as forward planetary protection (PP). Planetary protection probabilistic risk assessment (PRA) relies on two core methodologies-the contamination probability event tree analysis and statistical parameter estimation. Planetary protection engineers combine several techniques to estimate the bioburden present on spacecraft components. A direct assay to enumerate CFU (colony forming units) is the preferred methodology, but given a similar processing environment the bioburden present on certain components is inferred using: (1) a NASA defined bioburden estimate based upon the biological cleanliness of the manufacturing/assembly environment or (2) sampled data from a similar spacecraft component. The paper presents an empirical Bayesian framework to systematically treat bioburden estimation and its uncertainties on different levels starting with measurement procedures to combining different components to subsystems and whole spacecraft. It is shown that the Bayesian approach can effectively handle estimations and their uncertainties at different levels and produce a reliable estimate for bioburden to be used to evaluate the probability of contamination.

Seuylemezian, Arman↗

Radar-Based Bayesian Estimation of Ice Crystal Growth Parameters within a Microphysical Model

The potential for polarimetric Doppler radar measurements to improve predictions of ice microphysical processes within an idealized model–observational framework is examined. In an effort to more rigorously constrain ice growth processes (e.g., vapor deposition) with observations of natural clouds, a novel framework is developed to compare simulated and observed radar measurements, coupling a bulk adaptive-habit model of vapor growth to a polarimetric radar forward model. Bayesian inference on key microphysical model parameters is then used, via a Markov chain Monte Carlo sampler, to estimate the probability distribution of the model parameters. The statistical formalism of this method allows for robust estimates of the optimal parameter values, along with (non-Gaussian) estimates of their uncertainty. To demonstrate this framework, observations from Department of Energy radars in the Arctic during a case of pristine ice precipitation are used to constrain vapor deposition parameters in the adaptive habit model. The resulting parameter probability distributions provide physically plausible changes in ice particle density and aspect ratio during growth. A lack of direct constraint on the number concentration produces a range of possible mean particle sizes, with the mean size inversely correlated to number concentration. Consistency is found between the estimated inherent growth ratio and independent laboratory measurements, increasing confidence in the parameter PDFs and demonstrating the effectiveness of the radar measurements in constraining the parameters. The combined Doppler and polarimetric observations produce the highest-confidence estimates of the parameter PDFs, with the Doppler measurements providing a stronger constraint for this case.

Robert S. Schrom↗

Infusing Statistical Thinking into the NASA Quesst Community Test Campaign

Statistical thinking permeates many important decisions as NASA plans its Quesst mission, which will culminate in a series of community overflights using the X-59 aircraft to demonstrate low-noise supersonic flight. Month-long longitudinal surveys will be deployed to assess human perception and annoyance to this new acoustic phenomenon. NASA works with a large contractor team to develop systems and methodologies to estimate noise doses, to test and field socio-acoustic surveys, and to study the relationship between the two quantities, dose and response, through appropriate choices of statistical models. This latter dose-response relationship will serve as an important tool as national and international noise regulators debate whether overland supersonic flights could be permitted once again within permissible noise limits. In this presentation we highlight several areas where statistical thinking has come into play, including issues of sampling, classification and data fusion, and analysis of longitudinal survey data that are subject to rare events and the consequences of measurement error. We note several operational constraints that shape the appeal or feasibility of some decisions on statistical approaches, and we identify several important remaining questions to be addressed.

Bayesian model↗