Search NASASearch

SEARCH · Search NASA

Results for “Bayesian statistical modeling”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 55 records · Page 3

Surrogates for numerical simulations; optimization of eddy-promoter heat exchangers

Although the advent of fast and inexpensive parallel computers has rendered numerous previously intractable calculations feasible, many numerical simulations remain too resource-intensive to be directly inserted in engineering optimization efforts. An attractive alternative to direct insertion considers models for computational systems: the expensive simulation is evoked only to construct and validate a simplified, input-output model; this simplified input-output model then serves as a simulation surrogate in subsequent engineering optimization studies. A simple 'Bayesian-validated' statistical framework for the construction, validation, and purposive application of static computer simulation surrogates is presented. As an example, dissipation-transport optimization of laminar-flow eddy-promoter heat exchangers are considered: parallel spectral element Navier-Stokes calculations serve to construct and validate surrogates for the flowrate and Nusselt number; these surrogates then represent the originating Navier-Stokes equations in the ensuing design process.

Patera, Anthony T.

A Guide to the Literature on Learning Graphical Models

This literature review discusses different methods under the general rubric of learning Bayesian networks from data, and more generally, learning probabilistic graphical models. Because many problems in artificial intelligence, statistics and neural networks can be represented as a probabilistic graphical model, this area provides a unifying perspective on learning. This paper organizes the research in this area along methodological lines of increasing complexity.

Buntine, Wray L.

Studies in Astronomical Time Series Analysis. VI. Bayesian Block Representations

This paper addresses the problem of detecting and characterizing local variability in time series and other forms of sequential data. The goal is to identify and characterize statistically significant variations, at the same time suppressing the inevitable corrupting observational errors. We present a simple nonparametric modeling technique and an algorithm implementing it-an improved and generalized version of Bayesian Blocks [Scargle 1998]-that finds the optimal segmentation of the data in the observation interval. The structure of the algorithm allows it to be used in either a real-time trigger mode, or a retrospective mode. Maximum likelihood or marginal posterior functions to measure model fitness are presented for events, binned counts, and measurements at arbitrary times with known error distributions. Problems addressed include those connected with data gaps, variable exposure, extension to piece- wise linear and piecewise exponential representations, multivariate time series data, analysis of variance, data on the circle, other data modes, and dispersed data. Simulations provide evidence that the detection efficiency for weak signals is close to a theoretical asymptotic limit derived by [Arias-Castro, Donoho and Huo 2003]. In the spirit of Reproducible Research [Donoho et al. (2008)] all of the code and data necessary to reproduce all of the figures in this paper are included as auxiliary material.

signal detection

Separating Gravitational Wave Signals from Instrument Artifacts

Central to the gravitational wave detection problem is the challenge of separating features in the data produced by astrophysical sources from features produced by the detector. Matched filtering provides an optimal solution for Gaussian noise, but in practice, transient noise excursions or "glitches" complicate the analysis. Detector diagnostics and coincidence tests can be used to veto many glitches which may otherwise be misinterpreted as gravitational wave signals. The glitches that remain can lead to long tails in the matched filter search statistics and drive up the detection threshold. Here we describe a Bayesian approach that incorporates a more realistic model for the instrument noise allowing for fluctuating noise levels that vary independently across frequency bands, and deterministic "glitch fitting" using wavelets as "glitch templates", the number of which is determined by a trans-dimensional Markov chain Monte Carlo algorithm. We demonstrate the method's effectiveness on simulated data containing low amplitude gravitational wave signals from inspiraling binary black hole systems, and simulated non-stationary and non-Gaussian noise comprised of a Gaussian component with the standard LIGO/Virgo spectrum, and injected glitches of various amplitude, prevalence, and variety. Glitch fitting allows us to detect significantly weaker signals than standard techniques.

Littenberg, Tyson B.

Flood Hazard Assessment from Storm Tides, Rain and Sea Level Rise for a Tidal River Estuary

Cities and towns along the tidal Hudson River are highly vulnerable to flooding through the combination of storm tides and high streamflows, compounded by sea level rise. Here a three-dimensional hydrodynamic model, validated by comparing peak water levels for 76 historical storms, is applied in a probabilistic flood hazard assessment. In simulations, the model merges streamflows and storm tides from tropical cyclones (TCs), offshore extratropical cyclones (ETCs) and inland "wet extratropical" cyclones (WETCs). The climatology of possible ETC and WETC storm events is represented by historical events (1931-2013), and simulations include gauged streamflows and inferred ungauged streamflows (based on watershed area) for the Hudson River and its tributaries. The TC climatology is created using a stochastic statistical model to represent a wider range of storms than is contained in the historical record. TC streamflow hydrographs are simulated for tributaries spaced along the Hudson, modeled as a function of TC attributes (storm track, sea surface temperature, maximum wind speed) using a statistical Bayesian approach. Results show WETCs are important to flood risk in the upper tidal river (e.g., Albany, New York), ETCs are important in the estuary (e.g., New York City) and lower tidal river, and TCs are important at all locations due to their potential for both high surge and extreme rainfall. The raising of floods by sea level rise is shown to be reduced by approximately 30-60 percent at Albany due to the dominance of streamflow for flood risk. This can be explained with simple channel flow dynamics, in which increased depth throughout the river reduces frictional resistance, thereby reducing the water level slope and the upriver water level.

Tidal river

Uncertainty Management for Diagnostics and Prognostics of Batteries using Bayesian Techniques

Uncertainty management has always been the key hurdle faced by diagnostics and prognostics algorithms. A Bayesian treatment of this problem provides an elegant and theoretically sound approach to the modern Condition- Based Maintenance (CBM)/Prognostic Health Management (PHM) paradigm. The application of the Bayesian techniques to regression and classification in the form of Relevance Vector Machine (RVM), and to state estimation as in Particle Filters (PF), provides a powerful tool to integrate the diagnosis and prognosis of battery health. The RVM, which is a Bayesian treatment of the Support Vector Machine (SVM), is used for model identification, while the PF framework uses the learnt model, statistical estimates of noise and anticipated operational conditions to provide estimates of remaining useful life (RUL) in the form of a probability density function (PDF). This type of prognostics generates a significant value addition to the management of any operation involving electrical systems.

Saha, Bhaskar

The statistical analysis of circadian phase and amplitude in constant-routine core-temperature data

Accurate estimation of the phases and amplitude of the endogenous circadian pacemaker from constant-routine core-temperature series is crucial for making inferences about the properties of the human biological clock from data collected under this protocol. This paper presents a set of statistical methods based on a harmonic-regression-plus-correlated-noise model for estimating the phases and the amplitude of the endogenous circadian pacemaker from constant-routine core-temperature data. The methods include a Bayesian Monte Carlo procedure for computing the uncertainty in these circadian functions. We illustrate the techniques with a detailed study of a single subject's core-temperature series and describe their relationship to other statistical methods for circadian data analysis. In our laboratory, these methods have been successfully used to analyze more than 300 constant routines and provide a highly reliable means of extracting phase and amplitude information from core-temperature data.

NASA Discipline Regulatory Physiology

Bayesian Statistics and Uncertainty Quantification for Safety Boundary Analysis in Complex Systems

The analysis of a safety-critical system often requires detailed knowledge of safe regions and their highdimensional non-linear boundaries. We present a statistical approach to iteratively detect and characterize the boundaries, which are provided as parameterized shape candidates. Using methods from uncertainty quantification and active learning, we incrementally construct a statistical model from only few simulation runs and obtain statistically sound estimates of the shape parameters for safety boundaries.

Active Learning

Advanced Statistical Methods in Spacecraft Flight Software Cost Estimation: Bayesian Regression and Nonlinear Principal Components Analysis to Support System Engineering in the Early Project Lifecycle

This paper provides an overview of the new features and model updates in the upcoming release of the NASA Analogy Software Cost Tool (ASCoT). ASCoT, hosted within the Online NASA Space Estimation Tools (ONSET) on the One NASA Cost Engineering (ONCE) Database, is a web-based tool that provides a suite of estimation tools to support early lifecycle NASA flight software cost analysis. In addition to the traditional parametric flight software costing method COCOMO II, ASCoT contains a Bayesian linear regression to predict total flight software development cost as a function of total spacecraft cost, as well as four analogic methods: k-Nearest Neighbors (kNN) and Clustering models to predict Effort (in work-months) and total source lines of code (SLOC). These methods are designed to work primarily with system-level inputs such as mission type (orbiter, lander, etc.), mission destination (Earth, Inner Planetary, etc.), and the number of instruments and deployables. Nonlinear principal components analysis (NLPCA) is performed to find the principal features of the data composed of both categorical and numerical variables and is necessary prior to defining our analogic methods. Sensitivity analyses and in- and out-of-sample model performance results are presented for the Bayesian CER and the analogic models.

Johnson, James K.

Uncertainty Estimates of Psychoacoustic Thresholds Obtained from Group Tests

Adaptive psychoacoustic test methods, in which the next signal level depends on the response to the previous signal, are the most efficient for determining psychoacoustic thresholds of individual subjects. In many tests conducted in the NASA psychoacoustic labs, the goal is to determine thresholds representative of the general population. To do this economically, non-adaptive testing methods are used in which three or four subjects are tested at the same time with predetermined signal levels. This approach requires us to identify techniques for assessing the uncertainty in resulting group-average psychoacoustic thresholds. In this presentation we examine the Delta Method of frequentist statistics, the Generalized Linear Model (GLM), the Nonparametric Bootstrap, a frequentist method, and Markov Chain Monte Carlo Posterior Estimation and a Bayesian approach. Each technique is exercised on a manufactured, theoretical dataset and then on datasets from two psychoacoustics facilities at NASA. The Delta Method is the simplest to implement and accurate for the cases studied. The GLM is found to be the least robust, and the Bootstrap takes the longest to calculate. The Bayesian Posterior Estimate is the most versatile technique examined because it allows the inclusion of prior information.

Rathsam, Jonathan

Exploring the Connection Between Sampling Problems in Bayesian Inference and Statistical Mechanics

The Bayesian and statistical mechanical communities often share the same objective in their work - estimating and integrating probability distribution functions (pdfs) describing stochastic systems, models or processes. Frequently, these pdfs are complex functions of random variables exhibiting multiple, well separated local minima. Conventional strategies for sampling such pdfs are inefficient, sometimes leading to an apparent non-ergodic behavior. Several recently developed techniques for handling this problem have been successfully applied in statistical mechanics. In the multicanonical and Wang-Landau Monte Carlo (MC) methods, the correct pdfs are recovered from uniform sampling of the parameter space by iteratively establishing proper weighting factors connecting these distributions. Trivial generalizations allow for sampling from any chosen pdf. The closely related transition matrix method relies on estimating transition probabilities between different states. All these methods proved to generate estimates of pdfs with high statistical accuracy. In another MC technique, parallel tempering, several random walks, each corresponding to a different value of a parameter (e.g. "temperature"), are generated and occasionally exchanged using the Metropolis criterion. This method can be considered as a statistically correct version of simulated annealing. An alternative approach is to represent the set of independent variables as a Hamiltonian system. Considerab!e progress has been made in understanding how to ensure that the system obeys the equipartition theorem or, equivalently, that coupling between the variables is correctly described. Then a host of techniques developed for dynamical systems can be used. Among them, probably the most powerful is the Adaptive Biasing Force method, in which thermodynamic integration and biased sampling are combined to yield very efficient estimates of pdfs. The third class of methods deals with transitions between states described by rate constants. These problems are isomorphic with chemical kinetics problems. Recently, several efficient techniques for this purpose have been developed based on the approach originally proposed by Gillespie. Although the utility of the techniques mentioned above for Bayesian problems has not been determined, further research along these lines is warranted

Pohorille, Andrew

A Bayesian approach to microwave precipitation profile retrieval

A multichannel passive microwave precipitation retrieval algorithm is developed. Bayes theorem is used to combine statistical information from numerical cloud models with forward radiative transfer modeling. A multivariate lognormal prior probability distribution contains the covariance information about hydrometeor distribution that resolves the nonuniqueness inherent in the inversion process. Hydrometeor profiles are retrieved by maximizing the posterior probability density for each vector of observations. The hydrometeor profile retrieval method is tested with data from the Advanced Microwave Precipitation Radiometer (10, 19, 37, and 85 GHz) of convection over ocean and land in Florida. The CP-2 multiparameter radar data are used to verify the retrieved profiles. The results show that the method can retrieve approximate hydrometeor profiles, with larger errors over land than water. There is considerably greater accuracy in the retrieval of integrated hydrometeor contents than of profiles. Many of the retrieval errors are traced to problems with the cloud model microphysical information, and future improvements to the algorithm are suggested.

Evans, K. Franklin

A mathematical model of diurnal variations in human plasma melatonin levels

Studies in animals and humans suggest that the diurnal pattern in plasma melatonin levels is due to the hormone's rates of synthesis, circulatory infusion and clearance, circadian control of synthesis onset and offset, environmental lighting conditions, and error in the melatonin immunoassay. A two-dimensional linear differential equation model of the hormone is formulated and is used to analyze plasma melatonin levels in 18 normal healthy male subjects during a constant routine. Recently developed Bayesian statistical procedures are used to incorporate correctly the magnitude of the immunoassay error into the analysis. The estimated parameters [median (range)] were clearance half-life of 23.67 (14.79-59.93) min, synthesis onset time of 2206 (1940-0029), synthesis offset time of 0621 (0246-0817), and maximum N-acetyltransferase activity of 7.17(2.34-17.93) pmol x l(-1) x min(-1). All were in good agreement with values from previous reports. The difference between synthesis offset time and the phase of the core temperature minimum was 1 h 15 min (-4 h 38 min-2 h 43 min). The correlation between synthesis onset and the dim light melatonin onset was 0.93. Our model provides a more physiologically plausible estimate of the melatonin synthesis onset time than that given by the dim light melatonin onset and the first reliable means of estimating the phase of synthesis offset. Our analysis shows that the circadian and pharmacokinetics parameters of melatonin can be reliably estimated from a single model.

NASA Discipline Number 70-10

Bayesian Estimation of Earth’s Undiscovered Mineralogical Diversity Using Noninformative Priors

Recently, statistical distributions have been explored to provide estimates of the mineralogical diversity of Earth, and Earth-like planets. In this paper, a Bayesian approach is introduced to estimate Earth’s undiscovered mineralogical diversity. Samples are generated from a posterior distribution of the model parameters using Markov chain Monte Carlo simulations such that estimates and inference are directly obtained. It was previously shown that the mineral species frequency distribution conforms to a generalized inverse Gauss–Poisson (GIGP) large number of rare events model. Even though the model fit was good, the population size estimate obtained by using this model was found to be unreasonably low by mineralogists. In this paper, several zero-truncated, mixed Poisson distributions are fitted and compared, where the Poisson-lognormal distribution is found to provide the best fit. Subsequently, the population size estimates obtained by Bayesian methods are compared to the empirical Bayes estimates. Species accumulation curves are constructed and employed to estimate the population size as a function of sampling size. Finally, the relative abundances, and hence the occurrence probabilities of species in a random sample, are calculated numerically for all mineral species in Earth’s crust using the Poisson-lognormal distribution. These calculations are connected and compared to the calculations obtained in a previous paper using the GIGP model for which mineralogical criteria of an Earth-like planet were given.

Bayesian statistics

Structural model optimization using statistical evaluation

The results of research in applying statistical methods to the problem of structural dynamic system identification are presented. The study is in three parts: a review of previous approaches by other researchers, a development of various linear estimators which might find application, and the design and development of a computer program which uses a Bayesian estimator. The method is tried on two models and is successful where the predicted stiffness matrix is a proper model, e.g., a bending beam is represented by a bending model. Difficulties are encountered when the model concept varies. There is also evidence that nonlinearity must be handled properly to speed the convergence.

Collins, J. D.

Influence of Coronal Abundance Variations

The PI of this project was Jeff Scargle of NASA/Ames. Co-I's were Alma Connors of Eureka Scientific/Wellesley, and myself. Part of the work was subcontracted to Eureka Scientific via SAO, with Vinay Kashyap as PI. This project was originally assigned grant number NCC2-1206, and was later changed to NCC2-1350 for administrative reasons. The goal of the project was to obtain, derive, and develop statistical and data analysis tools that would be of use in the analyses of high-resolution, high-sensitivity data that are becoming available with new instruments. This is envisioned as a cross-disciplinary effort with a number of "collaborators" including some at SA0 (Aneta Siemiginowska, Peter Freeman) and at the Harvard Statistics department (David van Dyk, Rostislav Protassov, Xiao-li Meng, Epaminondas Sourlas, et al). We have developed a new tool to reliably measure the metallicities of thermal plasma. It is unfeasible to obtain high-resolution grating spectra for most stars, and one must make the best possible determination based on lower-resolution, CCD-type spectra. It has been noticed that most analyses of such spectra have resulted in measured metallicities that were significantly lower than when compared with analyses of high- resolution grating data where available (see, e.g., Brickhouse et al., 2000, ApJ 530,387). Such results have led to the proposal of the existence of so-called Metal Abundance Deficient, or "MAD" stars (e.g., Drake, J.J., 1996, Cool Stars 9, ASP Conf.Ser. 109, 203). We however find that much of these analyses may be systematically underestimating the metallicities, and using a newly developed method to correctly treat the low-counts regime at the high-energy tail of the stellar spectra (van Dyk et al. 2001, ApJ 548,224), have found that the metallicities of these stars are generally comparable to their photospheric values. The results were reported at the AAS (Sourlas, Yu, van Dyk, Kashyap, and Drake, 2000, BAAS 196, v32, #54.02), and at the conference on Statistical Challenges in Modem Astronomy (Sourlas, van Dyk, Kashyap, Drake, and Pease, 2003, SCMA 111, Eds. E.D.Feigelson, G.J.Babu, New York:Springer, p489-490). We also described the limitations of one of the most egregiously misused and misapplied statistical tests in astrophysical literature, the F-test for verifying model components (Protassov, van Dyk, Connors, Kashyap, and Siemiginowska, 2002, ApJ, 571,545). Indeed, a search through the ApJ archives turned up 170 papers in the 5 previous years that used the F-test explicitly in some form or the other, and with the vast majority of them not using it correctly! Indeed, looking at just 4 issues of the ApJ in 2001, we found 13 instances of its use, of which nine were demonstrably incorrect. Clearly, it is difficult to understate the importance of this issue. We also worked on speeding up Bayes Blocks and Sparse Bayes Blocks algorithms to make them more tractable for large searches. We also supported staistics students and postdocs in both explicit physics- model-based (spectra with tens of thousands of atomic lines) and "model-free" -- i.e. non-parametric or semi-parametric -- algorithms. Work on using more of the latter is just beginning; while using multi-scale methods for Poisson imaging has come to hition. In fact, "An Image Restoration Technique with Error Estimates", by D. Esch, A. Connors, M. Karovska, and D. van Dyk, was published by ApJ (Esch et a1.2004, ApJ, 610, 1213). The code has been delivered to M. Karovska for CXC; and is available for beta-testing upon request. The other large project we worked on was on the self-consistent modeling of logN-logs curves in the Poisson limit. logN-logs curves are a fundamental tool in the study of source populations, luminosity functions, and cosmological parameters. However, their determination is hampered by statistical effects such as the Eddington bias, incompleteness due to detection efficiency, faint source flux fluctuations, etc. We have develed a new and powerful method using the full Poisson machinery that allows us to model the logN-logs distribution of X-ray sources in a self-consistent manner. Because we properly account for all the above statistical effects, our modeling is valid over the full range of the data, and not just for strong sources, as is normally done. Using a Bayesian approach and modeling the fluxes with known functional forms such as simple or broken power-laws, and conditioning the expected photon counts on the fluxes, the background contamination, effective area, detector vignetting, and detection probability, we can delve deeply into the low counts regime and extend the usefulness of medium sensitivity surveys such as ChAMP by orders of magnitude. The built-in flexibility of the algorithm also allows a simultaneous analysis of multiple datasets. We have applied this analysis to a set a Chandra observations (Sourlas, Kashyap, Zezas, van Dyk, 2004, HEAD #8, #16.32)

Scargle, Jeffrey D.