Computing Sobol main effects indices with unstructured samples for discrete random variables and streaming data
Explore the source record for details and available documents.
SEARCH · Search NASA
Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.
Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.
Explore the source record for details and available documents.
We present a novel stochastic optimal control framework that accounts for various types of uncertainties, with application to reentry trajectory planning. The formulation of the optimal trajectory control problem is presented in the context of an indirect method where a functional objective associated with the terminal vehicle speed is to be minimized. Uncertain input parameters in the optimal trajectory control model, including aerodynamic parameters and initial and terminal conditions, are modeled as aleatory random variables, while the statistical parameters of these aleatory distributions are themselves random variables. The parametric and model uncertainties are simultaneously propagated through an extended polynomial chaos expansion (EPCE) formalism. Several metrics are described to evaluate response statistics and presented as insightful tools for robust decision making. Specifically, the response probability density function (PDF) reflecting influence of both epistemic and aleatory uncertainties is obtained. By sampling over the random variables representing model error, an ensemble of response PDFs is generated and the associated failure probability is estimated as a random variable with its own polynomial chaos expansion. Besides, the sensitivity index functions of response PDF with respect to the statistical parameters are evaluated. Coupling parametric and model uncertainties within the EPCE framework leads to a robust and efficient paradigm for multilevel uncertainty propagation and PDF characterization in general optimal control problems.
This paper reviews current theoretical and numerical approaches to optimization problems governed by partial differential equations (PDEs) that depend on random variables or random fields. Such problems arise in many engineering, science, economics and societal decision-making tasks. This paper focuses on problems in which the governing PDEs are parametrized by the random variables/fields, and the decisions are made at the beginning and are not revised once uncertainty is revealed. Examples of such problems are presented to motivate the topic of this paper, and to illustrate the impact of different ways to model uncertainty in the formulations of the optimization problem and their impact on the solution. A linear–quadratic elliptic optimal control problem is used to provide a detailed discussion of the set-up for the risk-neutral optimization problem formulation, study the existence and characterization of its solution, and survey numerical methods for computing it. Different ways to model uncertainty in the PDE-constrained optimization problem are surveyed in an abstract setting, including risk measures, distributionally robust optimization formulations, probabilistic functions and chance constraints, and stochastic orders. Furthermore, approximation-based optimization approaches and stochastic methods for the solution of the large-scale PDE-constrained optimization problems under uncertainty are described. Some possible future research directions are outlined.
Graphite is a quasi-brittle material, resulting in random variability in tensile strength distributions. To account for the random variability in strength, HHA-3000 of the ASME BPVC provides two semi-probabilistic methods for qualifying nuclear graphite components in the design stage, the simplified and full assessments. The full and simplified assessments apply statistical methods to engineering-based design problems. This is often referred to as reliability-based design. Reliability-based design (RBD) is a method to develop reliable designs by accounting for uncertainties and result in small chances of failure when also considering safety factors. RBDs provide reliability targets using semi-probabilistic approaches. RBD is implemented in ASME BPVC HHA-3000 for nuclear graphite components, but is not specific to that application. There has been much confusion around the methods implemented in ASME BPVC HHA-3000 for qualifying nuclear graphite components. To address the confusion, this paper takes a hierarchical approach. First, the general RBD framework is presented. Then, the semi-probabilistic methods and the underlying assumptions implemented in the assessments are presented. The semi-probabilistic methods are separated from the engineering modifications that have been made to the assessments. After building the framework and underlying assumptions, the specific methods in the full and simplified assessments are explained in three steps: inputs, methods, outputs. The methods are applied to an H-451 reflector block. Tensile strength properties for other graphite grades are provided.
Sensitivity analysis and reliability assessment are two important aspects of structural and system safety. Epistemic uncertainty with respect to probabilistic model of input parameters due to lack of knowledge is present in many scarce-data applications and complicates the characterization of uncertainty in model response. In this article, we present two importance measures to evaluate the impact of distribution parameters on the probability distribution function (PDF) of the output and the failure probability. The epistemic uncertainty associated with the distribution parameters is modeled as random variables. Additionally, a modified extended polynomial chaos expansion (MEPCE) approach is introduced in which aleatory and epistemic random variables are modeled and propagated simultaneously while allowing the separate assessment for any single epistemic variable. A MEPCE-based kernel density estimation (KDE) construction provides a composite map from each epistemic variable to the response PDF. The functional global sensitivity index of the PDF with respect to the distribution parameters is thus derived, as a function of output, which is both more informative and more efficient than standard scalar sensitivity measures. Reliability sensitivity indices can be readily evaluated by integrating the global sensitivity index function over the failure zone. Three illustrative examples are used to demonstrate the proposed methodology.
Herein we present a new modeling paradigm for optimization that we call random field optimization. Random fields are a powerful modeling abstraction that aims to capture the behavior of random variables that live on infinite-dimensional spaces (e.g., space and time) such as stochastic processes (e.g., time series, Gaussian processes, and Markov processes), random matrices, and random spatial fields. This paradigm involves sophisticated mathematical objects (e.g., stochastic differential equations and space-time kernel functions) and has been widely used in neuroscience, geoscience, physics, civil engineering, and computer graphics. Despite of this, however, random fields have seen limited use in optimization; specifically, existing optimization paradigms that involve uncertainty (e.g., stochastic programming and robust optimization) mostly focus on the use of finite random variables. This trend is rapidly changing with the advent of statistical optimization (e.g., Bayesian optimization) and multi-scale optimization (e.g., integration of molecular sciences and process engineering). Our work extends a recently-proposed abstraction for infinite-dimensional optimization problems by capturing more general uncertainty representations. Moreover, we discuss solution paradigms for this new class of problems based on finite transformations and sampling, and identify open questions and challenges.
We develop a method to estimate the sum of conditional means of a sequence of random variables given our access to only a subsequence by spot-checking. The method works with non-independent-and-identically-distributed (non-i.i.d.) random variables and can be applied for certifying ongoing quantum information tasks.
Stochastic modelling approaches are presented to capture random effects at multiple time and length scales. Random processes that occur at the microscale produce nondeterministic effects at the macroscale. Here we present three stochastic modeling approaches that describe random processes at microscopic length scales and map these processes to the macroscopic length scale. The first stochastic modeling approach is based upon a particle based numerical technique to solve a Stochastic Differential Equation (SDE) using an arbitrary diffusion process to capture random processes at the microstructural level. The second approach prescribes a Probability Density Function (PDF) for the drift and diffusion of the random variable derived using the forward and backward Kolmogorov equations. This method requires mean and drift evolution PDF transport equations. The third approach is the coupling of multiple random variables which are dependent on each other. The relationship of the PDFs and a coupling function, known as a copula, produces a Joint Probability Density Function (JPDF). These stochastic modeling approaches are implemented into a Multiple Component (MC) shock physics computational code and used to model statistical fracture and reactive flow applications.
Abstract Water vapor supersaturation in clouds is a random variable that drives activation and growth of cloud droplets. The Pi Convection–Cloud Chamber generates a turbulent cloud with a microphysical steady state that can be varied from clean to polluted by adjusting the aerosol injection rate. The supersaturation distribution and its moments, e.g., mean and variance, are investigated for varying cloud microphysical conditions. High-speed and collocated Eulerian measurements of temperature and water vapor concentration are combined to obtain the temporally resolved supersaturation distribution. This allows quantification of the contributions of variances and covariances between water vapor and temperature. Results are consistent with expectations for a convection chamber, with strong correlation between water vapor and temperature; departures from ideal behavior can be explained as resulting from dry regions on the warm boundary, analogous to entrainment. The saturation ratio distribution is measured under conditions that show monotonic increase of liquid water content and decrease of mean droplet diameter with increasing aerosol injection rate. The change in liquid water content is proportional to the change in water vapor concentration between no-cloud and cloudy conditions. Variability in the supersaturation remains even after cloud droplets are formed, and no significant buffering is observed. Results are interpreted in terms of a cloud microphysical Damköhler number (Da), under conditions corresponding to , i.e., the slow-microphysics regime. This implies that clouds with very clean regions, such that is satisfied, will experience supersaturation fluctuations without them being buffered by cloud droplet growth. Significance Statement The saturation ratio (humidity) in clouds controls the growth rate and formation of cloud droplets. When air in a turbulent cloud mixes, the humidity varies in space and time throughout the cloud. This is important because it means cloud droplets experience different growth histories, thereby resulting in broader size distributions. It is often assumed that growth and evaporation of cloud droplets buffers out some of the humidity variations. Measuring these variations has been difficult, especially in the field. The purpose of this study is to measure the saturation ratio distribution in clouds with a range of conditions. We measure the in-cloud saturation ratio using a convection cloud chamber with clean to polluted cloud properties. We found in clouds with low concentrations of droplets that the variations in the saturation ratio are not suppressed.
Quantum capacity, as the key figure of merit for a given quantum channel, upper bounds the channel's ability in transmitting quantum information. Identifying different types of channels, evaluating the corresponding quantum capacity, and finding the capacity-approaching coding scheme are the major tasks in quantum communication theory. Quantum channel in discrete variables has been discussed enormously based on various error models, while error model in the continuous variable channel has been less studied due to the infinite dimensional problem. In this paper, we investigate a general continuous variable quantum erasure channel. By defining an effective subspace of the continuous variable system, we find a continuous variable random coding model. We then derive the quantum capacity of the continuous variable erasure channel in the framework of decoupling theory. The discussion in this paper fills the gap of a quantum erasure channel in continuous variable setting and sheds light on the understanding of other types of continuous variable quantum channels.
The advantages offered by additive manufacturing over traditional processes has driven a great deal of industrial and academic interest in recent years. However, the process is relatively new and requires additional investigation to become sufficiently mature for wide scale industrial adoption. Electron beam melting powder bed fusion is one technology that has shown promise for fabricating high temperature resistant materials such as nickel based superalloys. The resulting microstructures typically exhibit a strong fiber texture in the build direction giving rise to anisotropic time-dependent deformation behavior. In order to accelerate the qualification of these materials for industrial adoption accurate numerical models are needed for simulating their behavior. In this work a crystal plasticity model including non-Schmid effects is presented for capturing creep anisotropy observed in additively manufactured IN738LC. The model is calibrated via a probabilistic framework where model parameters are treated as random variables. An iterative sequential design strategy is utilized to efficiently identify the probability density of the unknown model parameters. As a case study the model is utilized to investigate the behavior of randomly oriented equiaxed grain clusters sometimes observed embedded in the additively manufactured columnar structure. A synthetic realization is simulated and uncertainty is propagated through to the full-field response. Results indicate that these features are the source of significant creep relaxation and strain accumulation which partially explains observed grain boundary decohesion at these locations.
We consider a high-dimensional nonlinear computational model of a dynamical system, parameterized by a vector-valued control parameter, in the presence of uncertainties represented by an uncontrolled parameter modeled by a vector-valued random variable, and possibly with stochastic excitation. The objective is to construct a statistical surrogate model where the input is any deterministic value of the control parameter, and the output is a vector-valued observation of the computational model, which is a random vector whose probability measure is updated using a target dataset. To construct this statistical surrogate model, the stochastic response of the computational model must be built, which is a vector-valued time-discretized stochastic process in high dimension, depending on the control parameter. It is assumed that the computational cost of a single evaluation of the deterministic model is high. For the probabilistic updating, we consider a subset of the components of the observation of the computational model, defined as the “identification observation” of the computational model, for which a small target dataset is available. Therefore, the target dataset is associated with partial observability, corresponding to an incomplete data case. Given a prior probability model of the random control and uncontrolled parameters, a training dataset is constructed, consisting of realizations of the random triplet composed of the stochastic response, the random identification observation, and the random control parameter. Since the computational cost of a single evaluation of the deterministic model is assumed to be large, the training dataset is also of small size. The main challenges in this problem are the high dimensionality, partial observability leading to incomplete data in the target dataset for the identification observation of the computational model (which is not sufficient to identify the computational stochastic responses), and the availability of a small training dataset. To address these challenges, we propose a methodology based on statistical methods for constructing necessary reduced representations, direct probabilistic learning under constraints using probabilistic learning on manifolds (PLoM) constrained by the target dataset, and the use of a weak formulation of the Fourier transform of probability measures. Statistical conditioning is also employed to explore the learned dataset. The constructed predictive statistical surrogate model can be implemented in the context of online computation. Here, we apply this approach to a problem of nonlinear stochastic dynamics in high dimensions within the framework of deformable solids mechanics.
Learning Planar Ising Models is a software package written in Matlab for learning relationships among variable in a dataset using graphical models. The software package implements a generally-applicable algorithm for learning planar Ising models from any multivariate dataset. The code provides an algorithm for learning the best planar Ising model to approximate an arbitrary collection of binary random variables (possibly from sample data). Given the set of all pairwise correlations among variables, we select a planar graph and optimal planar Ising model defined on this graph to best approximate that set of correlations. The software includes demonstrations of the algorithm in simulations and for applications on publicly available datasets. Details of the algorithm, demonstration simulations, and applications are given in Johnson, et al; 2016. Reference: Johnson, J. K., Oyen, D., Chertkov, M., and Netrapalli, P. (2016). Learning planar Ising models. Journal of Machine Learning Research.
The Katz centrality of a node in a complex network is a measure of the node’s importance as far as the flow of information across the network is concerned. For ensembles of locally tree-like undirected random graphs, this observable is a random variable. Its full probability distribution is of interest but difficult to handle analytically because of its “global” character and its definition in terms of a matrix inverse. Leveraging a fast Gaussian Belief Propagation-Cavity algorithm to solve linear systems on tree-like structures, we show that i) the Katz centrality of a single instance can be computed recursively in a very fast way, and ii) the probability P ( K ) that a random node in the ensemble of undirected random graphs has centrality K satisfies a set of recursive distributional equations, which can be analytically characterized and efficiently solved using a population dynamics algorithm. We test our solution on ensembles of Erdős-Rényi and Scale Free networks in the locally tree-like regime, with excellent agreement. The analytical distribution of centrality for the configuration model conditioned on the degree of each node can be employed as a benchmark to identify nodes of empirical networks with over- and underexpressed centrality relative to a null baseline. We also provide an approximate formula based on a rank- 1 projection that works well if the network is not too sparse, and we argue that an extension of our method could be efficiently extended to tackle analytical distributions of other centrality measures such as PageRank for directed networks in a transparent and user-friendly way.
Abstract. Wind lidars are widespread and important tools in atmospheric observations. An intrinsic part of lidar measurement error is due to atmospheric variability in the remote-sensing scan volume. This study describes and quantifies the distribution of measurement error due to turbulence in varying atmospheric stability. While the lidar error model is general, we demonstrate the approach using large ensembles of virtual WindCube V2 lidar performing a profiling Doppler-beam-swinging scan in quasi-stationary large-eddy simulations (LESs) of convective and stable boundary layers. Error trends vary with the stability regime, time averaging of results, and observation height. A systematic analysis of the observation error explains dominant mechanisms and supports the findings of the empirical results. Treating the error under a random variable framework allows for informed predictions about the effect of different configurations or conditions on lidar performance. Convective conditions are most prone to large errors (up to 1.5 m s−1 in 1 Hz wind speed in strong convection), driven by the large vertical velocity variances in convective conditions and the high elevation angle of the scanning beams (62∘). Range-gate weighting induces a negative bias into the horizontal wind speeds near the surface shear layer (−0.2 m s−1 in the stable test case). Errors in the horizontal wind speed and direction computed from the wind components are sensitive to the background wind speed but have negligible dependence on the relative orientation of the instrument. Especially during low winds and in the presence of large errors in the horizontal velocity estimates, the reported wind speed is subject to a systematic positive bias (up to 0.4 m s−1 in 1 Hz measurements in strong convection). Vector time-averaged measurements can improve the behavior of the error distributions (reducing the 10 min wind speed error standard deviation to <0.3 m s−1 and the bias to <0.1 m s−1 in strong convection) with a predictable effectiveness related to the number of decorrelated samples in the time window. Hybrid schemes weighting the 10 min scalar- and vector-averaged lidar measurements are shown to be effective at reducing the wind speed biases compared to cup measurements in most of the simulated conditions, with time averages longer than 10 min recommended for best use in some unstable conditions. The approach in decomposing the error mechanisms with the help of the LES flow field could be extended to more complex measurement scenarios and scans.
SUMMARY We present a computationally efficient method to approximately propagate uncertainty when linearly inverting seismic data for point source, time variable moment tensor components. The method is based on the assumption that the data residual, given by the difference between the observed seismic data and the data predicated by a linear inversion, contains the effects of both data and model uncertainty. Our method uses a distribution of data residuals, added directly to the data, in a pseudo-Monte Carlo scheme. Using the assumption that the data residual is a stochastic process, we use the well-known Karhunen–Loève (KL) theorem to construct a distribution of data residuals, where the required basis functions are constructed using Fourier series. The Fourier series are scaled by a product of a random variable and the real-valued spectral amplitudes of the original data residual’s spectrum. Thus, the Fourier series and spectral amplitudes are eigenfunction-eigenvalue pairs used in the KL-based construction of data residual distribution. Using tests with synthetic data, we show that our method compares closely with a Finite Difference Monte Carlo (FDMC) method that we presented previously. More importantly, the method presented here is computationally several orders of magnitude faster than our previous FDMC method, and requires no a priori assumptions of model and/or data uncertainty.
The Johnson–Lindenstrauss (JL) lemma is a powerful tool for dimensionality reduction in modern algorithm design. The lemma states that any set of high-dimensional points in a Euclidean space can be projected into lower dimensions while approximately preserving pairwise Euclidean distances. Random matrices satisfying this lemma are called JL transforms (JLTs). Inspired by existing $s$-hashing JLTs with exactly $s$ nonzero elements on each column, the present work introduces an ensemble of sparse matrices encompassing so-called $s$-hashing-like matrices whose expected number of nonzero elements on each column is $s$. The independence of the sub-Gaussian entries of these matrices and the knowledge of their exact distribution play an important role in their analyses. Using properties of independent sub-Gaussian random variables, these matrices are demonstrated to be JLTs, and their smallest nontrivial singular values and largest singular values are estimated nonasymptotically using a technique from geometric functional analysis. As the dimensions of the matrix grow to infinity, these singular values are proved to converge almost surely to fixed quantities (by using the universal Bai–Yin law) and in distribution to the Gaussian orthogonal ensemble Tracy–Widom law after proper rescalings. Understanding the behaviors of extreme singular values is important in general because they are often used to define a measure of stability of matrix algorithms. For example, JLTs were recently used in derivative-free optimization algorithmic frameworks to select random subspaces in which are constructed random models or poll directions to achieve scalability, and hence estimating their smallest singular value in particular helps determine the dimension of these subspaces.
Understanding and accurately characterizing energy dissipation mechanisms in civil structures during earthquakes is an important element of seismic assessment and design. The most commonly used model is attributed to Rayleigh. This paper proposes a systematic approach to quantify the uncertainty associated with Rayleigh's damping model. Bayesian calibration with embedded model error is employed to treat the coefficients of the Rayleigh model as random variables using modal damping ratios. Through a numerical example, we illustrate how this approach works and how the calibrated model can address modeling uncertainty associated with the Rayleigh damping model.