Search NASA⌕ Search

SEARCH · Search NASA

Results for “Nonparametric kernel”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

Nonparametric, data-based kernel interpolation for particle-tracking simulations and kernel density estimation

Traditional interpolation techniques for particle tracking include binning and convolutional formulas that use pre-determined (i.e., closed-form, parameteric) kernels. In many instances, the particles are introduced as point sources in time and space, so the cloud of particles (either in space or time) is a discrete representation of the Green’s function of an underlying PDE. As such, each particle is a sample from the Green’s function; therefore, each particle should be distributed according to the Green’s function. In short, the kernel of a convolutional interpolation of the particle sample “cloud” should be a replica of the cloud itself. This idea gives rise to an iterative method by which the form of the kernel may be discerned in the process of interpolating the Green’s function. When the Green’s function is a density, this method is broadly applicable to interpolating a kernel density estimate based on random data drawn from a single distribution. We formulate and construct the algorithm and demonstrate its ability to perform kernel density estimation of skewed and/or heavy-tailed data including breakthrough curves.

42 ENGINEERING↗

Smoothing Lexis diagrams using kernel functions: A contemporary approach

Lexis diagrams are rectangular arrays of event rates indexed by age and period. Analysis of Lexis diagrams is a cornerstone of cancer surveillance research. Typically, population-based descriptive studies analyze multiple Lexis diagrams defined by sex, tumor characteristics, race/ethnicity, geographic region, etc. Inevitably the amount of information per Lexis diminishes with increasing stratification. Several methods have been proposed to smooth observed Lexis diagrams up front to clarify salient patterns and improve summary estimates of averages, gradients, and trends. In this article, we develop a novel bivariate kernel-based smoother that incorporates two key innovations. First, for any given kernel, we calculate its singular values decomposition, and select an optimal truncation point—the number of leading singular vectors to retain—based on the bias-corrected Akaike information criterion. Second, we model-average over a panel of candidate kernels with diverse shapes and bandwidths. The truncated model averaging approach is fast, automatic, has excellent performance, and provides a variance-covariance matrix that takes model selection into account. We present an in-depth case study (invasive estrogen receptor-negative breast cancer incidence among non-Hispanic white women in the United States) and simulate operating characteristics for 20 representative cancers. The truncated model averaging approach consistently outperforms any fixed kernel. Our results support the routine use of the truncated model averaging approach in descriptive studies of cancer.

60 APPLIED LIFE SCIENCES↗

Advances in statistical methods for cancer surveillance research: an age-period-cohort perspective

Background: Analysis of Lexis diagrams (population-based cancer incidence and mortality rates indexed by age group and calendar period) requires specialized statistical methods. However, existing methods have limitations that can now be overcome using new approaches. Methods: We assembled a “toolbox” of novel methods to identify trends and patterns by age group, calendar period, and birth cohort. We evaluated operating characteristics across 152 cancer incidence Lexis diagrams compiled from United States (US) Surveillance, Epidemiology and End Results Program data for 21 leading cancers in men and women in four race and ethnicity groups (the “cancer incidence panel”). Results: Nonparametric singular values adaptive kernel filtration (SIFT) decreased the estimated root mean squared error by 90% across the cancer incidence panel. A novel method for semi-parametric age-period-cohort analysis (SAGE) provided optimally smoothed estimates of age-period-cohort (APC) estimable functions and stabilized estimates of lack-of-fit (LOF). SAGE identified statistically significant birth cohort effects across the entire cancer panel; LOF had little impact. As illustrated for colon cancer, newly developed methods for comparative age-period-cohort analysis can elucidate cancer heterogeneity that would otherwise be difficult or impossible to discern using standard methods. Conclusions: Cancer surveillance researchers can now identify fine-scale temporal signals with unprecedented accuracy and elucidate cancer heterogeneity with unprecedented specificity. Birth cohort effects are ubiquitous modulators of cancer incidence in the US. The novel methods described here can advance cancer surveillance research.

60 APPLIED LIFE SCIENCES↗

Fiber Uncertainty Visualization for Bivariate Data With Parametric and Nonparametric Noise Models

Visualization and analysis of multivariate data and their uncertainty are top research challenges in data visualization. Constructing fiber surfaces is a popular technique for multivariate data visualization that generalizes the idea of level-set visualization for univariate data to multivariate data. Here, in this paper, we present a statistical framework to quantify positional probabilities of fibers extracted from uncertain bivariate fields. Specifically, we extend the state-of-the-art Gaussian models of uncertainty for bivariate data to other parametric distributions (e.g., uniform and Epanechnikov) and more general nonparametric probability distributions (e.g., histograms and kernel density estimation) and derive corresponding spatial probabilities of fibers. In our proposed framework, we leverage Green's theorem for closed-form computation of fiber probabilities when bivariate data are assumed to have independent parametric and nonparametric noise. Additionally, we present a nonparametric approach combined with numerical integration to study the positional probability of fibers when bivariate data are assumed to have correlated noise. For uncertainty analysis, we visualize the derived probability volumes for fibers via volume rendering and extracting level sets based on probability thresholds. We present the utility of our proposed techniques via experiments on synthetic and simulation datasets.

97 MATHEMATICS AND COMPUTING↗

Online MCMC Thinning with Kernelized Stein Discrepancy

A fundamental challenge in Bayesian inference is efficient representation of a target distribution. Many nonparametric approaches do so by sampling a large number of points using variants of Markov chain Monte Carlo (MCMC). Here, we propose an MCMC variant that retains only those posterior samples which exceed a kernelized Stein discrepancy (KSD) threshold, which we call KSD thinning. We establish the convergence and complexity trade-offs for several settings of KSD thinning as a function of the KSD threshold parameter, sample size, and other problem parameters. We provide experimental comparisons against other online nonparametric Bayesian methods that generate low-complexity posterior representations. We observe superior consistency/complexity trade-offs across a range of settings including MCMC sampling on two Bayesian inference problems from the biological sciences, and 10 × inference speedup and storage reduction for Bayesian neural networks with no loss of accuracy and no increase in training time. Our code is available at https://github.com/colehawkins/KSD-Thinning.

Bayesian inference↗

Quasar Identification Using Multivariate Probability Density Estimated from Nonparametric Conditional Probabilities

Nonparametric estimation for a probability density function that describes multivariate data has typically been addressed by kernel density estimation (KDE). A novel density estimator recently developed by Farmer and Jacobs offers an alternative high-throughput automated approach to univariate nonparametric density estimation based on maximum entropy and order statistics, improving accuracy over univariate KDE. This article presents an extension of the single variable case to multiple variables. The univariate estimator is used to recursively calculate a product array of one-dimensional conditional probabilities. In combination with interpolation methods, a complete joint probability density estimate is generated for multiple variables. Good accuracy and speed performance in synthetic data are demonstrated by a numerical study using known distributions over a range of sample sizes from 100 to 10 6 for two to six variables. Performance in terms of speed and accuracy is compared to KDE. The multivariate density estimate developed here tends to perform better as the number of samples and/or variables increases. As an example application, measurements are analyzed over five filters of photometric data from the Sloan Digital Sky Survey Data Release 17. The multivariate estimation is used to form the basis for a binary classifier that distinguishes quasars from galaxies and stars with up to 94% accuracy.

79 ASTRONOMY AND ASTROPHYSICS↗

Regularized inversion of aerosol hygroscopic growth factor probability density function: application to humidity-controlled fast integrated mobility spectrometer measurements

Abstract. Aerosol hygroscopic growth plays an important role in atmospheric particle chemistry and the effects of aerosol on radiation and hence climate. The hygroscopic growth is often characterized by a growth factor probability density function (GF-PDF), where the growth factor is defined as the ratio of the particle size at a specified relative humidity to its dry size. Parametric, least-squares methods are the most widely used algorithms for inverting the GF-PDF from measurements of the humidified tandem differential mobility analyzer (HTDMA) and have been recently applied to the GF-PDF inversion from measurements of the humidity-controlled fast integrated mobility spectrometer (HFIMS). However, these least-squares methods suffer from noise amplification due to the lack of regularization in solving the ill-posed problem, resulting in significant fluctuations in the retrieved GF-PDF and even occasional failures of convergence. In this study, we introduce nonparametric, regularized methods to invert the aerosol GF-PDF and apply them to HFIMS measurements. Based on the HFIMS kernel function, the forward convolution is transformed into a matrix-based form, which facilitates the application of the nonparametric inversion methods with regularizations, including Tikhonov regularization and Twomey's iterative regularization. Inversions of the GF-PDF using the nonparameteric methods with regularization are demonstrated using HFIMS measurements simulated from representative GF-PDFs of ambient aerosols. The characteristics of reconstructed GF-PDFs resulting from different inversion methods, including previously developed least-squares methods, are quantitatively compared. The result shows that Twomey's method generally outperforms other inversion methods. The capabilities of Twomey's method in reconstructing the pre-defined GF-PDFs and recovering the mode parameters are validated.

54 ENVIRONMENTAL SCIENCES↗

Physics-Informed Gaussian Process Inference of Liquid Structure from Scattering Data

We present a nonparametric Bayesian framework to infer radial distribution functions from experimental scattering measurements with uncertainty quantification using nonstationary Gaussian processes. The Gaussian process prior mean and kernel functions are designed to mitigate well-known numerical challenges with the Fourier transform, including discrete measurement binning and detector windowing, while encoding fundamental yet minimal physical knowledge of the liquid structure. We demonstrate uncertainty propagation of the Gaussian process posterior to unmeasured quantities of interest. Experimental radial distribution functions of liquid argon and water with uncertainty quantification are provided as both a proof of principle for the method and a benchmark for molecular models.

Chemical structure↗

Code for multi-shape Gaussian process (GP) fitting with uncertainty quantification (UQ)

This is code associated with the publication “Nonparametric Multi-shape Modeling with Uncertainty Quantification,” authored by Hengrui Luo (Lawrence Berkeley National Laboratory) and Justin Strait (Los Alamos National Laboratory). The code is used to fit multiple-output Gaussian process (GP) models of planar closed curves to collections of ordered point sets, allowing for flexible nonlinear prediction of the underlying curve under dense or sparse point set samplings and with or without noise, as well as tractable uncertainty quantification. To do this, we employ use of a periodic kernel to account for the nonlinear input space of closed curves, and combine with coregionalization models to account for dependence both (i) between curve coordinates, and (ii) between pairs of curves. Functions in the code are capable of fitting these models, as well as performing additional tasks with the fitted curves such as (a) shape registration and alignment, (b) shape averaging, and (c) fitting for curve sub-populations / clusters.

Strait, Justin↗