Retrofitting Decision Tree Classifiers Using Kernel Density Estimation
A novel method for combining decision trees and kernel density estimators is proposed. Standard.
SEARCH · Search NASA
Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.
Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.
A novel method for combining decision trees and kernel density estimators is proposed. Standard.
Kernel type density estimators calculated by the method of sieves. Proofs are presented for the characterization theorem: Let x(1), x(2),...x(n) be a random sample from a population with density f(0). Let sigma 0 and consider estimators f of f(0) defined by (1).
We present a new method for density estimation based on Mercer kernels. The density estimate can be understood as the density induced on a data manifold by a mixture of Gaussians fit in a feature space. As is usual, the feature space and data manifold are defined with any suitable positive-definite kernel function. We modify the standard EM algorithm for mixtures of Gaussians to infer the parameters of the density. One benefit of the approach is it's conceptual simplicity, and uniform applicability over many different types of data. Preliminary results are presented for a number of simple problems.
The objective of this challenge is to develop a data-based probabilistic model of uncertainty to predict the behavior of subsystems (payloads) by themselves and while coupled to a primary (target) system. Although this type of analysis is routinely performed and representative of issues faced in real-world system design and integration, there are still several key technical challenges that must be addressed when analyzing uncertain interconnected systems. For example, one key technical challenge is related to the fact that there is limited data on target configurations. Moreover, it is typical to have multiple data sets from experiments conducted at the subsystem level, but often samples sizes are not sufficient to compute high confidence statistics. In this challenge problem additional constraints are placed as ground rules for the participants. One such rule is that mathematical models of the subsystem are limited to linear approximations of the nonlinear physics of the problem at hand. Also, participants are constrained to use these models and the multiple data sets to make predictions about the target system response under completely different input conditions. Our approach involved initially the screening of several different methods. Three of the ones considered are presented herein. The first one is based on the transformation of the modal data to an orthogonal space where the mean and covariance of the data are matched by the model. The other two approaches worked solutions in physical space where the uncertain parameter set is made of masses, stiffnesses and damping coefficients; one matches confidence intervals of low order moments of the statistics via optimization while the second one uses a Kernel density estimation approach. The paper will touch on all the approaches, lessons learned, validation 1 metrics and their comparison, data quantity restriction, and assumptions/limitations of each approach. Keywords: Probabilistic modeling, model validation, uncertainty quantification, kernel density
A non-intrusive uncertainty quantification method is applied to computational analysis of supersonic, low-boom aircraft. The mean and standard deviation statistics of the pressure waveforms and loudness metrics are evaluated through use of numerical quadrature. The probability density function (p.d.f.) of these outputs is evaluated via kernel density estimation. The simulations use an inviscid, embedded-boundary Cartesian-mesh flow solver in the nearfield combined with an augmented Burgers’ equation solver for propagation in the farfield. The results show that the p.d.f. of the waveform is bimodal at shocks, which makes the mean and standard deviation statistics inappropriate. Despite this limitation, we show that the moment statistics can provide effective assessment of discrepancies when comparing with experimental data. This is demonstrated by presenting uncertainty analysis of a wind-tunnel test and showing that we significantly improve the predictions when we include the test uncertainties in the simulation. Normal distributions are obtained for the ground signature and loudness metrics, which is primarily due to the careful shaping of the low-boom waveform. Separation of variables and error control are used to reduce computational cost. We demonstrate that this is an efficient approach in the sense of balancing numerical errors in the statistics quadrature with discretization errors in the solvers.
Most work on preference learning has focused on pairwise preferences or rankings over individual items. In this paper, we present a method for learning preferences over sets of items. Our learning method takes as input a collection of positive examples--that is, one or more sets that have been identified by a user as desirable. Kernel density estimation is used to estimate the value function for individual items, and the desired set diversity is estimated from the average set diversity observed in the collection. Since this is a new learning problem, we introduce a new evaluation methodology and evaluate the learning method on two data collections: synthetic blocks-world data and a new real-world music data collection that we have gathered.
The design, analysis, and verification and validation of a spacecraft relies heavily on Monte Carlo simulations. Modern computational techniques are able to generate large amounts of Monte Carlo data but flight dynamics engineers lack the time and resources to analyze it all. The growing amounts of data combined with the diminished available time of engineers motivates the need to automate the analysis process. Pattern recognition algorithms are an innovative way of analyzing flight dynamics data efficiently. They can search large data sets for specific patterns and highlight critical variables so analysts can focus their analysis efforts. This work combines a few tractable pattern recognition algorithms with basic flight dynamics concepts to build a practical analysis tool for Monte Carlo simulations. Current results show that this tool can quickly and automatically identify individual design parameters, and most importantly, specific combinations of parameters that should be avoided in order to prevent specific system failures. The current version uses a kernel density estimation algorithm and a sequential feature selection algorithm combined with a k-nearest neighbor classifier to find and rank important design parameters. This provides an increased level of confidence in the analysis and saves a significant amount of time.
We introduce a new tool to planetary geology for quantifying the spatial arrangement of vent fields and volcanic provinces using non parametric kernel density estimation. Unlike parametricmethods where spatial density, and thus the spatial arrangement of volcanic vents, is simplified to fit a standard statistical distribution, non parametric methods offer more objective and data driven techniques to characterize volcanic vent fields. This method is applied to Syria Planum volcanic vent catalog data as well as catalog data for a vent field south of Pavonis Mons. The spatial densities are compared to terrestrial volcanic fields.
We study the orbital architectures of planetary systems orbiting within ∼1 AU of their stars by analyzing the ensemble of Kepler systems having two or more planet candidates. We use data from the entire Kepler mission, and in many cases we apply improved analysis techniques (e.g., replacing histograms by top-hat Kernel Density Estimators that avoid the loss of information resulting from choosing a particular phase for the bin boundaries) to extend and enhance the studies of Lissauer et al. (2011, ApJS 197, 8) and Fabrycky et al. (2014, ApJ 790, 146).These data show ~ 1700 transiting planet candidates in > 600 multiple-planet systems, far more than were available for our previous two studies. The increased numbers and better information about planetary radii and the properties of stellar hosts made possible by Gaia DR2 allow more statistically-robust analyses of the entire ensemble of Kepler multis as well as independent analyses of subsets of the population. We are thus able to contrast the dynamical configurations of small and large planets, short-period and longer-period planets, and planets orbiting various types of host stars. We reinforce our previous findings that most pairs of planets within the same system are neither in nor near low-order mean motion resonances and that there is a substantial excess of planets having period ratios slightly larger than those of first-order mean-motion resonances. However, neglecting three systems whose planets are locked in 3- body resonances and summing over all first-order mean motion resonances, the deficit of planet pairs with period ratios just narrow of resonance is as large as the excess of planets wide of resonance (within statistical uncertainties), suggesting that overall there is no overall excess of planet pairs in the vicinity of resonance. Other aspects of our study, including estimates of the typical relative inclinations of planetary orbits and their variations as functions of orbital period, planet sizes and stellar properties, are in progress, with results expected to be available for presentation by the time of the conference.
PHD general-purpose classifier computer program. Uses Bayesian methods to classify vectors of real numbers, based on combination of statistical techniques that include multivariate density estimation, Parzen density kernels, and EM (Expectation Maximization) algorithm. By means of simple graphical interface, user trains classifier to recognize two or more classes of data and then use it to identify new data. Written in ANSI C for Unix systems and optimized for online classification applications. Embedded in another program, or runs by itself using simple graphical-user-interface. Online help files makes program easy to use.
Two classes of nonparametric density estimators, the histogram and the kernel estimator, both require a choice of smoothing parameter, or 'window width'. The optimum choice of this parameter is in general very difficult. An upper bound to the choices that depends only on the standard deviation of the distribution is described.
Two nonparametric probability density estimators are considered. The first is the kernel estimator. The problem of choosing the kernel scaling factor based solely on a random sample is addressed. An interactive mode is discussed and an algorithm proposed to choose the scaling factor automatically. The second nonparametric probability estimate uses penalty function techniques with the maximum likelihood criterion. A discrete maximum penalized likelihood estimator is proposed and is shown to be consistent in the mean square error. A numerical implementation technique for the discrete solution is discussed and examples displayed. An extensive simulation study compares the integrated mean square error of the discrete and kernel estimators. The robustness of the discrete estimator is demonstrated graphically.
Various instruments are used to create images of the Earth and other objects in the universe in a diverse set of wavelength bands with the aim of understanding natural phenomena. These instruments are sometimes built in a phased approach, with some measurement capabilities being added in later phases. In other cases, there may not be a planned increase in measurement capability, but technology may mature to the point that it offers new measurement capabilities that were not available before. In still other cases, detailed spectral measurements may be too costly to perform on a large sample. Thus, lower resolution instruments with lower associated cost may be used to take the majority of measurements. Higher resolution instruments, with a higher associated cost may be used to take only a small fraction of the measurements in a given area. Many applied science questions that are relevant to the remote sensing community need to be addressed by analyzing enormous amounts of data that were generated from instruments with disparate measurement capability. This paper addresses this problem by demonstrating methods to produce high accuracy estimates of spectra with an associated measure of uncertainty from data that is perhaps nonlinearly correlated with the spectra. In particular, we demonstrate multi-layer perceptrons (MLPs), Support Vector Machines (SVMs) with Radial Basis Function (RBF) kernels, and SVMs with Mixture Density Mercer Kernels (MDMK). We call this type of an estimator a Virtual Sensor because it predicts, with a measure of uncertainty, unmeasured spectral phenomena.
The paper summarizes observations of selected solar flares made with a far-UV spectroheliograph (190-465 A) and a UV spectrograph (900-1900 A) aboard Skylab. The emission lines used in the present analysis are identified, and three events are described in detail: the flare of June 15, 1973, a small subflare observed on August 9, 1973, and the flare of January 21, 1974. Ultraviolet images of two other events are also presented in an attempt to sketch a general picture of a flare as seen in this spectral region. It is found that a small kernel seems to be the source of the primary energy release of a flare. The size, electron density, and ion temperature of a typical kernel are estimated, and it is noted that hot clouds of coronal gas at 20 million K surrounded the observed kernels. It is speculated that flare kernels might be very thin channels through which high-energy particles, originating in deep layers, are ejected into the corona.
Abrupt increases in the rate of magnetic reversals (magnetic reversal spurts) were first studied by many others. They hypothesized that spurts result from increased turbulence in the earth's core dynamo during episodes of intense bolide bombardment of the earth. Mechanisms for creating episodes of intense bombardment of the earth involve gravitational perturbation of the Oort cloud of comets, either by a hidden planet, a solar companion, or massive matter in the galactic plane. Herein, the time variation in reversal rate is analyzed using methods of statistical density estimation. A smooth, continuous estimate of reversal rate is obtained using an adaptive kernel method, in which the kernel width is adjusted as a function of reversal rate. The estimates near the ends of the data series (at 165 my ago and the present) are obtained by extending the data by reflection. The results show that the reversal spurts are not associated demonstrably with extinctions or well-dated impacts. If the spurts do record episodes of intense bombardment of the earth, then the mass extinctions do not, in general, occur at times of impacts. Furthermore, the large impact craters seen are not obviously related to the spurts, suggesting that the craters may have been caused by bolides of a different nature and with a different temporal pattern. However, the most simple explanation seems to be that the spurts do not record comet showers, either because the recording mechanism suggested by Muller and Morris is not effective or because comet showers are not triggered in the ways considered by Hut et al.
The characteristics of coronal emplacements preceding solar flares were investigated based on a comprehensive survey of Skylab soft X-ray images. A search interval of 30 min before flare was used in the X-ray observations. X-ray images with preflare enhancements were compared with high resolution H-alpha images and photospheric magnetograms and preflare enhancements were found in a statistically significant number of the observed preflare intervals. The enhancement events consisted of loops, kernels, and sinuous features with one to three separate preflare structures appearing in each interval. Typical gas pressures in the preflare X-ray features were estimated on the order of a few dyne per sq cm and densities were 4-10 x 10 to the -9th per cu cm for assumed average temperatures. H-alpha brightenings in the form of knots and patches were found in conjunction with the X-ray preflare features in nearly all of the intervals. It is concluded that H-alpha emission is characteristic of preflare emission processes. The observational data are interpreted within the framework of existing loop preheating models, and the results are discussed in detail.
When it is known a priori exactly to which finite dimensional manifold the probability density function gives rise to a set of samples, the parametric maximum likelihood estimation procedure leads to poor estimates and is unstable; while the nonparametric maximum likelihood procedure is undefined. A very general theory of maximum penalized likelihood estimation which should avoid many of these difficulties is presented. It is demonstrated that each reproducing kernel Hilbert space leads, in a very natural way, to a maximum penalized likelihood estimator and that a well-known class of reproducing kernel Hilbert spaces gives polynomial splines as the nonparametric maximum penalized likelihood estimates.
Energy-independent first-flight transport kernels are evaluated for a spherical region with an R(-2) density distribution. The uncollided angular-flux distribution is obtained and integrated for a source distribution that is proportional to the density to give the uncollided emitted particle flux and current density. These are useful for the calculation of mass, energy, and momentum carried away by fast particles born in the medium. The data are relevant to estimate escape from weakly bound atmospheres such as comet comae, dilute circumstellar envelopes, and some unconfined laboratory plasmas.