Search NASA⌕ Search

SEARCH · Search NASA

Results for “Kernel Density Estimation (KDE)”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

The effective number of parameters in kernel density estimation

We devise a new formula for measuring the effective degrees of freedom (EDoF) in kernel density estimation (KDE). Starting from the orthogonal polynomial sequence (OPS) expansion for the ratio of the empirical to the oracle density, we show how convolution with the kernel leads to a new OPS with respect to which one may express the resulting KDE. The expansion coefficients of the two OPS systems can then be related via a kernel sensitivity matrix, which leads to a natural oracle definition of EDoF through the trace operator. Asymptotic properties of the (empirical) plug-in EDoF are worked out through influence functions, and connections with other empirical EDoFs are established. Minimization of Kullback-Leibler divergence is investigated as an alternative to integrated squared error based bandwidth selection rules, yielding a new normal scale rule. The methodology, which arises from a proper oracle formulation and is not restricted to convolution kernels, suggests the possibility of a new bandwidth selection rule based on an information criterion such as AIC.

bandwidth selection↗

A functional global sensitivity measure and efficient reliability sensitivity analysis with respect to statistical parameters

Sensitivity analysis and reliability assessment are two important aspects of structural and system safety. Epistemic uncertainty with respect to probabilistic model of input parameters due to lack of knowledge is present in many scarce-data applications and complicates the characterization of uncertainty in model response. In this article, we present two importance measures to evaluate the impact of distribution parameters on the probability distribution function (PDF) of the output and the failure probability. The epistemic uncertainty associated with the distribution parameters is modeled as random variables. Additionally, a modified extended polynomial chaos expansion (MEPCE) approach is introduced in which aleatory and epistemic random variables are modeled and propagated simultaneously while allowing the separate assessment for any single epistemic variable. A MEPCE-based kernel density estimation (KDE) construction provides a composite map from each epistemic variable to the response PDF. The functional global sensitivity index of the PDF with respect to the distribution parameters is thus derived, as a function of output, which is both more informative and more efficient than standard scalar sensitivity measures. Reliability sensitivity indices can be readily evaluated by integrating the global sensitivity index function over the failure zone. Three illustrative examples are used to demonstrate the proposed methodology.

42 ENGINEERING↗

Next-Cycle Optimal Dilute Combustion Control via Online Learning of Cycle-to-Cycle Variability Using Kernel Density Estimators

Dilute combustion using exhaust gas recirculation (EGR) presents a cost-effective method for increasing the efficiency of spark-ignition (SI) engines. However, the maximum amount of EGR that can be used at a given condition is limited by a rapid increment of cycle-to-cycle variability (CCV). This study describes a methodology to design a model-based stochastic optimal controller to adjust the cycle-to-cycle fuel injection quantity in order to reduce CCV and further extend the dilute limit. Given the complexity and chaotic nature of combustion events, the controller was enhanced with online learning in order to identify the statistical properties of combustion efficiency, which are needed to generate predictions for next-cycle events. This study showed that a kernel density estimator (KDE) can be used to learn the combustion properties in real time and can be incorporated into the feedback policy in order to calculate the optimal control command. Experimental results suggested that the dilute limit can be extended from 18.5% to 21% EGR fraction at an operating condition relevant for highway cruising. Additionally, the proposed controller can achieve a large CCV reduction with less fuel enrichment compared to previous methods, overall contributing to an increase in 0.2% indicated fuel conversion efficiency.

33 ADVANCED PROPULSION SYSTEMS↗

A Multi-Fidelity Gaussian Process Regression Method for Probabilistic Wind Farm Power Curve Estimation

Accurate estimation of the power curve for wind turbines or wind farms is crucial to ensure their efficient operation and management. However, conventional methods for power curve estimation rely either on expensive and infrequent measurements or on low-quality numerical simulations. Moreover, the majority of previous studies on power curve estimation for wind turbines or wind farms focused on deterministic estimation, which provides a point estimate of the relationship between wind speed and power generation. Nevertheless, the deterministic approach fails to consider the inherent uncertainty associated with wind energy production resulting from varying turbine characteristics. This can lead to inaccurate power generation estimation and suboptimal decisions regarding energy management. In this paper, a kernel density estimation (KDE) based Multi-Fidelity Gaussian Process Regression (MFGPR) model is proposed to fuse theoretical power curve data and the ground true measurements to create a mapping of wind speed and wind power. By conducting a case study on an actual wind farm in China, the efficacy of the proposed MFGPR model was demonstrated in characterizing the variability of wind power. The probabilistic MFGPR model was also able to generate confidence intervals that encompassed the measured power, thereby improving the accuracy and confidence in wind power estimation or wind resource assessment. Overall, the proposed MFGPR model offers a reliable approach to integrate high-fidelity ground measurements and theoretical power curve data, resulting in precise wind resource assessment and power estimation.

Gaussian process regression↗

Efficient screening of rare large pit anomalies on polished surfaces using a minimalist sampling scheme

Lawrence Livermore National Laboratory (LLNL) has made significant strides in generating clean energy through its inertial confinement fusion (ICF) experiments. These experiments rely on high-density carbon (HDC) coated shells to encapsulate the fusion fuel. The success of these experiments is heavily dependent on the surface quality of these shells, as even minor imperfections, such as deep pits, can negatively impact fusion yield. Ensuring the required smoothness involves an extensive surface-finishing process that spans approximately 20 stages, making it both time-intensive and resource-demanding. A critical challenge in this process is the need for high-resolution scans to detect rare deep pits, which can be costly and impractical if performed on every shell. This highlights the necessity of developing more efficient scanning methods to optimize time and cost without compromising accuracy. To address these challenges, we introduce a novel approach that employs the multivariate Dvoretzky–Kiefer–Wolfowitz (DKW) inequality to provide a probabilistic upper bound on the error in estimating pit distribution characteristics via a Kernel Density Estimator (KDE). This error bound enables efficient and reliable estimation of pit distribution characteristics at a specified statistical confidence level using a minimal number of surface scans. The integrated DKW-KDE approach was validated through surface-finishing experiments across two batches of HDC-coated shells, demonstrating consistent and robust performance across multiple stages of the surface-finishing experiments. The validation studies suggest that the integrated DKW-KDE approach achieves comparable accuracy in estimating the risk of deleterious large pits with six scans, thus conserving time and resources. Further evaluations show that performance remains consistent across batches and over multiple polishing stages. In conclusion, based on these findings, one can leverage the minimal-scan insights to strategically improve the bottleneck inspection process, thus enhancing the productivity and quality of shell polishing and similar challenging manufacturing processes.

Inertial confinement fusion↗

Stochastic multiscale modeling for quantifying statistical and model errors with application to composite materials

This paper provides a coherent and efficient computational framework for stochastic multiscale analysis of material systems in the presence of parametric uncertainties and modeling errors. Uncertainty in those model parameters that are not deduced as upscaled quantities is attributed to an uncertainty “germ”. While such parameters can appear at any scale, they are predominant at the finest analysis scale. Additional uncertainties stemming from statistical estimation, attributed to lack of data and model error, are associated with each submodel contributing to the multiscale system. Here, a robust and efficient framework based on a generalized extended polynomial chaos expansion (gEPCE) is proposed to simultaneously propagate all these uncertainties in order to provide a probabilistic representation of specific quantities of interest (QoI). We characterize the full probability distribution of the QoI and the uncertainty in the failure probability pertaining to its tails. By combining gEPCE with kernel density estimation (KDE) and directional derivatives, we construct sensitivity measures that connect these statistical metrics of QoI to the various sources of uncertainty to assess their individual and combined impacts. An illustrative problem featuring three-point bending of a composite beam is investigated to demonstrate the presented approach.

36 MATERIALS SCIENCE↗

GalaxyFlow: upsampling hydrodynamical simulations for realistic mock stellar catalogues

ABSTRACT Cosmological N-body simulations of galaxies operate at the level of ‘star particles’ with a mass resolution on the scale of thousands of solar masses. Turning these simulations into stellar mock catalogues requires ‘upsampling’ the star particles into individual stars following the same phase-space density. In this paper, we introduce two new upsampling methods. First, we describe GalaxyFlow, a sophisticated upsampling method that utilizes normalizing flows to both estimate the stellar phase-space density and sample from it. Secondly, we improve on existing upsamplers based on adaptive kernel density estimation (KDE), using maximum likelihood estimation to fine-tune the bandwidth for such algorithms in a way that improves both the density estimation accuracy and upsampling results. We demonstrate our upsampling techniques on a neighbourhood of the Solar location in two simulated galaxies: Auriga 6 and h277. Both yield smooth stellar distributions that closely resemble the stellar densities seen in the Gaia DR3 catalogue. Furthermore, we introduce a novel multimodel classifier test to compare the accuracy of different upsampling methods quantitatively. This test confirms that GalaxyFlow more accurately estimates the density of the underlying star particles than methods based on KDE, at the cost of being more computationally intensive.

Lim, Sung Hak (ORCID:0000000330981092)↗

Detecting Anomalies in Time Series Using Kernel Density Approaches

This paper introduces a novel anomaly detection approach tailored for time series data with exclusive reliance on normal events during training. Our key innovation lies in the application of kernel-density estimation (KDE) to scrutinize reconstruction errors, providing an empirically derived probability distribution for normal events post-reconstruction. This non-parametric density estimation technique offers a nuanced understanding of anomaly detection, differentiating it from prevalent threshold-based mechanisms in existing methodologies. In post-training, events are encoded, decoded, and evaluated against the estimated density, providing a comprehensive notion of normality. In addition, we propose a data augmentation strategy involving variational autoencoder-generated events and a smoothing step for enhanced model robustness. The significance of our autoencoder-based approach is evident in its capacity to learn normal representation without prior anomaly knowledge. Through the KDE step on reconstruction errors, our method addresses the versatility of anomalies, departing from assumptions tied to larger reconstruction errors for anomalous events. Our proposed likelihood measure then distinguishes normal from anomalous events, providing a concise yet comprehensive anomaly detection solution. The extensive experimental results support the feasibility of our proposed method, yielding significantly improved classification performance by nearly 10% on the UCR benchmark data.

Frehner, Robin↗

Quasar Identification Using Multivariate Probability Density Estimated from Nonparametric Conditional Probabilities

Nonparametric estimation for a probability density function that describes multivariate data has typically been addressed by kernel density estimation (KDE). A novel density estimator recently developed by Farmer and Jacobs offers an alternative high-throughput automated approach to univariate nonparametric density estimation based on maximum entropy and order statistics, improving accuracy over univariate KDE. This article presents an extension of the single variable case to multiple variables. The univariate estimator is used to recursively calculate a product array of one-dimensional conditional probabilities. In combination with interpolation methods, a complete joint probability density estimate is generated for multiple variables. Good accuracy and speed performance in synthetic data are demonstrated by a numerical study using known distributions over a range of sample sizes from 100 to 10 6 for two to six variables. Performance in terms of speed and accuracy is compared to KDE. The multivariate density estimate developed here tends to perform better as the number of samples and/or variables increases. As an example application, measurements are analyzed over five filters of photometric data from the Sloan Digital Sky Survey Data Release 17. The multivariate estimation is used to form the basis for a binary classifier that distinguishes quasars from galaxies and stars with up to 94% accuracy.

79 ASTRONOMY AND ASTROPHYSICS↗

Data-driven Minimum Entropy Control for Stochastic Nonlinear Systems using the Cumulant-Generating Function

Here, we present a novel minimum entropy control algorithm for a class of stochastic nonlinear systems subjected to non-Gaussian noises. The entropy control can be considered as an optimization problem for the system randomness attenuation, but the mean value has to be considered separately. To overcome this disadvantage, a new representation of the system stochastic properties was given using the cumulant-generating function based on the moment-generating function, in which the mean value and the entropy was reflected by the shape of the cumulant-generating function. Based on the samples of the system output and control input, a time-variant linear model was identified, and the minimum entropy optimization was transformed to system stabilization. Then, an optimal control strategy was developed to achieve the randomness attenuation, and the boundedness of the controlled system output was analyzed. The effectiveness of the presented control algorithm was demonstrated by a numerical example. In this paper, a data-driven minimum entropy design is presented without pre-knowledge of the system model; entropy optimization is achieved by the system stabilization approach in which the stochastic distribution control and minimum entropy are unified using the same identified structure; and a potential framework is obtained since all the existing system stabilization methods can be adopted to achieve the minimum entropy objective.

42 ENGINEERING↗