Search NASA⌕ Search

SEARCH · Search NASA

Results for “kernel density estimation”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

The effective number of parameters in kernel density estimation

We devise a new formula for measuring the effective degrees of freedom (EDoF) in kernel density estimation (KDE). Starting from the orthogonal polynomial sequence (OPS) expansion for the ratio of the empirical to the oracle density, we show how convolution with the kernel leads to a new OPS with respect to which one may express the resulting KDE. The expansion coefficients of the two OPS systems can then be related via a kernel sensitivity matrix, which leads to a natural oracle definition of EDoF through the trace operator. Asymptotic properties of the (empirical) plug-in EDoF are worked out through influence functions, and connections with other empirical EDoFs are established. Minimization of Kullback-Leibler divergence is investigated as an alternative to integrated squared error based bandwidth selection rules, yielding a new normal scale rule. The methodology, which arises from a proper oracle formulation and is not restricted to convolution kernels, suggests the possibility of a new bandwidth selection rule based on an information criterion such as AIC.

bandwidth selection↗

Efficient screening of rare large pit anomalies on polished surfaces using a minimalist sampling scheme

Lawrence Livermore National Laboratory (LLNL) has made significant strides in generating clean energy through its inertial confinement fusion (ICF) experiments. These experiments rely on high-density carbon (HDC) coated shells to encapsulate the fusion fuel. The success of these experiments is heavily dependent on the surface quality of these shells, as even minor imperfections, such as deep pits, can negatively impact fusion yield. Ensuring the required smoothness involves an extensive surface-finishing process that spans approximately 20 stages, making it both time-intensive and resource-demanding. A critical challenge in this process is the need for high-resolution scans to detect rare deep pits, which can be costly and impractical if performed on every shell. This highlights the necessity of developing more efficient scanning methods to optimize time and cost without compromising accuracy. To address these challenges, we introduce a novel approach that employs the multivariate Dvoretzky–Kiefer–Wolfowitz (DKW) inequality to provide a probabilistic upper bound on the error in estimating pit distribution characteristics via a Kernel Density Estimator (KDE). This error bound enables efficient and reliable estimation of pit distribution characteristics at a specified statistical confidence level using a minimal number of surface scans. The integrated DKW-KDE approach was validated through surface-finishing experiments across two batches of HDC-coated shells, demonstrating consistent and robust performance across multiple stages of the surface-finishing experiments. The validation studies suggest that the integrated DKW-KDE approach achieves comparable accuracy in estimating the risk of deleterious large pits with six scans, thus conserving time and resources. Further evaluations show that performance remains consistent across batches and over multiple polishing stages. In conclusion, based on these findings, one can leverage the minimal-scan insights to strategically improve the bottleneck inspection process, thus enhancing the productivity and quality of shell polishing and similar challenging manufacturing processes.

Inertial confinement fusion↗

GalaxyFlow: upsampling hydrodynamical simulations for realistic mock stellar catalogues

ABSTRACT Cosmological N-body simulations of galaxies operate at the level of ‘star particles’ with a mass resolution on the scale of thousands of solar masses. Turning these simulations into stellar mock catalogues requires ‘upsampling’ the star particles into individual stars following the same phase-space density. In this paper, we introduce two new upsampling methods. First, we describe GalaxyFlow, a sophisticated upsampling method that utilizes normalizing flows to both estimate the stellar phase-space density and sample from it. Secondly, we improve on existing upsamplers based on adaptive kernel density estimation (KDE), using maximum likelihood estimation to fine-tune the bandwidth for such algorithms in a way that improves both the density estimation accuracy and upsampling results. We demonstrate our upsampling techniques on a neighbourhood of the Solar location in two simulated galaxies: Auriga 6 and h277. Both yield smooth stellar distributions that closely resemble the stellar densities seen in the Gaia DR3 catalogue. Furthermore, we introduce a novel multimodel classifier test to compare the accuracy of different upsampling methods quantitatively. This test confirms that GalaxyFlow more accurately estimates the density of the underlying star particles than methods based on KDE, at the cost of being more computationally intensive.

Lim, Sung Hak (ORCID:0000000330981092)↗

Experimental and Computational Evaluation of Lipidomic In-Source Fragmentation as a Result of Postionization with Matrix-Assisted Laser Desorption/Ionization

Matrix-assisted laser desorption/ionization mass spectrometry imaging (MALDI-MSI) can provide spatially resolved molecular information about a sample. Recently, a postionization approach (MALDI-2) has been commercially integrated with MALDI-MSI, allowing for bettered sensitivity and consequent improved spatial resolution. While advantages of MALDI-2 have previously been established, we demonstrate here statistically increased in-source fragmentation (ISF) results from postionization with a commercial instrument. Via lipid standard analyses, known MALDI ISF pathways (e.g., loss of trimethylamine) were statistically increased in MALDI-2 compared to MALDI-1 (65–172% increase in fragmentation). Gas phase molecular modeling with density functional theory estimated that the most-weighted virtual orbitals to excite within lipids involve ester and phosphate bonds. Protonated lipid excitation energies are furthermore red-shifted compared to those of other adduct types [e.g., 254 nm for protonated PC(16:0/18:1)] and approach the MALDI-2 laser energy (266 nm). Analysis of rat brain homogenate detected statistically more positive-ion mode peaks with MALDI-2 (1090) than that with MALDI-1 (719), where Kernel density estimations showed that the majority of this enhancement occurs with low m/z ions (i.e., m/z 75–500). Taken together with the lipid standard data, these observations may indicate ISF due to postionization. Finally, while artifact contributions from matrix blanks were also noted, both experimental and computational data sets suggest that the overall extent of ISF is statistically increased in MALDI-2 compared to MALDI-1.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Machine learning materials properties with accurate predictions, uncertainty estimates, domain guidance, and persistent online accessibility

One compelling vision of the future of materials discovery and design involves the use of machine learning (ML) models to predict materials properties and then rapidly find materials tailored for specific applications. However, realizing this vision requires both providing detailed uncertainty quantification (model prediction errors and domain of applicability) and making models readily usable. At present, it is common practice in the community to assess ML model performance only in terms of prediction accuracy (e.g. mean absolute error), while neglecting detailed uncertainty quantification and robust model accessibility and usability. Here, we demonstrate a practical method for realizing both uncertainty and accessibility features with a large set of models. We develop random forest ML models for 33 materials properties spanning an array of data sources (computational and experimental) and property types (electrical, mechanical, thermodynamic, etc). All models have calibrated ensemble error bars to quantify prediction uncertainty and domain of applicability guidance enabled by kernel-density-estimate-based feature distance measures. All data and models are publicly hosted on the Garden-AI infrastructure, which provides an easy-to-use, persistent interface for model dissemination that permits models to be invoked with only a few lines of Python code. We demonstrate the power of this approach by using our models to conduct a fully ML-based materials discovery exercise to search for new stable, highly active perovskite oxide catalyst materials.

domain of applicability↗

Enabling probabilistic learning on manifolds through double diffusion maps

Here, we present a generative learning framework for probabilistic sampling that extends Probabilistic Learning on Manifolds (PLoM), which is designed to generate statistically consistent realizations of a random vector in a finite-dimensional Euclidean space, informed by a (representative) set of observations. In its original form, PLoM constructs a reduced-order probabilistic model by combining three main components: (a) kernel density estimation to approximate the underlying probability measure, (b) Diffusion Maps to characterize the manifold of the data, and (c) a reduced-order Itô Stochastic Differential Equation (ISDE) to sample from the learned distribution. However, its sampling dynamics are posed in the ambient space and the retained number of reduced coordinates is chosen by projection-reconstruction error. In practice, this often (i) requires more coordinates than the data’s intrinsic dimension to achieve stable sampling and (ii) lacks a smooth, basis-independent lifting back to the data domain; moreover, standard Diffusion Maps emphasize harmonic eigenfunctions and can miss non-harmonic latent structure. We address these limitations by decoupling geometry learning from sampling: a first Diffusion Maps pass identifies non-harmonic coordinates on which we formulate a full-order ISDE directly in the latent space, while Double Diffusion Maps captures multiscale geometric features and Geometric Harmonics (GH) learns a smooth lifting map to the ambient variables that is independent of the particular diffusion basis. This hybrid design preserves the system’s dynamical richness with a compact geometric representation and enables principled out-of-sample inference. The effectiveness and robustness of the proposed method are illustrated through two numerical studies: one based on data generated from two-dimensional Hermite polynomial functions and another based on high-fidelity simulations of a detonation wave in a reactive flow.

Double diffusion maps↗

DONKEY: A Flexible and Accurate Algorithm for Clustering

We propose an accurate clustering algorithm suitable for the varied and multidimensional data sets that correspond to temporal snapshots from on-the-fly nonadiabatic trajectory-based simulations of photoexcited dynamics. The algorithm approximates the underlying probability density function using variable kernel density estimation, with local maxima corresponding to cluster centers. Each data point is then assigned to one of the maxima by employing a maximization procedure. Finally, clusters artificially separated by minor fluctuations in the probability density are merged. The algorithm does not require parameter tuning, which ensures flexibility and reduces the risk of bias. It is tested on several synthetic data sets, where it consistently outperforms conventional clustering algorithms. As a final example, the algorithm is applied to the excited dynamics of the norbornadiene ⇌ quadricyclane (C 7 H 8 ) molecular photoswitch, demonstrating how distinct reaction pathways can be identified.

algorithms↗

A meshless stochastic method for Poisson–Nernst–Planck equations

A plethora of biological, physical, and chemical phenomena involve transport of charged particles (ions). Its continuum-scale description relies on the Poisson–Nernst–Planck (PNP) system, which encapsulates the conservation of mass and charge. The numerical solution of these coupled partial differential equations is challenging and suffers from both the curse of dimensionality and difficulty in efficiently parallelizing. We present a novel particle-based framework to solve the full PNP system by simulating a drift–diffusion process with time- and space-varying drift. We leverage Green’s functions, kernel-independent fast multipole methods, and kernel density estimation to solve the PNP system in a meshless manner, capable of handling discontinuous initial states. The method is embarrassingly parallel, and the computational cost scales linearly with the number of particles and dimension. We use a series of numerical experiments to demonstrate both the method’s convergence with respect to the number of particles and computational cost vis-à-vis a traditional partial differential equation solver.

Chemistry↗

Polynomial Chaos Surrogate Construction for Random Fields with Parametric Uncertainty

Engineering and applied science rely on computational experiments to rigorously study physical systems. The mathematical models used to probe these systems are highly complex, and sampling-intensive studies often require prohibitively many simulations for acceptable accuracy. Surrogate models provide a means of circumventing the high computational expense of sampling such complex models. In particular, polynomial chaos expansions (PCEs) have been successfully used for uncertainty quantification studies of deterministic models where the dominant source of uncertainty is parametric. We discuss an extension to conventional PCE surrogate modeling to enable surrogate construction for stochastic computational models that have intrinsic noise in addition to parametric uncertainty. We develop a PCE surrogate on a joint space of intrinsic and parametric uncertainty, enabled by Rosenblatt transformations, which are evaluated via kernel density estimation of the associated conditional cumulative distributions. Furthermore, we extend the construction to random field data via the Karhunen–Loève expansion. We then take advantage of closed-form solutions for computing PCE Sobol indices to perform a global sensitivity analysis of the model which quantifies the intrinsic noise contribution to the overall model output variance. Additionally, the resulting joint PCE is generative in the sense that it allows generating random realizations at any input parameter setting that are statistically approximately equivalent to realizations from the underlying stochastic model. The method is demonstrated on a chemical catalysis example model and a synthetic example controlled by a parameter that enables a switch from unimodal to bimodal response distributions.

97 MATHEMATICS AND COMPUTING↗

PV-Finder: ML Based Algorithm for Primary Vertex Identification

he CMS detector at the High-Luminosity Large Hadron Collider (HL-LHC) will operate in challenging conditions with expected pile-up of up to 200 collisions per bunch crossing, necessitating the development of a more resilient primary vertex (PV) reconstruction method to ensure the integrity of data analysis and the efficiency of the CMS triggering system. This contribution describes preliminary studies on a new ML based PV-Finder method for PV identification. The method is based on a model trained using Kernel Density Estimations (KDEs) derived from the positions of reconstructed tracks at the beamline, incorporating uncertainties from track parameters. It also utilizes target histograms, modeled as Gaussian distributions centered on the actual ground truth values of specific primary vertices.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

Minimum entropy filtering for a single output non-Gaussian stochastic system using state transformation

This paper presents a novel filter design for the single-output stochastic non-linear systems subjected to non-Gaussian noises and the proposed assumptions. Based on a state transformation, the unmeasurable states of the systems can be estimated where non-linear terms in the systems have been eliminated. It has been shown that the estimation error is linearly dynamical regarding to the presented vector-valued filter gain which can be optimised by minimising the entropy-based performance criterion. In addition, the convergence of the presented algorithm is analysed in mean-square sense and a numerical example is given to verify the effectiveness of the presented filtering algorithm. Meanwhile, the extended Kalman filter, unscented particle filter and minimum entropy filter are given for the comparisons of the filtering performance. Following the presented framework, some extensions of the presented filtering algorithm are discussed to indicate the flexibility of the filter design. The contribution of this paper can be summarised as establishing a novel minimum entropy filtering framework which consists of model transformation, entropy optimisation and convergence analysis.

42 ENGINEERING↗

Computing the QRPA level density with the finite amplitude method

Here, we describe a new algorithm to calculate the vibrational nuclear level density of an atomic nucleus. Fictitious perturbation operators that probe the response of the system are generated by drawing their matrix elements from some probability distribution function. We use the Finite Amplitude Method to explicitly compute the response for each such sample. With the help of the Kernel Polynomial Method, we build an estimator of the vibrational level density and provide the upper bound of the relative error in the limit of infinitely many random samples. The new algorithm can give accurate estimates of the vibrational level density. Since it is based on drawing multiple samples of perturbation operators, its computational implementation is naturally parallel and scales like the number of available processing units.

73 NUCLEAR PHYSICS AND RADIATION PHYSICS↗

Training quantum neural networks using the quantum information bottleneck method

Abstract We provide in this paper a concrete method for training a quantum neural network to maximize the relevant information about a property that is transmitted through the network. This is significant because it gives an operationally well founded quantity to optimize when training autoencoders for problems where the inputs and outputs are fully quantum. We provide a rigorous algorithm for computing the value of the quantum information bottleneck quantity within error ε that requires O ( log 2 ⁡ ( 1 / ϵ ) + 1 / δ 2 ) queries to a purification of the input density operator if its spectrum is supported on { 0 } ⋃ [ δ , 1 − δ ] for δ > 0 and the kernels of the relevant density matrices are disjoint. We further provide algorithms for estimating the derivatives of the QIB function, showing that quantum neural networks can be trained efficiently using the QIB quantity given that the number of gradient steps required is polynomial.

Çatlı, Ahmet Burak (ORCID:0000000152294141)↗

Hierarchical Speed Planner for Automated Vehicles: A Framework for Lagrangian Variable Speed Limit in Mixed-Autonomy Traffic

Here, this article presents a novel hierarchical speed planning framework for variable speed limits in mixed-autonomy traffic environments, leveraging server-side macroscopic control and vehicle-side microscopic execution. The framework integrates real-time traffic state estimation (TSE) and reinforcement learning (RL)-based control to mitigate congestion and improve traffic flow. A TSE enhancement module combines macroscopic data from sources like INRIX with high-resolution observations from connected autonomous vehicles (CAVs), enabling predictive modeling to address latency and noise. The target speed design module employs kernel smoothing and a buffer zone strategy to optimize traffic density and flow around bottlenecks. The proposed system was validated in the largest open-road test to date with 100 CAVs, demonstrating an overall 8% traffic density decrease, with a specific decrease of 7% upstream, 10% downstream, and a 52% decrease during the congestion formation phase at bottlenecks.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗

Particle Filter Based Inference Testing

The primary intent of PAR-FIT (Particle Filter based Inference Testing) is to provide hard inductive evidence that a machine learning model is capable and proven for an individual test input. By examining training data used to form the underlying model functional correlation, an estimate of the reliability that a model will make the correct prediction can be made. The Sequential Probability Ratio Test is used to derive a qualitative evaluation for reliability based on hypothesis testing. The PAR-FIT framework achieves this by implementing a particle filter and the sequential probability ratio test algorithms on the machine learning model training data to determine relevancy of new individual test samples to the training dataset. The kernel function evaluates the local proximity and density of training data used to derive a prediction outcome. Particles are used to probabilistically determine which training data to evaluate for proximity. For test samples that are within a close proximity to and surrounded by multiple training data points, the evaluated reliability of the prediction is high. For test samples that are anomalies not represented by the training dataset, in low density data clusters, or are far from existing data points, the evaluated reliability is low as insufficient training evidence exists to suggest the model is capable of making the correct prediction. Sequential Probability Ratio Test is further used to determine when a hypothesis on whether a signal can be rejected or accepted for use. The ratio test collects sequence information from the particle filter to test whether the signal is anomalous or normal via hypothesis testing of the underlying distributions.

Chen, Edward [Idaho National Laboratory (INL), Ida↗

Conformal Hierarchical Simulation-Based Inference with Local Validity

Trustworthy and interpretable uncertainty quantification is a long-standing challenge in artificial intelligence. Simulation-based inference (SBI) comprises a broad swath of approaches for estimating latent parameters with uncertainties. Although flexible neural density estimators in SBI can be remark- ably expressive capturing highly structured, high-dimensional posteriors their credible regions can be badly mis-calibrated and are often only accompanied by heuristic coverage checks. We present the first SBI framework that delivers finite-sample local valid coverage guarantees that hold in the neighborhood of each observation. Our framework can couple any off-the-shelf hierarchical SBI engine with a confor- mal Bayesian post-processing step that operates on the posterior predictive density. A kernel-weighted conformity score adapts the conformal quantile to the local geometry of the data, yielding prediction sets that are simultaneously (i) marginally calibrated, (ii) locally valid, and (iii) hierarchical, handling global and observation-specific parameters in a single pass. Through experiments on synthetic data and benchmarks from neuroscience and physics, we show that our approach attains 1 − α coverage, where prior SBI methods under- or over-cover. Our approach also maintains a competitive, credible set size with minimal computational overhead. Finally, our approach can be used to make predictions on real data and give valid credible regions modulo weight-initialization-based model mis-specification.

Trivedi, Shubhendu [Fermilab]↗

Resonant inelastic x-ray scattering in warm-dense Fe compounds beyond the SASE FEL resolution limit

Resonant inelastic x-ray scattering (RIXS) is a widely used spectroscopic technique, providing access to the electronic structure and dynamics of atoms, molecules, and solids. However, RIXS requires a narrow bandwidth x-ray probe to achieve high spectral resolution. The challenges in delivering an energetic monochromated beam from an x-ray free electron laser (XFEL) thus limit its use in few-shot experiments, including for the study of high energy density systems. Here we demonstrate that by correlating the measurements of the self-amplified spontaneous emission (SASE) spectrum of an XFEL with the RIXS signal, using a dynamic kernel deconvolution with a neural surrogate, we can achieve electronic structure resolutions substantially higher than those normally afforded by the bandwidth of the incoming x-ray beam. We further show how this technique allows us to discriminate between the valence structures of Fe and Fe2O3, and provides access to temperature measurements as well as M-shell binding energies estimates in warm-dense Fe compounds.

74 ATOMIC AND MOLECULAR PHYSICS↗

The Cosmic Evolution of C IV Absorbers at 1.4 < z < 4.5: Insights from 100,000 Systems in DESI Quasars

We present the largest catalog to date of triply ionized carbon (C IV ) absorbers detected in quasar spectra from the Dark Energy Spectroscopic Instrument. Using an automated matched-kernel convolution method with adaptive signal-to-noise thresholds, we identify 101,487 C IV systems in the redshift range 1.4 < z < 4.5 from 300,637 quasar spectra. Completeness is estimated via Monte Carlo simulations, and the catalog is 50% complete at EW C IV ≥ 0.4 Å. The differential equivalent width frequency distribution declines exponentially and shows weak redshift evolution. The absorber incidence per unit comoving path increases by a factor of 2–5 from z ≈ 4.5 to z ≈ 1.4, with stronger redshift evolution for strong systems. Using column densities derived from the apparent optical depth method, we constrain the cosmic mass density of C IV , Ω C IV , which increases by a factor of ∼3.8 from (0.82 ± 0.05) × 10 −8 at z ≈ 4.5 to (3.16 ± 0.2) × 10 −8 at z ≈ 1.4. From Ω C IV , we estimate a lower limit on intergalactic medium metallicity ${\mathrm{log}}({Z}_{{\rm{IGM}}}/{Z}_{\odot })\gtrsim -3.25$ at z ∼ 2.3, with a smooth decline at higher redshifts. These trends trace the cosmic star formation history and He II photoheating rate, suggesting a link between C IV enrichment, star formation, and UV background over ∼3 Gyr. The catalog also provides a critical resource for future studies connecting circumgalactic metals to galaxy evolution, especially near cosmic noon.

79 ASTRONOMY AND ASTROPHYSICS↗