Search NASASearch

SEARCH · Search NASA

Results for “Kernel”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 37 records · Page 2

The effective number of parameters in kernel density estimation

We devise a new formula for measuring the effective degrees of freedom (EDoF) in kernel density estimation (KDE). Starting from the orthogonal polynomial sequence (OPS) expansion for the ratio of the empirical to the oracle density, we show how convolution with the kernel leads to a new OPS with respect to which one may express the resulting KDE. The expansion coefficients of the two OPS systems can then be related via a kernel sensitivity matrix, which leads to a natural oracle definition of EDoF through the trace operator. Asymptotic properties of the (empirical) plug-in EDoF are worked out through influence functions, and connections with other empirical EDoFs are established. Minimization of Kullback-Leibler divergence is investigated as an alternative to integrated squared error based bandwidth selection rules, yielding a new normal scale rule. The methodology, which arises from a proper oracle formulation and is not restricted to convolution kernels, suggests the possibility of a new bandwidth selection rule based on an information criterion such as AIC.

bandwidth selection

Physicochemical evolution of uranium nitride kernel microstructure with varying carbon distribution for advanced TRISO fuel forms

Uranium nitride (UN) has emerged as a fuel candidate for advanced nuclear reactor concepts due to its superior uranium density, thermal conductivity, and high melting temperature. However, the fabrication route for converting UO 2 to UN is complex and difficult to standardize. Although the chemistry of this conversion process is well-studied, more insight into the physicochemical dynamics of this conversion using advanced characterization techniques can help further our understanding of this material system. This work leveraged thermogravimetric analysis (TGA), X-ray diffraction (XRD), and nondestructive 3D X-ray computed tomography (XCT) to characterize dynamic microstructural changes in the UO 2 → UCO → UN fabrication pathway for two kernels with a varying carbon distribution in the starting composition. TGA and XRD were used to quantify changes in the mass, density, and chemical composition of the two kernels, while three-dimensional image processing and segmentation of XCT data were used to quantify the volume, surface area, and spatial distribution of features within each kernel for multiple steps along the fabrication pathway. The analysis indicates distinct differences between the two kernels that are correlated to downstream conversion efficiency. In conclusion, this work is among the first to perform 3D quantification of physicochemical evolution during UN conversion, providing quantitative correlation between processing, properties, and expected fuel performance.

Nuclear fuel

Collins-Soper kernel in the QCD instanton vacuum

We outline a general framework for evaluating the nonperturbative soft functions in the quantum chromodynamics (QCD) instanton vacuum. In particular, from the soft function we derive the Collins-Soper (CS) kernel, which drives the rapidity evolution of the transverse-momentum-dependent parton distributions. The resulting CS kernel, when supplemented with the perturbative contribution, agrees well with recent lattice results and some phenomenological parametrizations. Moreover, our CS kernel depends logarithmically on the large quark transverse separation, providing a key constraint on its phenomenological parametrization. Finally, a lattice calculation can be directly compared to our generic results in Euclidean signature, thus providing a new approach for evalulating the soft function and extracting the CS kernel by analytical continuation.

QCD phenomenology

A Study of Performance Portability of Low-bit Fused Matrix-Vector Multiplication Kernels in SYCL

Understanding the causes of performance gaps between a portable programming model and a vendor-specific programming model is important for improving performance portability. This paper studies performance portability of low-bit fused general matrix-vector multiplication kernels in SYCL on vendors’ graphics processing units (GPUs). This work introduces the use case, explains the kernel implementations in detail, evaluates the performance of the CUDA, HIP, and SYCL kernels on datacenter, desktop, and laptop GPUs, and investigates the causes of performance gaps. The results show that loop unrolling, kernel dispatch overhead, and sum reduction contribute to the gaps.

Jin, Zheming [ORNL] (ORCID:000000027197780X)

Boundary Corrections for Kernel Approximation to Differential Operators

The kernel-based approach to operator approximation for partial differential equations has been shown to be unconditionally stable for linear PDEs and numerically exhibit unconditional stability for non-linear PDEs. These methods have the same computational cost as an explicit finite difference scheme but can exhibit order reduction at boundaries. In previous work on periodic domains, order reduction was addressed, yielding high-order accuracy. The issue addressed in this work is the elimination of order reduction of the kernel-based approach for a more general set of boundary conditions. Further, we consider the case of both first and second order operators. To demonstrate the theory, we provide not only the mathematical proofs but also experimental results by applying various boundary conditions to different types of equations. The results agree with the theory, demonstrating a systematic path to high order for kernel-based methods on bounded domains.

97 MATHEMATICS AND COMPUTING

Transient anisotropic kernel for probabilistic learning on manifolds

PLoM (Probabilistic Learning on Manifolds) is a method introduced in 2016 for handling small training datasets by projecting an Itô equation from a stochastic dissipative Hamiltonian dynamical system, acting as the MCMC generator, for which the KDE-estimated probability measure with the training dataset is the invariant measure. PLoM performs a projection on a reduced-order vector basis related to the training dataset, using the diffusion maps (DMAPS) basis constructed with a time-independent isotropic kernel. In this paper, we propose a new ISDE projection vector basis built from a transient anisotropic kernel, providing an alternative to the DMAPS basis to improve statistical surrogates for stochastic manifolds with heterogeneous data. The construction ensures that for times near the initial time, the DMAPS basis coincides with the transient basis. For larger times, the differences between the two bases are characterized by the angle of their spanned vector subspaces. The optimal instant yielding the optimal transient basis is determined using an estimation of mutual information from Information Theory, which is normalized by the entropy estimation to account for the effects of the number of realizations used in the estimations. Consequently, this new vector basis better represents statistical dependencies in the learned probability measure for any dimension. Three applications with varying levels of statistical complexity and data heterogeneity validate the proposed theory, showing that the transient anisotropic kernel improves the learned probability measure.

Diffusion maps

A fractional calculus framework for open quantum dynamics: From Liouville to Lindblad to memory kernels

Open quantum systems exhibit dynamics ranging from unitary evolution to irreversible dissipation. While the Gorini–Kossakowski–Sudarshan–Lindblad equation uniquely characterizes Markovian completely positive and trace-preserving (CPTP) evolution, many physical platforms display non-Markovian features such as algebraic relaxation and coherence backflow. Fractional calculus provides a natural way to model such long-memory behavior through power-law temporal kernels introduced by fractional time derivatives. Here, we develop a unified framework that embeds fractional master equations within the broader hierarchy of open-system formalisms. The fractional equation forms a structured subclass of memory-kernel models, reduces to the Lindblad form at unit order, and, through Bochner–Phillips subordination, admits a CPTP representation as an average over Lindblad semigroups. Its resolvent structure further connects fractional dynamics to established non-Markovian approaches, including Nakajima–Zwanzig kernels and hierarchical equations of motion, providing a compact surrogate for long-memory effects. This formulation positions fractional calculus as a rigorous and practical language for modeling non-Markovian quantum dynamics in chemical physics and physical chemistry, providing a CPTP-preserving, computationally efficient surrogate for structured condensed-phase environments where long-time memory and dissipation play a central role.

71 CLASSICAL AND QUANTUM MECHANICS, GENERAL PHYSIC

Transverse-momentum-dependent pion structures from lattice QCD: Collins-Soper kernel, soft factor, TMDWF, and TMDPDF

We present the first lattice quantum chromodynamics (QCD) calculation of the pion valence-quark transverse-momentum-dependent parton distribution function (TMDPDF) within the framework of large-momentum effective theory (LaMET). Using correlators fixed in the Coulomb gauge (CG), we computed the quasi-TMD beam function for a pion with a mass of 300 MeV, a fine lattice spacing of 𝑎 =0.06 fm, and multiple large momenta up to 3 GeV. The intrinsic soft functions in the CG approach are extracted from form factors with large momentum transfer, and as a byproduct, we also obtain the corresponding Collins-Soper (CS) kernel. Our determinations of both the soft function and the CS kernel agree with perturbation theory at small transverse separations (𝑏 ⊥ ) between the quarks. At larger 𝑏 ⊥ , the CS kernel remains consistent with recent results obtained using both CG and gauge-invariant TMD correlators in the literature. By combining next-to-leading logarithmic factorization of the quasi-TMD beam function and the soft function, we obtain an 𝑥-dependent pion valence-quark TMDPDF for transverse separations 𝑏 ⊥ ≳1 fm. Interestingly, we find that the 𝑏 ⊥ dependence of the phenomenological parametrizations of TMDPDF for moderate values of 𝑥 are in reasonable agreement with our QCD determinations. In addition, we present results for the transverse-momentum-dependent wave function for a heavier pion with 670 MeV mass.

73 NUCLEAR PHYSICS AND RADIATION PHYSICS

Integral Kernel Methods for Nonlinear Parabolic-Elliptic Systems

Nonlinear parabolic-elliptic systems arise in many physical, biological, and chemical phenomena such as chemotaxis, ion transport, self-gravitating particles, and Brownian vortices. Existing methods struggle with the strong coupling and high nonlinearity and nonlocality of some of these systems, especially the ill-conditioned, convection-dominated problems. To overcome numerical difficulties, current approaches rely on initial guesses, preconditioning, or iterative techniques with no convergence guarantees. They might suffer from poor scalability, large memory usage, and difficulty to parallelize. Inspired by the connection of parabolic-elliptic systems to stochastic processes, we introduce a novel meshless, monolithic, and fully explicit method that naturally encapsulates the elliptic and parabolic operators into a single step which updates each node deterministically with global information. By being fully quadrature-based, it avoids solving systems of discretized equations and does not utilize initial guesses or preconditioning, while requiring little memory and being easy to parallelize. We first derive the method in an integral kernel formulation with quadratic complexity in the number of integration nodes and then leverage kernel-independent fast multipole methods (FMM) to present a scalable algorithm with linear complexity. We provide numerical examples for the Poisson-Nernst-Planck equations in one, two, and three dimensions, together with the derivation of the integral kernel for each case. Furthermore, the examples demonstrate the fast convergence and scalability of the FMM-accelerated algorithm, as well as its suitability for convection-dominated problems, making it competitive against traditional PDE solvers.

PDE systems

The Collins-Soper Kernel from Lattice QCD

I will present the first complete determination of the quark Collins-Soper kernel, which relates TMDs at different rapidity scales, using lattice QCD and including systematic control of quark mass, operator mixing, and discretization effects. Next-to-next-to-leading logarithmic matching is used to match lattice-calculable distributions to the corresponding TMDs. The continuum-extrapolated lattice QCD results are consistent with several recent phenomenological parametrizations of the Collins-Soper kernel and are precise enough to disfavor other parametrizations. I will also discuss a first exploration of the gluon Collins-Soper kernel.

Wagman, Michael [Fermilab]

Kernelized approaches to streaming compression of scientific data

In this paper three algorithms are developed for the streaming compression of scientific data. The algorithms presented are reliant on the theory of vector-valued reproducing kernel Hilbert spaces and operator valued kernel. Further, the scientific data is modeled as a snapshot of time dependent vector field F(x, t) over a manifold M and the recovery of the data is framed as a learning problem. These processes are then appropriately modified and ana lyzed for the streaming scenario in which data is generated without the ability to revisit past entries.

97 MATHEMATICS AND COMPUTING

Application of Portable Parallelization Strategies for GPUs on track reconstruction kernels

Utilizing the computational power of GPUs is one of the key ingredients to meet the computing challenges presented to the next generation of High-Energy Physics (HEP) experiments. Unlike CPUs, developing software for GPUs often involves using architecturespecific programming languages promoted by the GPU vendors and hence limits the platform that the code can run on. Various portability solutions have been developed to achieve portable, performant software across different GPU vendors. Given the rapid evolution of these portability solutions, an early adoption of them in simple HEP testbed applications will help us understand the strengths and weaknesses of respective approaches.We apply several portability solutions, including Alpaka, Kokkos, SYCL and std::execution::par, on kernels for track propagation extracted from the mkFit project. We report on the development experience of the same application with different portability solutions, as well as their performance on GPUs, measured as the throughput of the kernels, from different manufacturers such as NVIDIA, AMD and Intel.

Kwok, Martin [Fermilab] (ORCID:0000000286936146)

Unsupervised atomic data mining via multi-kernel graph autoencoders for machine learning force fields

Constructing a chemically diverse dataset while avoiding sampling bias is critical to training efficient and generalizable force fields. However, in computational chemistry and materials science, many common dataset generation techniques are prone to oversampling regions of the potential energy surface. Furthermore, these regions can be difficult to identify and isolate from each other or may not align well with human intuition, making it challenging to systematically remove bias in the dataset. While traditional clustering and pruning (down-sampling) approaches can be useful for this, they can often lead to information loss or a failure to properly identify distinct regions of the potential energy surface due to difficulties associated with the high dimensionality of atomic descriptors. In this work, we introduce the Multi-kernel Edge Attention-based Graph Autoencoder (MEAGraph) model, an unsupervised approach for analyzing atomic datasets. MEAGraph combines multiple linear kernel transformations with attention-based message passing to capture geometric sensitivity and enable effective dataset pruning without relying on labels or extensive training. Demonstrated applications on niobium, tantalum, and iron datasets show that MEAGraph efficiently groups similar atomic environments, allowing for the use of basic pruning techniques for removing sampling bias. This approach provides an effective method for representation learning and clustering that can be used for data analysis, outlier detection, and dataset optimization.

Materials science

Collins-Soper kernel and reduced soft function in lattice QCD

We evaluate the Collins-Soper kernel and the reduced soft function in lattice QCD, incorporating 𝒪⁡(𝛼 𝑠 ) matching corrections. The calculation relies on the evaluation of the quasitransverse momentum–dependent wave function with asymmetric staple-shaped quark bilinear operators and four-point meson form factors. These quantities are computed nonperturbatively using two 𝑁 𝑓 =2 + 1 + 1 twisted-mass fermion ensembles with the same lattice spacing of 𝑎 = 0.093 fm: the first ensemble has a lattice size of 24 3 × 48 and a pion mass of 346 MeV, and the second one has a lattice size of 32 3 × 64 and a pion mass of 261 MeV. The Collins-Soper kernel and the soft function are needed for the determination of the transverse momentum–dependent parton distribution functions.

Lattice QCD

Mojo: MLIR-based Performance-Portable HPC Science Kernels on GPUs for the Python Ecosystem

We explore the performance and portability of the novel Mojo language for scientific computing workloads on GPUs. As the first language based on the LLVM’s Multi-Level Intermediate Representation (MLIR) compiler infrastructure, Mojo aims to close performance and productivity gaps by combining Python’s interoperability and CUDA-like syntax for compile-time portable GPU programming. We target four scientific workloads: a seven-point stencil (memory-bound), BabelStream (memory-bound), miniBUDE (compute-bound), and Hartree–Fock (compute-bound with atomic operations); and compare their performance against vendor baselines on NVIDIA H100 and AMD MI300A GPUs. We show that Mojo’s performance is competitive with CUDA and HIP for memory-bound kernels, whereas gaps exist on AMD GPUs for atomic operations and for fast-math compute-bound kernels on both AMD and NVIDIA GPUs. Although the learning curve and programming requirements are still fairly low-level, Mojo can close significant gaps in the fragmented Python ecosystem in the convergence of scientific computing and AI.

Godoy, William [ORNL] (ORCID:0000000225905178)

Technical note: Recommendations for diagnosing cloud feedbacks and rapid cloud adjustments using cloud radiative kernels

Abstract. The cloud radiative kernel method is a popular approach to quantify cloud feedbacks and rapid cloud adjustments to increased CO2 concentrations and to partition contributions from changes in cloud amount, altitude, and optical depth. However, because this method relies on cloud property histograms derived from passive satellite sensors or produced by passive satellite simulators in models, changes in obscuration of lower-level clouds by upper-level clouds can cause apparent low-cloud feedbacks and adjustments, even in the absence of changes in lower-level cloud properties. Here, we provide a methodology for properly diagnosing the impact of changing obscuration on cloud feedbacks and adjustments and quantify these effects across climate models. Averaged globally and across global climate models, properly accounting for obscuration leads to weaker positive feedbacks from lower-level clouds and stronger positive feedbacks from upper-level clouds while simultaneously removing a mostly artificial anti-correlation between them. Given that the methodology for diagnosing cloud feedbacks and adjustments using cloud radiative kernels has evolved over several papers, and obscuration effects have only occasionally been considered in recent papers, this paper serves to establish recommended best practices and to provide a corresponding code base for community use.

54 ENVIRONMENTAL SCIENCES

First constraints on the nonperturbative gluon Collins-Soper kernel

The gluon Collins-Soper kernel, which encodes the rapidity evolution of transverse-momentum-dependent gluon distributions, is constrained for the first time in the nonperturbative regime, for transverse momentum scales $q_{T} \in [ 300\text{ MeV}, 1.3\text{ GeV}]$. The constraints are determined in lattice QCD at a close-to-physical pion mass $M_π= 172(3)\text{ MeV}$, a single lattice spacing $a=0.15\text{ fm}$, and next-to-next-to-leading logarithmic matching in Large-Momentum Effective Theory. These results represent the first step toward a controlled determination of the gluon Collins-Soper kernel in QCD, with eventual phenomenological import and relevance to present and future experiments sensitive to the gluon structure of hadronic matter.

Avkhadiev, Artur [Argonne; MIT, Cambridge, CTP]

Kernel fusion in atomistic spin dynamics simulations on Nvidia GPUs using tensor core

In atomistic spin dynamics simulations, the time cost of constructing the space- and time-displaced pair correlation function in real space increases quadratically as the number of spins N, leading to significant computational effort. The GEMM subroutine can be adopted to accelerate the calculation of the dynamical spin-spin correlation function, but the computational cost of simulating large spin systems (>40000 spins) on CPUs remains expensive. In this work, we perform the simulation on the graphics processing unit (GPU), a hardware solution widely used as an accelerator for scientific computing and deep learning. Here we show that GPUs can accelerate the simulation up to 25-fold compared to multi-core CPUs when using the GEMM subroutine on both. To hide memory latency, we fuse the element-wise operation into the GEMM kernel using CUTLASS that can improve the performance by 26% ~ 33% compared to implementation based on cuBLAS. Furthermore, we perform the on-the-fly calculation in the epilogue of the GEMM subroutine to avoid saving intermediate results on global memory, which makes the large-scale atomistic spin dynamics simulation feasible and affordable.

97 MATHEMATICS AND COMPUTING