Search NASA⌕ Search

SEARCH · Search NASA

Results for “Kernel”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 325 records · Page 18

Generating entangled steady states in multistable open quantum systems via initial state control

Entanglement underpins the power of quantum technologies, yet it is fragile and typically destroyed by dissipation. Paradoxically, the same dissipation, when carefully engineered, can drive a system toward robust entangled steady states. However, this engineering task is nontrivial, as dissipative many-body systems are complex, particularly when they support multiple steady states. Here, we derive analytic expressions that predict how the steady state of a system evolving under a Lindblad equation depends on the initial state, without requiring integration of the dynamics. These results extend Refs. [V. V. Albert and L. Jiang, Phys. Rev. A 89, 022118 (2014); V. V. Albert et al., Phys. Rev. X 6, 041031 (2016)], showing that while the steady-state manifold is determined by the Liouvillian kernel, the weights within it depend on both the Liouvillian and the initial state. We identify a special class of Liouvillians for which the steady state depends only on the initial overlap with the kernel. Our framework provides analytical insight and a computationally efficient tool for predicting steady states in open quantum systems. As an application, we propose schemes to generate metrologically useful entangled steady states in spin ensembles via balanced collective decay.

Dissipative dynamics↗

IRIS: Exploring Performance Scaling of the Intelligent Runtime System and its Dynamic Scheduling Policies

High-Performance Computing is becoming increasingly heterogeneous, relying on a diverse mix of hardware to achieve good performance. Paradoxically, current drivers and frameworks for these devices typically require separate languages and implementations for each vendor. Furthermore, there are few tools and little support to schedule codes between these devices in a truly heterogeneous manner-partly because of this fragmentation between vendors and the languages each supports. To overcome both limitations, the Intelligent Runtime System (IRIS) was developed. It allows a common task abstraction to automatically be shared among contemporary vendors and is run from a single host-side API. At runtime, IRIS queries the host system and registers which frameworks and drivers are available, these determine which kernels can be used by the scheduler-CPUs via OpenMP, Nvidia GPUs (CUDA), AMD GPUs (HIP), and Intel and Xilinx FPGAs with OpenCL. IRIS enables tasks to be scheduled to any heterogeneous device and resolves to the appropriate kernel binary at runtimeit only uses the devices supported by the system on which it is run. IRIS supports single-task and graph-based expressions of dependencies of tasks. Additionally, IRIS features a range of dynamic scheduling policies, allowing complex chains of tasks and interactions to be executed, relieving the programmer/user from considering the system to assign tasks to devices optimally. This paper presents the peak performance attainable by IRIS over a range of systems-each with different numbers and types of accelerator devices, it highlights the flexibility of IRIS since these devices are truly heterogeneous, relying on different backends (drivers, frameworks, and languages) which historically required unique implementations to utilize them. We then use this peak performance as a baseline to compare increasingly complex chains of tasks (with increasingly complex task dependencies) and evaluate how IRIS copes. Finally, we consider the performance of different IRIS scheduling policies on this range of task graphs.

Johnston, Beau↗

Scatter and Blur Corrections for High-Energy X-Ray Radiography

High-energy X-ray radiography is useful as a highly penetrating method for imaging through dense materials. However, the primary modes of interaction of X-rays at these energies involve scattering or the production of secondary high-energy photons, which can interfere with the image. In addition, detector blurring, often resulting from scatter within the detector, can reduce image sharpness. Both of these processes can be mitigated with the use of convolution kernels, with the main challenge being that the proper kernel to use is not known, particularly for the scatter contribution. By radiographing solid slabs of uniform attenuation, we show that point spread functions and material-specific point scatter functions can be determined to significantly reduce the effect of detector blurring and object scatter. Constraining the fits to the slabs and uniform transmission within the slabs is sufficient to recover these functions. A functional form that reproduces the angular distribution of high-energy bremsstrahlung X-rays is presented for recovering point scatter functions. In conclusion, the method is applied to radiographs of objects from bremsstrahlung X-ray sources operating at 4- and 7.5-MV endpoint energies and a significant increase in sharpness is observed.

Blind deconvolution↗

Discovery of Probabilistic Dirichlet-to-Neumann Maps on Graphs

Dirichlet-to-Neumann maps enable the coupling of multiphysics simulations across computational subdomains by ensuring continuity of state variables and fluxes at artificial interfaces. We present a novel method for learning Dirichlet-to-Neumann maps on graphs using Gaussian processes, specifically for problems where the data obey a conservation law arising from an underlying partial differential equation. Our approach combines discrete exterior calculus and nonlinear optimal recovery to infer relationships between vertex and edge values. This framework yields data-driven predictions with uncertainty quantification across the entire graph, even when observations are limited to a subset of vertices and edges. By minimizing the reproducing kernel Hilbert space norm while penalizing kernel complexity through maximum likelihood estimation, our method ensures that the resulting surrogate strictly enforces conservation laws without overfitting. We demonstrate our method on two representative applications: subsurface flow in fracture networks and arterial blood flow. Finally, the results demonstrate that the method maintains high accuracy and well-calibrated uncertainty estimates even under severe data scarcity, highlighting its potential for scientific applications where limited data and reliable uncertainty quantification are critical.

Dirichlet-to-Neumann map↗

IRIS-MASH: Efficient Multi-device Asynchronous Multi-Stream Heterogeneous Computing

In the rapidly evolving field of high-performance computing (HPC), effectively leveraging heterogeneous devices through asynchronous task programming is paramount. This paper presents a robust asynchronous task programming model tailored for a multi-device, multi-stream execution environment that incorporates a diverse array of heterogeneous computing units, including GPUs from various vendors and other accelerators. Current state-of-the-art task programming models provide methodologies to support asynchronous task executions, but they typically handle homogeneous devices using native programming languages, while support for heterogeneous devices is limited to frameworks like OpenCL. This gap presents significant challenges in abstracting heterogeneous devices to harness their true asynchronous capabilities effectively using their native programming languages. By implementing asynchronous task execution, our model significantly boosts the performance of tiled algorithm task graphs through overlapping data transfers with computation and enabling the simultaneous execution of multiple kernels. We integrate this approach into a heterogeneous Intelligent Runtime System (IRIS) and assess its performance using a suite of tiled algorithm benchmarks from the heterogeneous math kernels library (MatRIS) based on IRIS. Experimental results demonstrate a performance improvement ranging from 1.6 × to 2 × over IRIS without asynchronous support, and a notable 22% performance enhancement compared to established runtime systems such as StarPU and PaRSEC. This approach significantly improves computation efficiency of HPC workflows and provides a solid base for future exploration and development in the area of asynchronous task programming in heterogeneous systems.

Miniskar, Narasinga Rao [ORNL] (ORCID:000000018259↗

Fine-Grained Application Energy and Power Measurements on the Frontier Exascale System

The increasing complexity and power/energy demands of heterogeneous exascale systems, such as the Frontier supercomputer, present significant challenges for measuring and optimizing power consumption in applications. Current tools either lack the resolution to capture fine-grained power and energy measurements, fail to validate in-band measurements against out-of-band power sensors, or cannot integrate this information with application performance events in a scalable manner. This paper introduces a novel open-source performance toolkit that integrates extended PAPI components with Score-P plugins to enable in-band, fine-grained power and energy measurements, while also supporting validation using power meter measurements for both CPUs and GPUs. One key contribution is the ability to perform millisecond-level power and energy measurements for AMD MI250X GPUs, mapping them to application performance events within a single trace and measurement system that scales. Our toolkit combines coarse-grained measurements from cray_pm counters with high-resolution metrics from rocm_smi and RAPL, converting GPU instantaneous accumulated energy into power to capture both transient and steady-state power behavior, a capability often missed by out-of-band and monitoring tools. By mapping these metrics to specific application regions, developers can identify energy hotspots, address inefficiencies in GPU kernel execution, and validate in-band measurements against external measurements. We demonstrate the effectiveness of this approach through case studies using benchmarks such as GPU rocblas_sgemm, BLIS c_blas_dgemm, and rocHPL, highlighting the variability of the measurements and the impact of transient power spikes on kernel-level efficiency.

Hernandez Mendoza, Oscar [ORNL] (ORCID:00000002538↗

Code for the manuscript "Mori-Zwanzig Modal Decomposition"

We would like to create an open source repository in LANL's github on code written in Julia, in which we implement and extend the data-driven Mori-Zwanzig method for extracting large-scale spatio-temporal structures from data, which we call MZMD. This method is an extension of Dynamic Mode Decomposition (DMD) in which Mori-Zwanzig memory kernels are included into the associated companion matrix. In the code we would like to release, we apply MZMD to a flow over a cylinder with Reynolds number 100 rather than the much larger data set used in the associated manuscript. DMD is used extensively in the fluid dynamics community mainly for extracting large scale spatio-temporal structures (patters) from flow data. This is useful for understanding the key mechanisms that generate certain complex dynamical process relevant in engineering design. In MZMD, we improve upon DMD by adding the Mori-Zwanzig memory kernels, and show this improvement is especially important in strongly nonlinear regions of the flow.

Woodward, Michael↗

Temperature Inversions below 1 km from a V-Band Scanning Radiometer at the North Slope of Alaska

A single-channel (56.7 GHz) scanning radiometer was deployed in August 2022 at the Atmospheric Radiation Measurement (ARM) North Slope of Alaska site near Utqiaġvik. The radiometer is designed to provide temperature profiles between 0 and 1 km every 5 min. Averaging kernels show that this single-channel radiometer, taking observations at 10 discrete elevation angles, yields approximately the same information as a seven-channel V-band radiometer scanning three elevation angles. The instrument is able to reproduce the occurrence of temperature inversions between the surface and 1 km and their strength showing a correlation of 0.85, bias of −0.6 K, and slope of 1.04 with respect to radiosondes. Uncertainty in the inversion base height varies from 50 m near the surface to ∼300 m above 0.4 km when compared with radiosondes. Conversely, the inversion top height is overestimated and has higher uncertainty due to the degrading effects of the averaging kernels on the vertical resolution of the retrievals. Thanks to the high temporal resolution of the retrievals, the diurnal cycle of boundary layer temperature was evaluated showing that the radiometer can capture some aspects of the boundary layer thermal structure. The present analysis provides an overview of the capabilities of this simple observing configuration for selected applications.

Atmospheric profilers↗

Batched sparse direct solver design and evaluation in SuperLU_DIST

Over the course of interactions with various application teams, the need for batched sparse linear algebra functions has emerged in order to make more efficient use of the GPUs for many small and sparse linear algebra problems. In this paper, we present our recent work on a batched sparse direct solver for GPUs. The sparse LU factorization is computed by the levels of the elimination tree, leveraging the batched dense operations at each level and a new batched Scatter GPU kernel. The sparse triangular solve is computed by the level sets of the directed acyclic graph (DAG) of the triangular matrix. Batched operations overcome the large overhead associated with launching many small kernels. For medium sized matrix batches with not-so-small bandwidth, using an NVIDIA A100 GPU, our new batched sparse direct solver is orders of magnitude faster than a batched banded solver and uses less than one-tenth of the memory.

Boukaram, Wajih↗

PETSc/TAO developments for GPU-based early exascale systems

The Portable Extensible Toolkit for Scientific Computation (PETSc) library provides scalable solvers for nonlinear time-dependent differential and algebraic equations and for numerical optimization via the Toolkit for Advanced Optimization (TAO). PETSc is used in dozens of scientific fields and is an important building block for many simulation codes. During the U.S. Department of Energy’s Exascale Computing Project, the PETSc team has made substantial efforts to enable efficient utilization of the massive fine-grain parallelism present within exascale compute nodes and to enable performance portability across exascale architectures. We recap some of the challenges that designers of numerical libraries face in such an endeavor, and then discuss the many developments we have made, which include the addition of new GPU backends, features supporting efficient on-device matrix assembly, better support for asynchronicity and GPU kernel concurrency, and new communication infrastructure. In conclusion, we evaluate the performance of these developments on some pre-exascale systems as well as the early exascale systems Frontier and Aurora, using compute kernel, communication layer, solver, and mini-application benchmark studies, and then close with a few observations drawn from our experiences on the tension between portable performance and other goals of numerical libraries.

Exascale Computing Project (ECP)↗

On nonlocal problems with Neumann boundary conditions: scaling and convergence for nonlocal operators and solutions

Formulations of Neumann-type boundary conditions for boundary value problems in the nonlocal framework are beset with difficulties, some related to the choice of a proper scaling. Here we identify a space-dependent scaling for a nonlocal Neumann operator, for which we prove linear in δ (δ being the radius for the support for the kernel) convergence of the Neumann operator and $\mathcal{O}$(δ 2 ) convergence of solutions to their classical counterparts. The pointwise-like convergence of the nonlocal normal operator is cast as a new type of two-scale operator-point convergence, which we call condensated convergence . The results hold for general integrable kernels, a setting which is favored in numerical simulations. We support this analysis with numerical convergence studies using a piecewise linear discontinuous Galerkin discretization and show an $\mathcal{O}$(δ 2 ) rate of convergence of solutions, also exhibiting an $\mathcal{O}$(h 2 ) convergence, where h is the mesh size.

97 MATHEMATICS AND COMPUTING↗

Milestone 49 Report: Batched Sparse LA Phase 5 Implementation

Batched sparse linear algebra operations in general, and solvers in particular, have become the major algorithmic development activity and foremost performance engineering effort in the numerical software libraries work on modern hardware with accelerators such as GPUs. Many applications, ECP and non-ECP alike, require simultaneous solutions of many small linear systems of equations that are structurally sparse in one form or another. In order to move towards high hardware utilization levels, it is important to provide these applications with appropriate interface designs to be both functionally efficient and performance portable and give full access to the appropriate batched sparse solvers running on modern hardware accelerators prevalent across DOE supercomputing sites since the inception of ECP. To this end, we present here a summary of recent advances on the interface designs in use by HPC software libraries supporting batched sparse linear algebra and the development of sparse batched kernel codes for solvers and preconditioners. We also address the potential interoperability opportunities to keep the corresponding software portable between the major hardware accelerators from AMD, Intel, and NVIDIA, while maintaining the appropriate disclosure levels conforming to the active NDA agreements. The presented interface specifications include a mix of batched band, sparse iterative, and sparse direct solvers with their accompanying functionality that is already required by the application codes or we anticipated to be needed in the near future. This report summarizes progress in Kokkos Kernels and the xSDK libraries MAGMA, Ginkgo, hypre, PETSc, and SuperLU.

97 MATHEMATICS AND COMPUTING↗

Unifying Combinatorial and Graphical Methods in Artificial Intelligence

Recently, a new graph Laplacian, called the inner product Laplacian, was introduced which generalizes many existing Laplacians, including the normalized and combinatorial Laplacian and their weighted variants. The key observation behind the inner product Laplacian is that by defining appropriate inner product spaces on the vertices and edges, the standard Laplacians can be recovered as Hodge Laplacians over the simplicial complex formed by the edges and vertices. These inner product spaces form a natural way to incorporate non-combinatorial information into the definition of a domain-specific Laplacian. In particular, in contrast to current domain-specific weighting schemes which rely solely on edge weights, information regarding the similarity of non-adjacent vertices and arbitrary pairs of edges can be effectively incorporated into the Laplacian. In order to illustrate this approach we consider the problem of calculating the potential energy of an atomistic configuration using Graph Neural Networks. In comparison with start-of-the-art approaches, such as SchNet, our approach replaces a learned (via auto-encoder) representation of the atom types with an inner product space on atoms based on scientific knowledge (e.g., electronegativity). We will illustrate how this approach captures key chemical properties of the molecules and compare the energy calculations with state-of-the-art neural network approaches. However, to compute the resulting Laplacian involves a mixture of sparse and dense matrix computation and yields a dense matrix as the basis for the graph convolution. This dense convolutional kernel necessitates moving away from the standard message passing framework for graph neural networks and increases the computational cost of applying the kernel. In order to mitigate these costs we investigate means of leveraging the mixed sparse and dense computations to reduce the overall computational cost and how these approaches can be automatically transferred to energy efficient hardware (e.g., field programmable gate arrays (FPGAs)).

97 MATHEMATICS AND COMPUTING↗

Improved Weld Residual Stress Modeling System in BlackBear

This report presents enhancements to the MOOSE-based BlackBear application aimed at improving its capability to simulate welding and other thermo-mechanical manufacturing processes. Two primary avenues of improvement are pursued. First, to enhance user accessibility, we introduce a centralized default block restriction mechanism that ensures coverage checks are performed within user-specified default blocks. This default setting is applied consistently to all block-describable objects, such as variables, kernels, and more. In addition, we develop a modular action for moving heat source simulations, which integrates path file parsing, subdomain modification, and heat source kernel enforcement into a single, streamlined configuration. Second, to improve solver robustness, we implement an alternative method for assigning initial conditions to the updated active domain during the simulation, thereby enhancing convergence behavior. To validate the framework, we design and conduct several benchmark simulations, including heat conduction with progressive material addition, linear elasticity with time-dependent material deposition, and viscoplasticity model with isotropic hardening under similar conditions. Finally, we demonstrate the effectiveness of the proposed framework through large-scale thermo-mechanical welding simulations in both two and three dimensions.

42 ENGINEERING↗

Spectrally Stabilized Interface Capturing Formulation and Implementation in Nek5000/NekRS

This report documents the formulation of a novel level-set method for incompressible two-phase flows in the continuous Galerkin (CG) high order spectral element framework. The overall method hinges on a novel implementation of the spectral vanishing viscosity (SVV) operator for the stabilization of linear/non-linear hyperbolic problems. The multidimensional SVV convolution kernels, which in essence, have a similar effect as a high pass filter applied to the derivatives, are formulated by exploiting the tensor product form, analogous to the construction of the usual stiffness matrix system. The resulting kernels are directionally decoupled and ensure a linear, symmetric positive definite, elliptic matrix operator. The SVV formulation is demonstrated to provide a robust stabilizing mechanism through challenging linear and non-linear hyperbolic problems, including problems pertinent to the level-set formulation. The two-phase framework conceptualized herein is based on the conservative level-set (CLS) method which represents the interface between the fluids by the 0.5 iso-contour of the smoothed Heaviside function. The CLS method is augmented with a preconditioning procedure for interface normals using the signed distance function which precludes the manifestation of spurious oscillations in the vicinty of the interface. Further, the existing mixed explicit-implicit approach for the solution of Navier-Stokes equations in Nek5000, as described in Tomboulides et al, is augmented with a pressure coefficient splitting approach for the Poisson equation, which greatly accelerated the convergence of pressure solver for two-phase systems with large density ratio. The robustness and accuracy of the overall two-phase method is demonstrated through canonical challenging problems involving high density and viscosity ratios, with and without surface tension. The two-phase formulation is wholly implemented in Nek5000 and the SVV stabilization method is implemented in NekRS, which is the essential precursor to the two-phase framework, undergoing active development.

97 MATHEMATICS AND COMPUTING↗

Radiation Effects on Network on Chips (NoC) Laboratory Directed Research and Development (LDRD) project

This project was motivated by State-of-the-Art (SOTA) technology that incorporates Network on Chips (NOC) for efficient data communication across the various computer kernels. For example, on the AMD Versal Field Programmable Gate Arrays (FPGA), an NoC has been incorporated for fast data communication from the programmable logic and other computer kernels (processing system, adaptable intelligence engines, etc.). The radiation effects on the legacy technology of this FPGA, such as the programmable logic, are well understood, and established methods exist to measure cross-sections when new families/generations are released; however, newly incorporated technologies, such as the NoC, are not fully understood and could introduce new failure points into the mission space.

36 MATERIALS SCIENCE↗

Explosive Soot Challenge (Final Report)

This project assembled a broad ensemble of modeling and experimentation tools to study the morphological and optical properties of detonation soots in explosive fireballs. A gram-scale hemispherical high explosive was studied in a low-pressure controlled environment using in-situ experimentation with diffusely illuminated visible absorption spectroscopy, particle sizing through light scattering techniques, and post-test collections with subsequent morphological analysis. Hydrocode modeling was performed to replicate the detonation flow observations, and subsequent aerosol kinetics models provided particle size distributions and extinction coefficients from the hydrocode results. Experimentally observed soot morphologies agreed with expectation from the literature - a bimodal distribution was found, brought upon by the particles growing to a size where their inertia and fluid wakes are non-negligible. The aerosol kinetics model did not replicate the observed bimodal size distribution for lack of a coagulation kernel to represent the behavior. To recover particulate optical properties, a spectrally resolved absorption spectroscopy method termed Spectral diffuse back-illuminated extinction imaging (SBI-EI) was developed and implemented on two explosive types. Inverting the absorption spectra using a Kramers-Kronig consistent method yielded the complex index of refraction for the soots produced by the explosives. This method resulted in an unrealistic index of refraction for one of the two explosives, and this is suspected to be due to the model neglecting scattering brought upon by the large particle sizes observed. In addition to the core work, three additional studies were performed in parallel. These investigated the impact of scattering on diffuse absorption spectroscopy, studied how soots oxidate and sublimate in a well-controlled shock tube, and laid the theoretical groundwork for a new collision kernel to replicate the bimodal size distribution from the observations. Summaries of these efforts are included at the end of this report.

45 MILITARY TECHNOLOGY, WEAPONRY, AND NATIONAL DEF↗

Post-Irradiation Examination on MiniFuel UCO and UO 2 TRISO Particles Irradiated in HFIR at High Power

Post-irradiation examination (PIE) of MiniFuel compacts was conducted at Oak Ridge National Laboratory (ORNL) under the Nuclear Science User Facilities project in collaboration with Kairos Power (KP) to evaluate the performance of tristructural-isotropic (TRISO) particles under high particle power and fluoride-salt-cooled high-temperature reactor (FHR)-relevant conditions. MiniFuel compacts containing low-enriched uranium oxide-uranium carbide (LEUCO), low-enriched uranium dioxide (LEUO2), and natural UCO (NUCO) kernels were irradiated for four cycles at ORNL’s High Flux Isotope Reactor (HFIR) at target temperatures between 500°C and 900°C. Post irradiation, the experiment was disassembled at ORNL to recover the MiniFuel subcapsules, which were subsequently punctured to measure fission gas release. Subcapsule disassembly allowed the recovery of components of interest, such as silicon carbide (SiC) thermometry, fuel specimens, fission product sinks, and SiC spacers. The experimental irradiation temperature was confirmed by analyzing the SiC thermometry via dilatometry. PIE on the fuel specimens included gamma spectrometry and deconsolidation leach burn leach, which were complemented by imaging techniques such as x-ray computed tomography, optical microscopy, and electron microscopy. The PIE results provide insight into TRISO particle integrity, fission product retention, coating performance, and kernel migration, informing fuel qualification for application in KP’s FHR concept.

22 GENERAL STUDIES OF NUCLEAR REACTORS↗