Modeling Optimization of Stencil Computations Via Domain-level Properties
Explore the source record for details and available documents.
SEARCH · Search NASA
Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.
Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.
Explore the source record for details and available documents.
This project investigated new techniques for compiling stencil and stencil-like computations. Stencils computations are commonly found in applications such as image processing, physical simulations, image processing, and machine learning. In recent years, many high-performance domain-specific languages (DSLs) have been proposed to optimize stencil computations. To leverage such DSLs, however, existing codes often need to be rewritten. Such rewriting is manual, labor intensive, and error-prone. To alleviate such issues, this project investigated the application of program synthesis and artificial learning techniques to enable stencil computations to automatically leverage new high-performance DSLs. Rather than constructing syntax driven rules, verified lifting uses program synthesis to search for a target code fragment to compile the given input code into. In addition, it also searches for a proof that validates how the found target code fragment preserves the semantics of the original input. Thus, the target code fragment is guaranteed to be semantically equivalent to the input.
We present FPDetect, a low-overhead approach for detecting logical errors and soft errors affecting stencil computations without generating false positives. We develop an offline analysis that tightly estimates the number of floating-point bits preserved across stencil applications. This estimate rigorously bounds the values expected in the data space of the computation. Violations of this bound can be attributed with certainty to errors. FPDetect helps synthesize error detectors customized for user-specified levels of accuracy and coverage. FPDetect also enables overhead reduction techniques based on deploying these detectors coarsely in space and time. Experimental evaluations demonstrate the practicality of our approach.
With the rise of exascale systems and large, data-centric workflows, the need to observe and analyze high performance computing (HPC) applications during their execution is becoming increasingly important. HPC applications are typically not designed with online monitoring in mind, therefore, the observability challenge lies in being able to access and analyze interesting events with low overhead while seamlessly integrating such capabilities into existing and new applications. We explore how our service-based observation, monitoring, and analytics (SOMA) approach to collecting and aggregating both application-specific diagnostic data and performance data addresses these needs. Furthermore, we present our SOMA framework and demonstrate its viability with LULESH, a hydrodynamics proxy application. Then we focus on Astaroth, a multi-GPU library for stencil computations, highlighting the integration of the TAU and APEX performance tools and SOMA for application and performance data monitoring.
Five computationally expensive stencil evaluation routines from SW4(https://github.com/geodynamics/sw4 GPL license) have been extracted and packaged with a driver to create a mini-app for evaluating compiler and GPU performance. The kernels can executed on AMD and Nvidia GPUs with and without RAJA. The kernel driver generates synthetic inputs, checks for correctness and measures kernel run times.
We present a sixth order finite difference scheme for Poisson’s equation when discretized with the compact 27-point stencil based on Mehrstellen corrections of the forcing function term f. Our approach results in a sixth order accurate solution error as opposed to a fourth-order error imposed by the classical Mehrstellen correction for the 19-point and 27-point stencils. The present study is a continuation of former work of Spotz and Carey (1996) on compact finite difference schemes for Poisson’s equation where sixth order convergence may be obtained under the assumption that the fourth order derivatives of f are determined analytically. Specifically, we show that sixth order convergence can still be attained when only values of f at grid points are available. The sixth order Mehrstellen scheme is further coupled with a Method of Local Corrections (MLC) 3D Poisson solver improving to sixth order accuracy the results reported in Kavouklis and Colella (2019). The MLC test case considered involves an adaptive grid that comprises 7.5 billion cells.
The use of limiting methods for high-order numerical approximations of hyperbolic conservation laws generally requires defining an admissible region/bounds for the solution. In this work, we present a novel approach for computing solution bounds and limiting for the Euler equations through the kinetic representation provided by the Boltzmann equation, which allows for extending limiters designed for linear advection directly to the Euler equations. Given an arbitrary set of solution values to compute bounds over (e.g., numerical stencil) and a desired linear advection limiter, the proposed approach yields an analytic expression for the admissible region of particle distribution function values, which may be numerically integrated to yield a set of bounds for the density, momentum, and total energy. Further, these solution bounds are shown to preserve positivity of density/pressure/internal energy and, when paired with a limiting technique, can robustly resolve strong discontinuities while recovering high-order accuracy in smooth regions without any ad hoc corrections (e.g., relaxing the bounds). This approach is demonstrated in the context of an explicit unstructured high-order discontinuous Galerkin/flux reconstruction scheme for a variety of difficult problems in gas dynamics, including cases with extreme shocks and shock-vortex interactions. Furthermore, this work presents a foundation for limiting techniques for more complex macroscopic governing equations that can be derived from an underlying kinetic representation for which admissible solution bounds are not well-understood.
Iterative stencils are used widely across the spectrum of High Performance Computing (HPC) applications. Many efforts have been put into optimizing stencil GPU kernels, given the prevalence of GPU-accelerated supercomputers. To improve the data locality, temporal blocking is an optimization that combines a batch of time steps to process them together. Under the observation that GPUs are evolving to resemble CPUs in some aspects, we revisit temporal blocking optimizations for GPUs. We explore how temporal blocking schemes can be adapted to the new features in the recent Nvidia GPUs, including large scratchpad memory, hardware prefetching, and device-wide synchronization. We propose a novel temporal blocking method, EBISU, which champions low device occupancy to drive aggressive deep temporal blocking on large tiles that are executed tile-by-tile. We compare EBISU with state-of-the-art temporal blocking libraries: STENCILGEN and AN5D. We also compare with state-of-the-art stencil auto-tuning tools that are equipped with temporal blocking optimizations: ARTEMIS and DRSTENCIL. Over a wide range of stencil benchmarks, EBISU achieves speedups up to 2.53x and a geometric mean speedup of 1.49x over the best state-of-the-art performance in each stencil benchmark.
The paper deals with a new effective numerical technique on unfitted Cartesian meshes for simulations of heterogeneous elastic materials. Here, we develop the optimal local truncation error method (OLTEM) with 27- point stencils (similar to those for linear finite elements) for the 3-D time-independent elasticity equations with irregular interfaces. Only displacement unknowns at each internal Cartesian grid point are used. The interface conditions are added to the expression for the local truncation error and do not change the width of the stencils. The unknown stencil coefficients are calculated by the minimization of the local truncation error of the stencil equations and yield the optimal second order of accuracy for OLTEM with the 27-point stencils on unfitted Cartesian meshes. A new post-processing procedure for accurate stress calculations has been developed. Similar to basic computations it uses OLTEM with the 27-point stencils and the elasticity equations. The post-processing procedure can be easily extended to unstructured meshes and can be independently used with existing numerical techniques (e.g., with finite elements). Numerical experiments show that at an accuracy of 0.1% for stresses, OLTEM with the new post-processing procedure significantly (by 10 5 -10 9 times) reduces the number of degrees of freedom compared to linear finite elements. OLTEM with the 27-point stencils yields even more accurate results than high-order finite elements with wider stencils.
For the first time the optimal local truncation error method (OLTEM) with 125-point stencils and unfitted Cartesian meshes has been developed in the general 3-D case for the Poisson equation for heterogeneous materials with smooth irregular interfaces. The 125-point stencils equations that are similar to those for quadratic finite elements are used for OLTEM. The interface conditions for OLTEM are imposed as constraints at a small number of interface points and do not require the introduction of additional unknowns, i.e., the sparse structure of global discrete equations of OLTEM is the same for homogeneous and heterogeneous materials. The stencils coefficients of OLTEM are calculated by the minimization of the local truncation error of the stencil equations. These derivations include the use of the Poisson equation for the relationship between the different spatial derivatives. Such a procedure provides the maximum possible accuracy of the discrete equations of OLTEM. In contrast to known numerical techniques with quadratic elements and third order of accuracy on conforming and unfitted meshes, OLTEM with the 125-point stencils provides 11-th order of accuracy, i.e., an extremely large increase in accuracy by 8 orders for similar stencils. The numerical results show that OLTEM yields much more accurate results than high-order finite elements with much wider stencils. The increased numerical accuracy of OLTEM leads to an extremely large increase in computational efficiency. Additionally, a new post-processing procedure with the 125-point stencil has been developed for the calculation of the spatial derivatives of the primary function. The post-processing procedure includes the minimization of the local truncation error and the use of the Poisson equation. It is demonstrated that the use of the partial differential equation (PDE) for the 125-point stencils improves the accuracy of the spatial derivatives by 6 orders compared to post-processing without the use of PDE as in existing numerical techniques. At an accuracy of 0.1% for the spatial derivatives, OLTEM reduces the number of degrees of freedom by 900 - 4∙10 6 times compared to quadratic finite elements. The developed post-processing procedure can be easily extended to unstructured meshes and can be independently used with existing post-processing techniques (e.g., with finite elements).
We have developed the Optimal Local Truncation Error Method (OLTEM) with 10-th order of accuracy on unfitted Cartesian meshes for a system of 3-D elasticity equations with smooth irregular interfaces. 5 x 5 x 5 = 125-point stencils (similar to those for quadratic finite elements) for elastic heterogeneous materials are used for OLTEM. There are no unknowns at the interface points between different materials; the structure of the global discrete equations is the same for homogeneous and heterogeneous materials. The calculation of unknown stencil coefficients is based on the minimization of the local truncation error of the stencil equations and yields the optimal 10-th order of accuracy for OLTEM on unfitted Cartesian meshes, i.e., the increase by 7 orders in accuracy compared to quadratic finite elements on conformal meshes. A new post-processing procedure provides the 9-th order of accuracy for stresses in the 3-D case. Similar to basic computations it uses OLTEM with the 125-point stencils, the interface conditions and the elasticity equations. It was shown that the use of the elasticity equations for post-processing improves the accuracy of 0.1% stresses by 6 orders compared to post-processing without the use of PDEs. At an accuracy of for stresses, OLTEM with the new post-processing procedure reduces the number of degrees of freedom by 360 - 8000 times compared to quadratic finite elements with similar stencils. OLTEM with the 125-point stencils yields even more accurate results than high-order finite elements with much wider stencils. OLTEM provides accurate numerical results for compressible and nearly incompressible materials.
We evaluate AI-assisted generative capabilities on fundamental numerical kernels in high-performance computing (HPC), including AXPY, GEMV, GEMM, SpMV, Jacobi Stencil, and CG. We test the generated kernel codes for a variety of language-supported programming models, including (1) C++ (e.g., OpenMP [including offload], OpenACC, Kokkos, SyCL, CUDA, and HIP), (2) Fortran (e.g., OpenMP [including offload] and OpenACC), (3) Python (e.g., numpy, Numba, cuPy, and pyCUDA), and (4) Julia (e.g., Threads, CUDA.jl, AMDGPU.jl, and KernelAbstractions.jl). We use the GitHub Copilot capabilities powered by the GPT-based OpenAI Codex available in Visual Studio Code as of April 2023 to generate a vast amount of implementations given simple + + prompt variants. To quantify and compare the results, we propose a proficiency metric around the initial 10 suggestions given for each prompt. Results suggest that the OpenAI Codex outputs for C++ correlate with the adoption and maturity of programming models. For example, OpenMP and CUDA score really high, whereas HIP is still lacking. We found that prompts from either a targeted language such as Fortran or the more general purpose Python can benefit from adding code keywords, while Julia prompts perform acceptably well for its mature programming models (e.g., Threads and CUDA.jl). We expect for these benchmarks to provide a point of reference for each programming model's community. Overall, understanding the convergence of large language models, AI, and HPC is crucial due to its rapidly evolving nature and how it is redefining human-computer interactions.
We explore the performance and portability of the novel Mojo language for scientific computing workloads on GPUs. As the first language based on the LLVM’s Multi-Level Intermediate Representation (MLIR) compiler infrastructure, Mojo aims to close performance and productivity gaps by combining Python’s interoperability and CUDA-like syntax for compile-time portable GPU programming. We target four scientific workloads: a seven-point stencil (memory-bound), BabelStream (memory-bound), miniBUDE (compute-bound), and Hartree–Fock (compute-bound with atomic operations); and compare their performance against vendor baselines on NVIDIA H100 and AMD MI300A GPUs. We show that Mojo’s performance is competitive with CUDA and HIP for memory-bound kernels, whereas gaps exist on AMD GPUs for atomic operations and for fast-math compute-bound kernels on both AMD and NVIDIA GPUs. Although the learning curve and programming requirements are still fairly low-level, Mojo can close significant gaps in the fragmented Python ecosystem in the convergence of scientific computing and AI.
We present a higher-order finite volume method for solving elliptic PDEs with jump conditions on interfaces embedded in a 2D Cartesian grid. Second, fourth, and sixth order accuracy is demonstrated on a variety of tests including problems with high-contrast and spatially varying coefficients, large discontinuities in the source term, and complex interface geometries. We include a generalized truncation error analysis based on cell-centered Taylor series expansions, which then define stencils in terms of local discrete solution data and geometric information. In the process, we develop a simple method based on Green's theorem for computing exact geometric moments directly from an implicit function definition of the embedded interface. This approach produces stencils with a simple bilinear representation, where spatially-varying coefficients and jump conditions can be easily included and finite volume conservation can be enforced.
We evaluate Julia as a single language and ecosystem paradigm powered by LLVM to develop workflow components for high-performance computing. We run a Gray-Scott, 2-variable diffusion-reaction application using a memory-bound, 7-point stencil kernel on Frontier, the US Department of Energy’s first exascale supercomputer. We evaluate the performance, scaling, and trade-offs of (i) the computational kernel on AMD’s MI250x GPUs, (ii) weak scaling up to 4,096 MPI processes/GPUs or 512 nodes, (iii) parallel I/O writes using the ADIOS2 library bindings, and (iv) Jupyter Notebooks for interactive analysis. Results suggest that although Julia generates a reasonable LLVM-IR, a nearly 50% performance difference exists vs. native AMD HIP stencil codes when running on the GPUs. As expected, we observed near-zero overhead when using MPI and parallel I/O bindings for system-wide installed implementations. Consequently, Julia emerges as a compelling high-performance and high-productivity workflow composition language, as measured on the fastest supercomputer in the world.
Today, multiGPU nodes are widely used in high-performance computing and data centers. However, current programming models do not provide simple, transparent, and portable support for automatically targeting multiple GPUs within a node on application areas of array programming. In this paper, we describe a new application programming interface based on the Kokkos programming model to enable array computation on multiple GPUs in a transparent and portable way across both NVIDIA and AMD GPUs. We implement different variations of this technique to accommodate the exchange of stencils (array boundaries) among different GPU memory spaces, and we provide autotuning to select the proper number of GPUs, depending on the computational cost of the operations to be computed on arrays, that is completely transparent to the programmer. We evaluate our multiGPU extension on Summit (#5 TOP500), with six NVIDIA V100 Volta GPUs per node, and Crusher that contains identical hardware/software as Frontier (#1 TOP500), with four AMD MI250X GPUs, each with 2 Graphics Compute Dies (GCDs)for a total of 8 GCDs per node. We also compare the performance of this solution against the use of MPI + Kokkos, which is the cur-rent de facto solution for multiple GPUs in Kokkos. Our evaluation shows that the new Kokkos solution provides good scalability for many GPUs and a faster and simpler solution (from a programming productivity perspective) than MPI + Kokkos.
Explore the source record for details and available documents.
In this work, we propose an Exponential DG framework for partial differential equations. We decompose 7 governing equations into linear and nonlinear parts to which we apply the discontinuous Galerkin 8 (DG) spatial discretization. In particular, we construct the linear part using Jacobian that effectively 9 capture stiff characteristics in the system. The former is integrated analytically, whereas the latter 10 is approximated. This approach i) is stable with a large Courant number (Cr > 1); ii) supports 11 high-order solutions both in time and space; iii) is computationally favorable compared to IMEX 12 DG methods with no preconditioner; iv) becomes comparable to explicit RKDG methods on uniform 13 mesh and beneficial on non-uniform grid for Euler equations; v) is scalable in a modern massively 14 parallel computing architecture due to its explicit nature of exponential time integrators and com15 pact communication stencil of DG method. Numerical results demonstrate the performance of our 16 proposed methods through various examples. We also discuss the stability and convergence analysis 17 for our exponential DG scheme in the context of Burgers equation.