Search NASA⌕ Search

SEARCH · Search NASA

Results for “CUDA-aware”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

GPU-accelerated DNS of compressible turbulent flows

Here, this paper explores strategies to transform an existing CPU-based high-performance computational fluid dynamics solver, HyPar, for compressible flow simulations on emerging exascale heterogeneous (CPU+GPU) computing platforms. The scientific motivation for developing a GPU-enhanced version of HyPar is to simulate canonical turbulent flows at the highest resolution possible on such platforms. We show that optimizing memory operations and thread blocks results in 200x speedup of computationally intensive kernels compared with a CPU core. Using multiple GPUs and CUDA-aware MPI communication, we demonstrate both strong and weak scaling of our GPU-based HyPar implementation on the NVIDIA Volta V100 GPUs. We simulate the decay of homogeneous isotropic turbulence in a triply periodic box on grids with up to 1024 3 points (5.3 billion degrees of freedom) and on up to 1,024 GPUs. We compare the wall times for CPU-only and CPU+GPU simulations. The results presented in the paper are obtained on the Summit and Lassen supercomputers at Oak Ridge and Lawrence Livermore National Laboratories, respectively.

97 MATHEMATICS AND COMPUTING↗

Optimizing the hypre solver for manycore and GPU architectures

The solution of large-scale combustion problems with codes such as Uintah on modern computer architectures requires the use of multithreading and GPUs to achieve performance. Uintah uses a low-Mach number approximation that requires iteratively solving a large system of linear equations. The Hypre iterative solver has solved such systems in a scalable way for Uintah, but the use of OpenMP with Hypre leads to at least slowdown due to OpenMP overheads. The proposed solution uses the MPI Endpoints within Hypre, where each team of threads acts as a different MPI rank. This approach minimizes OpenMP synchronization overhead and performs as fast or (up to 1.44) faster than Hypre's MPI-only version, and allows the rest of Uintah to be optimized using OpenMP. The profiling of the GPU version of Hypre shows the bottleneck to be the launch overhead of thousands of micro-kernels. The GPU performance was improved by fusing these micro-kernels and was further optimized by using Cuda-aware MPI, resulting in an overall speedup of 1.16—1.44 compared to the baseline GPU implementation. The above optimization strategies were published in the International Conference on Computational Science 2020 [1]. This work extends the previously published research by carrying out the second phase of communication-centered optimizations in Hypre to improve its scalability on large-scale supercomputers. Additionally, this includes an efficient non-blocking inter-thread communication scheme, communication-reducing patch assignment, and expression of logical communication parallelism to a new version of the MPICH library that utilizes the underlying network parallelism [2]. The above optimizations avoid communication bottlenecks previously observed during strong scaling and improve performance by up to 2 on 256 nodes of Intel Knight's Landing processor.

97 MATHEMATICS AND COMPUTING↗

Accelerating the Lagrangian Particle Tracking in Hydrologic Modeling to Continental-Scale

Unprecedented climate change and anthropogenic activities have induced increasing ecohydrological problems, which have motivated the development of large-scale modeling for solutions. Water age/quality is as important as water quantity for understanding the water cycle. However, current scientific progress in tracking water parcels at large-scale with high spatiotemporal resolutions is far behind that in water balance/quantity owing to the lack of powerful tools. EcoSLIM is a particle tracking model that works with the hydrologic model ParFlow-CLM, which couples surface-subsurface hydrology with land surface processes. Here, we demonstrate a parallel framework to accelerate EcoSLIM to continental-scale on a distributed, multi-GPU platform with CUDA-Aware MPI. In tests from catchment-, to regional-, and then to continental-scale using 25-million to 1.6-billion particles, EcoSLIM shows significant speedup and excellent parallel performance. The parallel framework is portable to atmospheric and oceanic particle tracking models, where parallelization is inadequate and a standard parallel framework is absent. Parallelized EcoSLIM is a promising tool to accelerate our understanding of the terrestrial water cycle and the upscaling of subsurface hydrology to Earth system models.

54 ENVIRONMENTAL SCIENCES↗

TEMPI: An Interposed MPI Library with Canonical Representation of MPI Datatypes [Poster]

TEMPI provides a transparent non-contiguous data-handling layer compatible with various MPIs. MPI Datatypes are a powerful abstraction for allowing an MPI implementation to operate on non-contiguous data. CUDA-aware MPI implementations must also manage transfer of such data between the host system and GPU. The non-unique and recursive nature of MPI datatypes mean that providing fast GPU handling is a challenge. The same noncontiguous pattern may be described in a variety of ways, all of which should be treated equivalently by an implementation. This work introduces a novel technique to do this for strided datatypes. Methods for transferring non-contiguous data between the CPU and GPU depends on the properties of the data layout. This work shows that a simple performance model can accurately select the fastest method. Unfortunately, the combination of MPI software and system hardware available may not provide sufficient performance. The contributions of this work are deployed on OLCF Summit through an interposer library which does not require privileged access to the system to use

97 MATHEMATICS AND COMPUTING↗

Venado acceptance: results and tips [Slides]

Nvidia compiler support is not available through cray-mpich/compiler wrapper interface. Adjust CMAKE files to use the correct COMPILER_ID in conditionals and explicit variables to package flags. Use pinned host memory in cray-libsci_acc and cublasXt calls. Set a large blockDim for cublasXt calls. Try MPS and/or explicit numactl binding if performance is lackluster. Use CRAY_MALLOPT_OFF=1 if unexpected OOM errors appear using cce. Use MPICH_SMP_SINGLE_COPY_MODE=CMA for xpmem issues. Use MPICH_OPT_THREAD_SYNC=0 for MPI_THREAD issues. Poor CUDA-aware MPI performance remains an issue.

97 MATHEMATICS AND COMPUTING↗

Characterizing the performance of node-aware strategies for irregular point-to-point communication on heterogeneous architectures

Supercomputer architectures are trending toward higher computational throughput due to the inclusion of heterogeneous compute nodes. These multi-GPU nodes increase on-node computational efficiency, while also increasing the amount of data to be communicated and the number of potential data flow paths. In this work, we characterize the performance of irregular point-to-point communication with MPI on heterogeneous compute environments through performance modeling, demonstrating the limitations of standard communication strategies for both device-aware and staging-through-host communication techniques. Presented models suggest staging communicated data through host processes then using node-aware communication strategies for high inter-node message counts. Notably, the models also predict that node-aware communication utilizing all available CPU cores to communicate inter-node data leads to the most performant strategy when communicating with a high number of nodes. Furthermore, model validation is provided via a case study of irregular point-to-point communication patterns in distributed sparse matrix–vector products. Importantly, we include a discussion on the implications model predictions have on communication strategy design for emerging supercomputer architectures.

97 MATHEMATICS AND COMPUTING↗

Modeling Data Movement Performance on Heterogeneous Architectures

The cost of data movement on parallel systems varies greatly with machine architecture, job partition, and nearby jobs. Performance models that accurately capture the cost of data movement provide a tool for analysis, allowing for communication bottlenecks to be pinpointed. Modern heterogeneous architectures yield increased variance in data movement as there are a number of viable paths for inter-GPU communication. In this paper, we present performance models for the various paths of inter-node communication on modern heterogeneous architectures, including the trade-off between GPUDirect communication and copying to CPUs. Furthermore, we present a novel optimization for inter-node communication based on these models, utilizing all available CPU cores per node. Finally, we show associated performance improvements for MPI collective operations.

97 MATHEMATICS AND COMPUTING↗

Evaluation of PETSc on a Heterogeneous Architecture, the OLCF Summit System: Part II - Basic Communication Performance

Nearest-neighbor communication is at the heart of many high-performance parallel computations. We report on the performance of such communication on the Oak Ridge Leadership Computing Facility system Summit in the context of the PETSc communication module. The analysis in this report includes basic Ping-Pong point-to-point communication and regular and irregular nearest-neighbor communication.

97 MATHEMATICS AND COMPUTING↗