Search NASA⌕ Search

DOE OSTI · 1848079

Threaded Multi-Core GEMM with MoA and Cache-Blocking: Preprint

Abstract

A threaded multi-core implementation of the high performance dense linear algebra matrix-matrix multiply GEMM kernel is described. This kernel is widely implemented by vendors in the basic linear algebra subroutine BLAS library. The mathematics of arrays (MoA) paradigm due to Mullin (1988) results in contiguous memory accesses by employing outer-product forms. Our performance studies demonstrate that the MoA implementation of double precision DGEMM combined with optimal cache-blocking strategies results in at least a 25% performance gain on the Intel Xeon Skylake processor over the vendor supplied Intel MKL basic linear algebra libraries. Results are presented for the NREL Eagle supercomputer. The multi-core DGEMM achieves over 100 GigaFlops/sec with eight openMP threads.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Thomas, Stephen, Mullin, Lenore, Swirydowicz, Kasia, Khan, Rishi. 2022-03-01. Threaded Multi-Core GEMM with MoA and Cache-Blocking: Preprint. https://www.osti.gov/biblio/1848079

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related reports

Improving the Performance of DGEMM with MoA and Cache-Blocking: Preprint

The goal of this paper is to demonstrate performance enhancements of the high performance dense linear algebra matrix-matrix multiply DGEMM kernel, widely implemented by vendors in the basic linear algebra subroutine BLAS library. The mathematics of arrays (MoA) paradigm due to Mullin (1988) results in contiguous memory accesses in combination with Church-Rosser complete language constructs optimized for target processor architectures [3]. Our performance studies demonstrate that the MoA implementation of DGEMM combined with optimal cache-blocking strategies results in at least a 25% performance gain on both Intel Xeon Skylake and IBM Power-9 processors over the vendor supplied Intel MKL and IBM ESSL basic linear algebra libraries. Results are presented for the NREL Eagle and ORNL Summit supercomputers.

cache-blocking↗