Search NASA⌕ Search

Engineering topics

Zhang, Chengming

Publications and source records attributed to Zhang, Chengming.

H-GCN: A Graph Convolutional Network Accelerator on Versal ACAP Architecture

Recently Graph Neural Networks (GNNs) have drawn tremendous attentions due to their unique capability to extend the Machine Learning (ML) approaches to broadly defined applications with unstructured data, especially graphs. Comparing with other ML modalities, the acceleration of GNNs is as critical but even more challenging due to the irregularity and heterogeneity from graph typologies that together limit the performance. Existing efforts mainly focus on handling graphs’ irregularity, however, have not studied the heterogeneity. To this end, in this work, we propose H-GCN, a PL-AIE-based hybrid accelerator that leverages the emerging heterogeneity of Xilinx Versal ACAPs to achieve high-performance GNN inference. In particular, H-GCN partitions each graph into three subgraphs based on its inherent heterogeneity and processes them using PL and the newly emerged AIE respectively. To further improve the performance, we explore the sparsity support of AIE and develop an efficient density-aware method to map tiles of SpMM onto the systolic tensor array automatically. Compared with the current state-of-the-art GCN accelerator, HGCN achieves on average 1.5× speedups.

Zhang, Chengming↗

CEAZ: Accelerating Parallel I/O Via Hardware-Algorithm Co-Designed Adaptive Lossy Compression

As supercomputers continue to grow to exa-scale, the amount of data that needs to be saved or transmitted is exploding. To this end, many previous works have studied using error-bounded lossy compressors to reduce the data size and improve the I/O performance. However, little work has been done for effectively offloading lossy compression onto FPGA-based SmartNICs to reduce the compression overhead. In this paper, we propose a hardware-algorithm co-design of efficient and adaptive lossy compressor for scientific data on FPGAs (called CEAZ) to accelerate parallel I/O. Our contribution is fourfold: (1) We propose an efficient Huffman coding approach that can adaptively update Huffman codewords online based on codewords generated offline (from a variety of representative scientific datasets). (2) We derive a theoretical analysis to support a precise control of compression ratio under an error-bounded compression mode, enabling accurate offline Huffman codewords generation. This also help us create a fixed-ratio compression mode for consistent throughput. (3) We develop an efficient compression pipeline by adopting cuSZ’s dual-quantization algorithm to our hardware use case. (4) We evaluate CEAC on five real-world datasets with both a single FPGA board and 256 nodes from Bridges2 supercomputer. Experiments show that CEAZ outperforms the second-best FPGA-based lossy compressor by 2× of throughput and 9.6× of compression ratio. It also improves MPI_File_write and MPI_Gather throughputs by up to 32.7× and 31.4×, respectively.

Zhang, Chengming↗