Search NASA⌕ Search

Engineering topics

Shi, Runbin

Publications and source records attributed to Shi, Runbin.

O3BNN-R: An Out-Of-Order Architecture for High-Performance and Regularized BNN Inference

Binarized Neural Networks (BNN) have drawn tremendous attention due to significantly reduced computational complexity and memory demand. They have especially shown great potential in cost- and power-restricted domains, such as IoT and smart edge-devices, where reaching a certain accuracy bar is often sufficient, and real-time is highly desired.In this work, we demonstrate that the highly-condensed BNN model can be shrunk significantly further by dynamically pruning irregular redundant edges. Based on two new observations on BNN-specific properties, an out-of-order (OoO) architecture – O3BNN-R, can curtail edge evaluation in cases where the binary output of a neuron can be determined early. Similar to Instruction-Level-Parallelism(ILP), these fine-grained, irregular, runtime pruning opportunities are traditionally presumed to be difficult to exploit. In order to increase the pruning opportunities, we also optimize the training process by adding 2 regularization items in the loss function (1) for pooling pruning and (2) for threshold pruning. We evaluate our design on an FPGA platform using three well-known networks, including VggNet-16, AlexNet for ImageNet, and a VGG-like network for Cifar-10.

Geng, Tong↗

AWB-GCN: A Graph Convolutional Network Accelerator with Runtime Workload Rebalancing

The recent development of deep learning has been mostly focusing on Euclidean data, such as images, videos, audios, etc. However, most real-world information and relation are often expressed as graphs. To efficiently learn from graph data, graph convolutional networks (GCNs) emerge as a promising approach, showing advantages in several practical applications such as social network analysis, knowledge discovery, 3D modeling, motion capturing, etc. Real-world graphs are usually extremely large and imbalanced, posting significant performance demand and design challenges on the hardware dedicated for GCN inference. In this paper, we propose an architecture design called UW-GCN to accelerate graph convolutional network inference. To tackle the major performance bottleneck from workload imbalance, we propose dynamic neighborhood stealing and remote chunk shuffling techniques, relying on hardware flexibility to achieve hardware auto-tuning under negligible area or delay overhead. Specifically, UW-GCN is able to smartly profile the sparse graph pattern while continuously adjusting the workload distribution via routing reconfiguration among parallel processing elements (PEs). The ideal configuration is then reused in the remaining iterations. To the best of our knowledge, this is the first accelerator design particularly for GCN and the first work relying on hardware auto-tuning, which is normally based on software, to achieve near-optimal workload balance in processing sparse structures.

Geng, Tong↗