Engineering topics
Wang, Tianqi
Publications and source records attributed to Wang, Tianqi.
O3BNN-R: An Out-Of-Order Architecture for High-Performance and Regularized BNN Inference
Binarized Neural Networks (BNN) have drawn tremendous attention due to significantly reduced computational complexity and memory demand. They have especially shown great potential in cost- and power-restricted domains, such as IoT and smart edge-devices, where reaching a certain accuracy bar is often sufficient, and real-time is highly desired.In this work, we demonstrate that the highly-condensed BNN model can be shrunk significantly further by dynamically pruning irregular redundant edges. Based on two new observations on BNN-specific properties, an out-of-order (OoO) architecture – O3BNN-R, can curtail edge evaluation in cases where the binary output of a neuron can be determined early. Similar to Instruction-Level-Parallelism(ILP), these fine-grained, irregular, runtime pruning opportunities are traditionally presumed to be difficult to exploit. In order to increase the pruning opportunities, we also optimize the training process by adding 2 regularization items in the loss function (1) for pooling pruning and (2) for threshold pruning. We evaluate our design on an FPGA platform using three well-known networks, including VggNet-16, AlexNet for ImageNet, and a VGG-like network for Cifar-10.
Self-Assembled Periodic Nanostructures Using Martensitic Phase Transformations
We describe a novel approach for the rational design and synthesis of self-assembled periodic nanostructures using martensitic phase transformations. We demonstrate this approach in a thin film of perovskite SrSnO 3 with reconfigurable periodic nanostructures consisting of regularly spaced regions of sharply contrasted dielectric properties. The films can be designed to have different periodicities and relative phase fractions via chemical doping or strain engineering. The dielectric contrast within a single film can be tuned using temperature and laser wavelength, effectively creating a variable photonic crystal. Our results show the realistic possibility of designing large-area self-assembled periodic structures using martensitic phase transformations with the potential of implementing "built-to-order" nanostructures for tailored optoelectronic functionalities.
Precursor selection in hybrid molecular beam epitaxy of alkaline-earth stannates
One of the challenges of oxide molecular beam epitaxy (MBE) is the synthesis of oxides containing metals with high electronegativity (metals that are hard to oxidize). The use of reactive organometallic precursors can potentially address this issue. To investigate the formation of radicals in MBE, we explored three carefully chosen metal-organic precursors of tin for SnO 2 and BaSnO 3 growth: tetramethyltin (TMT), tetraethyltin (TET), and hexamethylditin (HMDT). All three precursors produced single-crystalline, atomically smooth, and epitaxial SnO 2 (101) films on r-Al 2 O 3 (101¯2) in the presence of oxygen plasma. The study of growth kinetics revealed reaction-limited and flux-limited regimes except for TET, which also exhibited a decrease in the deposition rate with increasing temperature above ~800 °C. Contrary to these similarities, the performance of these precursors was dramatically different for BaSnO 3 growth. TMT and TET were ineffective in supplying adequate tin, whereas HMDT yielded phase-pure, stoichiometric BaSnO 3 films. Significantly, HMDT resulted in phase-pure and stoichiometric BaSnO 3 films even without the use of an oxygen plasma (i.e., with molecular oxygen alone). Furthermore, these results are discussed using the ability of HMDT to form tin radicals and therefore assisting with Sn → Sn 4+ oxidation reaction. Structural and electronic transport properties of films grown using HMDT with and without oxygen plasma are compared. This study provides guideline for the choice of precursors that will enable the synthesis of metal oxides containing hard-to-oxidize metals using reactive radicals in MBE.
AWB-GCN: A Graph Convolutional Network Accelerator with Runtime Workload Rebalancing
The recent development of deep learning has been mostly focusing on Euclidean data, such as images, videos, audios, etc. However, most real-world information and relation are often expressed as graphs. To efficiently learn from graph data, graph convolutional networks (GCNs) emerge as a promising approach, showing advantages in several practical applications such as social network analysis, knowledge discovery, 3D modeling, motion capturing, etc. Real-world graphs are usually extremely large and imbalanced, posting significant performance demand and design challenges on the hardware dedicated for GCN inference. In this paper, we propose an architecture design called UW-GCN to accelerate graph convolutional network inference. To tackle the major performance bottleneck from workload imbalance, we propose dynamic neighborhood stealing and remote chunk shuffling techniques, relying on hardware flexibility to achieve hardware auto-tuning under negligible area or delay overhead. Specifically, UW-GCN is able to smartly profile the sparse graph pattern while continuously adjusting the workload distribution via routing reconfiguration among parallel processing elements (PEs). The ideal configuration is then reused in the remaining iterations. To the best of our knowledge, this is the first accelerator design particularly for GCN and the first work relying on hardware auto-tuning, which is normally based on software, to achieve near-optimal workload balance in processing sparse structures.