Publications (7)
ECI: a Customizable Cache Coherency Stack for Hybrid FPGA-CPU Architectures
Abishek Ramdas, Michael Giardino, Runbin Shi +4
Unlike other accelerators, FPGAs are capable of supporting cache coherency, thereby turning them into a more powerful architectural option than just a peripheral accelerator. Howev…
Mix and Match: A Novel FPGA-Centric Deep Neural Network Quantization Framework
Sung-En Chang, Yanyu Li, Mengshu Sun +5
Deep Neural Networks (DNNs) have achieved extraordinary performance in various application domains. To support diverse DNN models, efficient implementations of DNN inference on edg…
CSB-RNN: A Faster-than-Realtime RNN Acceleration Framework with Compressed Structured Blocks
Runbin Shi, Peiyan Dong, Tong Geng +6
Recurrent neural networks (RNNs) have been widely adopted in temporal sequence analysis, where realtime performance is often in demand. However, RNNs suffer from heavy computationa…
AWB-GCN: A Graph Convolutional Network Accelerator with Runtime Workload Rebalancing
Tong Geng, Ang Li, Runbin Shi +8
Deep learning systems have been successfully applied to Euclidean data such as images, video, and audio. In many applications, however, information and their relationships are bett…
Co-design Hardware and Algorithm for Vector Search
Wenqi Jiang, Shigang Li, Yu Zhu +8
Vector search has emerged as the foundation for large-scale information retrieval and machine learning systems, with search engines like Google and Bing processing tens of thousand…
MSP: An FPGA-Specific Mixed-Scheme, Multi-Precision Deep Neural Network Quantization Framework
Sung-En Chang, Yanyu Li, Mengshu Sun +4
With the tremendous success of deep learning, there exists imminent need to deploy deep learning models onto edge devices. To tackle the limited computing and storage resources in…