activity
20182021
most citedAPQ: Joint Search for Network Architecture, Pruning and Quantization Policy

22 citations · 59 across the 4 of their papers we have counts for

collaborators

5 papers

cs.LG20214 cited

How to Design Sample and Computationally Efficient VQA Models

Karan Samel, Zelin Zhao, Binghong Chen +3

In multi-modal reasoning tasks, such as visual question answering (VQA), there have been many modeling and training paradigms tested. Previous models propose different methods for…

cs.CV202016 cited

Hardware-Centric AutoML for Mixed-Precision Quantization

Kuan Wang, Zhijian Liu, Yujun Lin +2

Model quantization is a widely used technique to compress and accelerate deep neural network (DNN) inference. Emergent DNN hardware accelerators begin to support mixed precision (1…

cs.LG202022 cited

APQ: Joint Search for Network Architecture, Pruning and Quantization Policy

Tianzhe Wang, Kuan Wang, Han Cai +3

We present APQ for efficient deep learning inference on resource-constrained hardware. Unlike previous methods that separately search the neural architecture, pruning policy, and q…

cs.LG201917 cited

Design Automation for Efficient Deep Learning Computing

Song Han, Han Cai, Ligeng Zhu +4

Efficient deep learning computing requires algorithm and hardware co-design to enable specialization: we usually need to change the algorithm to reduce memory footprint and improve…

cs.CV2018

HAQ: Hardware-Aware Automated Quantization with Mixed Precision

Kuan Wang, Zhijian Liu, Yujun Lin +2

Model quantization is a widely used technique to compress and accelerate deep neural network (DNN) inference. Emergent DNN hardware accelerators begin to support mixed precision (1…