2 papers
cs.CL2022
Fast DistilBERT on CPUs
Haihao Shen, Ofir Zafrir, Bo Dong +7
Transformer-based language models have become the standard approach to solving natural language processing tasks. However, industry adoption usually requires the maximum throughput…
cs.CV2018
Highly Efficient 8-bit Low Precision Inference of Convolutional Neural Networks with IntelCaffe
Jiong Gong, Haihao Shen, Guoming Zhang +6
High throughput and low latency inference of deep neural networks are critical for the deployment of deep learning applications. This paper presents the efficient inference techniq…