220 citations · 301 across the 4 of their papers we have counts for
Showing cs.LGShow all
2 papers · 1 filter
cs.LG2022★ 50 cited
FP8 Formats for Deep Learning
Paulius Micikevicius, Dusan Stosic, Neil Burgess +12
FP8 is a natural progression for accelerating deep learning training inference beyond the 16-bit formats common in modern processors. In this paper we propose an 8-bit floating poi…
cs.LG2020★ 220 cited
Integer Quantization for Deep Learning Inference: Principles and Empirical Evaluation
Hao Wu, Patrick Judd, Xiaojie Zhang +2
Quantization techniques can reduce the size of Deep Neural Networks and improve inference latency and throughput by taking advantage of high throughput integer instructions. In thi…