13 citations · 13 across the 2 of their papers we have counts for
2 papers
cs.LG2021
Understanding and Overcoming the Challenges of Efficient Transformer Quantization
Yelysei Bondarenko, Markus Nagel, Tijmen Blankevoort
Transformer-based architectures have become the de-facto standard models for a wide range of Natural Language Processing tasks. However, their memory footprint and high latency are…
cs.LG2021★ 13 cited
A White Paper on Neural Network Quantization
Markus Nagel, Marios Fournarakis, Rana Ali Amjad +3
While neural networks have advanced the frontiers in many applications, they often come at a high computational cost. Reducing the power and latency of neural network inference is…