activity
20192022
most citedNeural Network Quantization with AI Model Efficiency Toolkit (AIMET)

24 citations · 69 across the 6 of their papers we have counts for

collaborators

10 papers

cs.LG2022

Cyclical Pruning for Sparse Neural Networks

Suraj Srinivas, Andrey Kuzmin, Markus Nagel +3

Current methods for pruning neural network weights iteratively apply magnitude-based pruning on the model weights and re-train the resulting model to recover lost accuracy. In this…

cs.LG202224 cited

Neural Network Quantization with AI Model Efficiency Toolkit (AIMET)

Sangeetha Siddegowda, Marios Fournarakis, Markus Nagel +3

While neural networks have advanced the frontiers in many machine learning applications, they often come at a high computational cost. Reducing the power and latency of neural netw…

cs.LG2021

Understanding and Overcoming the Challenges of Efficient Transformer Quantization

Yelysei Bondarenko, Markus Nagel, Tijmen Blankevoort

Transformer-based architectures have become the de-facto standard models for a wide range of Natural Language Processing tasks. However, their memory footprint and high latency are…

cs.LG202113 cited

A White Paper on Neural Network Quantization

Markus Nagel, Marios Fournarakis, Rana Ali Amjad +3

While neural networks have advanced the frontiers in many applications, they often come at a high computational cost. Reducing the power and latency of neural network inference is…

cs.LG2021

In-Hindsight Quantization Range Estimation for Quantized Training

Marios Fournarakis, Markus Nagel

Quantization techniques applied to the inference of deep neural networks have enabled fast and efficient execution on resource-constraint devices. The success of quantization durin…

cs.LG2020

Bayesian Bits: Unifying Quantization and Pruning

Mart van Baalen, Christos Louizos, Markus Nagel +4

We introduce Bayesian Bits, a practical method for joint mixed precision quantization and pruning through gradient based optimization. Bayesian Bits employs a novel decomposition o…