24 citations · 69 across the 6 of their papers we have counts for
10 papers
Cyclical Pruning for Sparse Neural Networks
Suraj Srinivas, Andrey Kuzmin, Markus Nagel +3
Current methods for pruning neural network weights iteratively apply magnitude-based pruning on the model weights and re-train the resulting model to recover lost accuracy. In this…
Neural Network Quantization with AI Model Efficiency Toolkit (AIMET)
Sangeetha Siddegowda, Marios Fournarakis, Markus Nagel +3
While neural networks have advanced the frontiers in many machine learning applications, they often come at a high computational cost. Reducing the power and latency of neural netw…
Understanding and Overcoming the Challenges of Efficient Transformer Quantization
Yelysei Bondarenko, Markus Nagel, Tijmen Blankevoort
Transformer-based architectures have become the de-facto standard models for a wide range of Natural Language Processing tasks. However, their memory footprint and high latency are…
A White Paper on Neural Network Quantization
Markus Nagel, Marios Fournarakis, Rana Ali Amjad +3
While neural networks have advanced the frontiers in many applications, they often come at a high computational cost. Reducing the power and latency of neural network inference is…
In-Hindsight Quantization Range Estimation for Quantized Training
Marios Fournarakis, Markus Nagel
Quantization techniques applied to the inference of deep neural networks have enabled fast and efficient execution on resource-constraint devices. The success of quantization durin…
Bayesian Bits: Unifying Quantization and Pruning
Mart van Baalen, Christos Louizos, Markus Nagel +4
We introduce Bayesian Bits, a practical method for joint mixed precision quantization and pruning through gradient based optimization. Bayesian Bits employs a novel decomposition o…