24 citations · 204 across the 35 of their papers we have counts for
4 papers · 2 filters
FP8 Quantization: The Power of the Exponent
Andrey Kuzmin, Mart Van Baalen, Yuwei Ren +3
When quantizing neural networks for efficient inference, low-bit integers are the go-to format for efficiency. However, low-bit floating point numbers have an extra degree of freed…
Overcoming Oscillations in Quantization-Aware Training
Markus Nagel, Marios Fournarakis, Yelysei Bondarenko +1
When training neural networks with simulated quantization, we observe that quantized weights can, rather unexpectedly, oscillate between two grid-points. The importance of this eff…
Cyclical Pruning for Sparse Neural Networks
Suraj Srinivas, Andrey Kuzmin, Markus Nagel +3
Current methods for pruning neural network weights iteratively apply magnitude-based pruning on the model weights and re-train the resulting model to recover lost accuracy. In this…
Neural Network Quantization with AI Model Efficiency Toolkit (AIMET)
Sangeetha Siddegowda, Marios Fournarakis, Markus Nagel +3
While neural networks have advanced the frontiers in many machine learning applications, they often come at a high computational cost. Reducing the power and latency of neural netw…