24 citations · 37 across the 2 of their papers we have counts for
2 papers
cs.LG2023★ 24 cited
FP8 versus INT8 for efficient deep learning inference
Mart van Baalen, Andrey Kuzmin, Suparna S Nair +8
Recently, the idea of using FP8 as a number format for neural network training has been floating around the deep learning world. Given that most training is currently conducted wit…
cs.LG2022★ 13 cited
FP8 Quantization: The Power of the Exponent
Andrey Kuzmin, Mart Van Baalen, Yuwei Ren +3
When quantizing neural networks for efficient inference, low-bit integers are the go-to format for efficiency. However, low-bit floating point numbers have an extra degree of freed…