1 paper
Jeremias Bohn, Tizian Dippold, Mahdi Koubaa +2
Quantization has become an invaluable tool to reduce memory requirements and inference speed of modern language models, in particular to make them available for consumer setups and…