1 paper
Ismail Hossain, Nafi Ullah Shafin, Mohammad Abdullah Al Mumin
Post-training quantization lowers the memory footprint of Large Language Models (LLMs) and speeds up inference, which is why it is now common for on-device deployment. Most of what…