1 paper · 1 filter
Zhe Ding, Su Pan, Duowei Pan
Post-training quantization (PTQ) has become an important technique for reducing the inference cost of Large Language Models (LLMs). While recent mixed-precision methods improve ult…