1 paper
Zhe Ding, Su Pan, Duowei Pan
Post-training quantization (PTQ) has become an important technique for reducing the inference cost of Large Language Models (LLMs). While recent mixed-precision methods improve ult…