1 paper
Xin Nie, Haicheng Zhang, Liang Dong +3
Mixed-precision quantization is a promising approach for compressing large language models under tight memory budgets. However, existing mixed-precision methods typically suffer fr…