1 paper
Fei Wang, Chao Xue, Taoran Liu +3
Mixed-precision quantization (MPQ) has become a key technique for deploying large language models under stringent memory and compute constraints. We first identify a phenomenon tha…