1 paper
Deokjae Lee, Sihun Chu, Hyun Oh Song
Mixed-precision quantization (MPQ) assigns a different bitwidth to each linear layer of a large language model (LLM) to minimize the quantization-induced quality loss under a fixed…