1 paper
Gunjun Lee, Sehwan Son, Younjoo Lee +2
Weight-only post-training quantization (PTQ) enables the deployment of large language models under tight memory budgets, but accuracy often collapses at 2-3 bits. Existing backprop…