2 papers
cs.CL2026
UniSVQ: 2-bit Unified Scalar-Vector Quantization
Haoyu Wang, Haiyan Zhao, Xingyu Yu +4
Post-training quantization at the 2-bit level enables low-cost deployment and inference acceleration for large language models (LLMs). Scalar quantization (SQ) and vector quantizat…
cs.LG2026
GAMMA: Global Bit Allocation for Mixed-Precision Models under Arbitrary Budgets
Zhangyang Yao, Haiyan Zhao, Haoyu Wang +3
Mixed-precision quantization improves the budget--accuracy trade-off for large language models (LLMs) by allocating more bits to sensitive modules. However, automating this allocat…