2 papers
cs.LG2025
First-Order Error Matters: Accurate Compensation for Quantized Large Language Models
Xingyu Zheng, Haotong Qin, Yuye Li +5
Post-training quantization (PTQ) offers an efficient approach to compressing large language models (LLMs), significantly reducing memory access and computational costs. Existing co…
cs.LG2025
An Empirical Study of Qwen3 Quantization
Xingyu Zheng, Yuye Li, Haoran Chu +7
The Qwen series has emerged as a leading family of open-source Large Language Models (LLMs), demonstrating remarkable capabilities in natural language understanding tasks. With the…