Showing cs.CLShow all
2 papers · 1 filter
cs.CL2026
Attend to Your Own Thoughts: Breaking the Barrier for Post-Training Quantization of Reasoning LLMs through the Lens of 1.58-Bit Quantization
Shigeng Wang, Chao Li, Yangyuxuan Kang +2
We propose ScaleQ-1.58, a scalable ternary post-training quantization (PTQ) framework for reasoning LLMs. Its core insight stems from an empirical finding: although modern LLMs are…
cs.CL2026
CAT-Q: Cost-efficient and Accurate Ternary Quantization for LLMs
Shigeng Wang, Chao Li, Yangyuxuan Kang +2
In this paper, we present CAT-Q, Cost-efficient and Accurate Ternary Quantization, for compressing and accelerating LLMs. Unlike existing state-of-the-art ternary quantization meth…