3 papers
cs.CL2026
CAT-Q: Cost-efficient and Accurate Ternary Quantization for LLMs
Shigeng Wang, Chao Li, Yangyuxuan Kang +2
In this paper, we present CAT-Q, Cost-efficient and Accurate Ternary Quantization, for compressing and accelerating LLMs. Unlike existing state-of-the-art ternary quantization meth…
cs.AI2026
SliderQuant: Accurate Post-Training Quantization for LLMs
Shigeng Wang, Chao Li, Yangyuxuan Kang +3
In this paper, we address post-training quantization (PTQ) for large language models (LLMs) from an overlooked perspective: given a pre-trained high-precision LLM, the predominant…
cs.CV2024
ScaleKD: Strong Vision Transformers Could Be Excellent Teachers
Jiawei Fan, Chao Li, Xiaolong Liu +1
In this paper, we question if well pre-trained vision transformer (ViT) models could be used as teachers that exhibit scalable properties to advance cross architecture knowledge di…