From the 1 of 18 linked papers with an AI index.
1 citations · 1 across the 7 of their papers we have counts for
3 papers · 1 filter
OSAQ: Outlier Self-Absorption for Accurate Low-bit LLM Quantization
Zhikai Li, Zhen Dong, Xuewen Liu +2
Large Language Models (LLMs) have demonstrated remarkable capabilities. However, their massive parameter scale leads to significant resource consumption and latency during inferenc…
TTAQ: Towards Stable Post-training Quantization in Continuous Domain Adaptation
Junrui Xiao, Zhikai Li, Lianwei Yang +2
Post-training quantization (PTQ) reduces excessive hardware cost by quantizing full-precision models into lower bit representations on a tiny calibration set, without retraining. D…
RepQuant: Towards Accurate Post-Training Quantization of Large Transformer Models via Scale Reparameterization
Zhikai Li, Xuewen Liu, Jing Zhang +1
Large transformer models have demonstrated remarkable success. Post-training quantization (PTQ), which requires only a small dataset for calibration and avoids end-to-end retrainin…