1 paper · 1 filter
Jiatong Ding, Bingxin Xing, Yu Zhang +9
Quantization has been widely adopted in LLM training and inference to reduce cost and improve efficiency. However, low-bit quantization of the \emph{attention} module often introdu…