1 paper
Mengzhao Chen, Yi Liu, Jiahao Wang +3
Existing weight-activation quantization methods for Large Language Models (LLMs) primarily address channel-wise outliers but often neglect token-wise outliers, which limits the acc…