1 paper · 1 filter
Ke Li, Dong An, Xiaoling Zang +6
Low-bit activation quantization remains a major bottleneck in efficient large language model (LLM) deployment. The difficulty is not only that activations contain outliers, but tha…