1 paper
Jinguang Wang, Yuexi Yin, Haifeng Sun +5
Quantizing the activations of large language models (LLMs) has been a significant challenge due to the presence of structured outliers. Most existing methods focus on the per-token…