1 paper · 1 filter
Wonsuk Jang, Thierry Tambe
The rapidly increasing size of large language models (LLMs) presents significant challenges in memory usage and computational costs. Quantizing both weights and activations can add…