1 paper · 1 filter
Yao Fu, Xianxuan Long, Runchao Li +5
Quantization enables efficient deployment of large language models (LLMs) in resource-constrained environments by significantly reducing memory and computation costs. While quantiz…