1 paper · 1 filter
Qingcheng Zhu, Yangyang Ren, Linlin Yang +6
Deploying large language models (LLMs) is challenging due to their massive parameters and high computational costs. Ultra low-bit quantization can significantly reduce storage and…