1 paper · 1 filter
Hongyuan Liu, Yawei Li, Zhiqiang Que +3
Efficient large language model (LLM) serving is increasingly constrained by deployment cost. Quantization is a key technique for reducing serving cost, yet even state-of-the-art 4-…