5 citations · 13 across the 14 of their papers we have counts for
1 paper · 1 filter
Yijia Zhang, Lingran Zhao, Shijie Cao +6
Efficient deployment of large language models (LLMs) necessitates low-bit quantization to minimize model size and inference cost. While low-bit integer formats (e.g., INT8/INT4) ha…