1 paper
Phong Nam Huu Nguyen, Khoi M. Le, Cong-Duy T Nguyen +3
Quantization is an effective approach to reduce the memory footprint and inference cost of large language models (LLMs), yet maintaining performance in the ultra-low-bit regime rem…