1 paper
Siqing Song, Chuang Wang, Ruiqi Wang +2
Quantizing large language models (LLMs) to 1-bit precision significantly reduces computational costs, but existing quantization techniques suffer from noticeable performance degrad…