1 paper · 1 filter
Qingyao Yang, Runming Yang, He Xiao +7
While binary quantization theoretically promises extreme compression and acceleration for Large Language Models (LLMs), existing research often overlooks the necessity of specializ…