1 paper
Jahyun Koo, Dahoon Park, Sangwoo Jung +1
To overcome the burden on the memory size and bandwidth due to ever-increasing size of large language models (LLMs), aggressive weight quantization has been recently studied, while…