1 paper
Junho Yoon, Geom Lee, Donghyeon Jeon +2
Quantization has been widely studied as an effective technique for reducing the memory requirement of large language models (LLMs), potentially improving the latency time as well.…