1 paper
Sangwoo Kwon, Seong Hoon Seo, Jae W. Lee +1
How can we effectively handle queries for on-device large language models (LLMs) with varying runtime constraints, such as latency and accuracy? Multi-scale quantization addresses…