1 paper · 1 filter
Ting-Yun Chang, Muru Zhang, Jesse Thomason +1
Low-bit weight-only quantization significantly reduces the memory footprint of large language models (LLMs), but disproportionately affects certain examples. We analyze diverse 3-4…