1 paper · 1 filter
Jung Hyun Lee, June Yong Yang, Jungwook Choi +1
As large language models continue to scale, low-bit weight-only post-training quantization (PTQ) offers a practical solution to their memory-efficient deployment. Although block-wi…