2 citations · 2 across the 7 of their papers we have counts for
1 paper · 1 filter
Xiangbo Qi, Chaoyi Jiang, Murali Annavaram
Large language models (LLMs) have grown beyond the memory capacity of single GPU devices, necessitating quantization techniques for practical deployment. While NF4 (4-bit NormalFlo…