2 citations · 3 across the 3 of their papers we have counts for
1 paper · 1 filter
Shu-Hao Zhang, Le-Tong Huang, Xiang-Sheng Deng +5
Quantization has emerged as a mainstream approach for deploying Large Language Models (LLMs) on resource-constrained devices, yet compressing precision below 4-bit typically causes…