1 paper
Haodong Wang, Junjie Liu, Zicong Hong +4
4-bit quantization reduces the memory footprint and latency of large language model inference, but its aggressive precision reduction can severely degrade accuracy. Prior methods a…