1 citations · 1 across the 1 of their papers we have counts for
1 paper
Ying Zhang, Peng Zhang, Mincong Huang +7
Quantization is a proven effective method for compressing large language models. Although popular techniques like W8A8 and W4A16 effectively maintain model performance, they often…