1 paper
Yuzhuang Xu, Xu Han, Zonghan Yang +5
Model quantification uses low bit-width values to represent the weight matrices of existing models to be quantized, which is a promising approach to reduce both storage and computa…