1 paper · 1 filter
Weijia Han, Lisha Qu
Reported serving speedups from quantized kernels typically bundle the weight format, the kernel, and the inference runtime into one number. We present an attribution study on four…