2 citations · 2 across the 3 of their papers we have counts for
1 paper · 1 filter
Hong Guo, Nianhui Guo, Weixing Wang +3
W4A4 quantization promises full utilization of INT4 Tensor Cores, yet group dequantization overhead on CUDA Cores has driven existing systems to mixed-precision fallbacks. We prese…