1 paper · 1 filter
Chun-Ting Chen, Dongmin Han, Hangyeol Mun +6
Block Quantization (BQ) enables efficient LLM inference by quantizing both weights and activations, but its design space remains underexplored. Through hardware-accuracy design spa…