1 paper
Haohui Han, Yuming Wan, Hongni Wang +4
Low-bit attention accelerates Transformer inference by moving the QK⊤ and PV matrix multiplications to FP8 or FP4 matrix engines. However, the softmax path often evaluates…