1 paper
Chao Zhang, Yifan Ji, Ziyan Zhang +2
Dynamic sparse attention can reduce the quadratic cost of long-context prefilling without changing model weights. MInference assigns each attention head one pattern offline and est…