1 paper · 1 filter
Wenshuai Yao, Wenyong Zhou, Hanyong Shao +5
The paper proposes ReTopK, a training‑free technique that speeds up dynamic top‑K sparse attention for long‑context language models by reusing supports from historically similar qu…