1 paper · 1 filter
Xu Yang, Jiapeng Zhang, Zhangke +6
Sparse long-context inference requires efficient token retrieval in both prefill and decode. Existing methods often use different retrieval strategies for the two stages, preventin…