1 paper
Xu Yang, Jiapeng Zhang, Zhangke +6
Sparse long-context inference requires efficient token retrieval in both prefill and decode. Existing methods often use different retrieval strategies for the two stages, preventin…