1 paper
Zhibin Wang, Ziyu Zhong, Nuo Shen +3
Speculative decoding and dynamic sparse attention are two complementary approaches for accelerating long-context LLM inference: the former amortizes target-model execution across m…