25 citations · 35 across the 14 of their papers we have counts for
1 paper · 2 filters
Xin Wang, Ziming Miao, Yi Zhu +4
Large language model (LLM)-based agent applications often incur high response time. Speculative decoding is a promising solution to improve the inference efficiency of LLM agents w…