25 citations · 35 across the 14 of their papers we have counts for
Showing 2026Show all
2 papers · 1 filter
cs.CL2026
AgentSpec: Speculative Decoding for Batch Inference of LLM Agents
Xin Wang, Ziming Miao, Yi Zhu +4
Large language model (LLM)-based agent applications often incur high response time. Speculative decoding is a promising solution to improve the inference efficiency of LLM agents w…
cs.CV2026
MMSpec: Benchmarking Speculative Decoding for Vision-Language Models
Hui Shen, Xin Wang, Ping Zhang +11
Vision-language models (VLMs) achieve strong performance on multimodal tasks but suffer from high inference latency due to large model sizes and long multimodal contexts. Speculati…