1 paper · 1 filter
Yunhe Han, Yunqi Gao, Bing Hu +4
Speculative decoding can significantly accelerate LLM inference, especially given that its cloud-edge collaborative deployment offers cloud workload offloading, offline robustness,…