1 paper
Xiangchen Li, Saeid Ghafouri, Jiakun Fan +3
Speculative decoding enables collaborative Large Language Model (LLM) inference across cloud and edge by separating lightweight token drafting from heavyweight verification. While…