1 paper · 1 filter
Junjie Ma, Jinlong Li, Jiajun Luo
Inference with modern Large Language Models (LLMs) is expensive and slow, and speculative sampling has emerged as an effective solution to this problem. However, the number of call…