1 paper
Seongjin Cha, Gyuwan Kim, Dongsu Han +2
Self-speculative decoding (SSD) accelerates LLM inference by skipping layers to create an efficient draft model, yet existing methods often rely on static heuristics that ignore th…