1 paper
Kumari Nishu, Han-Byul Kim, Santosh Chilkunda +13
Mixture-of-Experts (MoE) models are increasingly deployed alongside Speculative Decoding (SD) to accelerate inference, but combining the two is challenging. SD improves the inferen…