1 paper
Haochen Huang, Shengxuan Qiu, Meng Li
Deploying large language models (LLMs) for inference on edge devices is challenging due to severe memory and bandwidth constraints. While speculative decoding and Mixture-of-Expert…