1 paper · 1 filter
Jiawen Qi, Chang Gao, Zhaochun Ren +1
Deploying Large Language Models (LLMs) on edge devices remains challenging due to their quadratically increasing computations with the sequence length. Existing studies for dynamic…