1 paper · 1 filter
Yufei Li, Yu Fu, Yue Dong +1
Large language models (LLMs) deployed on edge servers are increasingly used in latency-sensitive applications such as personalized assistants, recommendation, and content moderatio…