1 paper · 1 filter
Cheng Li, Jiexiong Liu, Yixuan Chen +1
On-device LLM inference faces a trilemma of response latency, limited hardware resources and user privacy. Full cloud inference delivers strong computing power but exposes user pro…