1 paper
Ruiyi Tao, Xiaolong Tu, Haoxin Wang
Large Language Models (LLMs) are increasingly deployed on mobile and embedded devices to improve privacy and reduce network latency. Yet on-device inference faces a fundamental con…