1 paper
Guanyu Cai, Ruiming Tian, Lang Yang +4
Deploying Large Language Models (LLMs) on mobile devices enhances privacy and reduces latency, but is severely bottlenecked by hardware inefficiency. We present the first comprehen…