1 paper
Renyuan Liu, Yuyang Leng, Kaiyan Liu +7
On-device LLM inference is attractive for privacy and responsiveness, but remains challenging on mobile and embedded devices because model weights far exceed available DRAM. Prior…