1 paper
Sunjung Lee, Sanghoon Cha, Hyeonsu Kim +8
On-device deployments of large language models (LLMs) are rapidly proliferating across mobile and edge platforms. LLM inference comprises a compute-intensive prefill phase and a me…