1 paper
Peirong Zheng, Wenchao Xu, Haozhao Wang +2
The deployment of large language models' (LLMs) inference at the edge can facilitate prompt service responsiveness while protecting user privacy. However, it is critically challeng…