1 paper
Bingjie Zhu, Zhixiong Chen, Liqiang Zhao +2
Large language model (LLM) inference at the network edge is a promising serving paradigm that leverages distributed edge resources to run inference near users and enhance privacy.…