1 paper
Mulei Ma, Xinyi Xu, Minrui Xu +3
LLMs are increasingly executed in edge where limited GPU memory and heterogeneous computation jointly constrain deployment which motivates model partitioning and request scheduling…