1 citations · 1 across the 2 of their papers we have counts for
1 paper · 1 filter
Mulei Ma, Xinyi Xu, Minrui Xu +3
LLMs are increasingly executed in edge where limited GPU memory and heterogeneous computation jointly constrain deployment which motivates model partitioning and request scheduling…