1 citations · 1 across the 7 of their papers we have counts for
1 paper · 1 filter
Mingyuan Fan, Yu Liu, Fuyi Wang +1
The deployment of large language models (LLMs) on resource-constrained devices remains challenging, spurring interest in split inference, where models are partitioned between clien…