1 paper
Mingyu Sun, Xiao Zhang, Shen Qu +5
Providing lossless inference services of LLMs on edge devices remains challenging, especially given the extremely tight memory budgets. The existing offloading techniques inevitabl…