1 paper · 1 filter
Wenbo Sun, Qiming Guo, Wenlu Wang +1
Deploying Large Language Models (LLMs) on resource-constrained devices remains challenging due to limited memory, lack of GPUs, and the complexity of existing runtimes. In this pap…