1 paper · 1 filter
Xiangwen Zhuge, Xu Shen, Zeyu Wang +6
Efficient LLM inference on resource-constrained devices presents significant challenges in compute and memory utilization. Due to limited GPU memory, existing systems offload model…