1 paper
Xiangwen Zhuge, Xu Shen, Zeyu Wang +6
Efficient LLM inference on resource-constrained devices presents significant challenges in compute and memory utilization. Due to limited GPU memory, existing systems offload model…