1 citations · 1 across the 8 of their papers we have counts for
1 paper · 1 filter
Xiangwen Zhuge, Xu Shen, Zeyu Wang +6
Efficient LLM inference on resource-constrained devices presents significant challenges in compute and memory utilization. Due to limited GPU memory, existing systems offload model…