2 papers
cs.DC2025
Opt4GPTQ: Co-Optimizing Memory and Computation for 4-bit GPTQ Quantized LLM Inference on Heterogeneous Platforms
Yaozheng Zhang, Wei Wang, Jie Kong +5
The increasing adoption of large language models (LLMs) on heterogeneous computing platforms poses significant challenges to achieving high inference efficiency. To address these e…
cs.DC2025
Boosting LLM Serving through Spatial-Temporal GPU Resource Sharing
Zejia Lin, Hongxin Xu, Guanyi Chen +3
Modern LLM serving systems confront inefficient GPU utilization due to the fundamental mismatch between compute-intensive prefill and memory-bound decode phases. While current prac…