1 paper
Junhan Liao, Minxian Xu, Wanyi Zheng +4
To meet strict Service-Level Objectives (SLOs),contemporary Large Language Models (LLMs) decouple the prefill and decoding stages and place them on separate GPUs to mitigate the di…