1 paper
Chongpeng Liu, Xiaojian Liao, Hancheng Liu +2
This paper presents PipeBoost, a low-latency LLM serving system for multi-GPU (serverless) clusters, which can rapidly launch inference services in response to bursty requests with…