2 papers
cs.DC2026
KVServe: Service-Aware KV Cache Compression for Communication-Efficient Disaggregated LLM Serving
Zedong Liu, Xinyang Ma, Dejun Luo +9
LLMs are widely adopted in production, pushing inference systems to their limits. Disaggregated LLM serving (e.g., PD separation and KV state disaggregation) improves scalability a…
cs.DC2026
TACO: Efficient Communication Compression of Intermediate Tensors for Scalable Tensor-Parallel LLM Training
Man Liu, Xingchen Liu, Xingjian Tian +8
Handling communication overhead in large-scale tensor-parallel training remains a critical challenge due to the dense, near-zero distributions of intermediate tensors, which exacer…