3 papers
cs.DC2026
DisagFusion: Asynchronous Pipeline Parallelism and Elastic Scheduling for Disaggregated Diffusion Serving
Hantian Zha, Teng Ma, Yang Yong +7
Diffusion-based generation is increasingly powering production content pipelines; however, deploying these models at scale remains a significant challenge. Model weights frequently…
cs.AR2026
Adaptive Multi-Objective Tiered Storage Configuration for KV Cache in LLM Service
Xianzhe Zheng, Zhengheng Wang, Ruiyan Ma +17
The memory-for-computation paradigm of KV caching is essential for accelerating large language model (LLM) inference service, but limited GPU high-bandwidth memory (HBM) capacity m…
cs.DC2026
LatencyPrism: Online Non-intrusive Latency Sculpting for SLO-Guaranteed LLM Inference
Yin Du, Jiayi Ren, Xiayu Sun +4
LLM inference latency critically determines user experience and operational costs, directly impacting throughput under SLO constraints. Even brief latency spikes degrade service qu…