3 papers
cs.DC2026
RAC: Reference-Aware Activation Compression for Communication-Efficient Split LLM Inference
Guotao Yang, Mingxi Zhao, Haopeng Li +4
Large language model (LLM) agents repeatedly process long, privacy-sensitive contexts, while cloud-only deployment exposes user data beyond the trusted endpoint and fully local dep…
cs.DC2026
AsymSpec: Efficient Cloud-Edge Speculative Decoding over Asymmetric Networks
Guotao Yang, Hao Chen, Rui Guo +5
Cloud-edge speculative decoding places a lightweight draft model at an edge gateway and a higher-quality target model in the cloud, but inserts communication into every speculative…
cs.DC2024
Harpagon: Minimizing DNN Serving Cost via Efficient Dispatching, Scheduling and Splitting
Zhixin Zhao, Yitao Hu, Ziqi Gong +5
Advances in deep neural networks (DNNs) have significantly contributed to the development of real-time video processing applications. Efficient scheduling of DNN workloads in cloud…