2 papers
cs.DC2026
FlowSpec: Continuous Pipelined Speculative Decoding for Efficient Distributed LLM Inference
Xing Liu, Lizhuo Luo, Ming Tang +2
Distributed inference serves as a promising approach to enabling the inference of large language models (LLMs) at the network edge. It distributes the inference process to multiple…
cs.DC2025
CoCoI: Distributed Coded Inference System for Straggler Mitigation
Xing Liu, Chao Huang, Ming Tang
Convolutional neural networks (CNNs) are widely applied in real-time applications on resource-constrained devices. To accelerate CNN inference, prior works proposed to distribute t…