2 papers
cs.DC2026
PipeSD: An Efficient Cloud-Edge Collaborative Pipeline Inference Framework with Speculative Decoding
Yunhe Han, Yunqi Gao, Bing Hu +4
Speculative decoding can significantly accelerate LLM inference, especially given that its cloud-edge collaborative deployment offers cloud workload offloading, offline robustness,…
cs.DC2025
FlowMoE: A Scalable Pipeline Scheduling Framework for Distributed Mixture-of-Experts Training
Yunqi Gao, Bing Hu, Mahdi Boloursaz Mashhadi +5
The parameter size of modern large language models (LLMs) can be scaled up via the sparsely-activated Mixture-of-Experts (MoE) technique to avoid excessive increase of the computat…