2 papers
cs.DC2025
Collaborative Speculative Inference for Efficient LLM Inference Serving
Luyao Gao, Jianchun Liu, Hongli Xu +3
Speculative inference is a promising paradigm employing small speculative models (SSMs) as drafters to generate draft tokens, which are subsequently verified in parallel by the tar…
cs.DC2024
Accelerating End-Cloud Collaborative Inference via Near Bubble-free Pipeline Optimization
Luyao Gao, Jianchun Liu, Hongli Xu +3
End-cloud collaboration offers a promising strategy to enhance the Quality of Service (QoS) in DNN inference by offloading portions of the inference workload from end devices to cl…