3 papers
cs.IT2026
DiP-SD: Distributed Pipelined Speculative Decoding for Efficient LLM Inference at the Edge
Yaodan Xu, Sheng Zhou, Zhisheng Niu
Speculative decoding has emerged as a promising technique for large language model (LLM) inference by accelerating autoregressive decoding via draft-then-verify. This paper studies…
cs.DC2025
Joint Optimization of Offloading, Batching and DVFS for Multiuser Co-Inference
Yaodan Xu, Sheng Zhou, Zhisheng Niu
With the growing integration of artificial intelligence in mobile applications, a substantial number of deep neural network (DNN) inference requests are generated daily by mobile d…
cs.DC2025
SMDP-Based Dynamic Batching for Improving Responsiveness and Energy Efficiency of Batch Services
Yaodan Xu, Sheng Zhou, Zhisheng Niu
For servers incorporating parallel computing resources, batching is a pivotal technique for providing efficient and economical services at scale. Parallel computing resources exhib…