3 papers
cs.CL2025
When, What, and How: Rethinking Retrieval-Enhanced Speculative Decoding
Min Fang, Zhihui Fu, Qibin Zhao +1
Speculative decoding (SD) has emerged as an effective technique to accelerate large language model (LLM) inference without compromising output quality. However, the achievable spee…
cs.CL2025
DiffuSpec: Unlocking Diffusion Language Models for Speculative Decoding
Guanghao Li, Zhihui Fu, Min Fang +4
As large language models (LLMs) scale up, accuracy improves, but the autoregressive (AR) nature of decoding increases latency since each token requires a serial forward pass. Specu…
cs.DC2025
SimDC: A High-Fidelity Device Simulation Platform for Device-Cloud Collaborative Computing
Ruiguang Pei, Junjie Wu, Dan Peng +4
The advent of edge intelligence and escalating concerns for data privacy protection have sparked a surge of interest in device-cloud collaborative computing. Large-scale device dep…