3 papers
cs.CL2026
DIVERSED: Relaxed Speculative Decoding via Dynamic Ensemble Verification
Ziyi Wang, Siva Rajesh Kasa, Ankith M S +6
Speculative decoding is an effective technique for accelerating large language model inference by drafting multiple tokens in parallel. In practice, its speedup is often bottleneck…
cs.CL2026
Ultra-Fast Language Generation via Discrete Diffusion Divergence Instruct
Haoyang Zheng, Xinyang Liu, Cindy Xiangrui Kong +5
Fast and high-quality language generation is the holy grail that people pursue in the age of AI. In this work, we introduce Discrete Diffusion Divergence Instruct (DiDi-Instruct),…
cs.LG2025
SQS: Bayesian DNN Compression through Sparse Quantized Sub-distributions
Ziyi Wang, Nan Jiang, Guang Lin +1
Compressing large-scale neural networks is essential for deploying models on resource-constrained devices. Most existing methods adopt weight pruning or low-bit quantization indivi…