2 papers
cs.LG2026
Dispersion Loss Counteracts Embedding Condensation and Improves Generalization in Small Language Models
Chen Liu, Xingzhi Sun, Xi Xiao +8
Large language models (LLMs) achieve remarkable performance through ever-increasing parameter counts, but scaling incurs steep computational costs. To better understand LLM scaling…
cs.CL2026
Making Every Verified Token Count: Adaptive Verification for MoE Speculative Decoding
Lehan Pan, Ziyang Tao, Ruoyu Pang +3
Tree-based speculative decoding accelerates autoregressive generation by verifying multiple draft candidates in parallel, but this advantage weakens for sparse Mixture-of-Experts (…