2 papers
cs.CL2026
OUTLETS: Output-Length Prediction from Speculative Decoding Backbones
Weihuang Wen, Yingying Liu, Yichuan Liu +5
The heavy-tailed distribution of output lengths in Large Language Model (LLM) serving poses major challenges for resource provisioning and cluster scheduling. Although output-lengt…
cs.DC2025
Synera: Synergistic LLM Serving across Device and Cloud at Scale
Genglin Wang, Liekang Zeng, Bufang Yang +6
Large Language Models (LLMs) are becoming key components in various mobile operating systems, driving smart applications like interactive chatbots and personal assistants. While br…