2 papers
cs.DC2026
FASER: Fine-Grained Phase Management for Speculative Decoding in Dynamic LLM Serving
Wenyan Chen, Chengzhi Lu, Yanying Lin +1
Speculative decoding (SD) is a widely used approach for accelerating decode-heavy LLM inference workloads. While online inference workloads are highly dynamic, existing SD systems…
cs.DC2026
CodecSight: Leveraging Video Codec Signals for Efficient Streaming VLM Inference
Yulin Zou, Yan Chen, Wenyan Chen +5
Video streaming analytics is a crucial workload for vision-language model serving, but the high cost of multimodal inference limits scalability. Prior systems reduce inference cost…