2 citations · 2 across the 2 of their papers we have counts for
2 papers
cs.DC2026★ 2 cited
FlexPipe: Adapting Dynamic LLM Serving Through Inflight Pipeline Refactoring in Fragmented Serverless Clusters
Yanying Lin, Shijie Peng, Chengzhi Lu +2
Serving Large Language Models (LLMs) in production faces significant challenges from highly variable request patterns and severe resource fragmentation in serverless clusters. Curr…
cs.DC2026
FASER: Fine-Grained Phase Management for Speculative Decoding in Dynamic LLM Serving
Wenyan Chen, Chengzhi Lu, Yanying Lin +1
Speculative decoding (SD) is a widely used approach for accelerating decode-heavy LLM inference workloads. While online inference workloads are highly dynamic, existing SD systems…