2 papers
cs.DC2026
Benchmarking LLM Serving Systems for Agentic AI Workloads with XPerf
Michael Wang, Yikang Yue, Shaobo Li +3
We present XPerf, a benchmarking framework that load-tests LLM serving systems with diverse agentic AI workloads. It provides detailed profiling of the serving system and hardware,…
cs.LG2026
Vegas: Self-Speculative Decoding with Verification-Guided Sparse Attention
Yikang Yue, Yuqi Xue, Jian Huang
Long-context large language model (LLM) inference has become the norm for today's AI applications. However, it is severely bottlenecked by the increasing memory demands of its KV c…