2 papers
cs.DC2026
PEEK: Predictive Queue-Informed KV Cache Management for LLM Serving
Bing Xie, Zhipeng Wang, Masahiro Tanaka +1
We present PEEK, a lightweight scheduling and eviction framework for both online (streaming) and offline (batch) LLM serving; this paper focuses on the online regime. PEEK maintain…
cs.LG2025
Scaling Up Data Parallelism in Decentralized Deep Learning
Bing Xie, Junqi Yin, Zhenyu Zhou +2
Although it has been extensively explored in theory, decentralized learning is not yet green-lighted for production use, largely due to a lack of stability, scalability, and genera…