2 papers
cs.DC2026
Lynx: Progressive Speculative Quantization for accelerating KV Transfer in Long-Context Inference
Wenchen Han, Gingfung Matthew Yeung, Marco Barletta +3
Long-context inference is increasingly common in large language model (LLM) serving, driven by retrieval-augmented generation and agentic systems. In disaggregated inference, these…
cs.LG2026
Learning to Score: Tuning Cluster Schedulers through Reinforcement Learning
Martin Asenov, Qiwen Deng, Gingfung Yeung +1
Efficiently allocating incoming jobs to nodes in large-scale clusters can lead to substantial improvements in both cluster utilization and job performance. In order to allocate inc…