2 citations · 5 across the 14 of their papers we have counts for
Showing 2026Show all
3 papers · 1 filter
cs.DC2026
SMetric: Rethink LLM Scheduling for Serving Agents with Balanced Session-centric Scheduling
Jiahao Wang, Kaizhan Lin, Kaixi Zhang +7
LLM scheduling is critical to serving, yet how well existing designs fit agentic serving--where agents, not humans, issue the requests--remains unclear. Agents shift the workload i…
cs.DB2026
Efficient Vector Search in the Wild: One Model for Multi-K Queries
Yifan Peng, Jiafei Fan, Xingda Wei +7
Learned top-K search is a promising approach for serving vector queries with both high accuracy and performance. However, current models trained for a specific K value fail to gene…
cs.OS2026
Towards Fully-fledged GPU Multitasking via Proactive Memory Scheduling
Weihang Shen, Yinqiu Chen, Rong Chen +1
The limited HBM capacity has become the primary bottleneck for hosting an increasing number of larger-scale GPU tasks. While demand paging extends capacity via host DRAM, it incurs…