1 citations · 1 across the 2 of their papers we have counts for
3 papers
cs.LG2025★ 1 cited
RouterArena: An Open Platform for Comprehensive Comparison of LLM Routers
Yifan Lu, Rixin Liu, Jiayi Yuan +4
Today's LLM ecosystem comprises a wide spectrum of models that differ in size, capability, and cost. No single model is optimal for all scenarios; hence, LLM routers have become es…
cs.DC2025
From Models to Operators: Rethinking Autoscaling Granularity for Large Generative Models
Xingqi Cui, Chieh-Jan Mike Liang, Jiarong Xing +1
Serving large generative models such as LLMs and multi- modal transformers requires balancing user-facing SLOs (e.g., time-to-first-token, time-between-tokens) with provider goals…
cs.OS2025
Towards Efficient and Practical GPU Multitasking in the Era of LLM
Jiarong Xing, Yifan Qiao, Simon Mo +5
GPU singletasking is becoming increasingly inefficient and unsustainable as hardware capabilities grow and workloads diversify. We are now at an inflection point where GPUs must em…