5 papers
Compass: SLO-aware Query Planner for Compound AI Serving at Scale
Banruo Liu, Wei-Yu Lin, Minghao Fang +2
The rise of compound AI serving that integrates multiple operators in a pipeline enables end-user applications such as generative AI-powered meeting companions, autonomous driving,…
CacheFlow: Efficient LLM Serving with 3D-Parallel KV Cache Restoration
Sean Nian, Jiahao Fang, Qilong Feng +2
KV cache restoration has emerged as a dominant bottleneck in serving long-context LLM workloads, including multi-turn conversations, retrieval-augmented generation, and agentic pip…
JITServe: SLO-aware LLM Serving with Imprecise Request Information
Wei Zhang, Zhiyu Wu, Yi Mu +5
The integration of Large Language Models (LLMs) into applications ranging from interactive chatbots to multi-agent systems has introduced a wide spectrum of service-level objective…
Dora: QoE-Aware Hybrid Parallelism for Distributed Edge AI
Jianli Jin, Ziyang Lin, Qianli Dong +5
With the proliferation of edge AI applications, satisfying user quality of experience (QoE) requirements, such as model inference latency, has become a first class objective, as th…
Single-agent or Multi-agent Systems? Why Not Both?
Mingyan Gao, Yanzi Li, Banruo Liu +4
Multi-agent systems (MAS) decompose complex tasks and delegate subtasks to different large language model (LLM) agents and tools. Prior studies have reported the superior accuracy…