Showing cs.CLShow all
2 papers · 1 filter
cs.CL2026
TOPAS: Workflow-Aware Prefix-State Scheduling for Multi-Agent LLM Serving
Hongqiu Ni, Han Tian, Chi Zhang +2
Prefix caching introduces a fundamental tradeoff in multi-agent large language model (LLM) serving: retaining a long system-prompt key-value (KV) cache for an agent accelerates fut…
cs.CL2025
Astraea: A State-Aware Scheduling Engine for LLM-Powered Agents
Hongqiu Ni, Jiabao Zhang, Guopeng Li +4
Large Language Models (LLMs) are increasingly being deployed as intelligent agents. Their multi-stage workflows, which alternate between local computation and calls to external net…