3 papers
cs.CL2026
TOPAS: Workflow-Aware Prefix-State Scheduling for Multi-Agent LLM Serving
Hongqiu Ni, Han Tian, Chi Zhang +2
Prefix caching introduces a fundamental tradeoff in multi-agent large language model (LLM) serving: retaining a long system-prompt key-value (KV) cache for an agent accelerates fut…
cs.DC2025
Embedding Samples Dispatching for Recommendation Model Training in Edge Environments
Guopeng Li, Haisheng Tan, Chi Zhang +5
Training deep learning recommendation models (DLRMs) on edge workers brings several benefits, particularly in terms of data privacy protection, low latency and personalization. How…
cs.CL2025
Astraea: A State-Aware Scheduling Engine for LLM-Powered Agents
Hongqiu Ni, Jiabao Zhang, Guopeng Li +4
Large Language Models (LLMs) are increasingly being deployed as intelligent agents. Their multi-stage workflows, which alternate between local computation and calls to external net…