4 papers
Latency-Aware Orchestration for Multi-Agent LLM Workflows on Heterogeneous GPUs
Jinghao Wang, Yifeng Zhang, Xiao Zhou +7
Concurrent multi-agent workflows expose future dependencies and serving-state requirements while running on heterogeneous GPU pools with time-varying load, model residency, and res…
ElastiCo: Elastic Configuration and Interference-Aware Orchestration for GPU Clusters
Jinghao Wang, Yihang Zhou, Xiaoyang Sun +5
Modern GPU clusters must simultaneously serve deep learning training and offline large language model inference workloads, yet existing schedulers treat these as isolated resource…
SpecBox: Speculative Sandbox Scheduling for Efficient LLM Agent Serving
Yihui Zhang, Tianyu Wo, Jinghao Wang +7
As LLM agents increasingly rely on the Model Context Protocol (MCP) to invoke isolated external sandboxes, disaggregated sandbox deployment introduces a fundamental tension between…
Maestro: Workload-Aware Cross-Cluster Scheduling for LLM-Based Multi-Agent Systems
Jinghao Wang, Xiao Zhou, Xiaoyang Sun +6
Large Language Model based Multi-Agent Systems (LLM-MAS) have emerged as a powerful paradigm for tackling complex tasks by breaking them into collaborative workflows of specialized…