2 papers
cs.DC2026
Latency-Aware Orchestration for Multi-Agent LLM Workflows on Heterogeneous GPUs
Jinghao Wang, Yifeng Zhang, Xiao Zhou +7
Concurrent multi-agent workflows expose future dependencies and serving-state requirements while running on heterogeneous GPU pools with time-varying load, model residency, and res…
cs.MA2026
Efficient LLM Serving for Agentic Workflows: A Data Systems Perspective
Noppanat Wadlom, Junyi Shen, Yao Lu
Agentic workflows are composed of sequences of interdependent Large Language Model (LLM) calls, and they have become a dominant workload in modern AI systems. These workflows exhib…