2 papers
cs.DC2026
TokenDance: Scaling Multi-Agent LLM Serving via Collective KV Cache Sharing
Zhuohang Bian, Feiyang Wu, Chengrui Zhang +3
Multi-agent LLM applications organize execution in synchronized rounds where a central scheduler gathers outputs from all agents and redistributes the combined context. This All-Ga…
cs.OS2025
Taming and Controlling Performance and Energy Trade-offs Automatically in Network Applications
Han Dong, Yara Awad, Sanjay Arora +2
In this paper, we demonstrate that a server running a single latency-sensitive application can be treated as a black box to reduce energy consumption while meeting an SLA target. W…