2 papers
cs.LG2026
WAR: Workload-Aware Rollouts for Synchronous Agentic Reinforcement Learning
Ryan Xu, Atlas Zhao, David Bao +1
Long-horizon rollout generation has become the dominant systems bottleneck in agentic reinforcement learning (RL). As agents interact with environments over many turns, trajectorie…
cs.DC2026
Talaria: Session-Aware Serverless Serving of Hundred-Billion-Parameter LLMs
Utopia Meng, Unicornt Zhao, Derek Li +2
Serverless multi-model LLM systems multiplex popularity-skewed model catalogs over shared GPU pools, yet typically schedule each request independently. Tool-using agents break this…