1 paper · 1 filter
Dongxin Guo, Jikun Wu, Siu Ming Yiu
AI agents execute tens to hundreds of chained LLM calls per task, yet GPU schedulers treat each call as independent, discarding gigabytes of intermediate state between steps and in…