1 paper · 1 filter
Yongtong Wu, Shaoyuan Chen, Yinmin Zhong +10
The performance of multi-turn, agentic LLM inference is increasingly dominated by KV-Cache storage I/O rather than computation. In prevalent disaggregated architectures, loading th…