1 citations · 1 across the 3 of their papers we have counts for
1 paper · 1 filter
Yongtong Wu, Shaoyuan Chen, Yinmin Zhong +10
The performance of multi-turn, agentic LLM inference is increasingly dominated by KV-Cache storage I/O rather than computation. In prevalent disaggregated architectures, loading th…