1 paper · 1 filter
Yue Zhu, Hao Yu, Chen Wang +2
The increasing adoption of large language models (LLMs) with extended context windows necessitates efficient Key-Value Cache (KVC) management to optimize inference performance. Inf…