1 citations · 1 across the 1 of their papers we have counts for
1 paper · 1 filter
Jing Zou, Shangyu Wu, Hancong Duan +2
Efficiently serving Large Language Models (LLMs) with persistent Prefix Key-Value (KV) Cache is critical for applications like conversational search and multi-turn dialogue. Servin…