1 paper
Jing Zou, Shangyu Wu, Hancong Duan +2
Efficiently serving Large Language Models (LLMs) with persistent Prefix Key-Value (KV) Cache is critical for applications like conversational search and multi-turn dialogue. Servin…