1 paper
Hyungmin Kim, Minsoo Kim, Hongseok Kim +1
Multi-turn LLM serving accumulates dialogue history whose Key-Value (KV) cache grows with every turn and every user, quickly exceeding the model weights themselves and making memor…