1 paper
Luchang Li, Shuaishuai Wang, Zhao Ruan +2
Prefix caching is critical for efficient large language model (LLM) serving, particularly for agentic workloads that repeatedly invoke the model with a growing conversation and too…