1 paper
Shaoke Fang, Ziang Li, Wenfei Wu +3
Prefix caching is a key optimization in Large Language Model (LLM) serving, reusing attention Key-Value (KV) states across requests with shared prompt prefixes to reduce expensive…