1 citations · 1 across the 4 of their papers we have counts for
1 paper · 1 filter
Hanchen Li, Yuhan Liu, Yihua Cheng +2
Across large language model (LLM) applications, we observe an emerging trend for reusing KV caches to save the prefill delays of processing repeated input texts in different LLM in…