1 paper
Ziyi Cao, Qingyi Si, Jingbin Zhang +1
Large language models face significant cost challenges in long-sequence inference. To address this, reusing historical Key-Value (KV) Cache for improved inference efficiency has be…