1 paper
Payman Behnam, Yaosheng Fu, Ritchie Zhao +3
Transformer-based Large Language Models rely critically on the KV cache to efficiently handle extended contexts during the decode phase. Yet, the size of the KV cache grows proport…