1 paper
Bo Lv, Quan Zhou, Xuanang Ding +2
The bottleneck associated with the key-value(KV) cache presents a significant challenge during the inference processes of large language models. While depth pruning accelerates inf…