1 paper · 1 filter
Yu Li, Binxu Li, Tian Lan
Autoregressive decoding in Transformer-based language models relies on the KV cache, whose memory footprint grows linearly with sequence length and becomes the primary bottleneck f…