1 paper
Yiming Yao, Chenyang Lyu, Xuanfan Ni +4
Long-form audio inputs make the KV cache the dominant memory cost of speech LLMs. Prefill-only KV compression methods permanently discard audio KV positions once evicted, with no p…