1 paper
Jinsong Shu, Chenyang Wu, Zhongle Xie +2
Key-Value (KV) caching is essential for efficient inference in multimodal large language models (MLLMs), yet its memory footprint grows linearly with context length and becomes a m…