1 paper · 1 filter
Cheng-Yu Yang, Shao-Yuan Lo, Yu-Lun Liu
Vision-language models (VLMs) project images into hundreds to thousands of visual tokens, making decoder inference expensive in both attention computation and KV-cache memory. Exis…