1 paper
Yilin Feng, Ahmed Burak Gulhan, Mahmut Taylan Kandemir
Vision-Language Models (VLMs) process thousands of visual tokens per image alongside comparatively few text tokens, yet existing compression methods treat both modalities uniformly…