2 papers
cs.CV2026
KVCapsule: Efficient Sequential KV Cache Compression for Vision-Language Models with Asymmetric Redundancy
Yingbing Huang, Tharun Adithya Srikrishnan, Steven K. Reinhardt +1
Vision-Language Models (VLMs) have emerged as a critical and fast-growing extension of Large Language Models (LLMs) that enable multimodal reasoning through both text and image inp…
cs.CV2026
BlindSight: Harnessing Sparsity for Efficient Vision-Language Models
Tharun Adithya Srikrishnan, Deval Shah, Timothy Hein +3
Large vision-language models (VLMs) enable joint processing of text and images. However, incorporating vision data significantly increases the prompt length, resulting in a longer…