2 papers
cs.CV2026
EvoCut: Multi-Layer Evolution-Aware Visual Token Compression for Efficient Large Vision-Language Models
Hongyu Lu, Feng Zhang, Wenwei Jin +5
Large vision-language models (LVLMs) achieve strong performance on image and video understanding tasks, but their inference efficiency is constrained by the large number of visual…
cs.CV2024
P4Q: Learning to Prompt for Quantization in Visual-language Models
Huixin Sun, Runqi Wang, Yanjing Li +4
Large-scale pre-trained Vision-Language Models (VLMs) have gained prominence in various visual and multimodal tasks, yet the deployment of VLMs on downstream application platforms…