2 papers
cs.CV2026
Token-Budget Distillation: Transferring Full-Token Semantics to Compressed Video Vision-Language Models
Xiaoyang Guo, Guoping Luo, Jusheng Zhang +2
Adapting video vision-language models (VLMs) is computationally expensive because video inputs produce a large number of visual tokens, making both fine-tuning and inference costly…
cs.CV2026
MissingBench-Verified: Probing Vision-Language Models' Inability to Detect Missing Object Parts
Wenqi Marshall Guo, Qingyun Qian, Shiyu Zhou +2
Vision Language Models (VLMs) are well known for hallucinating non-existent objects in images. Objects with missing parts present a unique challenge for VLMs, stemming from both re…