2 papers
cs.CV2025
Self-Improving VLM Judges Without Human Annotations
Inna Wanyin Lin, Yushi Hu, Shuyue Stella Li +5
Effective judges of Vision-Language Models (VLMs) are crucial for model development. Current methods for training VLM judges mainly rely on large-scale human preference annotations…
cs.LG2024
Exploring How Generative MLLMs Perceive More Than CLIP with the Same Vision Encoder
Siting Li, Pang Wei Koh, Simon Shaolei Du
Recent research has shown that CLIP models struggle with visual reasoning tasks that require grounding compositionality, understanding spatial relationships, or capturing fine-grai…