1 citations · 1 across the 12 of their papers we have counts for
4 papers · 1 filter
Diagnosing Visual Ignorance in Vision-Language Models
Runyu Zhou, Qi Zhang, Qixun Wang +1
Vision-Language Models (VLMs) frequently rely on language priors, producing confident answers that are weakly grounded in visual evidence. While this behavior is widely observed, i…
UniG2U-Bench: Do Unified Models Advance Multimodal Understanding?
Zimo Wen, Boxiu Li, Wanbo Zhang +11
Unified multimodal models have recently demonstrated strong generative capabilities, yet whether and when generation improves understanding remains unclear. Existing benchmarks lac…
SSL4RL: Revisiting Self-supervised Learning as Intrinsic Reward for Visual-Language Reasoning
Xiaojun Guo, Runyu Zhou, Yifei Wang +8
Vision-language models (VLMs) have shown remarkable abilities by integrating large language models with visual inputs. However, they often fail to utilize visual evidence adequatel…
Identifiable Contrastive Learning with Automatic Feature Importance Discovery
Qi Zhang, Yifei Wang, Yisen Wang
Existing contrastive learning methods rely on pairwise sample contrast to learn data representations, but the learned features often lack clear interpretability f…