Showing cs.CVShow all
2 papers · 1 filter
cs.CV2025
VLRMBench: A Comprehensive and Challenging Benchmark for Vision-Language Reward Models
Jiacheng Ruan, Wenzhen Yuan, Xian Gao +6
Although large visual-language models (LVLMs) have demonstrated strong performance in multimodal tasks, errors may occasionally arise due to biases during the reasoning process. Re…
cs.CV2023
Improving Vision-and-Language Reasoning via Spatial Relations Modeling
Cheng Yang, Rui Xu, Ye Guo +5
Visual commonsense reasoning (VCR) is a challenging multi-modal task, which requires high-level cognition and commonsense reasoning ability about the real world. In recent years, l…