4 papers
MV-STRIDE: Enabling MLLMs to Master Multi-View Spatial Reasoning via Hierarchical Capability Modeling
Jin Xu, Xiaojian Huang, Zhuodong Luo +6
Despite the rapid progress of Multimodal Large Language Models (MLLMs) in 2D vision-language tasks, robust multi-view spatial reasoning remains a fundamental bottleneck due to the…
DT2IT-MRM: Debiased Preference Construction and Iterative Training for Multimodal Reward Modeling
Zhihong Zhang, Jie Zhao, Xiaojian Huang +5
Multimodal reward models (MRMs) play a crucial role in aligning Multimodal Large Language Models (MLLMs) with human preferences. Training a good MRM requires high-quality multimoda…
MER 2026: From Discriminative Emotion Recognition to Generative Emotion Understanding
Zheng Lian, Xiaojiang Peng, Kele Xu +16
MER2026 marks the fourth edition of the MER series of challenges. The MER series provides valuable data resources to the research community and offers tasks centered on recent rese…
NTIRE 2025 Challenge on Image Super-Resolution (x4): Methods and Results
Zheng Chen, Kai Liu, Jue Gong +108
This paper presents the NTIRE 2025 image super-resolution (4) challenge, one of the associated competitions of the 10th NTIRE Workshop at CVPR 2025. The challenge aims to r…