2 papers
cs.CV2026
ROVER: Routing Object-Centric Visual Evidence for Grounded Multi-Image Reasoning
Guannan Lv, Ren Nie, Hongjian Dou +1
Multimodal Large Language Models (MLLMs) have increasingly localized and interleaved visual evidence for deliberative reasoning. Grounding-based approaches typically focus on regio…
cs.AI2026
DIVA-GRPO: Enhancing Multimodal Reasoning through Difficulty-Adaptive Variant Advantage
Haowen Gao, Zhenyu Zhang, Liang Pang +7
Reinforcement learning (RL) with group relative policy optimization (GRPO) has become a widely adopted approach for enhancing the reasoning capabilities of multimodal large languag…