1 paper · 1 filter
Xinhao Zhong, Yuxia Qiao, Junhao Li +3
Reinforcement-learning (RL) post-training equips multimodal large reasoning models (MLRMs) with exploratory chains of thought (CoT), substantially improving visual reasoning. Howev…