1 paper · 1 filter
Ziqi Miao, Haonan Jia, Lijun Li +4
Reinforcement learning with verifiable rewards (RLVR) has substantially enhanced the reasoning capabilities of multimodal large language models (MLLMs). However, existing RLVR appr…