1 paper
Lin Qiu, Hanqing Zeng, Yao Liu +3
Reinforcement learning (RL) has emerged as an effective paradigm for improving the reasoning capability of vision-language models (VLMs). However, RL-based optimization typically d…