1 paper · 1 filter
Yizhou Liu, Dingkang Yang, Zizhi Chen +5
Reinforcement learning (RL) with rule-based reward functions has recently shown great promise in enhancing the reasoning depth and generalization ability of vision-language models…