1 paper · 1 filter
Wenhui Zhu, Xuanzhao Dong, Xin Li +6
Recently, reinforcement learning (RL)-based tuning has shifted the trajectory of Multimodal Large Language Models (MLLMs), particularly following the introduction of Group Relative…