1 paper
Yunhao Wang, Ziting Li, Shuai Chen +6
Aligning large-scale vision-language models (VLMs) for complex reasoning via reinforcement learning is often hampered by the limitations of existing policy optimization algorithms,…