1 paper · 1 filter
Michael Jerge, Joseph Pelczar, Justin Downes
Reinforcement learning with verifiable rewards (RLVR) improves the reasoning ability of vision-language models (VLMs), and diversifying the rollouts within each optimization group…