1 paper · 1 filter
Omar Sharif, Eftekhar Hossain, Nikhil Singh +1
Reinforcement learning with verifiable rewards has driven major gains in LLM reasoning, and it is intuitive to assume this recipe will transfer well to multimodal models. However,…