103 citations · 154 across the 36 of their papers we have counts for
1 paper · 2 filters
Reuben Tan, Baolin Peng, Zhengyuan Yang +15
Agentic reasoning models trained with multimodal reinforcement learning (MMRL) have become increasingly capable, yet they are almost universally optimized using sparse, outcome-bas…