1 paper · 1 filter
Shitian Zhao, Shaoheng Lin, Ming Li +4
Reinforcement learning for agentic multimodal models often suffers from interaction collapse, where models learn to reduce tool usage and multi-turn reasoning, limiting the benefit…