1 paper
Chuming Shen, Wei Wei, Xiaoye Qu +1
DeepSeek-R1 has demonstrated powerful reasoning capabilities in the text domain through stable reinforcement learning (RL). Recently, in the multimodal domain, works have begun to…