1 paper
Yi Chen, Yuying Ge, Rui Wang +4
Recent reinforcement learning approaches, such as outcome-supervised GRPO, have advanced Chain-of-Thought reasoning in large language models (LLMs), yet their adaptation to multimo…