most citedA Survey of Reinforcement Learning for Large Reasoning Models

2 citations · 3 across the 7 of their papers we have counts for

collaborators

9 papers

cs.CV2025

DiffThinker: Towards Generative Multimodal Reasoning with Diffusion Models

Zefeng He, Xiaoye Qu, Yafu Li +3

While recent Multimodal Large Language Models (MLLMs) have attained significant strides in multimodal reasoning, their reasoning processes remain predominantly text-centric, leadin…

cs.CV2025

VideoSSR: Video Self-Supervised Reinforcement Learning

Zefeng He, Xiaoye Qu, Yafu Li +3

Reinforcement Learning with Verifiable Rewards (RLVR) has substantially advanced the video understanding capabilities of Multimodal Large Language Models (MLLMs). However, the rapi…

cs.CL20252 cited

A Survey of Reinforcement Learning for Large Reasoning Models

Kaiyan Zhang, Yuxin Zuo, Bingxiang He +36

In this paper, we survey recent advances in Reinforcement Learning (RL) for reasoning with Large Language Models (LLMs). RL has achieved remarkable success in advancing the frontie…

cs.CV2025

FrameThinker: Learning to Think with Long Videos via Multi-Turn Frame Spotlighting

Zefeng He, Xiaoye Qu, Yafu Li +3

While Large Vision-Language Models (LVLMs) have achieved substantial progress in video understanding, their application to long video reasoning is hindered by uniform frame samplin…

cs.LG20251 cited

Rethinking Entropy Regularization in Large Reasoning Models

Yuxian Jiang, Yafu Li, Guanxu Chen +3

Reinforcement learning with verifiable rewards (RLVR) has shown great promise in enhancing the reasoning abilities of large reasoning models (LRMs). However, it suffers from a crit…

cs.AI2025

Conditional Advantage Estimation for Reinforcement Learning in Large Reasoning Models

Guanxu Chen, Yafu Li, Yuxian Jiang +6

Reinforcement Learning with Verifiable Rewards (RLVR) for large language models (LLMs) has achieved remarkable progress in enhancing LLMs' reasoning capabilities on tasks with clea…