3 papers
cs.CV2025
DiffThinker: Towards Generative Multimodal Reasoning with Diffusion Models
Zefeng He, Xiaoye Qu, Yafu Li +3
While recent Multimodal Large Language Models (MLLMs) have attained significant strides in multimodal reasoning, their reasoning processes remain predominantly text-centric, leadin…
cs.CV2025
VideoSSR: Video Self-Supervised Reinforcement Learning
Zefeng He, Xiaoye Qu, Yafu Li +3
Reinforcement Learning with Verifiable Rewards (RLVR) has substantially advanced the video understanding capabilities of Multimodal Large Language Models (MLLMs). However, the rapi…
cs.CV2025
FrameThinker: Learning to Think with Long Videos via Multi-Turn Frame Spotlighting
Zefeng He, Xiaoye Qu, Yafu Li +3
While Large Vision-Language Models (LVLMs) have achieved substantial progress in video understanding, their application to long video reasoning is hindered by uniform frame samplin…