8 papers
Co-RL: Unsupervised Reasoning Emerges from Diverse Cohort in Multi-agent RL
Yunhao Yang, Yuexin Bian, Yunjie Tian +6
Reinforcement learning (RL) has emerged as a powerful approach for improving reasoning in language and vision-language models, yet its strongest successes still depend heavily on g…
Self-Supervised Visual On-Policy Distillation
Yijiang Li, Yijun Liang, Yunjie Tian +6
Visual on-policy distillation relies heavily on an informative teacher-student asymmetry, through either a larger, stronger teacher or privileged supervision, such as reference ans…
On-Policy Self-Distillation without Any Supervision
Yijiang Li, Bingyang Wang, Yijun Liang +3
On-policy (Self-)Distillation (OPD / OPSD) has shown strong potential for post-training large language models (LLMs). However, existing methods still rely heavily on external super…
Increasing Computation Resolves Conflicts in Vision Language Models
Bingyang Wang, Yijiang Li, Yitong Qiao +7
Cognitive control, the ability to coordinate competing information sources in pursuit of goals, is fundamental to intelligent behavior. We systematically investigate whether Vision…
A Very Big Video Reasoning Suite
Maijunxian Wang, Ruisi Wang, Juyi Lin +53
Rapid progress in video models has largely focused on visual quality, leaving their reasoning capabilities underexplored. Video reasoning grounds intelligence in spatiotemporally c…
Core Knowledge Deficits in Multi-Modal Language Models
Yijiang Li, Qingying Gao, Tianwei Zhao +8
While Multi-modal Large Language Models (MLLMs) demonstrate impressive abilities over high-level perception and reasoning, their robustness in the wild remains limited, often falli…