collaborators

15 papers

cs.CV2026

Switch-Reasoner: Learn When to Think in Multitask Mixtures via Reinforcement Learning

Yiyang Fang, Pei Fu, Jinjie Li +7

Multimodal Large Language Models (MLLMs) often follow a fixed Think-then-Answer paradigm, which is inefficient in heterogeneous multitask settings because simple inputs may not req…

cs.LG2026

On the Geometry of On-Policy Distillation

Zhennan Shen, Yanshu Li, Qingyu Yin +6

On-policy distillation (OPD) is increasingly used to improve large language model reasoning, but its training dynamics remain poorly understood. We characterize the trajectory of O…

cs.CL2026

Reinforcement Learning from Denoising Feedback

Qi He, Huan Chen, Ya Guo +3

Policy loss estimation remains a fundamental and long-standing challenge in reinforcement learning (RL) for diffusion language models (DLMs). We introduce Reinforcement Learning fr…

cs.CL2026

Code2Math: Can Your Code Agent Effectively Evolve Math Problems Through Exploration?

Dadi Guo, Yuejin Xie, Qingyu Liu +11

As large language models (LLMs) advance their mathematical capabilities toward the IMO and research level, the scarcity of challenging, high-quality problems has become a significa…

cs.CL2026

ClinTutor-R1: Advancing Scalable and Robust One-to-Many Alignment in Clinical Socratic Education

Zhitao He, Haolin Yang, Zeyu Qin +1

While Large Language Models (LLMs) have achieved remarkable success in dyadic (one-on-one) instruction, they face significant challenges in One-to-Many alignment, such as clinical…

cs.CL2026

MARS-SQL: A multi-agent reinforcement learning framework for Text-to-SQL

Haolin Yang, Jipeng Zhang, Zhitao He +2

Large Language Models (LLMs) often struggle with the precise logic and schema alignment required for complex Text-to-SQL tasks. While current methods rely heavily on static prompti…