15 papers
Switch-Reasoner: Learn When to Think in Multitask Mixtures via Reinforcement Learning
Yiyang Fang, Pei Fu, Jinjie Li +7
Multimodal Large Language Models (MLLMs) often follow a fixed Think-then-Answer paradigm, which is inefficient in heterogeneous multitask settings because simple inputs may not req…
On the Geometry of On-Policy Distillation
Zhennan Shen, Yanshu Li, Qingyu Yin +6
On-policy distillation (OPD) is increasingly used to improve large language model reasoning, but its training dynamics remain poorly understood. We characterize the trajectory of O…
Reinforcement Learning from Denoising Feedback
Qi He, Huan Chen, Ya Guo +3
Policy loss estimation remains a fundamental and long-standing challenge in reinforcement learning (RL) for diffusion language models (DLMs). We introduce Reinforcement Learning fr…
On Stable Long-Form Generation: Benchmarking and Mitigating Length Volatility
Zhitao He, Haolin Yang, Rui Min +2
Large Language Models (LLMs) excel at long-context understanding but exhibit significant limitations in long-form generation. Existing studies primarily focus on single-generation…
Entropy Centroids as Intrinsic Rewards for Test-Time Scaling
Wenshuo Zhao, Qi Zhu, Xingshan Zeng +4
An effective way to scale up test-time compute of large language models is to sample multiple responses and then select the best one, as in Grok Heavy and Gemini Deep Think. Existi…
Code2Math: Can Your Code Agent Effectively Evolve Math Problems Through Exploration?
Dadi Guo, Yuejin Xie, Qingyu Liu +11
As large language models (LLMs) advance their mathematical capabilities toward the IMO and research level, the scarcity of challenging, high-quality problems has become a significa…