4 papers
Distill Where You Fail: Recovering Learning Signals of Negative RL-Groups from Adaptive Teacher Guidance
Zhuowen Han, Jinwei Xiao, Zhengxi Lu +9
Reinforcement learning with verifiable rewards (RLVR) has become a standard paradigm for post-training large language models (LLMs). While Group Relative Policy Optimization (GRPO)…
Finding the Evidence: Discovering Decision-Supporting Tokens for On-Policy Reasoning Distillation
Jinwei Xiao, Zhuowen Han, Yueqing Sun +6
On-policy distillation transfers reasoning ability through dense token-level supervision, yet the nature of the transferable signal remains unclear. We discover that reasoning chai…
Emergent Slow Thinking in LLMs as Inverse Tree Freezing
Sihan Hu, Xiansheng Cai, Yuan Huang +5
Reinforcement learning with verifiable rewards (RLVR) enables large language models to acquire slow, multi-step reasoning from sparse final-answer signals. We provide a statistical…
Inverse Knowledge Search over Verifiable Reasoning: Synthesizing a Scientific Encyclopedia from a Long Chains-of-Thought Knowledge Base
Yu Li, Yuan Huang, Tao Wang +19
Most scientific materials compress reasoning, presenting conclusions while omitting the derivational chains that justify them. This compression hinders verification by lacking expl…