6 papers · 1 filter
ADWIN: Adaptive Windows for Horizon-Aware On-Policy Distillation
Kun Liang, Chenming Tang, Clive Bai +3
On-policy distillation (OPD) transfers reasoning behavior by training a student on teacher feedback along student-generated trajectories, but standard full-rollout training ties ev…
RLVR Datasets and Where to Find Them: Tracing Data Lineage for Better Training Data
Hsiu-Yuan Huang, Weijie Liu, Chenming Tang +5
The proliferation of Reinforcement Learning from Verifiable Rewards (RLVR) datasets has exacerbated provenance collapse due to unclear lineage among existing datasets. To bridge th…
Democratizing Tool Learning with Environments Fully Simulated by a Free 8B Language Model
Chenming Tang, Hsiu-Yuan Huang, Weijie Liu +3
Reinforcement learning (RL) has become a prevalent paradigm for training tool calling agents, which typically requires online interactive environments. Existing approaches either r…
ORBIT: On-policy Exploration-Exploitation for Controllable Multi-Budget Reasoning
Kun Liang, Clive Bai, Xin Xu +5
Recent Large Reasoning Models (LRMs) achieve strong performance by leveraging long-form Chain-of-Thought (CoT) reasoning, but uniformly applying overlong reasoning at inference tim…
Do Not Step Into the Same River Twice: Learning to Reason from Trial and Error
Chenming Tang, Hsiu-Yuan Huang, Weijie Liu +3
Reinforcement learning with verifiable rewards (RLVR) has significantly boosted the reasoning capability of language models (LMs). However, existing RLVR approaches train LMs based…
Think Outside the Policy: In-Context Steered Policy Optimization
Hsiu-Yuan Huang, Chenming Tang, Weijie Liu +3
Existing Reinforcement Learning from Verifiable Rewards (RLVR) methods, such as Group Relative Policy Optimization (GRPO), have achieved remarkable progress in improving the reason…