4 papers · 1 filter
Branch2Skill: Efficient Skill Evolution Through Reasoning Trees
Yanwei Ren, Haotian Zhang, Likang Xiao +5
Skill evolution improves agent skills through feedback over time, with failed trajectories often providing informative signals by revealing incomplete or misleading behaviors. Howe…
Recycling Failures: Salvaging Exploration in RLVR via Fine-Grained Off-Policy Guidance
Yanwei Ren, Haotian Zhang, Likang Xiao +6
Reinforcement Learning from Verifiable Rewards (RLVR) has emerged as a powerful paradigm for enhancing the complex reasoning capabilities of Large Reasoning Models. However, standa…
ContextPRM: Leveraging Contextual Coherence for multi-domain Test-Time Scaling
Haotian Zhang, Liu Liu, Baosheng Yu +5
Process reward models (PRMs) have demonstrated significant efficacy in enhancing the mathematical reasoning capabilities of large language models (LLMs) by leveraging test-time sca…
SIGMA: Refining Large Language Model Reasoning via Sibling-Guided Monte Carlo Augmentation
Yanwei Ren, Haotian Zhang, Fuxiang Wu +4
Enhancing large language models by simply scaling up datasets has begun to yield diminishing returns, shifting the spotlight to data quality. Monte Carlo Tree Search (MCTS) has eme…