6 papers
Branch2Skill: Efficient Skill Evolution Through Reasoning Trees
Yanwei Ren, Haotian Zhang, Likang Xiao +5
Skill evolution improves agent skills through feedback over time, with failed trajectories often providing informative signals by revealing incomplete or misleading behaviors. Howe…
FBOS-RL: Feedback-Driven Bi-Objective Synergistic Reinforcement Learning
Xikai Zhang, Yongzhi Li, Likang Xiao +6
Reinforcement learning has become a cornerstone for aligning and unlocking the reasoning capabilities of large-scale models. At its core, the training loop of GRPO and its variants…
Recycling Failures: Salvaging Exploration in RLVR via Fine-Grained Off-Policy Guidance
Yanwei Ren, Haotian Zhang, Likang Xiao +6
Reinforcement Learning from Verifiable Rewards (RLVR) has emerged as a powerful paradigm for enhancing the complex reasoning capabilities of Large Reasoning Models. However, standa…
SPOGW: a Score-based Preference Optimization method via Group-Wise comparison for workflows
Yitong Cui, Liu Liu, Baosheng Yu +5
Large language models (LLMs) have exhibited significant capabilities in addressing challenging problems throughout various fields, often through the use of agentic workflows that a…
IMAGINE: Integrating Multi-Agent System into One Model for Complex Reasoning and Planning
Xikai Zhang, Bo Wang, Likang Xiao +4
Although large language models (LLMs) have made significant strides across various tasks, they still face significant challenges in complex reasoning and planning. For example, eve…
ContextPRM: Leveraging Contextual Coherence for multi-domain Test-Time Scaling
Haotian Zhang, Liu Liu, Baosheng Yu +5
Process reward models (PRMs) have demonstrated significant efficacy in enhancing the mathematical reasoning capabilities of large language models (LLMs) by leveraging test-time sca…