collaborators

6 papers

cs.AI2026

Branch2Skill: Efficient Skill Evolution Through Reasoning Trees

Yanwei Ren, Haotian Zhang, Likang Xiao +5

Skill evolution improves agent skills through feedback over time, with failed trajectories often providing informative signals by revealing incomplete or misleading behaviors. Howe…

cs.LG2026

FBOS-RL: Feedback-Driven Bi-Objective Synergistic Reinforcement Learning

Xikai Zhang, Yongzhi Li, Likang Xiao +6

Reinforcement learning has become a cornerstone for aligning and unlocking the reasoning capabilities of large-scale models. At its core, the training loop of GRPO and its variants…

cs.AI2026

Recycling Failures: Salvaging Exploration in RLVR via Fine-Grained Off-Policy Guidance

Yanwei Ren, Haotian Zhang, Likang Xiao +6

Reinforcement Learning from Verifiable Rewards (RLVR) has emerged as a powerful paradigm for enhancing the complex reasoning capabilities of Large Reasoning Models. However, standa…

cs.AI2025

SPOGW: a Score-based Preference Optimization method via Group-Wise comparison for workflows

Yitong Cui, Liu Liu, Baosheng Yu +5

Large language models (LLMs) have exhibited significant capabilities in addressing challenging problems throughout various fields, often through the use of agentic workflows that a…

cs.AI2025

IMAGINE: Integrating Multi-Agent System into One Model for Complex Reasoning and Planning

Xikai Zhang, Bo Wang, Likang Xiao +4

Although large language models (LLMs) have made significant strides across various tasks, they still face significant challenges in complex reasoning and planning. For example, eve…

cs.AI2025

ContextPRM: Leveraging Contextual Coherence for multi-domain Test-Time Scaling

Haotian Zhang, Liu Liu, Baosheng Yu +5

Process reward models (PRMs) have demonstrated significant efficacy in enhancing the mathematical reasoning capabilities of large language models (LLMs) by leveraging test-time sca…