8 papers
Branch2Skill: Efficient Skill Evolution Through Reasoning Trees
Yanwei Ren, Haotian Zhang, Likang Xiao +5
Skill evolution improves agent skills through feedback over time, with failed trajectories often providing informative signals by revealing incomplete or misleading behaviors. Howe…
SearchSwarm: Towards Delegation Intelligence in Agentic LLMs for Long-Horizon Deep Research
Xiaochong Lan, Pu Ning, Quan Chen +8
Large language models are increasingly expected to handle complex, long-horizon real-world tasks whose context demands can grow without bound, yet model context windows remain inhe…
Recycling Failures: Salvaging Exploration in RLVR via Fine-Grained Off-Policy Guidance
Yanwei Ren, Haotian Zhang, Likang Xiao +6
Reinforcement Learning from Verifiable Rewards (RLVR) has emerged as a powerful paradigm for enhancing the complex reasoning capabilities of Large Reasoning Models. However, standa…
IMAGINE: Integrating Multi-Agent System into One Model for Complex Reasoning and Planning
Xikai Zhang, Bo Wang, Likang Xiao +4
Although large language models (LLMs) have made significant strides across various tasks, they still face significant challenges in complex reasoning and planning. For example, eve…
Long-Horizon Visual Imitation Learning via Plan and Code Reflection
Quan Chen, Chenrui Shi, Qi Chen +6
Learning from long-horizon demonstrations with complex action sequences presents significant challenges for visual imitation learning, particularly in understanding temporal relati…
SPOGW: a Score-based Preference Optimization method via Group-Wise comparison for workflows
Yitong Cui, Liu Liu, Baosheng Yu +5
Large language models (LLMs) have exhibited significant capabilities in addressing challenging problems throughout various fields, often through the use of agentic workflows that a…