6 papers
SkillForge: Evolving Verifiable Skills for Reinforcement Learning Agents
Shidong Yang, Ziyu Ma, Tongwen Huang +5
Large language model (LLM) agents are trained with reinforcement learning (RL) for complex decision-making tasks. However, most RL-trained agents remain episodic and cannot accumul…
DiDPO: Diff-in-Diff Policy Optimization for Coding Agent Training
Xucong Wang, Zhe Zhao, Liheng Yu +3
Reinforcement learning with Verifiable Reward (RLVR) has emerged as a powerful paradigm for training coding agents, where the execution feedback from compilation and tests provides…
ReSum: Synergizing LLM Reasoning and Summarization with Reinforcement Learning
Xucong Wang, Ziyu Ma, Yong Wang +5
Reinforcement Learning with Verifiable Rewards (RLVR) is a central technique for improving long-horizon reasoning in Large Language Models (LLMs). However, existing RLVR methods of…
APPO: Agentic Procedural Policy Optimization
Xucong Wang, Ziyu Ma, Yong Wang +5
Recent advances in agentic Reinforcement Learning (RL) have substantially improved the multi-turn tool-use capabilities of large language model agents. However, most existing metho…
Role-Agent: Bootstrapping LLM Agents via Dual-Role Evolution
Xucong Wang, Ziyu Ma, Shidong Yang +4
Although Large Language Model (LLM) agents have demonstrated strong performance on complex tasks, their learning is often limited by inefficient interaction feedback and static tra…
Rethinking Crystal Symmetry Prediction: A Decoupled Perspective
Liheng Yu, Zhe Zhao, Xucong Wang +2
Efficiently and accurately determining the symmetry is a crucial step in the structural analysis of crystalline materials. Existing methods usually mindlessly apply deep learning m…