collaborators

6 papers

cs.CL2026

SkillForge: Evolving Verifiable Skills for Reinforcement Learning Agents

Shidong Yang, Ziyu Ma, Tongwen Huang +5

Large language model (LLM) agents are trained with reinforcement learning (RL) for complex decision-making tasks. However, most RL-trained agents remain episodic and cannot accumul…

cs.AI2026

DiDPO: Diff-in-Diff Policy Optimization for Coding Agent Training

Xucong Wang, Zhe Zhao, Liheng Yu +3

Reinforcement learning with Verifiable Reward (RLVR) has emerged as a powerful paradigm for training coding agents, where the execution feedback from compilation and tests provides…

cs.AI2026

ReSum: Synergizing LLM Reasoning and Summarization with Reinforcement Learning

Xucong Wang, Ziyu Ma, Yong Wang +5

Reinforcement Learning with Verifiable Rewards (RLVR) is a central technique for improving long-horizon reasoning in Large Language Models (LLMs). However, existing RLVR methods of…

cs.LG2026

APPO: Agentic Procedural Policy Optimization

Xucong Wang, Ziyu Ma, Yong Wang +5

Recent advances in agentic Reinforcement Learning (RL) have substantially improved the multi-turn tool-use capabilities of large language model agents. However, most existing metho…

cs.AI2026

Role-Agent: Bootstrapping LLM Agents via Dual-Role Evolution

Xucong Wang, Ziyu Ma, Shidong Yang +4

Although Large Language Model (LLM) agents have demonstrated strong performance on complex tasks, their learning is often limited by inefficient interaction feedback and static tra…

cs.LG2025

Rethinking Crystal Symmetry Prediction: A Decoupled Perspective

Liheng Yu, Zhe Zhao, Xucong Wang +2

Efficiently and accurately determining the symmetry is a crucial step in the structural analysis of crystalline materials. Existing methods usually mindlessly apply deep learning m…