collaborators

16 papers

cs.LG2026

SMOPD: Multi-Reward Reinforcement Learning via Specialize-and-Merge Online Policy Distillation

Wen Wang, Jiahua Bao, Tu Yongsiqi +8

We aim to improve model performance in multi-reward reinforcement learning training process. Existing Group reward-Decoupled Normalization Policy Optimization (GDPO) has mitigated…

cs.CL2026

Skill Self-Play: Pushing the Frontier of LLM Capability with Co-Evolving Skills

Siyuan Huang, Pengyu Cheng, Haotian Liu +10

LLM training is shifting from manual design and annotation to interaction-driven self-evolution. However, existing self-evolutionary methods face a fundamental dilemma between task…

cs.AI2026

From Passive Retrieval to Active Memory Navigation: Learning to Use Memory as a Structured Action Space

Yue Xu, Yutao Sun, Yihao Liu +7

Long-term user memory is essential for personalized conversational agents, yet many memory systems still expose memory through passive retrieval interfaces, making the model a cons…

cs.AI2026

Dynamo: Dynamic Skill-Tool Evolution for Vision-Language Agents

Yutao Sun, Yanting Miao, Hao-Xuan Ma +8

Improving vision-language models (VLMs) on visual reasoning typically requires retraining or hand-designed prompts and tools. We present Dynamo, a training-free framework that adap…

cs.LG2026

GDPO: Mitigating Multi-Reward Conflicts via Group-Dynamic reward-Decoupled Policy Optimization

Haotian Liu, Yihao Liu, Jingwei Ni +11

As LLMs advance, post-training reinforcement learning (RL) increasingly relies on multi-dimensional rewards to cultivate comprehensive capabilities. This shift demands new algorith…

cs.AI2026

Trace2Skill: Distill Trajectory-Local Lessons into Transferable Agent Skills

Jingwei Ni, Yihao Liu, Xinpeng Liu +7

Large Language Model (LLM) agents increasingly rely on domain-specific skills, yet manually authoring such skills does not scale, and skills generated purely from parametric knowle…