Showing cs.LGShow all
3 papers · 1 filter
cs.LG2026
RLCSD: Reinforcement Learning with Contrastive On-Policy Self-Distillation
Leyi Pan, Shuchang Tao, Yunpeng Zhai +5
On-policy self-distillation (OPSD) provides dense, token-level supervision for reasoning models by aligning a model's own distribution with the distribution it produces under privi…
cs.LG2026
Auto-Rubric: Learning From Implicit Weights to Explicit Rubrics for Reward Modeling
Lipeng Xie, Sen Huang, Zhuo Zhang +9
Conventional reward modeling relies on gradient descent over neural weights, creating opaque, data-hungry "black boxes." We propose a paradigm shift from implicit to explicit rewar…
cs.LG2025
AgentEvolver: Towards Efficient Self-Evolving Agent System
Yunpeng Zhai, Shuchang Tao, Cheng Chen +10
Autonomous agents powered by large language models (LLMs) have the potential to significantly enhance human productivity by reasoning, using tools, and executing complex tasks in d…