agent training 1large language models 1outcome verification 1reinforcement learning 1self-distillation 1
From the 1 of 3 linked papers with an AI index.
Showing cs.AIShow all
3 papers · 1 filter
cs.AI2026
From Scoring to Acting: Outcome-Verified Comparative Self-Distillation for LLM Agents
Xu Xia, Jinghua Piao, Min Yang +3
The paper introduces Outcome-Verified Comparative Self-Distillation (OVCSD), a method that lets large language model agents internalize skills by supervising them with teachers who…
cs.AI2026
SkillMaster: Toward Autonomous Skill Mastery in LLM Agents
Min Yang, Jinghua Piao, Xu Xia +4
Skills provide an effective mechanism for improving LLM agents on complex tasks, yet in existing agent frameworks, their creation, refinement, and selection are typically governed…
cs.AI2024
Partially Observable Mean Field Multi-Agent Reinforcement Learning Based on Graph-Attention
Min Yang, Guanjun Liu, Ziyuan Zhou
Traditional multi-agent reinforcement learning algorithms are difficultly applied in a large-scale multi-agent environment. The introduction of mean field theory has enhanced the s…