agent training 1large language models 1outcome verification 1reinforcement learning 1self-distillation 1
From the 1 of 22 linked papers with an AI index.
Showing cs.AIShow all
3 papers · 1 filter
cs.AI2026
From Scoring to Acting: Outcome-Verified Comparative Self-Distillation for LLM Agents
Xu Xia, Jinghua Piao, Min Yang +3
The paper introduces Outcome-Verified Comparative Self-Distillation (OVCSD), a method that lets large language model agents internalize skills by supervising them with teachers who…
cs.AI2026
SkillMaster: Toward Autonomous Skill Mastery in LLM Agents
Min Yang, Jinghua Piao, Xu Xia +4
Skills provide an effective mechanism for improving LLM agents on complex tasks, yet in existing agent frameworks, their creation, refinement, and selection are typically governed…
cs.AI2026
The First Impression Problem: Internal Bias Triggers Overthinking in Reasoning Models
Renfei Dang, Zhening Li, Shujian Huang +1
Reasoning models often exhibit overthinking, characterized by redundant reasoning steps. We identify \emph{internal bias} elicited by the input question as a key trigger of such be…