From the 1 of 19 linked papers with an AI index.
19 papers
From Scoring to Acting: Outcome-Verified Comparative Self-Distillation for LLM Agents
Xu Xia, Jinghua Piao, Min Yang +3
The paper introduces Outcome-Verified Comparative Self-Distillation (OVCSD), a method that lets large language model agents internalize skills by supervising them with teachers who…
A Data-Efficient Path to Multilingual LLMs: Language Expansion via Post-training PARAM Integration into Upcycled MoE
Hao Zhou, Tianhao Li, Zhijun Wang +6
Expanding Large Language Models~(LLMs) to new languages is a costly endeavor, demanding extensive Continued Pre-Training~(CPT) and data-intensive alignment. While recent data-free…
SkillMaster: Toward Autonomous Skill Mastery in LLM Agents
Min Yang, Jinghua Piao, Xu Xia +4
Skills provide an effective mechanism for improving LLM agents on complex tasks, yet in existing agent frameworks, their creation, refinement, and selection are typically governed…
How Does Alignment Enhance LLMs' Multilingual Capabilities? A Language Neurons Perspective
Shimao Zhang, Zhejian Lai, Xiang Liu +5
Multilingual Alignment is an effective and representative paradigm to enhance LLMs' multilingual capabilities, which transfers the capabilities from the high-resource languages to…
TAPO: Translation Augmented Policy Optimization for Multilingual Mathematical Reasoning
Xu Huang, Zhejian Lai, Zixian Huang +2
Large Language Models (LLMs) have demonstrated remarkable proficiency in English mathematical reasoning, yet a significant performance disparity persists in multilingual contexts,…
The First Impression Problem: Internal Bias Triggers Overthinking in Reasoning Models
Renfei Dang, Zhening Li, Shujian Huang +1
Reasoning models often exhibit overthinking, characterized by redundant reasoning steps. We identify \emph{internal bias} elicited by the input question as a key trigger of such be…