6 papers
OpenSkillEval: Automatically Auditing the Open Skill Ecosystem for LLM Agents
Jiahao Ying, Boxian Ai, Wei Tang +2
Skills, i.e., structured workflow instructions distilled for large language models (LLMs), are becoming an increasingly important mechanism for improving agent performance on real-…
In-Context Learning Operates as Concept Subspace Learning
Wei Tang, Xinyan Jiang, Fakhri Karray +1
Regression and Bayesian accounts of in-context learning (ICL) explain how demonstrations can induce predictors, while mechanistic analyses often identify compact activation directi…
SetPO: Set-Level Policy Optimization for Diversity-Preserving LLM Reasoning
Chenyi Li, Yuan Zhang, Bo Wang +4
Reinforcement learning with verifiable rewards has shown notable effectiveness in enhancing large language models (LLMs) reasoning performance, especially in mathematics tasks. How…
The Rise of Parameter Specialization for Knowledge Storage in Large Language Models
Yihuai Hong, Yiran Zhao, Wei Tang +3
Over time, a growing wave of large language models from various series has been introduced to the community. Researchers are striving to maximize the performance of language models…
Disentangling Language and Culture for Evaluating Multilingual Large Language Models
Jiahao Ying, Wei Tang, Yiran Zhao +3
This paper introduces a Dual Evaluation Framework to comprehensively assess the multilingual capabilities of LLMs. By decomposing the evaluation along the dimensions of linguistic…
EvoWiki: Evaluating LLMs on Evolving Knowledge
Wei Tang, Yixin Cao, Yang Deng +8
Knowledge utilization is a critical aspect of LLMs, and understanding how they adapt to evolving knowledge is essential for their effective deployment. However, existing benchmarks…