2 citations · 3 across the 16 of their papers we have counts for
17 papers · 1 filter
Qwen-AgentWorld: Language World Models for General Agents
Yuxin Zuo, Zikai Xiao, Li Sheng +30
A world model predicts environment dynamics based on current observations and actions, serving as a core cognitive mechanism for reasoning and planning. In this work, we investigat…
ARES: Automated Rubric Synthesis for Scalable LLM Reinforcement Learning
Xiaoyuan Li, Keqin Bao, Moxin Li +5
Rubric-based rewards offer a promising way to extend reinforcement learning (RL) for large language models beyond tasks with automatically verifiable answers. However, scaling rubr…
Unified Data Selection for LLM Reasoning
Xiaoyuan Li, Yubo Ma, Chengpeng Li +6
Effectively training Large Language Models (LLMs) for complex, long-CoT reasoning is often bottlenecked by the need for massive high-quality reasoning data. Existing methods are ei…
SkillGraph: Skill-Augmented Reinforcement Learning for Agents via Evolving Skill Graphs
Xiaoyuan Li, Moxin Li, Keqin Bao +4
Skill libraries enable large language model agents to reuse experience from past interactions, but most existing libraries store skills as isolated entries and retrieve them only b…
On Predicting the Post-training Potential of Pre-trained LLMs
Xiaoyuan Li, Yubo Ma, Kexin Yang +5
The performance of Large Language Models (LLMs) on downstream tasks is fundamentally constrained by the capabilities acquired during pre-training. However, traditional benchmarks l…
Qwen-Scope: Turning Sparse Features into Development Tools for Large Language Models
Boyi Deng, Xu Wang, Yaoning Wang +15
Large language models have achieved remarkable capabilities across diverse tasks, yet their internal decision-making processes remain largely opaque, limiting our ability to inspec…