6 papers
AgentPanel: Toward a New Paradigm for Human--AI Collaboration in Exploring Scientific Questions
Zhiyao Cui, Qianyi Wang, Haoyang Yan +26
Identifying promising scientific ideas remains an important challenge in research practice. Researchers commonly rely on small-group discussions or one-to-one interactions with a s…
SKT: Skill-Use Training at Scale via Verified Synthetic Data Generation
Zelin Tan, Yiqun Zhang, Hao Li +11
Agent skills have become an important mechanism for equipping language-model agents with reusable procedural knowledge. However, providing skills alone does not guarantee that curr…
Self-Harness: Harnesses That Improve Themselves
Hangfan Zhang, Shao Zhang, Kangcong Li +5
The performance of LLM-based agents is jointly shaped by their base models and the harnesses that mediate their interaction with the environment. Because different models exhibit d…
The Landscape of Agentic Reinforcement Learning for LLMs: A Survey
Guibin Zhang, Hejia Geng, Xiaohang Yu +22
The emergence of agentic reinforcement learning (Agentic RL) marks a paradigm shift from conventional reinforcement learning applied to large language models (LLM RL), reframing LL…
PAPO: Stabilizing Rubric Integration Training via Decoupled Advantage Normalization
Zelin Tan, Zhouliang Yu, Bohan Lin +9
We propose Process-Aware Policy Optimization (PAPO), a method that integrates process-level evaluation into Group Relative Policy Optimization (GRPO) through decoupled advantage no…
Single-Agent Scaling Fails Multi-Agent Intelligence: Towards Foundation Models with Native Multi-Agent Intelligence
Shuyue Hu, Haoyang Yan, Yiqun Zhang +3
Foundation models (FMs) are increasingly assuming the role of the ''brain'' of AI agents. While recent efforts have begun to equip FMs with native single-agent abilities -- such as…