most citedThe Landscape of Agentic Reinforcement Learning for LLMs: A Survey

1 citations · 2 across the 16 of their papers we have counts for

collaborators
Showing cs.AIShow all

7 papers · 1 filter

cs.AI2026

AgentPanel: Toward a New Paradigm for Human--AI Collaboration in Exploring Scientific Questions

Zhiyao Cui, Qianyi Wang, Haoyang Yan +26

Identifying promising scientific ideas remains an important challenge in research practice. Researchers commonly rely on small-group discussions or one-to-one interactions with a s…

cs.AI2026

SKT: Skill-Use Training at Scale via Verified Synthetic Data Generation

Zelin Tan, Yiqun Zhang, Hao Li +11

Agent skills have become an important mechanism for equipping language-model agents with reusable procedural knowledge. However, providing skills alone does not guarantee that curr…

cs.AI20261 cited

The Landscape of Agentic Reinforcement Learning for LLMs: A Survey

Guibin Zhang, Hejia Geng, Xiaohang Yu +22

The emergence of agentic reinforcement learning (Agentic RL) marks a paradigm shift from conventional reinforcement learning applied to large language models (LLM RL), reframing LL…

cs.AI2026

PAPO: Stabilizing Rubric Integration Training via Decoupled Advantage Normalization

Zelin Tan, Zhouliang Yu, Bohan Lin +9

We propose Process-Aware Policy Optimization (PAPO), a method that integrates process-level evaluation into Group Relative Policy Optimization (GRPO) through decoupled advantage no…

cs.AI2026

Embodied Science: Closing the Discovery Loop with Agentic Embodied AI

Xiang Zhuang, Chenyi Zhou, Kehua Feng +10

Artificial intelligence has demonstrated remarkable capability in predicting scientific properties, yet scientific discovery remains an inherently physical, long-horizon pursuit go…

cs.AI2026

VIKI-R: Coordinating Embodied Multi-Agent Cooperation via Reinforcement Learning

Li Kang, Xiufeng Song, Heng Zhou +6

Coordinating multiple embodied agents in dynamic environments remains a core challenge in artificial intelligence, requiring both perception-driven reasoning and scalable cooperati…