10 papers
CAPO: Critic-Guided Action-Aligned Policy Optimization for Advancing LLM Agent Capabilities
Daoyu Wang, Qingchuan Li, Mingyue Cheng +6
Reinforcement learning (RL) has become a key technique for improving the agentic capabilities of large language models (LLMs). Although critic-free methods such as GRPO are increas…
EvoArena: Tracking Memory Evolution for Robust LLM Agents in Dynamic Environments
Jundong Xu, Qingchuan Li, Jiaying Wu +11
Large language model (LLM) agents have achieved strong performance on a wide range of benchmarks, yet most evaluations assume static environments. In contrast, real-world deploymen…
TabClaw: An Interactive and Self-Evolving Agent for Spreadsheet Manipulation and Table Reasoning
Mingyue Cheng, Shuo Yu, Daoyu Wang +5
Spreadsheets and tables are widely used representations for structured data analysis, but effective analysis still requires substantial manual effort and domain expertise. Recent l…
Claw-R1: A Step-Level Data Middleware System for Agentic Reinforcement Learning
Daoyu Wang, Mingyue Cheng, Qingchuan Li +3
Agentic reinforcement learning (RL) has become an important post-training paradigm for turning LLMs from static chatbots into interactive agents, giving rise to representative appl…
SocraticPO: Policy Optimization via Interactive Guidance
Zirui Liu, Jie Ouyang, Qi Liu +8
Reinforcement learning (RL) for large language models usually supervises reasoning with scalar outcome rewards, such as binary correctness. Such rewards provide an optimization dir…
Agent-R1: A Unified and Modular Framework for Agentic Reinforcement Learning
Mingyue Cheng, Shuo Yu, Daoyu Wang +7
Large language models (LLMs) have rapidly evolved from single-turn text generators into the foundation of increasingly capable agents. As these agents take on more complex reasonin…