collaborators

10 papers

cs.CL2026

CAPO: Critic-Guided Action-Aligned Policy Optimization for Advancing LLM Agent Capabilities

Daoyu Wang, Qingchuan Li, Mingyue Cheng +6

Reinforcement learning (RL) has become a key technique for improving the agentic capabilities of large language models (LLMs). Although critic-free methods such as GRPO are increas…

cs.CL2026

EvoArena: Tracking Memory Evolution for Robust LLM Agents in Dynamic Environments

Jundong Xu, Qingchuan Li, Jiaying Wu +11

Large language model (LLM) agents have achieved strong performance on a wide range of benchmarks, yet most evaluations assume static environments. In contrast, real-world deploymen…

cs.CL2026

TabClaw: An Interactive and Self-Evolving Agent for Spreadsheet Manipulation and Table Reasoning

Mingyue Cheng, Shuo Yu, Daoyu Wang +5

Spreadsheets and tables are widely used representations for structured data analysis, but effective analysis still requires substantial manual effort and domain expertise. Recent l…

cs.LG2026

Claw-R1: A Step-Level Data Middleware System for Agentic Reinforcement Learning

Daoyu Wang, Mingyue Cheng, Qingchuan Li +3

Agentic reinforcement learning (RL) has become an important post-training paradigm for turning LLMs from static chatbots into interactive agents, giving rise to representative appl…

cs.LG2026

SocraticPO: Policy Optimization via Interactive Guidance

Zirui Liu, Jie Ouyang, Qi Liu +8

Reinforcement learning (RL) for large language models usually supervises reasoning with scalar outcome rewards, such as binary correctness. Such rewards provide an optimization dir…

cs.CL2026

Agent-R1: A Unified and Modular Framework for Agentic Reinforcement Learning

Mingyue Cheng, Shuo Yu, Daoyu Wang +7

Large language models (LLMs) have rapidly evolved from single-turn text generators into the foundation of increasingly capable agents. As these agents take on more complex reasonin…