3 papers
cs.LG2026
Claw-R1: A Step-Level Data Middleware System for Agentic Reinforcement Learning
Daoyu Wang, Mingyue Cheng, Qingchuan Li +3
Agentic reinforcement learning (RL) has become an important post-training paradigm for turning LLMs from static chatbots into interactive agents, giving rise to representative appl…
cs.LG2026
SocraticPO: Policy Optimization via Interactive Guidance
Zirui Liu, Tingyue Pan, Jie Ouyang +8
Reinforcement learning (RL) for large language models usually supervises reasoning with scalar outcome rewards, such as binary correctness. Such rewards provide an optimization dir…
cs.CL2025
Agent-R1: A Unified and Modular Framework for Agentic Reinforcement Learning
Mingyue Cheng, Shuo Yu, Daoyu Wang +7
Large language models (LLMs) have rapidly evolved from single-turn text generators into the foundation of increasingly capable agents. As these agents take on more complex reasonin…