1 citations · 1 across the 10 of their papers we have counts for
Showing cs.LGShow all
2 papers · 1 filter
cs.LG2026
Start Classifying: Categorical Critics for LLM Reinforcement Learning
Zhijian Zhou, Long Li, Xuan Zhang +7
Proximal Policy Optimization (PPO) for large language models typically trains its critic by mean-squared-error (MSE) regression on scalar value targets. Although scalar MSE is stat…
cs.LG2025★ 1 cited
Learn the Ropes, Then Trust the Wins: Self-imitation with Progressive Exploration for Agentic Reinforcement Learning
Yulei Qin, Xiaoyu Tan, Zhengbao He +13
Reinforcement learning (RL) is the dominant paradigm for sharpening strategic tool use capabilities of LLMs on long-horizon, sparsely-rewarded agent tasks, yet it faces a fundament…