1 citations · 1 across the 8 of their papers we have counts for
Showing cs.LGShow all
2 papers · 1 filter
cs.LG2026
Entropy Pacing Policy Optimization for Multi-Task Agentic Reinforcement Learning
Zetian Hu, Shunyu Liu, Junjie Zhang +4
Recent breakthroughs of Reinforcement Learning (RL) have highlighted its potential for complex agentic Large Language Model (LLM) tasks. However, existing efforts largely focus on…
cs.LG2025
Intra-Trajectory Consistency for Reward Modeling
Chaoyang Zhou, Shunyu Liu, Zengmao Wang +4
Reward models are critical for improving large language models (LLMs), particularly in reinforcement learning from human feedback (RLHF) or inference-time verification. Current rew…