Showing cs.LGShow all
3 papers · 1 filter
cs.LG2026
The Dark Room in the Reward Channel: Dense Prediction Rewards Collapse GRPO-Trained LLM Agents -- and The Channel, Not the Content, Decides What Works
Yu Wang
Dense per-step supervision is the standard remedy for sparse-reward long-horizon LLM agents: reward the policy for predicting its next observation, which looks provably safe under…
cs.LG2025
Extending Test-Time Scaling: A 3D Perspective with Context, Batch, and Turn
Chao Yu, Qixin Tan, Jiaxuan Gao +7
Reasoning reinforcement learning (RL) has recently revealed a new scaling effect: test-time scaling. Thinking models such as R1 and o1 improve their reasoning accuracy at test time…
cs.LG2024
Reward-Robust RLHF in LLMs
Yuzi Yan, Xingzhou Lou, Jialian Li +6
As Large Language Models (LLMs) continue to progress toward more advanced forms of intelligence, Reinforcement Learning from Human Feedback (RLHF) is increasingly seen as a key pat…