Showing cs.LGShow all
3 papers · 1 filter
cs.LG2025
Token Hidden Reward: Steering Exploration-Exploitation in Group Relative Deep Reinforcement Learning
Wenlong Deng, Yi Ren, Yushu Li +4
Reinforcement learning with verifiable rewards has significantly advanced the reasoning capabilities of large language models, yet how to explicitly steer training toward explorati…
cs.LG2025
Efficient kernelized bandit algorithms via exploration distributions
Bingshan Hu, Zheng He, Danica J. Sutherland
We consider a kernelized bandit problem with a compact arm set and a fixed but unknown reward function with a finite norm in some Reproducing Kern…
cs.LG2025
On the Effect of Negative Gradient in Group Relative Deep Reinforcement Optimization
Wenlong Deng, Yi Ren, Muchen Li +3
Reinforcement learning (RL) has become popular in enhancing the reasoning capabilities of large language models (LLMs), with Group Relative Policy Optimization (GRPO) emerging as a…