4 papers
Token Hidden Reward: Steering Exploration-Exploitation in Group Relative Deep Reinforcement Learning
Wenlong Deng, Yi Ren, Yushu Li +4
Reinforcement learning with verifiable rewards has significantly advanced the reasoning capabilities of large language models, yet how to explicitly steer training toward explorati…
On the Hardness of Conditional Independence Testing In Practice
Zheng He, Roman Pogodin, Yazhe Li +3
Tests of conditional independence (CI) underpin a number of important problems in machine learning and statistics, from causal discovery to evaluation of predictor fairness and out…
Efficient kernelized bandit algorithms via exploration distributions
Bingshan Hu, Zheng He, Danica J. Sutherland
We consider a kernelized bandit problem with a compact arm set and a fixed but unknown reward function with a finite norm in some Reproducing Kern…
On the Effect of Negative Gradient in Group Relative Deep Reinforcement Optimization
Wenlong Deng, Yi Ren, Muchen Li +3
Reinforcement learning (RL) has become popular in enhancing the reasoning capabilities of large language models (LLMs), with Group Relative Policy Optimization (GRPO) emerging as a…