3 citations · 3 across the 4 of their papers we have counts for
3 papers · 1 filter
Part I: Tricks or Traps? A Deep Dive into RL for LLM Reasoning
Zihe Liu, Jiashun Liu, Yancheng He +13
Reinforcement learning for LLM reasoning has rapidly emerged as a prominent research area, marked by a significant surge in related studies on both algorithmic innovations and prac…
Open RL Benchmark: Comprehensive Tracked Experiments for Reinforcement Learning
Shengyi Huang, Quentin Gallouédec, Florian Felten +30
In many Reinforcement Learning (RL) papers, learning curves are useful indicators to measure the effectiveness of RL algorithms. However, the complete raw data of the learning curv…
Reward Scale Robustness for Proximal Policy Optimization via DreamerV3 Tricks
Ryan Sullivan, Akarsh Kumar, Shengyi Huang +2
Most reinforcement learning methods rely heavily on dense, well-normalized environment rewards. DreamerV3 recently introduced a model-based method with a number of tricks that miti…