188 citations · 301 across the 43 of their papers we have counts for
1 paper · 2 filters
Dingwei Zhu, Zhiheng Xi, Shihan Dou +17
Training reinforcement learning (RL) systems in real-world environments remains challenging due to noisy supervision and poor out-of-domain (OOD) generalization, especially in LLM…