2 citations · 2 across the 3 of their papers we have counts for
1 paper · 1 filter
Yiran Guo, Zhongjian Qiao, Yingqi Xie +5
Effective exploration is a key challenge in reinforcement learning for large language models: discovering high-quality trajectories within a limited sampling budget from the vast n…