15 citations · 16 across the 9 of their papers we have counts for
1 paper · 2 filters
Zhicheng Zhang, Ziyan Wang, Yali Du +1
Exploration remains a key bottleneck for reinforcement learning (RL) post-training of large language models (LLMs), where sparse feedback and large action spaces can lead to premat…