From the 1 of 6 linked papers with an AI index.
1 paper · 1 filter
Shenghua He, Tian Xia, Xuan Zhou +1
We study a common challenge in reinforcement learning for large language models (LLMs): the Zero-Reward Assumption, where non-terminal actions (i.e., intermediate token generations…