1 paper
Tiantian Zhang, Jierui Zuo, Michael Chen +1
Recent theory suggests that reward-model-first methods can be more sample-efficient than direct policy fitting when the reward function is statistically simpler than the induced po…