1 citations · 1 across the 1 of their papers we have counts for
1 paper
Jiaxin Guo, Zewen Chi, Li Dong +4
Reward models play a critical role in guiding large language models toward outputs that align with human expectations. However, an open challenge remains in effectively utilizing t…