1 paper · 1 filter
Bin Chen, Xinzge Gao, Chuanrui Hu +3
Generative Reward Models (GRMs) provide greater flexibility than scalar reward models in capturing human preferences, but their effectiveness is limited by poor reasoning capabilit…