1 paper
Jianxiang Zang, Yongda Wei, Ruxue Bai +5
Reliable reward models (RMs) are critical for ensuring the safe alignment of large language models (LLMs). However, current RM evaluation methods focus solely on preference percept…