1 paper · 1 filter
Renhao Li, Jianhong Tu, Yang Su +6
Reward models (RMs) play a critical role in aligning large language models (LLMs) with human preferences. Yet in the domain of tool learning, the lack of RMs specifically designed…