1 paper
Lichen Li, Hengguang Zhou, Yijun Liang +2
Reward hacking in code generation, where models exploit evaluation loopholes to obtain high reward without correctly solving the intended task, poses a critical challenge for Reinf…