1 paper
Yunsheng Lu, Zijiang Yang, Licheng Pan +1
Reward models are central to aligning large language models, yet they often overfit to spurious cues such as response length and overly agreeable tone. Most prior work weakens thes…