54 citations · 125 across the 25 of their papers we have counts for
1 paper · 2 filters
Atticus Wang, Iván Arcuschin, Arthur Conmy
Reward models are central to large language model (LLM) post-training. However, past work has shown that they can reward spurious or undesirable attributes such as length, format,…