1 paper · 1 filter
Indraneil Paul, Goran GlavaÅ¡, Goran Glavaš +1
Reward models (RMs) have become an indispensable fixture of the language model (LM) post-training playbook, enabling policy alignment and test-time scaling. Research on the applica…