Showing cs.LGShow all
2 papers · 1 filter
cs.LG2025
Rethinking Reward Model Evaluation Through the Lens of Reward Overoptimization
Sunghwan Kim, Dongjin Kang, Taeyoon Kwon +3
Reward models (RMs) play a crucial role in reinforcement learning from human feedback (RLHF), aligning model behavior with human preferences. However, existing benchmarks for rewar…
cs.LG2024
Evaluating Robustness of Reward Models for Mathematical Reasoning
Sunghwan Kim, Dongjin Kang, Taeyoon Kwon +4
Reward models are key in reinforcement learning from human feedback (RLHF) systems, aligning the model behavior with human preferences. Particularly in the math domain, there have…