3 papers
cs.LG2026
Mitigating Reward Hacking in RLHF via Advantage Sign Robustness
Shinnosuke Ono, Johannes Ackermann, Soichiro Nishimori +2
Reward models (RMs) used in reinforcement learning from human feedback (RLHF) are vulnerable to reward hacking: as the policy maximizes a learned proxy reward, true quality plateau…
cs.LG2025
Recursive Reward Aggregation
Yuting Tang, Yivan Zhang, Johannes Ackermann +3
In reinforcement learning (RL), aligning agent behavior with specific objectives typically requires careful design of the reward function, which can be challenging when the desired…
cs.LG2025
On Symmetric Losses for Robust Policy Optimization with Noisy Preferences
Soichiro Nishimori, Yu-Jie Zhang, Thanawat Lodkaew +1
Optimizing policies based on human preferences is key to aligning language models with human intent. This work focuses on reward modeling, a core component in reinforcement learnin…