4 papers
Proxy Reward Internalization and Mechanistic Exploitation: A Learned Precursor to Reward Hacking and Its Generalization
Mohammad Beigi, Ming Jin, Lifu Huang
Reward hacking is usually studied after it becomes visible, once a model earns high proxy reward while failing the intended task. We instead study what proxy RL teaches before that…
IR: Contrastive Inverse Reinforcement Learning for Interpretable Detection and Mitigation of Reward Hacking
Mohammad Beigi, Ming Jin, Junshan Zhang +3
Reinforcement Learning from Human Feedback (RLHF) enables powerful LLM alignment but can introduce reward hacking - models exploit spurious correlations in proxy rewards without ge…
Adversarial Reward Auditing for Active Detection and Mitigation of Reward Hacking
Mohammad Beigi, Ming Jin, Junshan Zhang +2
Reinforcement Learning from Human Feedback (RLHF) remains vulnerable to reward hacking, where models exploit spurious correlations in learned reward models to achieve high scores w…
DiPT: Enhancing LLM reasoning through diversified perspective-taking
Hoang Anh Just, Mahavir Dabas, Lifu Huang +2
Existing work on improving language model reasoning typically explores a single solution path, which can be prone to errors. Inspired by perspective-taking in social studies, this…