1 paper
Paria Rashidinejad, Yuandong Tian
Aligning AI systems with human preferences typically suffers from the infamous reward hacking problem, where optimization of an imperfect reward model leads to undesired behaviors.…