1 paper · 1 filter
Joshua Clymer, Garrett Baker, Rohan Subramani +1
As AI systems become more intelligent and their behavior becomes more challenging to assess, they may learn to game the flaws of human feedback instead of genuinely striving to fol…