39 citations · 45 across the 3 of their papers we have counts for
1 paper · 1 filter
Robin Young
We propose an information-theoretic formalization of the distinction between two fundamental AI safety failure modes: deceptive alignment and goal drift. While both can lead to sys…