24 citations · 35 across the 3 of their papers we have counts for
1 paper · 1 filter
Ryan Greenblatt, Carson Denison, Benjamin Wright +17
We present a demonstration of a large language model engaging in alignment faking: selectively complying with its training objective in training to prevent modification of its beha…