1 citations · 1 across the 3 of their papers we have counts for
1 paper · 1 filter
Kartik Garg, Shourya Mishra, Kartikeya Sinha +8
Alignment faking is a form of strategic deception in AI in which models selectively comply with training objectives when they infer that they are in training, while preserving diff…