1 paper
Kartik Garg, Shourya Mishra, Kartikeya Sinha +8
Alignment faking is a form of strategic deception in AI in which models selectively comply with training objectives when they infer that they are in training, while preserving diff…