6 citations · 7 across the 3 of their papers we have counts for
1 paper · 1 filter
Abhay Sheshadri, John Hughes, Julian Michael +4
Alignment faking in large language models presented a demonstration of Claude 3 Opus and Claude 3.5 Sonnet selectively complying with a helpful-only training objective to prevent m…