Showing cs.AIShow all
2 papers · 1 filter
cs.AI2026
Behavioural Analysis of Alignment Faking
Nathaniel Mitrani Hadida, Rhea Karty, David Williams-King +1
Alignment faking (AF) refers to a model strategically complying with a training objective to avoid behavioural modification while preserving its deployment preferences. Understandi…
cs.AI2026
Chain-of-thought obfuscation learned from output supervision can generalise to unseen tasks
Nathaniel Mitrani Hadida, Sassan Bhanji, Cameron Tice +1
Chain-of-thought (CoT) reasoning provides a significant performance uplift to LLMs by enabling planning, exploration, and deliberation of their actions. CoT is also a powerful tool…