Showing cs.AIShow all
2 papers · 1 filter
cs.AI2026
"Did you lie?" Evaluating Lie Detectors across Model Scale and Belief-Verified Model Organisms
Alan Cooney, David Africa, Geoffrey Irving
Robust lie detectors for language models could enable powerful techniques for auditing, monitoring, and post-hoc investigation of model behaviour, but evaluating them requires test…
cs.AI2026
Behavioural Analysis of Alignment Faking
Nathaniel Mitrani Hadida, Rhea Karty, David Williams-King +1
Alignment faking (AF) refers to a model strategically complying with a training objective to avoid behavioural modification while preserving its deployment preferences. Understandi…