9 papers
Reading Between the Dots: Decoding Hidden Computation across Filler Tokens
Kaley Brauer, Claudio Mayrink Verdun, Samuel Marks
Frontier LLMs can perform multi-step reasoning over content-free filler tokens like dots or counting sequences, producing correct answers with no visible chain-of-thought (CoT). Th…
Model Spec Midtraining: Improving How Alignment Training Generalizes
Chloe Li, Nevan Wichers, Sara Price +2
Some frontier AI developers aim to align language models to a Model Spec or Constitution that describes the intended model behavior. However, standard alignment fine-tuning -- trai…
Censored LLMs as a Natural Testbed for Secret Knowledge Elicitation
Helena Casademunt, Bartosz CywiÅski, Khoi Tran +3
Large language models sometimes produce false or misleading responses. Two approaches to this problem are honesty elicitation -- modifying prompts or weights so that the model answ…
Steering Evaluation-Aware Language Models to Act Like They Are Deployed
Tim Tian Hua, Andrew Qin, Samuel Marks +1
Large language models (LLMs) can sometimes detect when they are being evaluated and adjust their behavior to appear more aligned, compromising the reliability of safety evaluations…
Liars' Bench: Evaluating Lie Detectors for Language Models
Kieron Kretschmar, Walter Laurito, Sharan Maiya +1
Prior work has introduced techniques for detecting when large language models (LLMs) lie, that is, generate statements they believe are false. However, these techniques are typical…
Auditing Games for Sandbagging
Jordan Taylor, Sid Black, Dillon Bowen +10
Future AI systems could conceal their capabilities ('sandbagging') during evaluations, potentially misleading developers and auditors. We stress-tested sandbagging detection techni…