collaborators

9 papers

cs.CL2026

Reading Between the Dots: Decoding Hidden Computation across Filler Tokens

Kaley Brauer, Claudio Mayrink Verdun, Samuel Marks

Frontier LLMs can perform multi-step reasoning over content-free filler tokens like dots or counting sequences, producing correct answers with no visible chain-of-thought (CoT). Th…

cs.AI2026

Model Spec Midtraining: Improving How Alignment Training Generalizes

Chloe Li, Nevan Wichers, Sara Price +2

Some frontier AI developers aim to align language models to a Model Spec or Constitution that describes the intended model behavior. However, standard alignment fine-tuning -- trai…

cs.LG2026

Censored LLMs as a Natural Testbed for Secret Knowledge Elicitation

Helena Casademunt, Bartosz Cywiński, Khoi Tran +3

Large language models sometimes produce false or misleading responses. Two approaches to this problem are honesty elicitation -- modifying prompts or weights so that the model answ…

cs.CL2026

Steering Evaluation-Aware Language Models to Act Like They Are Deployed

Tim Tian Hua, Andrew Qin, Samuel Marks +1

Large language models (LLMs) can sometimes detect when they are being evaluated and adjust their behavior to appear more aligned, compromising the reliability of safety evaluations…

cs.CL2026

Liars' Bench: Evaluating Lie Detectors for Language Models

Kieron Kretschmar, Walter Laurito, Sharan Maiya +1

Prior work has introduced techniques for detecting when large language models (LLMs) lie, that is, generate statements they believe are false. However, these techniques are typical…

cs.AI2025

Auditing Games for Sandbagging

Jordan Taylor, Sid Black, Dillon Bowen +10

Future AI systems could conceal their capabilities ('sandbagging') during evaluations, potentially misleading developers and auditors. We stress-tested sandbagging detection techni…