Showing cs.CLShow all
3 papers · 1 filter
cs.CL2026
Reading Between the Dots: Decoding Hidden Computation across Filler Tokens
Kaley Brauer, Claudio Mayrink Verdun, Samuel Marks
Frontier LLMs can perform multi-step reasoning over content-free filler tokens like dots or counting sequences, producing correct answers with no visible chain-of-thought (CoT). Th…
cs.CL2026
Steering Evaluation-Aware Language Models to Act Like They Are Deployed
Tim Tian Hua, Andrew Qin, Samuel Marks +1
Large language models (LLMs) can sometimes detect when they are being evaluated and adjust their behavior to appear more aligned, compromising the reliability of safety evaluations…
cs.CL2026
Liars' Bench: Evaluating Lie Detectors for Language Models
Kieron Kretschmar, Walter Laurito, Sharan Maiya +1
Prior work has introduced techniques for detecting when large language models (LLMs) lie, that is, generate statements they believe are false. However, these techniques are typical…