From the 1 of 5 linked papers with an AI index.
5 papers
A False Average: Chain-of-Thought Monitors Collapse Where They Are the Only Defense
Shikhar Shiromani, Leo Richter
Chain-of-thought (CoT) monitoring is meant to catch the reward hacks that look clean in the actions and betray themselves only in the reasoning. We show that this is exactly where…
Plausible Deniability Guarantees for Whistleblowers
Leo Richter, Matt J. Kusner
The paper proposes formal privacy guarantees for whistleblowers by applying per-report (0,δ)-differential privacy to audit selection transcripts, using a reduction to private conti…
ContextBench: Modifying Contexts for Targeted Latent Activation
Robert Graham, Edward Stevinson, Leo Richter +3
Identifying inputs that trigger specific behaviours or latent features in language models could have a wide range of safety use cases. We investigate a class of methods capable of…
Agentic Uncertainty Reveals Agentic Overconfidence
Jean Kaddour, Srijan Patel, Gbètondji Dovonon +3
Can AI agents predict whether they will succeed at a task? We study agentic uncertainty by eliciting success probability estimates before, during, and after task execution. All res…
An Auditing Test To Detect Behavioral Shift in Language Models
Leo Richter, Xuanli He, Pasquale Minervini +1
As language models (LMs) approach human-level performance, a comprehensive understanding of their behavior becomes crucial. This includes evaluating capabilities, biases, task perf…