works on

From the 1 of 5 linked papers with an AI index.

collaborators

5 papers

cs.CR2026

A False Average: Chain-of-Thought Monitors Collapse Where They Are the Only Defense

Shikhar Shiromani, Leo Richter

Chain-of-thought (CoT) monitoring is meant to catch the reward hacks that look clean in the actions and betray themselves only in the reasoning. We show that this is exactly where…

cs.CR2026

Plausible Deniability Guarantees for Whistleblowers

Leo Richter, Matt J. Kusner

The paper proposes formal privacy guarantees for whistleblowers by applying per-report (0,δ)-differential privacy to audit selection transcripts, using a reduction to private conti…

cs.AI2026

ContextBench: Modifying Contexts for Targeted Latent Activation

Robert Graham, Edward Stevinson, Leo Richter +3

Identifying inputs that trigger specific behaviours or latent features in language models could have a wide range of safety use cases. We investigate a class of methods capable of…

cs.AI2026

Agentic Uncertainty Reveals Agentic Overconfidence

Jean Kaddour, Srijan Patel, Gbètondji Dovonon +3

Can AI agents predict whether they will succeed at a task? We study agentic uncertainty by eliciting success probability estimates before, during, and after task execution. All res…

cs.LG2025

An Auditing Test To Detect Behavioral Shift in Language Models

Leo Richter, Xuanli He, Pasquale Minervini +1

As language models (LMs) approach human-level performance, a comprehensive understanding of their behavior becomes crucial. This includes evaluating capabilities, biases, task perf…