collaborators

9 papers

cs.AI2026

One Probe Won't Catch Them All: Towards Targeted Deception Detection

Vikram Natarajan, Devina Jain, Shivam Arora +2

Linear probes are a promising approach for monitoring AI systems for deceptive behaviour. Previous work has shown that a linear classifier trained on a contrastive instruction pair…

cs.CL2026

ABBEL: Learning Natural-Language Belief States for Memory-Efficient Interaction

Aly Lidayan, Jakob Bjorner, Satvik Golechha +2

As the time horizons of sequential decision-making tasks grow, keeping full interaction histories in model context becomes increasingly costly. Recent work reduces context lengths…

cs.AI2026

Among Us: A Sandbox for Measuring and Detecting Agentic Deception

Satvik Golechha, Adrià Garriga-Alonso

Prior studies on deception in language-based AI agents typically assess whether the agent produces a false statement about a topic, or makes a binary choice prompted by a goal, rat…

cs.AI2025

Auditing Games for Sandbagging

Jordan Taylor, Sid Black, Dillon Bowen +10

Future AI systems could conceal their capabilities ('sandbagging') during evaluations, potentially misleading developers and auditors. We stress-tested sandbagging detection techni…

cs.LG2025

Who's the Evil Twin? Differential Auditing for Undesired Behavior

Ishwar Balappanawar, Venkata Hasith Vattikuti, Greta Kintzley +2

Detecting hidden behaviors in neural networks poses a significant challenge due to minimal prior knowledge and potential adversarial obfuscation. We explore this problem by framing…

cs.LG2025

Training Neural Networks for Modularity aids Interpretability

Satvik Golechha, Dylan Cope, Nandi Schoots

An approach to improve network interpretability is via clusterability, i.e., splitting a model into disjoint clusters that can be studied independently. We find pretrained models t…