activity
20242026
collaborators

5 papers

cs.AI2026

When can we trust untrusted monitoring? A safety case sketch across collusion strategies

Nelson Gardner-Challis, Jonathan Bostock, Georgiy Kozhevnikov +4

AIs are increasingly being deployed with greater autonomy and capabilities, which increases the risk that a misaligned AI may be able to cause catastrophic harm. Untrusted monitori…

cs.CL2025

Hidden in Plain Text: Emergence & Mitigation of Steganographic Collusion in LLMs

Yohan Mathew, Ollie Matthews, Robert McCarthy +4

The rapid proliferation of frontier model agents promises significant societal advances but also raises concerns about systemic risks arising from unsafe interactions. Collusion to…

cs.LG2025

Studying Cross-cluster Modularity in Neural Networks

Satvik Golechha, Maheep Chaudhary, Joan Velja +2

An approach to improve neural network interpretability is via clusterability, i.e., splitting a model into disjoint clusters that can be studied independently. We define a measure…

cs.AI2024

'Explaining RL Decisions with Trajectories': A Reproducibility Study

Karim Abdel Sadek, Matteo Nulli, Joan Velja +1

This work investigates the reproducibility of the paper 'Explaining RL decisions with trajectories'. The original paper introduces a novel approach in explainable reinforcement lea…

cs.CL2024

Dynamic Vocabulary Pruning in Early-Exit LLMs

Jort Vincenti, Karim Abdel Sadek, Joan Velja +2

Increasing the size of large language models (LLMs) has been shown to lead to better performance. However, this comes at the cost of slower and more expensive inference. Early-exit…