5 papers
When can we trust untrusted monitoring? A safety case sketch across collusion strategies
Nelson Gardner-Challis, Jonathan Bostock, Georgiy Kozhevnikov +4
AIs are increasingly being deployed with greater autonomy and capabilities, which increases the risk that a misaligned AI may be able to cause catastrophic harm. Untrusted monitori…
Hidden in Plain Text: Emergence & Mitigation of Steganographic Collusion in LLMs
Yohan Mathew, Ollie Matthews, Robert McCarthy +4
The rapid proliferation of frontier model agents promises significant societal advances but also raises concerns about systemic risks arising from unsafe interactions. Collusion to…
Studying Cross-cluster Modularity in Neural Networks
Satvik Golechha, Maheep Chaudhary, Joan Velja +2
An approach to improve neural network interpretability is via clusterability, i.e., splitting a model into disjoint clusters that can be studied independently. We define a measure…
'Explaining RL Decisions with Trajectories': A Reproducibility Study
Karim Abdel Sadek, Matteo Nulli, Joan Velja +1
This work investigates the reproducibility of the paper 'Explaining RL decisions with trajectories'. The original paper introduces a novel approach in explainable reinforcement lea…
Dynamic Vocabulary Pruning in Early-Exit LLMs
Jort Vincenti, Karim Abdel Sadek, Joan Velja +2
Increasing the size of large language models (LLMs) has been shown to lead to better performance. However, this comes at the cost of slower and more expensive inference. Early-exit…