activity
20242026
collaborators

5 papers

cs.AI2026

A Decision-Theoretic Formalisation of Steganography With Applications to LLM Monitoring

Usman Anwar, Julianna Piskorz, David D. Baek +6

Large language models are beginning to show steganographic capabilities. Such capabilities could allow misaligned models to evade oversight mechanisms. Yet principled methods to de…

cs.LG2025

Understanding In-Context Learning of Linear Models in Transformers Through an Adversarial Lens

Usman Anwar, Johannes Von Oswald, Louis Kirsch +2

In this work, we make two contributions towards understanding of in-context learning of linear models by transformers. First, we investigate the adversarial robustness of in-contex…

cs.LG2025

Mitigating Goal Misgeneralization via Minimax Regret

Karim Abdel Sadek, Matthew Farrugia-Roberts, Usman Anwar +4

Safe generalization in reinforcement learning requires not only that a learned policy acts capably in new situations, but also that it uses its capabilities towards the pursuit of…

cs.LG2024

Learning to Forget using Hypernetworks

Jose Miguel Lara Rangel, Stefan Schoepf, Jack Foster +2

Machine unlearning is gaining increasing attention as a way to remove adversarial data poisoning attacks from already trained models and to comply with privacy and AI regulations.…

cs.LG2024

Noisy Zero-Shot Coordination: Breaking The Common Knowledge Assumption In Zero-Shot Coordination Games

Usman Anwar, Ashish Pandian, Jia Wan +2

Zero-shot coordination (ZSC) is a popular setting for studying the ability of reinforcement learning (RL) agents to coordinate with novel partners. Prior ZSC formulations assume th…