1 paper · 1 filter
Simon Lermen, Mateusz Dziemian, Natalia Pérez-Campanero AntolÃn
We demonstrate how AI agents can coordinate to deceive oversight systems using automated interpretability of neural networks. Using sparse autoencoders (SAEs) as our experimental f…