5 citations · 6 across the 3 of their papers we have counts for
1 paper · 1 filter
Simon Lermen, Mateusz Dziemian, Natalia Pérez-Campanero Antolín
We demonstrate how AI agents can coordinate to deceive oversight systems using automated interpretability of neural networks. Using sparse autoencoders (SAEs) as our experimental f…