20 citations · 20 across the 2 of their papers we have counts for
2 papers
cs.AI2025★ 20 cited
Frontier Models are Capable of In-context Scheming
Alexander Meinke, Bronson Schoen, Jérémy Scheurer +3
Frontier models are increasingly trained and deployed as autonomous agent. One safety concern is that AI agents might covertly pursue misaligned goals, hiding their true capabiliti…
cs.CL2024
TracrBench: Generating Interpretability Testbeds with Large Language Models
Hannes Thurnherr, Jérémy Scheurer
Achieving a mechanistic understanding of transformer-based language models is an open challenge, especially due to their large number of parameters. Moreover, the lack of ground tr…