1 citations · 1 across the 1 of their papers we have counts for
1 paper
Simon Storf, Rich Barton-Cooper, James Peters-Gill +1
Safe deployment of Large Language Model (LLM) agents in autonomous settings requires reliable oversight mechanisms. A central challenge is detecting scheming, where agents covertly…