9 citations · 9 across the 6 of their papers we have counts for
1 paper · 1 filter
Kexin Chen, Yi Liu, Haonan Zhang +3
As LLMs acquire stronger reasoning capabilities, deceptive behavior becomes an increasingly serious safety concern. Existing deception monitors either score visible transcripts or…