1 paper · 1 filter
Ely Hahami, Ishaan Sinha, Lavik Jain +2
Can large language models introspect, that is, accurately detect perturbations to their own internal states? We systematically investigate this question using activation steering i…