1 paper
Robert Graham, Edward Stevinson, Leo Richter +3
Identifying inputs that trigger specific behaviours or latent features in language models could have a wide range of safety use cases. We investigate a class of methods capable of…