1 paper
Gerard Boxo, Aman Neelappa, Shivam Raval
White-box monitors are a popular technique for detecting potentially harmful behaviours in language models. While they perform well in general, their effectiveness in detecting tex…