1 paper
Hanrui Luo, Shreyank N Gowda
Detecting jailbreak behaviour in large language models remains challenging, particularly when strongly aligned models produce harmful outputs only rarely. In this work, we present…