1 paper
Mark Russinovich, Blake Bullwinkel, Giorgio Severi +2
Language model safety is typically evaluated one interaction at a time. We show that a weaker, unaligned model can split a harmful task into benign-looking subproblems, consult a s…