374 citations · 602 across the 19 of their papers we have counts for
1 paper · 1 filter
David Glukhov, Ziwen Han, Ilia Shumailov +2
Vulnerability of Frontier language models to misuse and jailbreaks has prompted the development of safety measures like filters and alignment training in an effort to ensure safety…