1 citations · 1 across the 1 of their papers we have counts for
1 paper
Asa Cooper Stickland, Alexander Lyzhov, Jacob Pfau +2
Language models (LMs) have been shown to behave unexpectedly post-deployment. For example, new jailbreaks continually arise, allowing model misuse, despite extensive red-teaming an…