1 citations · 1 across the 1 of their papers we have counts for
1 paper
Zifan Wang, Christina Q. Knight, Jeremy Kritz +2
Large Language Model (LLM) safeguards, which implement request refusals, have become a widely adopted mitigation strategy against misuse. At the intersection of adversarial machine…