2 citations · 2 across the 1 of their papers we have counts for
1 paper
Nevan Wichers, Carson Denison, Ahmad Beirami
Red teaming is a common strategy for identifying weaknesses in generative language models (LMs), where adversarial prompts are produced that trigger an LM to generate unsafe respon…