4 citations · 4 across the 3 of their papers we have counts for
1 paper · 1 filter
Suyu Ge, Chunting Zhou, Rui Hou +5
Red-teaming is a common practice for mitigating unsafe behaviors in Large Language Models (LLMs), which involves thoroughly assessing LLMs to identify potential flaws and addressin…