3 citations · 6 across the 2 of their papers we have counts for
1 paper · 1 filter
Boyi Deng, Wenjie Wang, Fuli Feng +3
Large language models (LLMs) are susceptible to red teaming attacks, which can induce LLMs to generate harmful content. Previous research constructs attack prompts via manual or au…