1 citations · 1 across the 7 of their papers we have counts for
1 paper · 1 filter
Aman Goel, Xian Carrie Wu, Zhe Wang +2
Jailbreaking large-language models (LLMs) involves testing their robustness against adversarial prompts and evaluating their ability to withstand prompt attacks that could elicit u…