4 citations · 4 across the 3 of their papers we have counts for
1 paper · 1 filter
Muhammad Ahmed Shah, Roshan Sharma, Hira Dhamyal +10
It has been shown that Large Language Model (LLM) alignments can be circumvented by appending specially crafted attack suffixes with harmful queries to elicit harmful responses. To…