2 citations · 5 across the 30 of their papers we have counts for
1 paper · 2 filters
Ora Nova Fandina, Leshem Choshen, Eitan Farchi +3
Consider a scenario where a harmfulness evaluation metric intended to filter unsafe responses from a Large Language Model. When applied to individual harmful prompt-response pairs,…