2 citations · 2 across the 3 of their papers we have counts for
1 paper · 1 filter
Ang Li, Yin Zhou, Vethavikashini Chithrra Raghuram +2
A high volume of recent ML security literature focuses on attacks against aligned large language models (LLMs). These attacks may extract private information or coerce the model in…