10 citations · 34 across the 22 of their papers we have counts for
1 paper · 1 filter
Salman Rahman, Liwei Jiang, James Shiffer +7
Multi-turn interactions with language models (LMs) pose critical safety risks, as harmful intent can be strategically spread across exchanges. Yet, the vast majority of prior work…