164 citations · 499 across the 10 of their papers we have counts for
1 paper · 1 filter
Eric Wallace, Kai Xiao, Reimar Leike +3
Today's LLMs are susceptible to prompt injections, jailbreaks, and other attacks that allow adversaries to overwrite a model's original instructions with their own malicious prompt…