8 citations · 8 across the 1 of their papers we have counts for
1 paper · 1 filter
Eric Wallace, Kai Xiao, Reimar Leike +3
Today's LLMs are susceptible to prompt injections, jailbreaks, and other attacks that allow adversaries to overwrite a model's original instructions with their own malicious prompt…