91 citations · 114 across the 13 of their papers we have counts for
1 paper · 1 filter
Zixuan Shangguan, Yanjie Dong, Lanjun Wang +3
Large language models (LLMs) have demonstrated exceptional proficiency in language understanding. However, when LLMs align their outputs with deceptive and/or misleading prompts, t…