2 papers
cs.AI2026
PsychJail: Exploring Psychological Jailbreaks via Multi-Turn Persuasion of LLM Policies
Zeyu Feng, Qingyu Wu, Yuzhe Luo +1
Large language models (LLMs) are increasingly deployed in education, healthcare, policy advising, and other interactive settings, where users engage them as sustained social interl…
cs.CL2026
Refusing Intent, Not Form: Wrapper-Based Intent-Group Supervision for LLM Safety
Ping Wu, Haibo Tong, Feifei Zhao +7
Safety tuning can improve harmful refusal, but models may learn surface-form shortcuts: wrapped harmful prompts bypass safety, while similarly wrapped benign prompts are over-refus…