1 paper · 1 filter
Buyun Liang, Kwan Ho Ryan Chan, Darshan Thaker +2
Jailbreak attacks exploit specific prompts to bypass LLM safeguards, causing the LLM to generate harmful, inappropriate, and misaligned content. Current jailbreaking methods rely h…