1 paper · 1 filter
Amirhossein Farzam, Majid Behabahani, Mani Malek +2
Large language models (LLMs) remain vulnerable to jailbreak prompts that are fluent and semantically coherent, and therefore difficult to detect with standard heuristics. A particu…