1 paper
Astha Mehta, Niruthiha Selvanayagam, Cedric Lam +10
An attacker can split a malicious goal into sub-prompts that each look benign on their own and only become harmful in combination. Existing LLM safety benchmarks evaluate prompts o…