1 paper
Nilanjana Das, Mathew Dawit, Aman Chadha +1
Jailbreak attacks expose a persistent failure mode in safety-aligned LLMs: models can be pushed into harmful behavior, but the internal representations enabling this shift remain p…