2 papers
cs.CR2026
Robust Harmful Features Under Jailbreak Attacks: Mechanistic Evidence from Attention Head Specialization in Large Language Models
Yanchen Yin, Dongqi Han, Linghui Li
Jailbreak attacks bypass LLM safety alignment, yet their mechanisms remain poorly understood. We provide evidence that attacks do not comprehensively eliminate safety features, but…
cs.IR2024
Boosting the Targeted Transferability of Adversarial Examples via Salient Region & Weighted Feature Drop
Shanjun Xu, Linghui Li, Kaiguo Yuan +1
Deep neural networks can be vulnerable to adversarially crafted examples, presenting significant risks to practical applications. A prevalent approach for adversarial attacks relie…