4 papers
Adaptive Probe-based Steering for Robust LLM Jailbreaking
Junxi Chen, Junhao Dong, Xiaohua Xie
Recent work has demonstrated the potential of contrastive steering for jailbreaking Large Language Models (LLMs). However, existing methods rely on limited and inherently biased co…
Releasing Inequality Phenomenon in -norm Adversarial Training via Input Gradient Distillation
Junxi Chen, Junhao Dong, Xiaohua Xie +1
Adversarial training (AT) is considered the most effective defense against adversarial attacks. However, a recent study revealed that \(\ell_{\infty}\)-norm adversarial training (\…
Mind the Trojan Horse: Image Prompt Adapter Enabling Scalable and Deceptive Jailbreaking
Junxi Chen, Junhao Dong, Xiaohua Xie
Recently, the Image Prompt Adapter (IP-Adapter) has been increasingly integrated into text-to-image diffusion models (T2I-DMs) to improve controllability. However, in this paper, w…
Survey on Adversarial Attack and Defense for Medical Image Analysis: Methods and Challenges
Junhao Dong, Junxi Chen, Xiaohua Xie +2
Deep learning techniques have achieved superior performance in computer-aided medical image analysis, yet they are still vulnerable to imperceptible adversarial attacks, resulting…