6 citations · 12 across the 4 of their papers we have counts for
4 papers
Safety-Utility Conflicts Are Not Global: Surgical Alignment via Head-Level Diagnosis
Wang Cai, Yilin Wen, Jinchang Hou +5
Safety alignment in Large Language Models (LLMs) inherently presents a multi-objective optimization conflict, often accompanied by an unintended degradation of general capabilities…
Enhancing Transferability of Adversarial Examples with Spatial Momentum
Guoqiu Wang, Huanqian Yan, Xingxing Wei
Many adversarial attack methods achieve satisfactory attack success rates under the white-box setting, but they usually show poor transferability when attacking other DNN models. M…
Unrestricted Adversarial Attacks on ImageNet Competition
Yuefeng Chen, Xiaofeng Mao, Yuan He +34
Many works have investigated the adversarial attacks or defenses under the settings where a bounded and imperceptible perturbation can be added to the input. However in the real-wo…
Improving Adversarial Transferability with Gradient Refining
Guoqiu Wang, Huanqian Yan, Ying Guo +1
Deep neural networks are vulnerable to adversarial examples, which are crafted by adding human-imperceptible perturbations to original images. Most existing adversarial attack meth…