3 papers
cs.CL2025
Semantic Representation Attack against Aligned Large Language Models
Jiawei Lian, Jianhong Pan, Lefan Wang +3
Large Language Models (LLMs) increasingly employ alignment techniques to prevent harmful outputs. Despite these safeguards, attackers can circumvent them by crafting prompts that i…
cs.CL2025
Revealing the Intrinsic Ethical Vulnerability of Aligned Large Language Models
Jiawei Lian, Jianhong Pan, Lefan Wang +3
Large language models (LLMs) are foundational explorations to artificial general intelligence, yet their alignment with human values via instruction tuning and preference learning…
cs.CV2024
Attack Anything: Blind DNNs via Universal Background Adversarial Attack
Jiawei Lian, Shaohui Mei, Xiaofei Wang +5
It has been widely substantiated that deep neural networks (DNNs) are susceptible and vulnerable to adversarial perturbations. Existing studies mainly focus on performing attacks b…