4 papers
LLM-Agnostic Semantic Representation Attack
Jiawei Lian, Jianhong Pan, Lefan Wang +4
Large Language Models (LLMs) increasingly employ alignment techniques to prevent harmful outputs. Despite these safeguards, attackers can circumvent them by crafting adversarial pr…
Semantic Representation Attack against Aligned Large Language Models
Jiawei Lian, Jianhong Pan, Lefan Wang +3
Large Language Models (LLMs) increasingly employ alignment techniques to prevent harmful outputs. Despite these safeguards, attackers can circumvent them by crafting prompts that i…
Revealing the Intrinsic Ethical Vulnerability of Aligned Large Language Models
Jiawei Lian, Jianhong Pan, Lefan Wang +3
Large language models (LLMs) are foundational explorations to artificial general intelligence, yet their alignment with human values via instruction tuning and preference learning…
PADetBench: Towards Benchmarking Physical Attacks against Object Detection
Jiawei Lian, Jianhong Pan, Lefan Wang +3
Physical attacks against object detection have gained increasing attention due to their significant practical implications. However, conducting physical experiments is extremely ti…