backdoor attacks 1black-box attacks 1data poisoning 1representation alignment 1trigger generalization 1
From the 1 of 3 linked papers with an AI index.
3 papers
cs.CR2026
Lilith: Backdoor Generalization under Training-Inference Trigger Shift
Zhou Feng, Jiahao Chen, Chunyi Zhou +6
The paper studies how backdoor attacks can remain effective when the trigger used at inference time differs from the one seen during training, and proposes Lilith, a black‑box meth…
cs.CR2026
Compiling Activation Steering into Weights via Null-Space Constraints for Stealthy Backdoors
Rui Yin, Tianxu Han, Naen Xu +8
Safety-aligned large language models (LLMs) are increasingly deployed in real-world pipelines, yet this deployment also enlarges the supply-chain attack surface: adversaries can di…
cs.CL2026
"I See What You Did There": Can Large Vision-Language Models Understand Multimodal Puns?
Naen Xu, Jiayi Sheng, Changjiang Li +7
Puns are a common form of rhetorical wordplay that exploits polysemy and phonetic similarity to create humor. In multimodal puns, visual and textual elements synergize to ground th…