3 papers
cs.CL2026
Towards Context-Invariant Safety Alignment for Large Language Models
Yixu Wang, Yang Yao, Xin Wang +4
Preference-based post-training aligns LLMs with human intent, yet safety behavior often remains brittle. A model may refuse a harmful request in a standard prompt but comply when t…
cs.CV2026
TAME: Test-Time Adversarial Prompt Tuning via Mixture-of-Experts for Vision-Language Models
Xin Wang, Yixu Wang, Jiaming Zhang +6
Large-scale pre-trained Vision-Language models (VLMs), such as CLIP, exhibit strong zero-shot generalization, yet remain highly vulnerable to imperceptible adversarial perturbation…
cs.CR2026
DarkLLM: Learning Language-Driven Adversarial Attacks with Large Language Models
Ye Sun, Xin Wang, Jiaming Zhang +7
While vision and multimodal foundation models underpin critical tasks from perception to complex reasoning, they remain highly vulnerable to adversarial attacks. However, tradition…