4 papers
ReACT-CLIP: Response-Aware Test-Time Defense for Vision--Language Models
Hashmat Shadab Malik, Toluwani Aremu, Samuele Poppi +2
Training-free test-time defenses offer a practical way to improve the adversarial robustness of CLIP-style vision--language models without modifying the pretrained model. However,…
Beyond False Stability: High-Noise Drift Gating for Test-Time Adversarial Defenses in Vision-Language Models
Hashmat Shadab Malik, Muzammal Naseer, Salman Khan
Vision-language models (VLMs) such as CLIP show strong zero-shot generalization but remain highly vulnerable to adversarial attacks. Adversarial training improves robustness but is…
Exploring Adversarial Robustness and Safety Alignment in Multilingual Multi-Modal Large Language Models
Hashmat Shadab Malik, Muzammal Naseer, Salman Khan
Multimodal Large Language Models integrate visual perception into language reasoning, introducing a continuous attack surface susceptible to adversarial attacks. Prior work on MLLM…
Investigating Adversarial Robustness of Multi-modal Large Language Models
Hashmat Shadab Malik, Muzammal Naseer, Salman Khan
Multi-modal Large Language Models (MLLMs) achieve strong performance on vision-language tasks, but incorporating visual inputs through a vision encoder (e.g., CLIP) substantially e…