14 citations · 26 across the 11 of their papers we have counts for
9 papers · 1 filter
ReACT-CLIP: Response-Aware Test-Time Defense for Vision--Language Models
Hashmat Shadab Malik, Toluwani Aremu, Samuele Poppi +2
Training-free test-time defenses offer a practical way to improve the adversarial robustness of CLIP-style vision--language models without modifying the pretrained model. However,…
Beyond False Stability: High-Noise Drift Gating for Test-Time Adversarial Defenses in Vision-Language Models
Hashmat Shadab Malik, Muzammal Naseer, Salman Khan
Vision-language models (VLMs) such as CLIP show strong zero-shot generalization but remain highly vulnerable to adversarial attacks. Adversarial training improves robustness but is…
Investigating Adversarial Robustness of Multi-modal Large Language Models
Hashmat Shadab Malik, Muzammal Naseer, Salman Khan
Multi-modal Large Language Models (MLLMs) achieve strong performance on vision-language tasks, but incorporating visual inputs through a vision encoder (e.g., CLIP) substantially e…
Sentence-level Prompts Benefit Composed Image Retrieval
Yang Bai, Xinxing Xu, Yong Liu +5
Composed image retrieval (CIR) is the task of retrieving specific images by using a query that involves both a reference image and a relative caption. Most existing CIR models adop…
3D Indoor Instance Segmentation in an Open-World
Mohamed El Amine Boudjoghra, Salwa K. Al Khatib, Jean Lahoud +4
Existing 3D instance segmentation methods typically assume that all semantic classes to be segmented would be available during training and only seen categories are segmented at in…
Diverse Data Augmentation with Diffusions for Effective Test-time Prompt Tuning
Chun-Mei Feng, Kai Yu, Yong Liu +2
Benefiting from prompt tuning, recent years have witnessed the promising performance of pre-trained vision-language models, e.g., CLIP, on versatile downstream tasks. In this paper…