activity
20222026
most citedCLIP model is an Efficient Continual Learner

14 citations · 26 across the 11 of their papers we have counts for

collaborators
Showing cs.CVShow all

9 papers · 1 filter

cs.CV2026

ReACT-CLIP: Response-Aware Test-Time Defense for Vision--Language Models

Hashmat Shadab Malik, Toluwani Aremu, Samuele Poppi +2

Training-free test-time defenses offer a practical way to improve the adversarial robustness of CLIP-style vision--language models without modifying the pretrained model. However,…

cs.CV2026

Beyond False Stability: High-Noise Drift Gating for Test-Time Adversarial Defenses in Vision-Language Models

Hashmat Shadab Malik, Muzammal Naseer, Salman Khan

Vision-language models (VLMs) such as CLIP show strong zero-shot generalization but remain highly vulnerable to adversarial attacks. Adversarial training improves robustness but is…

cs.CV2026

Investigating Adversarial Robustness of Multi-modal Large Language Models

Hashmat Shadab Malik, Muzammal Naseer, Salman Khan

Multi-modal Large Language Models (MLLMs) achieve strong performance on vision-language tasks, but incorporating visual inputs through a vision encoder (e.g., CLIP) substantially e…

cs.CV2023★ 5 cited

Sentence-level Prompts Benefit Composed Image Retrieval

Yang Bai, Xinxing Xu, Yong Liu +5

Composed image retrieval (CIR) is the task of retrieving specific images by using a query that involves both a reference image and a relative caption. Most existing CIR models adop…

cs.CV2023

3D Indoor Instance Segmentation in an Open-World

Mohamed El Amine Boudjoghra, Salwa K. Al Khatib, Jean Lahoud +4

Existing 3D instance segmentation methods typically assume that all semantic classes to be segmented would be available during training and only seen categories are segmented at in…

cs.CV2023★ 2 cited

Diverse Data Augmentation with Diffusions for Effective Test-time Prompt Tuning

Chun-Mei Feng, Kai Yu, Yong Liu +2

Benefiting from prompt tuning, recent years have witnessed the promising performance of pre-trained vision-language models, e.g., CLIP, on versatile downstream tasks. In this paper…