activity
20242026
most citedRobust-LLaVA: On the Effectiveness of Large-Scale Robust Image Encoders for Multi-modal Large Language Models

1 citations · 1 across the 8 of their papers we have counts for

collaborators

10 papers

cs.CV2026

ReACT-CLIP: Response-Aware Test-Time Defense for Vision--Language Models

Hashmat Shadab Malik, Toluwani Aremu, Samuele Poppi +2

Training-free test-time defenses offer a practical way to improve the adversarial robustness of CLIP-style vision--language models without modifying the pretrained model. However,…

cs.CV2026

CORTEX: A Structured Reasoning Benchmark for Trustworthy 3D Chest CT MLLMs

Hashmat Shadab Malik, Anees Ur Rehman Hashmi, Numan Saeed +3

Reasoning in multimodal large language models (MLLMs) has shown strong promise in medical imaging. However, this reasoning is usually free-form text judged only by its final answer…

cs.CV2026

Beyond False Stability: High-Noise Drift Gating for Test-Time Adversarial Defenses in Vision-Language Models

Hashmat Shadab Malik, Muzammal Naseer, Salman Khan

Vision-language models (VLMs) such as CLIP show strong zero-shot generalization but remain highly vulnerable to adversarial attacks. Adversarial training improves robustness but is…

cs.CV2026

Hierarchical Self-Supervised Adversarial Training for Robust Vision Models in Histopathology

Hashmat Shadab Malik, Shahina Kunhimon, Muzammal Naseer +2

Adversarial attacks pose significant challenges for vision models in critical fields like healthcare, where reliability is essential. Although adversarial training has been well st…

cs.CV20261 cited

Robust-LLaVA: On the Effectiveness of Large-Scale Robust Image Encoders for Multi-modal Large Language Models

Hashmat Shadab Malik, Fahad Shamshad, Muzammal Naseer +3

Multi-modal Large Language Models (MLLMs) excel in vision-language tasks but remain vulnerable to visual adversarial perturbations that can induce hallucinations, manipulate respon…

cs.CV2026

Towards Evaluating the Robustness of Visual State Space Models

Hashmat Shadab Malik, Fahad Shamshad, Muzammal Naseer +3

Vision State Space Models (VSSMs), a novel architecture that combines the strengths of recurrent neural networks and latent variable models, have demonstrated remarkable performanc…