1 citations · 1 across the 3 of their papers we have counts for
10 papers
Hierarchical Self-Supervised Adversarial Training for Robust Vision Models in Histopathology
Hashmat Shadab Malik, Shahina Kunhimon, Muzammal Naseer +2
Adversarial attacks pose significant challenges for vision models in critical fields like healthcare, where reliability is essential. Although adversarial training has been well st…
Robust-LLaVA: On the Effectiveness of Large-Scale Robust Image Encoders for Multi-modal Large Language Models
Hashmat Shadab Malik, Fahad Shamshad, Muzammal Naseer +3
Multi-modal Large Language Models (MLLMs) excel in vision-language tasks but remain vulnerable to visual adversarial perturbations that can induce hallucinations, manipulate respon…
Towards Evaluating the Robustness of Visual State Space Models
Hashmat Shadab Malik, Fahad Shamshad, Muzammal Naseer +3
Vision State Space Models (VSSMs), a novel architecture that combines the strengths of recurrent neural networks and latent variable models, have demonstrated remarkable performanc…
VANE-Bench: Video Anomaly Evaluation Benchmark for Conversational LMMs
Rohit Bharadwaj, Hanan Gani, Muzammal Naseer +2
The recent developments in Large Multi-modal Video Models (Video-LMMs) have significantly enhanced our ability to interpret and analyze video data. Despite their impressive capabil…
How Good is my Histopathology Vision-Language Foundation Model? A Holistic Benchmark
Roba Al Majzoub, Hashmat Malik, Muzammal Naseer +4
Recently, histopathology vision-language foundation models (VLMs) have gained popularity due to their enhanced performance and generalizability across different downstream tasks. H…
VURF: A General-purpose Reasoning and Self-refinement Framework for Video Understanding
Ahmad Mahmood, Ashmal Vayani, Muzammal Naseer +2
Recent studies have demonstrated the effectiveness of Large Language Models (LLMs) as reasoning modules that can deconstruct complex tasks into more manageable sub-tasks, particula…