1 citations · 1 across the 7 of their papers we have counts for
16 papers · 1 filter
Robust-LLaVA: On the Effectiveness of Large-Scale Robust Image Encoders for Multi-modal Large Language Models
Hashmat Shadab Malik, Fahad Shamshad, Muzammal Naseer +3
Multi-modal Large Language Models (MLLMs) excel in vision-language tasks but remain vulnerable to visual adversarial perturbations that can induce hallucinations, manipulate respon…
Towards Evaluating the Robustness of Visual State Space Models
Hashmat Shadab Malik, Fahad Shamshad, Muzammal Naseer +3
Vision State Space Models (VSSMs), a novel architecture that combines the strengths of recurrent neural networks and latent variable models, have demonstrated remarkable performanc…
SafeDiffusion-R1: Online Reward Steering for Safe Diffusion Post-Training
Komal Kumar, Ankan Deria, Abhishek Basu +3
Diffusion models have been widely studied for removing unsafe content learned during pre-training. Existing methods require expensive supervised data, either unsafe-text paired wit…
Towards Calibrating Prompt Tuning of Vision-Language Models
Ashshak Sharifdeen, Fahad Shamshad, Muhammad Akhtar Munir +6
Prompt tuning of large-scale vision-language models such as CLIP enables efficient task adaptation without updating model weights. However, it often leads to poor confidence calibr…
VFace: A Training-Free Approach for Diffusion-Based Video Face Swapping
Sanoojan Baliah, Yohan Abeysinghe, Rusiru Thushara +4
We present a training-free, plug-and-play method, namely VFace, for high-quality face swapping in videos. It can be seamlessly integrated with image-based face swapping approaches…
RAVEN: Erasing Invisible Watermarks via Novel View Synthesis
Fahad Shamshad, Nils Lukas, Karthik Nandakumar
Invisible watermarking has become a critical mechanism for authenticating AI-generated image content, with major platforms deploying watermarking schemes at scale. However, evaluat…