Publications (116)
Towards Multimodal Domain Generalization with Few Labels
Hongzhao Li, Hao Dong, Hualei Wan +3
Multimodal models ideally should generalize to unseen domains while remaining data-efficient to reduce annotation costs. To this end, we introduce and study a new problem, Semi-Sup…
XPlainVerse: A Million-Scale Benchmark for Explainable Deepfake Detection
Abhijeet Narang, Kartik Kuckreja, Shreya Ghosh +3
As deepfake detection models increasingly produce natural language explanations, their reasoning often remains weakly grounded in visual artifacts, limiting reliability and user tr…
TLAC: Two-stage LMM Augmented CLIP for Zero-Shot Classification
Ans Munir, Faisal Z. Qureshi, Muhammad Haris Khan +1
Contrastive Language-Image Pretraining (CLIP) has shown impressive zero-shot performance on image classification. However, state-of-the-art methods often rely on fine-tuning techni…
GeoDisaster: Benchmarking Orchestrated Agents for Operational Disaster Geo-Intelligence
Maram Hasan, Aman Verma, Savitra Roy +5
Remote-sensing vision-language models (RS-VLMs) have advanced Earth-observation analysis toward visual interpretation and instruction-following, yet fall short of operational geo-i…
ThinkGeo: Evaluating Tool-Augmented Agents for Remote Sensing Tasks
Akashah Shabbir, Muhammad Akhtar Munir, Akshay Dudhane +6
Recent progress in large language models (LLMs) has enabled tool-augmented agents capable of solving complex real-world tasks through step-by-step reasoning. However, existing eval…
Towards PerSense++: Advancing Training-Free Personalized Instance Segmentation in Dense Images
Muhammad Ibraheem Siddiqui, Muhammad Umer Sheikh, Hassan Abid +2
Segmentation in dense visual scenes poses significant challenges due to occlusions, background clutter, and scale variations. To address this, we introduce PerSense, an end-to-end,…