papers

Publications (116)

cs.CV2026

Towards Multimodal Domain Generalization with Few Labels

Hongzhao Li, Hao Dong, Hualei Wan +3

Multimodal models ideally should generalize to unseen domains while remaining data-efficient to reduce annotation costs. To this end, we introduce and study a new problem, Semi-Sup…

cs.CV2026

XPlainVerse: A Million-Scale Benchmark for Explainable Deepfake Detection

Abhijeet Narang, Kartik Kuckreja, Shreya Ghosh +3

As deepfake detection models increasingly produce natural language explanations, their reasoning often remains weakly grounded in visual artifacts, limiting reliability and user tr…

cs.CV2025

TLAC: Two-stage LMM Augmented CLIP for Zero-Shot Classification

Ans Munir, Faisal Z. Qureshi, Muhammad Haris Khan +1

Contrastive Language-Image Pretraining (CLIP) has shown impressive zero-shot performance on image classification. However, state-of-the-art methods often rely on fine-tuning techni…

cs.CV2026

GeoDisaster: Benchmarking Orchestrated Agents for Operational Disaster Geo-Intelligence

Maram Hasan, Aman Verma, Savitra Roy +5

Remote-sensing vision-language models (RS-VLMs) have advanced Earth-observation analysis toward visual interpretation and instruction-following, yet fall short of operational geo-i…

cs.CV2026

ThinkGeo: Evaluating Tool-Augmented Agents for Remote Sensing Tasks

Akashah Shabbir, Muhammad Akhtar Munir, Akshay Dudhane +6

Recent progress in large language models (LLMs) has enabled tool-augmented agents capable of solving complex real-world tasks through step-by-step reasoning. However, existing eval…

cs.CV2025

Towards PerSense++: Advancing Training-Free Personalized Instance Segmentation in Dense Images

Muhammad Ibraheem Siddiqui, Muhammad Umer Sheikh, Hassan Abid +2

Segmentation in dense visual scenes poses significant challenges due to occlusions, background clutter, and scale variations. To address this, we introduce PerSense, an end-to-end,…