12 papers
OpenEarthAgent: A Unified Framework for Tool-Augmented Geospatial Agents
Akashah Shabbir, Muhammad Umer Sheikh, Muhammad Akhtar Munir +8
Recent progress in multimodal reasoning has enabled agents that interpret imagery, connect it with language, and execute structured analytical tasks. Extending these capabilities t…
Real-Time Source-Free Object Detection
Sairam VCR, Varun Gopal, Poornima Jain +2
Real-world detectors for autonomous driving, surveillance, and robotics must handle domain-shifts under strict latency and memory constraints, yet existing source-free object detec…
CountZES: Counting via Zero-Shot Exemplar Selection
Muhammad Ibraheem Siddiqui, Muhammad Haris Khan
Object counting in complex scenes is particularly challenging in the zero-shot (ZS) setting, where instances of unseen categories are counted using only a class name. Existing ZS c…
Pixels Don't Lie (But Your Detector Might): Bootstrapping MLLM-as-a-Judge for Trustworthy Deepfake Detection and Reasoning Supervision
Kartik Kuckreja, Parul Gupta, Muhammad Haris Khan +1
Deepfake detection models often generate natural-language explanations, yet their reasoning is frequently ungrounded in visual evidence, limiting reliability. Existing evaluations…
Towards Calibrating Prompt Tuning of Vision-Language Models
Ashshak Sharifdeen, Fahad Shamshad, Muhammad Akhtar Munir +6
Prompt tuning of large-scale vision-language models such as CLIP enables efficient task adaptation without updating model weights. However, it often leads to poor confidence calibr…
Foundation Model Priors Enhance Object Focus in Feature Space for Source-Free Object Detection
Sairam VCR, Rishabh Lalla, Aveen Dayal +4
Current state-of-the-art approaches in Source-Free Object Detection (SFOD) typically rely on Mean-Teacher self-labeling. However, domain shift often reduces the detector's ability…