2 citations · 3 across the 10 of their papers we have counts for
7 papers · 1 filter
VisualOverload: Probing Visual Understanding of VLMs in Really Dense Scenes
Paul Gavrikov, Wei Lin, M. Jehanzeb Mirza +6
Is basic visual understanding really solved in state-of-the-art VLMs? We present VisualOverload, a slightly different visual question answering (VQA) benchmark comprising 2,720 que…
EFSA: Episodic Few-Shot Adaptation for Text-to-Image Retrieval
Muhammad Huzaifa, Yova Kementchedjhieva
Text-to-image retrieval is a critical task for managing diverse visual content, but common benchmarks for the task rely on small, single-domain datasets that fail to capture real-w…
Towards Energy-Efficiency by Navigating the Trilemma of Energy, Latency, and Accuracy
Boyuan Tian, Yihan Pang, Muhammad Huzaifa +2
Extended Reality (XR) enables immersive experiences through untethered headsets but suffers from stringent battery and resource constraints. Energy-efficient design is crucial to e…
Test-Time Low Rank Adaptation via Confidence Maximization for Zero-Shot Generalization of Vision-Language Models
Raza Imam, Hanan Gani, Muhammad Huzaifa +1
The conventional modus operandi for adapting pre-trained vision-language models (VLMs) during test-time involves tuning learnable prompts, ie, test-time prompt tuning. This paper i…
ObjectCompose: Evaluating Resilience of Vision-Based Models on Object-to-Background Compositional Changes
Hashmat Shadab Malik, Muhammad Huzaifa, Muzammal Naseer +2
Given the large-scale multi-modal training of recent vision-based models and their generalization capabilities, understanding the extent of their robustness is critical for their r…
Domain Adaptable Fine-Tune Distillation Framework For Advancing Farm Surveillance
Raza Imam, Muhammad Huzaifa, Nabil Mansour +2
In this study, we propose an automated framework for camel farm monitoring, introducing two key contributions: the Unified Auto-Annotation framework and the Fine-Tune Distillation…