4 papers
VisualOverload: Probing Visual Understanding of VLMs in Really Dense Scenes
Paul Gavrikov, Wei Lin, M. Jehanzeb Mirza +6
Is basic visual understanding really solved in state-of-the-art VLMs? We present VisualOverload, a slightly different visual question answering (VQA) benchmark comprising 2,720 que…
EFSA: Episodic Few-Shot Adaptation for Text-to-Image Retrieval
Muhammad Huzaifa, Yova Kementchedjhieva
Text-to-image retrieval is a critical task for managing diverse visual content, but common benchmarks for the task rely on small, single-domain datasets that fail to capture real-w…
An Integrated Approach to Robotic Object Grasping and Manipulation
Owais Ahmed, M Huzaifa, M Areeb +1
In response to the growing challenges of manual labor and efficiency in warehouse operations, Amazon has embarked on a significant transformation by incorporating robotics to assis…
ObjectCompose: Evaluating Resilience of Vision-Based Models on Object-to-Background Compositional Changes
Hashmat Shadab Malik, Muhammad Huzaifa, Muzammal Naseer +2
Given the large-scale multi-modal training of recent vision-based models and their generalization capabilities, understanding the extent of their robustness is critical for their r…