3 papers
cs.CV2026
VisualOverload: Probing Visual Understanding of VLMs in Really Dense Scenes
Paul Gavrikov, Wei Lin, M. Jehanzeb Mirza +6
Is basic visual understanding really solved in state-of-the-art VLMs? We present VisualOverload, a slightly different visual question answering (VQA) benchmark comprising 2,720 que…
cs.CV2025
EFSA: Episodic Few-Shot Adaptation for Text-to-Image Retrieval
Muhammad Huzaifa, Yova Kementchedjhieva
Text-to-image retrieval is a critical task for managing diverse visual content, but common benchmarks for the task rely on small, single-domain datasets that fail to capture real-w…
cs.RO2025
An Integrated Approach to Robotic Object Grasping and Manipulation
Owais Ahmed, M Huzaifa, M Areeb +1
In response to the growing challenges of manual labor and efficiency in warehouse operations, Amazon has embarked on a significant transformation by incorporating robotics to assis…