most citedBridging the Gap between Object and Image-level Representations for Open-Vocabulary Detection

70 citations · 89 across the 5 of their papers we have counts for

collaborators

17 papers

cs.LG2024

Efficient Localized Adaptation of Neural Weather Forecasting: A Case Study in the MENA Region

Muhammad Akhtar Munir, Fahad Shahbaz Khan, Salman Khan

Accurate weather and climate modeling is critical for both scientific advancement and safeguarding communities against environmental risks. Traditional approaches rely heavily on N…

cs.CV2024

Composed Video Retrieval via Enriched Context and Discriminative Embeddings

Omkar Thawakar, Muzammal Naseer, Rao Muhammad Anwer +4

Composed video retrieval (CoVR) is a challenging problem in computer vision which has recently highlighted the integration of modification text with visual queries for more sophist…

cs.CL20241 cited

PALO: A Polyglot Large Multimodal Model for 5B People

Muhammad Maaz, Hanoona Rasheed, Abdelrahman Shaker +6

In pursuit of more inclusive Vision-Language Models (VLMs), this study introduces a Large Multilingual Multimodal Model called PALO. PALO offers visual reasoning capabilities in 10…

cs.CV2024

Video-GroundingDINO: Towards Open-Vocabulary Spatio-Temporal Video Grounding

Syed Talal Wasim, Muzammal Naseer, Salman Khan +2

Video grounding aims to localize a spatio-temporal section in a video corresponding to an input text query. This paper addresses a critical limitation in current video grounding me…

cs.CV20234 cited

Cal-DETR: Calibrated Detection Transformer

Muhammad Akhtar Munir, Salman Khan, Muhammad Haris Khan +2

Albeit revealing impressive predictive performance for several computer vision tasks, deep neural networks (DNNs) are prone to making overconfident predictions. This limits the ado…

cs.CV2023

Multi-grained Temporal Prototype Learning for Few-shot Video Object Segmentation

Nian Liu, Kepan Nan, Wangbo Zhao +7

Few-Shot Video Object Segmentation (FSVOS) aims to segment objects in a query video with the same category defined by a few annotated support images. However, this task was seldom…