70 citations · 78 across the 4 of their papers we have counts for
4 papers
VideoMathQA: Benchmarking Mathematical Reasoning via Multimodal Understanding in Videos
Hanoona Rasheed, Abdelrahman Shaker, Anqi Tang +4
Mathematical reasoning in real-world video settings presents a fundamentally different challenge than in static images or text. It requires interpreting fine-grained visual informa…
VideoGPT+: Integrating Image and Video Encoders for Enhanced Video Understanding
Muhammad Maaz, Hanoona Rasheed, Salman Khan +1
Building on the advances of language models, Large Multimodal Models (LMMs) have contributed significant improvements in video understanding. While the current video LMMs utilize a…
PALO: A Polyglot Large Multimodal Model for 5B People
Muhammad Maaz, Hanoona Rasheed, Abdelrahman Shaker +6
In pursuit of more inclusive Vision-Language Models (VLMs), this study introduces a Large Multilingual Multimodal Model called PALO. PALO offers visual reasoning capabilities in 10…
Bridging the Gap between Object and Image-level Representations for Open-Vocabulary Detection
Hanoona Rasheed, Muhammad Maaz, Muhammad Uzair Khattak +2
Existing open-vocabulary object detectors typically enlarge their vocabulary sizes by leveraging different forms of weak supervision. This helps generalize to novel objects at infe…