3 papers
cs.CV2025
Video-CoM: Interactive Video Reasoning via Chain of Manipulations
Hanoona Rasheed, Mohammed Zumri, Muhammad Maaz +3
Recent multimodal large language models (MLLMs) have advanced video understanding, yet most still "think about videos" ie once a video is encoded, reasoning unfolds entirely in tex…
cs.CV2025
GeoPixel: Pixel Grounding Large Multimodal Model in Remote Sensing
Akashah Shabbir, Mohammed Zumri, Mohammed Bennamoun +2
Recent advances in large multimodal models (LMMs) have recognized fine-grained grounding as an imperative factor of visual understanding and dialogue. However, the benefits of such…
cs.CV2025
LlamaV-o1: Rethinking Step-by-step Visual Reasoning in LLMs
Omkar Thawakar, Dinura Dissanayake, Ketan More +12
Reasoning is a fundamental capability for solving complex multi-step problems, particularly in visual contexts where sequential step-wise understanding is essential. Existing appro…