activity
20242026
collaborators
Showing cs.CVShow all

12 papers · 1 filter

cs.CV2026

PixFoundation: Are We Heading in the Right Direction with Pixel-level Vision Foundation Models?

Mennatullah Siam

Multiple works have emerged to push the boundaries of multi-modal large language models (MLLMs) towards pixel-level understanding. The current trend is to train MLLMs with pixel-le…

cs.CV2025

PixFoundation 2.0: Do Video Multi-Modal LLMs Use Motion in Visual Grounding?

Mennatullah Siam

Multi-modal large language models (MLLMs) have shown impressive generalization across tasks using images and text modalities. While their extension to video has enabled tasks such…

cs.CV2025

Multiscale Video Transformers for Class Agnostic Segmentation in Autonomous Driving

Leila Cheshmi, Mennatullah Siam

Ensuring safety in autonomous driving is a complex challenge requiring handling unknown objects and unforeseen driving scenarios. We develop multiscale video transformers capable o…

cs.CV2025

A Vision Centric Remote Sensing Benchmark

Abduljaleel Adejumo, Faegheh Yeganli, Clifford Broni-bediako +3

Multimodal Large Language Models (MLLMs) have achieved remarkable success in vision-language tasks but their remote sensing (RS) counterpart are relatively under explored. Unlike n…

cs.CV2025

The Power of One: A Single Example is All it Takes for Segmentation in VLMs

Mir Rayat Imtiaz Hossain, Mennatullah Siam, Leonid Sigal +1

Large-scale vision-language models (VLMs), trained on extensive datasets of image-text pairs, exhibit strong multimodal understanding capabilities by implicitly learning associatio…

cs.CV2024

Dynamics Based Neural Encoding with Inter-Intra Region Connectivity

Mai Gamal, Mohamed Rashad, Eman Ehab +2

Extensive literature has drawn comparisons between recordings of biological neurons in the brain and deep neural networks. This comparative analysis aims to advance and interpret d…