2 papers
cs.CV2026
Mobile-O: Unified Multimodal Understanding and Generation on Mobile Device
Abdelrahman Shaker, Ahmed Heakl, Jaseel Muhammad +8
Unified multimodal models can both understand and generate visual content within a single architecture. Existing models, however, remain data-hungry and too heavy for deployment on…
cs.CV2025
VideoGLaMM: A Large Multimodal Model for Pixel-Level Visual Grounding in Videos
Shehan Munasinghe, Hanan Gani, Wenqi Zhu +4
Fine-grained alignment between videos and text is challenging due to complex spatial and temporal dynamics in videos. Existing video-based Large Multimodal Models (LMMs) handle bas…