4 papers
Not All Modalities Are Equal: Instruction-Aware Gating for Multimodal Videos
Bonan Ding, Umair Nawaz, Ufaq Khan +5
Pre-trained video large language models excel at visual reasoning. However, they struggle when videos arrive with auxiliary streams, such as audio, depth map, or dense temporal evi…
OpenSeg-R: Improving Open-Vocabulary Segmentation via Step-by-Step Visual Reasoning
Zongyan Han, Jiale Cao, Shuo Chen +3
Open-Vocabulary Segmentation (OVS) has drawn increasing attention for its capacity to generalize segmentation beyond predefined categories. However, existing methods typically pred…
VideoGLaMM: A Large Multimodal Model for Pixel-Level Visual Grounding in Videos
Shehan Munasinghe, Hanan Gani, Wenqi Zhu +4
Fine-grained alignment between videos and text is challenging due to complex spatial and temporal dynamics in videos. Existing video-based Large Multimodal Models (LMMs) handle bas…
DB-SAM: Delving into High Quality Universal Medical Image Segmentation
Chao Qin, Jiale Cao, Huazhu Fu +2
Recently, the Segment Anything Model (SAM) has demonstrated promising segmentation capabilities in a variety of downstream segmentation tasks. However in the context of universal m…