5 papers
VideoMaMa: Mask-Guided Video Matting via Generative Prior
Sangbeom Lim, Seoung Wug Oh, Jiahui Huang +3
Generalizing video matting models to real-world videos remains a significant challenge due to the scarcity of labeled data. To address this, we present Video Mask-to-Matte Model (V…
Visual Representation Alignment for Multimodal Large Language Models
Heeji Yoon, Jaewoo Jung, Junwan Kim +10
Multimodal large language models (MLLMs) trained with visual instruction tuning have achieved strong performance across diverse tasks, yet they remain limited in vision-centric tas…
URECA: Unique Region Caption Anything
Sangbeom Lim, Junwan Kim, Heeji Yoon +2
Region-level captioning aims to generate natural language descriptions for specific image regions while highlighting their distinguishing features. However, existing methods strugg…
Referring Video Object Segmentation via Language-aligned Track Selection
Seongchan Kim, Woojeong Jin, Sangbeom Lim +3
Referring video object segmentation (RVOS) requires tracking and segmenting an object throughout a video according to a given natural language expression, demanding both complex mo…
Multi-Granularity Video Object Segmentation
Sangbeom Lim, Seongchan Kim, Seungjun An +3
Current benchmarks for video segmentation are limited to annotating only salient objects (i.e., foreground instances). Despite their impressive architectural designs, previous work…