1 citations · 1 across the 8 of their papers we have counts for
8 papers · 1 filter
Report of the 8th LSVOS Challenge: Complex and Multimodal Video Object Segmentation
Chang Liu, Henghui Ding, Lingyi Hong +36
This report summarizes the 8th Large-scale Video Object Segmentation (LSVOS) Challenge, held in conjunction with ECCV 2026. The challenge evaluates video segmentation in three comp…
SAM3Dual: A 3rd Place Solution to the MOSEv2 Track, 8th LSVOS Challenge
JeongRae Kim, Chaehyun Kim, Changwon Lim
We present SAM3Dual, our third-place solution to the MOSEv2 track of the 8th Large-scale Video Object Segmentation (LSVOS) Challenge at ECCV 2026. SAM3Dual is a training-free infer…
Probing and Leveraging Video Diffusion Transformer Features for Robust Point Tracking
Soowon Son, Honggyu An, Jisu Nam +7
Despite achieving strong results on standard benchmarks, current point tracking methods rely on feature backbones that are rarely designed with the temporal coherence needed for ro…
C3G: Learning Compact 3D Representations with 2K Gaussians
Honggyu An, Jaewoo Jung, Mungyeom Kim +10
Reconstructing and understanding 3D scenes from unposed sparse views in a feed-forward manner remains as a challenging task in 3D computer vision. Recent approaches use per-pixel 3…
Visual Representation Alignment for Multimodal Large Language Models
Heeji Yoon, Jaewoo Jung, Junwan Kim +10
Multimodal large language models (MLLMs) trained with visual instruction tuning have achieved strong performance across diverse tasks, yet they remain limited in vision-centric tas…
Seg4Diff: Unveiling Open-Vocabulary Segmentation in Text-to-Image Diffusion Transformers
Chaehyun Kim, Heeseong Shin, Eunbeen Hong +5
Text-to-image diffusion models excel at translating language prompts into photorealistic images by implicitly grounding textual concepts through their cross-modal attention mechani…