activity
20242026
most citedTowards Open-Vocabulary Semantic Segmentation Without Semantic Labels

1 citations · 1 across the 8 of their papers we have counts for

collaborators
Showing cs.CVShow all

8 papers · 1 filter

cs.CV2026

Report of the 8th LSVOS Challenge: Complex and Multimodal Video Object Segmentation

Chang Liu, Henghui Ding, Lingyi Hong +36

This report summarizes the 8th Large-scale Video Object Segmentation (LSVOS) Challenge, held in conjunction with ECCV 2026. The challenge evaluates video segmentation in three comp…

cs.CV2026

SAM3Dual: A 3rd Place Solution to the MOSEv2 Track, 8th LSVOS Challenge

JeongRae Kim, Chaehyun Kim, Changwon Lim

We present SAM3Dual, our third-place solution to the MOSEv2 track of the 8th Large-scale Video Object Segmentation (LSVOS) Challenge at ECCV 2026. SAM3Dual is a training-free infer…

cs.CV2025

Probing and Leveraging Video Diffusion Transformer Features for Robust Point Tracking

Soowon Son, Honggyu An, Jisu Nam +7

Despite achieving strong results on standard benchmarks, current point tracking methods rely on feature backbones that are rarely designed with the temporal coherence needed for ro…

cs.CV2025

C3G: Learning Compact 3D Representations with 2K Gaussians

Honggyu An, Jaewoo Jung, Mungyeom Kim +10

Reconstructing and understanding 3D scenes from unposed sparse views in a feed-forward manner remains as a challenging task in 3D computer vision. Recent approaches use per-pixel 3…

cs.CV2025

Visual Representation Alignment for Multimodal Large Language Models

Heeji Yoon, Jaewoo Jung, Junwan Kim +10

Multimodal large language models (MLLMs) trained with visual instruction tuning have achieved strong performance across diverse tasks, yet they remain limited in vision-centric tas…

cs.CV2025

Seg4Diff: Unveiling Open-Vocabulary Segmentation in Text-to-Image Diffusion Transformers

Chaehyun Kim, Heeseong Shin, Eunbeen Hong +5

Text-to-image diffusion models excel at translating language prompts into photorealistic images by implicitly grounding textual concepts through their cross-modal attention mechani…