activity
20242026
collaborators

5 papers

cs.CV2026

Video-Oasis: Rethinking Evaluation of Video Understanding

Geuntaek Lim, Sungjune Park, Jaeyun Lee +5

The inherent complexity of video understanding makes it difficult to determine whether Video-LLM benchmark performance stems from visual perception, linguistic reasoning, or knowle…

cs.CV2026

CLIP Tricks You: Training-free Token Pruning for Efficient Pixel Grounding in Large VIsion-Language Models

Sangin Lee, Yukyung Choi

In large vision-language models, visual tokens typically constitute the majority of input tokens, leading to substantial computational overhead. To address this, recent studies hav…

cs.CV2026

Multi-Modal Guided Multi-Source Domain Adaptation for Object Detection

Sangin Lee, Seokjun Kwon, Jeongmin Shin +2

General object detection (OD) struggles to detect objects in the target domain that differ from the training distribution. To address this, recent studies demonstrate that training…

cs.CV2025

Boosting Cross-spectral Unsupervised Domain Adaptation for Thermal Semantic Segmentation

Seokjun Kwon, Jeongmin Shin, Namil Kim +2

In autonomous driving, thermal image semantic segmentation has emerged as a critical research area, owing to its ability to provide robust scene understanding under adverse visual…

cs.CV2024

Probabilistic Vision-Language Representation for Weakly Supervised Temporal Action Localization

Geuntaek Lim, Hyunwoo Kim, Joonsoo Kim +1

Weakly supervised temporal action localization (WTAL) aims to detect action instances in untrimmed videos using only video-level annotations. Since many existing works optimize WTA…