2 papers
cs.CV2026
CLIP Tricks You: Training-free Token Pruning for Efficient Pixel Grounding in Large VIsion-Language Models
Sangin Lee, Yukyung Choi
In large vision-language models, visual tokens typically constitute the majority of input tokens, leading to substantial computational overhead. To address this, recent studies hav…
cs.CV2026
Multi-Modal Guided Multi-Source Domain Adaptation for Object Detection
Sangin Lee, Seokjun Kwon, Jeongmin Shin +2
General object detection (OD) struggles to detect objects in the target domain that differ from the training distribution. To address this, recent studies demonstrate that training…