8 citations · 13 across the 3 of their papers we have counts for
3 papers
cs.CV2023★ 2 cited
IFSeg: Image-free Semantic Segmentation via Vision-Language Model
Sukmin Yun, Seong Hyeon Park, Paul Hongsuck Seo +1
Vision-language (VL) pre-training has recently gained much attention for its transferability and flexibility in novel concepts (e.g., cross-modality transfer) across various visual…
cs.CV2022★ 8 cited
Time Is MattEr: Temporal Self-supervision for Video Transformers
Sukmin Yun, Jaehyung Kim, Dongyoon Han +3
Understanding temporal dynamics of video is an essential aspect of learning better video representations. Recently, transformer-based architectural designs have been extensively ex…
cs.CV2022★ 3 cited
Patch-level Representation Learning for Self-supervised Vision Transformers
Sukmin Yun, Hankook Lee, Jaehyung Kim +1
Recent self-supervised learning (SSL) methods have shown impressive results in learning visual representations from unlabeled images. This paper aims to improve their performance f…