activity
20172024
most citedCharacter Region Awareness for Text Detection

58 citations · 229 across the 12 of their papers we have counts for

collaborators
Showing cs.CVShow all

26 papers · 1 filter

cs.CV2024

Rotary Position Embedding for Vision Transformer

Byeongho Heo, Song Park, Dongyoon Han +1

Rotary Position Embedding (RoPE) performs remarkably on language models, especially for length extrapolation of Transformers. However, the impacts of RoPE on computer vision domain…

cs.CV2023

Language-only Efficient Training of Zero-shot Composed Image Retrieval

Geonmo Gu, Sanghyuk Chun, Wonjae Kim +2

Composed image retrieval (CIR) task takes a composed query of image and text, aiming to search relative images for both conditions. Conventional CIR approaches need a training data…

cs.CV2023

Match me if you can: Semi-Supervised Semantic Correspondence Learning with Unpaired Images

Jiwon Kim, Byeongho Heo, Sangdoo Yun +2

Semantic correspondence methods have advanced to obtaining high-quality correspondences employing complicated networks, aiming to maximize the model capacity. However, despite the…

cs.CV2023

Masking meets Supervision: A Strong Learning Alliance

Byeongho Heo, Taekyung Kim, Sangdoo Yun +1

Pre-training with random masked inputs has emerged as a novel trend in self-supervised training. However, supervised learning still faces a challenge in adopting masking augmentati…

cs.CV202316 cited

What Do Self-Supervised Vision Transformers Learn?

Namuk Park, Wonjae Kim, Byeongho Heo +2

We present a comparative study on how and why contrastive learning (CL) and masked image modeling (MIM) differ in their representations and in their performance of downstream tasks…

cs.CV2023

Three Recipes for Better 3D Pseudo-GTs of 3D Human Mesh Estimation in the Wild

Gyeongsik Moon, Hongsuk Choi, Sanghyuk Chun +2

Recovering 3D human mesh in the wild is greatly challenging as in-the-wild (ITW) datasets provide only 2D pose ground truths (GTs). Recently, 3D pseudo-GTs have been widely used to…