most citedWhat Do Self-Supervised Vision Transformers Learn?

16 citations · 28 across the 5 of their papers we have counts for

collaborators

5 papers

cs.CL2023

MPCHAT: Towards Multimodal Persona-Grounded Conversation

Jaewoo Ahn, Yeda Song, Sangdoo Yun +1

In order to build self-consistent personalized dialogue agents, previous research has mostly focused on textual persona that delivers personal facts or personalities. However, to f…

cs.CV202316 cited

What Do Self-Supervised Vision Transformers Learn?

Namuk Park, Wonjae Kim, Byeongho Heo +2

We present a comparative study on how and why contrastive learning (CL) and masked image modeling (MIM) differ in their representations and in their performance of downstream tasks…

cs.CV2023

Three Recipes for Better 3D Pseudo-GTs of 3D Human Mesh Estimation in the Wild

Gyeongsik Moon, Hongsuk Choi, Sanghyuk Chun +2

Recovering 3D human mesh in the wild is greatly challenging as in-the-wild (ITW) datasets provide only 2D pose ground truths (GTs). Recently, 3D pseudo-GTs have been widely used to…

cs.LG20227 cited

A Unified Analysis of Mixed Sample Data Augmentation: A Loss Function Perspective

Chanwoo Park, Sangdoo Yun, Sanghyuk Chun

We propose the first unified theoretical analysis of mixed sample data augmentation (MSDA), such as Mixup and CutMix. Our theoretical results show that regardless of the choice of…

cs.CV20225 cited

Exploring Temporally Dynamic Data Augmentation for Video Recognition

Taeoh Kim, Jinhyung Kim, Minho Shim +4

Data augmentation has recently emerged as an essential component of modern training recipes for visual recognition tasks. However, data augmentation for video recognition has been…